These are RAG questions of the kind Amazon actually asks — the patterns reported from Amazon's AI-engineering rounds, where design trade-offs, scale and failure modes matter as much as definitions. Treat this page as a mock interview: say every answer out loud before revealing it. If one surprises you, the lesson behind it is linked at the bottom.
Amazon RAG concept questions
Design retrieval for a customer-facing assistant over 50 million product documents at 3,000 QPS with a p95 latency budget of 1.5 seconds. What dominates, and what do you cut?
Asked in
Explain the difference between a bi-encoder and a cross-encoder, and why RAG systems use both.
Asked in
How would you measure whether a change to your RAG system is an improvement, given 3,000 QPS of live traffic?
Asked in
Amazon RAG applied & hands-on questions
Hybrid search returns a vector ranking and a BM25 ranking with incomparable score scales. Implement Reciprocal Rank Fusion and explain the role of the constant k.
Asked in
How to use this page: Amazon rarely asks something you've never seen — they ask a standard RAG concept and then push one level deeper ("why?", "what would you do if..."). Master the concept in the RAG course lessons, and the follow-up stops being scary.
Keep practising: What Is RAG?, Chunking, Retrieval Techniques and Evaluating RAG cover what most Amazon RAG rounds test.

