Skip to content

Design a RAG system: a system design interview walkthrough

Design a question-answering service over private, changing documents. The hard parts are keeping retrieval inside the caller's permissions, publishing updates without mixed versions, and showing when the evidence does not support an answer. This is an interview design with explicit assumptions, not a claim about any provider's private architecture.

Define what a useful answer must guarantee

Assume employees ask questions about internal policies, product documentation, and support tickets. The system returns a short answer with links to supporting passages. Before choosing a vector database, agree on the contract:

  • Users see only documents their current identity is allowed to read.
  • Answers cite the source passages; when evidence is weak or conflicting, the system says so.
  • Document edits, deletions, and permission changes become effective within an agreed freshness window.
  • Ordinary text questions meet a stated latency target, including retrieval and generation.

Ask about corpus size, document formats, update and deletion rates, tenant count, query rate, languages, access-control rules, answer length, and the cost and latency budget. For this walkthrough, assume 10 million documents, about 2,000 tokens per document, mostly English text, and an explicitly versioned corpus. These are example inputs for reasoning, not universal sizing values.

Separate ingestion from query serving

Keep two paths. Ingestion converts source changes into a searchable, authorized index. Query serving authenticates the caller, retrieves evidence within that scope, and generates a cited answer. They share versioned document and permission metadata, but a slow embedding job should not block reads from the currently active index.

  1. 1

    Source connectors

    Read document changes and tombstones from authoritative systems. Attach a stable source ID and source version.

  2. 2

    Ingestion workers

    Parse and normalize content, preserve section boundaries, split it into chunks, and attach document, version, tenant, and access metadata.

  3. 3

    Index builder

    Generate embeddings in batches, write lexical and vector indexes, and record a durable status for every source version.

  4. 4

    Version publisher

    Make a new document version active only after all required chunks are queryable; retire the previous version as one logical change.

  5. 5

    Query API

    Authenticate the user, derive the permitted tenant and document scope, and enforce that scope before constructing model context.

  6. 6

    Retriever and reranker

    Combine semantic and lexical candidates when the corpus needs both, then spend extra ranking work only on the smaller candidate set.

  7. 7

    Answer service

    Send bounded, authorized passages to the model, validate returned citation IDs, and return source links or an abstention.

Make document updates safe to retry

Treat indexing as a versioned state transition, not a series of unrelated vector writes. A useful idempotency key is the source, document, and source version together. A retry for that identity must not create a second logical copy. Store chunk IDs that include the document version, plus the source URI and section path needed to build citations.

  1. Record the observed source version and enqueue it durably.
  2. Parse, normalize, and chunk the content while preserving headings and other useful boundaries.
  3. Write chunks and metadata to a pending version; retry failed batches without exposing partial content.
  4. Verify the expected chunks are indexed, then atomically switch the active-version pointer.
  5. Remove the old version and invalidate affected caches after the switch.

Deletion is its own indexed operation. A tombstone should prevent a lagging connector or retry from resurrecting deleted content. If the source system revokes a user's permission, the query path must stop returning the document before a slower full re-embedding job completes.

Retrieve evidence before asking the model

First derive an authorization scope from the verified identity. Apply it in the retrieval service, not in a browser parameter and not as an instruction to the language model. Check returned document IDs against that scope again before passing text to the model. A model cannot repair a permission leak after private text is already in its context.

For the baseline, embed the query and find nearby chunks, with metadata filters for tenant, document version, and access scope. Semantic retrieval handles paraphrases; lexical retrieval can rescue exact names, error codes, and policy terms that embeddings rank poorly. If evaluation shows the combination helps, merge the candidate lists, rerank a small set against the original query, then pass only the best passages within a fixed token budget. Reranking improves ordering; it cannot recover evidence the first-stage retriever never found.

Do not choose chunk size by habit. Test different boundaries against representative questions. Small chunks sharpen matches but can lose context; large chunks preserve context but consume more prompt tokens and may dilute relevance. Keep stable source IDs and offsets so citations still point to the exact passage after ranking.

In multi-turn conversations, resolve references such as “that policy” into a search query, but keep the original user question for answer generation. Query rewriting is an optimization to measure; it is not a substitute for retrieval evaluation.

Estimate index size before choosing a store

With the example assumptions, 10 million documents at 2,000 tokens each contain about 20 billion source tokens. At roughly 200 tokens per chunk, that is about 100 million chunks before overlap, metadata, or replicas. If an embedding has 1,536 dimensions stored at two bytes per dimension, its raw vector is about 3 KB; 100 million such vectors are roughly 300 GB of raw vector values alone.

This is a lower-bound illustration, not a storage quote. Chunk overlap increases vector count; the ANN structure, text index, metadata, source copies, and replication add storage. Measure the actual embedding dimensions, index implementation, compression, and recall target. Then compare the cost of a managed vector service with a relational or search engine that can also handle the required filtering and lexical search.

Query capacity has a different shape: count searches per second, candidates returned per query, rerank calls, prompt tokens, generated tokens, and their p95 latency. Use the capacity estimator for request, storage, and network arithmetic; benchmark embeddings, retrieval, reranking, and generation on the chosen workload.

Trace the failures that change the design

A document update is only partly indexed
Keep the new version pending and continue serving the prior complete version. Alert on indexing age; never publish a partial set as current.
A user loses access after a result was cached
Recheck authorization before returning answer text and citation links. Include the permission version in cache keys, or bypass shared answer caches when permissions cannot be represented safely.
A deleted document returns in search
Propagate a tombstone keyed by source and document version. Make connector checkpoints and retries respect the tombstone, and test deletion against both lexical and vector indexes.
Retrieval finds no adequate evidence
Return a clear abstention or ask a clarifying question. Do not let the model fill the gap from general training data while presenting the answer as document-backed.
Two current sources disagree
Return both citations and state the conflict, or prefer an explicitly authoritative and newer source if the product contract defines that policy. Do not silently collapse the disagreement.
The embedding or model provider is slow
Bound retries and queues separately for ingestion and query traffic. Keep reads on the active index during ingestion delays; apply a timeout or a documented degraded response to queries.

Evaluate retrieval and answers separately

Build a versioned test set from real question shapes, expected source passages, permission scopes, and expected abstentions. Evaluate retrieval before generation: recall at k asks whether required evidence appears in the retrieved set; precision and ranking metrics show how much irrelevant context is included and where it appears. Slice results by document type, tenant, query length, and update age.

Evaluate answers separately for claim support, correctness against a trusted reference, citation accuracy, and appropriate abstention. Include adversarial access tests: a user must not retrieve a document by guessing its title, URL, ID, or another tenant's metadata. Track p50 and p95 latency and cost for each stage, so retrieval regressions are distinguishable from model delays. A high answer score cannot compensate for a permission failure.

The Cohere RAG evaluation guide separates retrieval metrics from claim-level generation checks. The Cohere RAG documentation also notes that retrieved material can itself be outdated or wrong, so grounding does not guarantee a correct answer.

A concise interview answer

I would first agree on corpus size, freshness, latency, and who may read each document. I would keep ingestion separate from query serving, version every source, and publish a version only after its chunks are indexed. At query time I would derive the access scope from the authenticated identity, retrieve semantic and lexical candidates within that scope, rerank a bounded set, and generate only from those passages with validated citations. If the evidence is weak or conflicting, the system abstains or says what conflicts. I would test retrieval recall, grounded claims, citation accuracy, permission isolation, update lag, latency, and cost as separate signals.

Primary sources and scope

  • Cohere RAG documentation describes retrieval, reranking, generation with citations, and caveats about source quality. This walkthrough uses general mechanisms, not a vendor-specific architecture.
  • Cohere RAG output evaluation demonstrates retrieval precision, recall, and ranking metrics, then evaluates generated claims separately.
  • OpenAI vector-store file batches documents batch indexing, per-file attributes, and configurable chunking in one hosted retrieval API.

RAG system design interview questions

How do you design a RAG system in a system design interview?
Start with the answer contract, corpus size, freshness target, latency, and access rules. Separate document ingestion from query serving. Version and chunk source documents, attach authorization metadata, retrieve candidates using lexical and semantic search, optionally rerank them, then generate only from authorized evidence with citations or abstention. Measure retrieval quality and grounded answer quality separately.
How should a RAG system enforce document permissions?
Derive the user and tenant scope from a trusted identity, apply that scope inside the retrieval service before documents enter the prompt, and recheck authorization before returning citations. Do not trust a client-supplied tenant filter or rely on the language model to hide unauthorized text.
Does retrieval-augmented generation prevent hallucinations?
No. Retrieved context can be stale, irrelevant, incomplete, or wrong, and a model can still make unsupported claims. Return source-linked citations, measure whether claims are supported, and abstain when evidence is missing or conflicting.
How do you keep RAG answers fresh when documents change?
Track immutable document versions and indexing state. Process updates idempotently, publish a new version only after its chunks are indexed, and make the old version unavailable after the switch. Deletions and permission revocations must invalidate retrieval and caches promptly.
Can I practise this design on SysGeeks?
Use the walkthrough as a reference, then rehearse the scope, retrieval trade-offs, access boundaries, and failure cases in a timed system design mock interview.

Now explain how a permission revocation takes effect while an old answer is cached.