Skip to content

AI system design interview practice

Practise three AI system design problems: serving language models, answering from private documents, and letting an agent take actions. The same interview framework applies, but token workloads, retrieval quality, permissions, and model-driven side effects change what a strong design must guarantee.

Choose a system to design

  1. 1

    Design an LLM inference service

    Estimate prompt and generated-token throughput, schedule GPU work, stream within a latency target, and handle overload or worker failure.

    Main design pressure: Throughput and latency.

  2. 2

    Design a RAG system for private documents

    Separate ingestion from query serving, keep retrieval inside the caller's permissions, publish fresh versions safely, and return evidence-backed answers.

    Main design pressure: Retrieval, permissions, and freshness.

  3. 3

    Design a customer-support AI agent

    Bound tool access, persist runs and actions, make external effects safe to retry, and decide what to do when the model or a tool fails.

    Main design pressure: Authority and side effects.

What changes when a system uses AI?

Capacity is more than requests per second
Prompt length, generated tokens, context size, and concurrent generations shape inference cost and latency. Estimate them separately, then benchmark the model and hardware instead of guessing a GPU count.
Quality needs its own measurement
Retrieval can be stale or irrelevant, and a model can make unsupported claims. Measure retrieval quality separately from answer quality; use citations or abstention when the available evidence is weak.
Permissions must hold before data reaches the model
Derive access scope from trusted identity and enforce it in application code before building model context. A prompt instruction is not an authorization boundary.
Model decisions and real-world actions need different controls
Keep policy checks, durable action state, and idempotency in deterministic services. Give the model only the tools and arguments the application has authorized.

A repeatable way to answer

  1. 1. Define the contract. Identify who uses the system, what the model may do, what quality means, and which actions must never happen without authorization.
  2. 2. Put numbers on the workload. Estimate request rate, token and context distributions, corpus size or update rate, latency targets, and cost limits.
  3. 3. Separate the paths. Draw latency-sensitive serving separately from ingestion and evaluation, and keep orchestration separate from permission checks and external writes.
  4. 4. Trace a failure. Work through overload, stale data, model or worker failure, and uncertain tool outcomes. Explain what the caller sees and how the system recovers safely.
  5. 5. Explain how you will know it works. Pick measures for latency and cost alongside task success, retrieval quality, or action correctness.

AI system design interview questions

What is different about an AI system design interview?
The core interview skills stay the same: clarify requirements, estimate capacity, compare designs, and trace failures. AI systems add constraints such as token-length distributions, model quality, retrieval freshness, evaluation, and the boundary between model suggestions and authorized actions.
Should I memorize an AI system design architecture?
No. Start from the task, quality target, data permissions, latency, and cost. Then choose the simplest architecture that meets them and explain how it behaves when the model, index, GPU worker, or tool is unavailable.
Are these questions asked by specific companies?
No. These are original practice exercises based on public engineering material and documented mechanisms. They do not claim to reproduce any employer's interview questions or private architecture.

These are practice designs, not a list of employer interview questions. The walkthroughs explain their assumptions and link to the public engineering sources behind the design choices.