惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Secure Thoughts
月光博客
月光博客
Y
Y Combinator Blog
量子位
J
Java Code Geeks
The GitHub Blog
The GitHub Blog
MyScale Blog
MyScale Blog
aimingoo的专栏
aimingoo的专栏
Microsoft Azure Blog
Microsoft Azure Blog
Apple Machine Learning Research
Apple Machine Learning Research
博客园_首页
罗磊的独立博客
Google DeepMind News
Google DeepMind News
大猫的无限游戏
大猫的无限游戏
M
MIT News - Artificial intelligence
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
Visual Studio Blog
S
Schneier on Security
S
Security Affairs
Project Zero
Project Zero
L
LINUX DO - 热门话题
H
Hacker News: Front Page
Google Online Security Blog
Google Online Security Blog
L
Lohrmann on Cybersecurity
Latest news
Latest news
P
Palo Alto Networks Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
www.infosecurity-magazine.com
www.infosecurity-magazine.com
MongoDB | Blog
MongoDB | Blog
Blog — PlanetScale
Blog — PlanetScale
The Last Watchdog
The Last Watchdog
Help Net Security
Help Net Security
I
Intezer
The Register - Security
The Register - Security
小众软件
小众软件
C
Check Point Blog
NISL@THU
NISL@THU
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
T
Tor Project blog
D
Docker
Hugging Face - Blog
Hugging Face - Blog
WordPress大学
WordPress大学
Forbes - Security
Forbes - Security
H
Hackread – Cybersecurity News, Data Breaches, AI and More
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
T
Tailwind CSS Blog
Security Latest
Security Latest
博客园 - 司徒正美
IT之家
IT之家

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
AI System Design Interview Questions: ChatGPT, RAG, LLM Inference, and Agents
Arslan Ahmad · 2026-06-25 · via DEV Community

System design interviews are changing.

Traditional questions such as “Design Twitter,” “Design Uber,” and “Design YouTube” are still important. They test whether you understand databases, caching, partitioning, replication, messaging, and high availability.

But engineers working on modern platforms now encounter a different category of problem:

  • Design a ChatGPT-like conversational assistant.
  • Design a retrieval-augmented generation system.
  • Design an LLM inference platform.
  • Design an AI agent that can call external tools.
  • Design an enterprise AI assistant for private documents.
  • Design an evaluation platform for generative AI applications.

These questions still require classical distributed-systems knowledge. An AI product needs APIs, queues, storage, authentication, observability, rate limiting, and reliable deployment.

The difference is that it also introduces expensive accelerators, probabilistic output, long-running requests, model routing, vector retrieval, prompt construction, safety controls, and quality evaluation.

This guide explains the most important AI system design interview questions and what a strong candidate should discuss for each.

For a broader preparation roadmap covering traditional and modern problems, see 64 System Design Interview Questions, Ranked From Easiest to Hardest.


Why AI System Design Is Different

A conventional service usually transforms an input into a deterministic output.

If a user requests order number 123, the service should retrieve order 123. Two identical requests should usually return the same underlying information.

Generative AI systems behave differently.

A model may produce different responses to the same prompt. A response can be grammatically convincing while being factually wrong. Latency depends on the number of generated tokens. Serving capacity is constrained by accelerator memory, not merely CPU utilization. Product quality may depend on prompts, retrieved context, model versions, safety filters, and external tools.

This creates several new design dimensions.

1. Quality is part of the architecture

Traditional systems are often measured using availability, latency, throughput, and error rate.

AI systems need those metrics, but they also need measures such as:

  • Answer correctness
  • Relevance
  • Groundedness
  • Retrieval quality
  • Hallucination rate
  • Tool-use success
  • Safety-policy compliance
  • User satisfaction

A system that returns a response in 200 milliseconds is not useful if that response is wrong.

2. Requests are computationally expensive

An ordinary API server may process thousands of lightweight requests per second.

An LLM request can occupy expensive GPU memory while processing a long prompt and generating hundreds of tokens. The architecture must therefore optimize batching, memory utilization, model placement, and request scheduling.

3. Latency is experienced as a stream

Users do not normally wait for an entire answer before seeing anything. Tokens are streamed as they are generated.

This introduces at least two important latency measurements:

  • Time to first token: How quickly generation begins.
  • Inter-token latency: How smoothly subsequent tokens arrive.

A system may have acceptable total latency but still feel slow if the first token takes too long.

4. Data enters the system in several ways

An AI application may depend on:

  • Model-training data
  • User prompts
  • Conversation history
  • Retrieved documents
  • Tool results
  • Feedback
  • Evaluation datasets
  • Safety policies

Each data type has distinct requirements for retention, privacy, freshness, and consistency.

5. Failure is not always binary

A traditional request may succeed or fail.

An AI request can technically succeed but produce a low-quality answer, retrieve the wrong documents, call the wrong tool, exceed a cost budget, or violate a safety rule.

The architecture must detect and respond to these softer failure modes.


A Framework for Answering Any AI System Design Question

Before considering individual questions, use a consistent interview structure.

Step 1: Clarify the product

Ask what the system is expected to do.

For example:

  • Is the assistant general-purpose or domain-specific?
  • Does it need private enterprise data?
  • Can it take actions or only provide answers?
  • Does it support text only, or also images, audio, and files?
  • Are responses expected in real time?
  • Does the system need citations?
  • Which decisions require human approval?

Without this clarification, “Design an AI assistant” is too broad.

Step 2: Define scale and service-level objectives

Estimate:

  • Daily and peak requests
  • Average prompt size
  • Average output length
  • Concurrent users
  • Required time to first token
  • Model size
  • GPU-memory requirements
  • Availability target
  • Cost per request

AI systems are often constrained by cost as much as by technical capacity.

Step 3: Separate the application layer from the model layer

The application layer may include:

  • Authentication
  • Billing
  • Conversation history
  • File management
  • User preferences
  • Rate limiting
  • Analytics

The AI layer may include:

  • Prompt construction
  • Retrieval
  • Model routing
  • Inference scheduling
  • Safety checks
  • Tool execution
  • Evaluation

Keeping these concerns separate makes the design easier to explain and evolve.

Step 4: Trace the complete request path

Describe what happens from the moment a user submits a prompt until the final result is displayed.

A typical path may be:

  1. Authenticate the request.
  2. Enforce quotas.
  3. Load conversation state.
  4. Retrieve relevant context.
  5. Construct the model prompt.
  6. Run input-safety checks.
  7. Select a model.
  8. Schedule inference.
  9. Stream tokens.
  10. Run output-safety checks.
  11. Store the response.
  12. Record metrics and feedback.

Step 5: Discuss failure, quality, and cost

A strong answer should explain:

  • What happens when the primary model is overloaded?
  • What happens when retrieval returns no useful documents?
  • How are duplicate tool calls prevented?
  • How does the system degrade gracefully?
  • How are model changes evaluated?
  • How is tenant data isolated?
  • How are expensive requests controlled?

These discussions distinguish a production design from a demo.


Question 1: Design ChatGPT

A ChatGPT-like system is one of the most comprehensive AI system design questions.

The functional requirements may include:

  • Starting a conversation
  • Sending prompts
  • Receiving streamed responses
  • Viewing conversation history
  • Regenerating an answer
  • Uploading files
  • Choosing among models
  • Enforcing free and paid usage limits

A useful high-level architecture contains the following components.

API gateway

The gateway handles authentication, request routing, rate limiting, quotas, and basic validation.

Long-running generation requests may use Server-Sent Events or WebSockets to stream tokens to clients.

Conversation service

This service manages:

  • Conversations
  • Messages
  • User preferences
  • Message ordering
  • Conversation titles
  • Retention and deletion

Conversation metadata can live in a transactional database, while large attachments may be placed in object storage.

Context builder

Models have finite context windows. The context builder decides what information should be included in the next request.

It may combine:

  • The system prompt
  • Recent conversation messages
  • A summary of older messages
  • Retrieved documents
  • User preferences
  • Tool outputs

Simply sending the entire conversation forever is expensive and eventually impossible. Older content may need summarization or selective retrieval.

Model gateway

The model gateway provides a single interface to multiple model backends.

It can route requests based on:

  • Task type
  • Required quality
  • User subscription
  • Context length
  • Latency target
  • Current capacity
  • Cost budget
  • Model availability

A simple request may use a smaller, faster model, while complex reasoning may be routed to a more capable one.

Inference scheduler

The scheduler assigns requests to model replicas running on accelerators.

It should consider:

  • Available GPU memory
  • Model placement
  • Prompt length
  • Output-token budget
  • Priority
  • Batch compatibility
  • Tenant quotas

A naive first-in, first-out scheduler can allow a few extremely long prompts to delay many short requests.

Streaming layer

Generated tokens should be forwarded incrementally to the user.

The system must also handle:

  • Client disconnections
  • User cancellation
  • Partial responses
  • Network retries
  • Moderation during generation
  • Final persistence after streaming completes

Safety layer

Input and output policies may detect:

  • Prompt injection
  • Sensitive information
  • Disallowed requests
  • Malicious files
  • Unsafe tool instructions
  • Data leakage

Safety should not be treated as one filter placed at the end. Different checks may be required before retrieval, before tool execution, before inference, and before returning the final response.

Important deep dives

An interviewer may ask:

  • How would you reduce time to first token?
  • How would you support 100 million users?
  • How would you prevent one tenant from consuming all GPU capacity?
  • How would you summarize long conversations?
  • How would you route between multiple models?
  • How would you preserve availability during GPU shortages?
  • How would you limit the cost for free users?

The Design ChatGPT walkthrough provides a structured example of this problem.


Question 2: Design a RAG System

Retrieval-augmented generation, or RAG, allows a model to answer using information retrieved from external sources.

A common interview prompt is:

Design an enterprise assistant that answers employee questions using internal documents and provides citations.

A RAG system has two major paths:

  1. The ingestion path
  2. The query path

The ingestion path

Documents may come from file uploads, internal wikis, cloud drives, databases, or support systems.

The ingestion pipeline performs several stages.

Document extraction

Files must be converted into usable text.

The system may need parsers for:

  • PDFs
  • Word documents
  • Presentations
  • HTML pages
  • Spreadsheets
  • Scanned images

The extraction process should preserve useful metadata such as titles, headings, page numbers, owners, and access permissions.

Chunking

Long documents are divided into smaller segments.

Chunks that are too large may contain irrelevant text and consume excessive context. Chunks that are too small may lose meaning.

Possible strategies include:

  • Fixed token windows
  • Paragraph-based chunking
  • Heading-aware chunking
  • Overlapping windows
  • Semantic chunking

There is no universally correct chunk size. It should be tested against representative questions.

Embedding generation

Each chunk is converted into a numerical vector using an embedding model.

The embedding service should be versioned because changing models can require re-embedding the entire corpus.

Indexing

The system stores:

  • Embeddings
  • Original text
  • Document metadata
  • Access-control information
  • Source location
  • Embedding version
  • Update timestamp

A vector index enables semantic retrieval. A traditional inverted index can support keyword retrieval. Many production systems combine both.

The query path

When a user submits a question:

  1. Authenticate the user.
  2. Generate an embedding for the query.
  3. Retrieve candidate chunks.
  4. Apply access-control filtering.
  5. Rerank the candidates.
  6. Select the best context.
  7. Construct the prompt.
  8. Generate the answer.
  9. Attach citations.
  10. Evaluate or log the result.

Hybrid retrieval

Semantic retrieval is useful when the query and source use different words with similar meanings.

Keyword retrieval is useful for exact terms such as:

  • Product codes
  • Error messages
  • Names
  • Dates
  • Identifiers

Combining both methods often produces better coverage.

Reranking

Vector similarity may retrieve documents that are generally related but not directly useful.

A reranker can score the top candidates more accurately before they are sent to the LLM. This improves answer quality while keeping the final prompt small.

Access control

Security is one of the most important parts of enterprise RAG.

A user should never retrieve a document they are not authorized to view. Filtering after the model has already received the document is too late.

Permissions should be enforced during retrieval, with tenant and user identity included in the query path.

Freshness and deletion

The system must react when:

  • A document changes.
  • A document is deleted.
  • Permissions change.
  • A user loses access.
  • A newer policy replaces an older one.

The ingestion pipeline may use event-driven updates, periodic crawling, or both.

RAG evaluation

A RAG system should separately evaluate:

  • Retrieval quality: Did the system find the relevant document?
  • Generation quality: Did the model use the retrieved context correctly?
  • Citation quality: Do the cited sources actually support the answer?

This separation is important. A poor answer can result from failed retrieval even when the model behaves correctly.


Question 3: Design an LLM Inference Platform

This question focuses less on the product interface and more on the infrastructure that serves models.

A possible prompt is:

Design a multi-tenant platform that serves several large language models to millions of requests.

The platform may need to support:

  • Multiple model families
  • Different model sizes
  • Streaming generation
  • Priority tiers
  • Autoscaling
  • Usage accounting
  • Model versioning
  • Regional deployment
  • Fine-tuned adapters

Inference gateway

The gateway exposes a consistent API and performs:

  • Authentication
  • Quota enforcement
  • Request validation
  • Model selection
  • Token-limit checks
  • Admission control
  • Cost estimation

Admission control is critical. Accepting unlimited work and allowing it to queue indefinitely creates poor latency and can destabilize the system.

Model registry

The registry tracks:

  • Model version
  • Artifact location
  • Supported hardware
  • Memory requirements
  • Context length
  • Quantization format
  • Deployment status
  • Safety and evaluation results

Rollouts should use immutable versions so requests and incidents can be traced to the exact model that served them.

Model placement

Loading a large model into GPU memory can take substantial time. The scheduler cannot treat models like lightweight stateless application containers.

It must decide:

  • Which models remain loaded
  • How many replicas each model receives
  • Where fine-tuned adapters are placed
  • When models should be unloaded
  • How capacity is distributed across regions

Popular models may remain warm, while rarely used models may accept a cold-start delay.

Prefill and decode

LLM inference contains two different computational phases.

Prefill processes the input prompt and can often benefit from parallel computation.

Decode generates tokens sequentially and is usually memory-bandwidth intensive.

Separating or independently scheduling these phases can improve utilization, but it also adds network and orchestration complexity.

Continuous batching

Instead of waiting for a fixed group of requests to finish together, continuous batching adds and removes requests dynamically as generation progresses.

This improves GPU utilization, especially when responses have different lengths.

The scheduler must still prevent long requests from starving shorter ones.

KV cache

The key-value cache stores intermediate attention state so the model does not recompute the entire prompt for every generated token.

KV-cache management affects:

  • Maximum concurrency
  • Memory pressure
  • Long-context support
  • Prefix reuse
  • Request eviction

A shared prompt prefix—such as a large system prompt—may sometimes be cached and reused across compatible requests.

Scaling

GPU utilization alone may not be sufficient for autoscaling.

Useful signals include:

  • Queue length
  • Time to first token
  • Tokens generated per second
  • KV-cache pressure
  • Number of active sequences
  • Predicted token demand
  • Model-specific backlog

Because accelerator provisioning may be slow, the platform may need reserved capacity and predictive scaling.

Graceful degradation

When capacity is limited, the system may:

  • Route to a smaller model.
  • Reduce the maximum output length.
  • Reject low-priority requests.
  • Disable expensive features.
  • Queue batch workloads.
  • Move traffic to another region.
  • Use a third-party model provider.

A strong interview answer discusses the quality and cost consequences of each fallback.


Question 4: Design an AI Agent Platform

An AI agent does more than produce text. It can plan a sequence of actions, call tools, observe results, update its state, and continue until a goal is completed.

A typical prompt might be:

Design an enterprise agent that can search internal documents, update tickets, send emails, and request human approval for sensitive actions.

Core components

Agent orchestrator

The orchestrator controls the execution loop:

  1. Receive a goal.
  2. Construct the current context.
  3. Ask the model for the next action.
  4. Validate the proposed action.
  5. Execute the selected tool.
  6. Store the result.
  7. Decide whether to continue.
  8. Produce the final response.

The orchestrator—not the model—should enforce hard limits such as maximum steps, timeouts, budgets, and approval requirements.

Tool registry

The tool registry describes each available capability:

  • Tool name
  • Purpose
  • Input schema
  • Required permissions
  • Timeout
  • Retry policy
  • Risk level
  • Whether human approval is required

Tool definitions should be versioned because changing their schemas can break existing agent behavior.

Tool execution service

Tool calls should run through controlled executors rather than allowing the model unrestricted access to internal systems.

The executor handles:

  • Authentication
  • Input validation
  • Secrets
  • Network policy
  • Timeouts
  • Retries
  • Audit logging
  • Output normalization

High-risk operations should use narrow, purpose-built APIs.

State and memory

Agents may need several kinds of memory.

Working memory contains the current task, observations, and intermediate steps.

Session memory preserves information during one user interaction.

Long-term memory stores information across sessions.

External memory may contain documents retrieved from databases or vector indexes.

Not everything should be stored forever. Memory needs explicit retention, privacy, and deletion policies.

Human approval

Actions such as sending payments, deleting data, publishing content, or modifying production systems should not be executed solely because a model requested them.

The agent can create a proposed action, pause its workflow, and wait for authorized approval.

The approval record should contain:

  • The intended action
  • The relevant parameters
  • Why it was proposed
  • The expected effect
  • The identity of the approver
  • An expiration time

Idempotency

Agents may retry actions after timeouts.

Without idempotency, a retry could send the same email twice, create duplicate tickets, or repeat a transaction.

Every state-changing tool call should include a stable execution identifier or idempotency key.

Agent-specific failure modes

A strong candidate should discuss:

  • Infinite planning loops
  • Repeated tool calls
  • Prompt injection inside retrieved content
  • Tool hallucination
  • Stale observations
  • Excessive cost
  • Partial workflow completion
  • Conflicting actions
  • Unauthorized data access

The system should impose:

  • Maximum step counts
  • Token budgets
  • Time limits
  • Per-tool permissions
  • Human checkpoints
  • Detailed audit logs
  • Recovery or compensation workflows

The Grokking Modern AI Fundamentals course can provide additional background on agentic AI, planning, memory, and tool-based behavior.


Additional AI System Design Questions to Practice

The four core questions cover much of the modern AI stack, but interviewers can frame the same concepts in narrower ways.

Design an enterprise AI copilot

Focus on tenant isolation, document permissions, RAG, conversation history, model routing, auditability, and data retention.

Design a coding assistant

Discuss repository indexing, code-aware chunking, low-latency suggestions, context selection, IDE integration, private-code protection, and evaluation of generated code.

Design an AI evaluation platform

Cover dataset versioning, offline evaluation, human review, model comparison, prompt experiments, regression detection, and production feedback.

Design a multi-model gateway

Explain routing between internal and third-party models based on cost, quality, latency, privacy, context length, and availability.

Design a semantic-search platform

Focus on ingestion, embeddings, vector indexes, hybrid retrieval, filtering, reranking, index updates, and relevance metrics.

Design an AI safety and guardrails service

Discuss policy versioning, input and output classification, prompt-injection detection, tool restrictions, personally identifiable information, appeals, and false-positive handling.

Design a prompt-management platform

Cover prompt templates, version control, experiments, rollout, rollback, tenant overrides, caching, and compatibility with changing model versions.

Design a multimodal assistant

Add image, audio, and document ingestion, media storage, preprocessing, modality-specific models, content safety, and larger payload management.


What Interviewers Look for in AI System Design Answers

A weak answer places an LLM box in the center of a diagram and connects it to an API.

A strong answer explains the system around the model.

Interviewers want to see whether you can reason about:

End-to-end architecture

Can you connect the client, application services, retrieval layer, model platform, storage, and observability systems?

Trade-offs

Can you compare:

  • Larger models versus smaller models
  • Quality versus latency
  • Quality versus cost
  • Long context versus retrieval
  • Hosted APIs versus self-hosting
  • Semantic retrieval versus keyword search
  • Autonomy versus human control

Reliability

Can the system continue operating when a model, vector index, tool, region, or third-party provider is unavailable?

Evaluation

Can you tell whether a new model or prompt actually improved the product?

Security

Can you protect tenant data, prevent unauthorized retrieval, constrain tool use, and manage sensitive prompts?

Cost

Can you estimate and control token usage, accelerator capacity, retrieval cost, storage, and third-party API spending?

The model is only one component. Production readiness comes from the architecture surrounding it.


How to Prepare

Candidates new to large-scale architecture should first learn the traditional foundations.

Grokking System Design Fundamentals introduces the core building blocks behind scalable systems, including databases, caches, queues, replication, partitioning, and load balancing.

The original Grokking the System Design Interview applies those concepts to common interview problems and teaches a structured way to move from requirements to architecture and trade-offs.

The System Design Interview Crash Course is useful for practicing a consistent interview framework across modern case studies, including a complete ChatGPT design problem.

Engineers preparing for senior and staff-level discussions can continue with Advanced System Design Interview, Volume II, which emphasizes open-ended problems, failures, and defensible architectural decisions.

Grokking Scalable Systems for Interviews is a useful next step for strengthening scalability, observability, fault tolerance, and performance reasoning.

For every AI design problem, practice three times:

  1. Design the happy path.
  2. Design for failure and overload.
  3. Defend the quality, safety, and cost trade-offs.

That third pass is where most of the valuable interview discussion occurs.


Final Takeaway

AI system design is not a replacement for traditional system design.

It is a traditional system design combined with a new set of constraints.

You still need to understand APIs, storage, caching, queues, partitioning, replication, security, observability, and fault tolerance.

But you must now apply those concepts to systems with:

  • Probabilistic outputs
  • Expensive inference
  • Streaming generation
  • Vector retrieval
  • Dynamic prompts
  • Long-lived context
  • External tools
  • Model evaluation
  • Safety requirements
  • Human approval

Start with four foundational problems:

  1. Design ChatGPT.
  2. Design a RAG platform.
  3. Design an LLM inference service.
  4. Design an AI agent platform.

Master the request flow, deep dives, failure modes, and trade-offs behind each one.

Once you can explain those systems clearly, most other AI system design questions become variations of the same underlying building blocks.