Design an LLM API Service (ChatGPT-scale)
DifficultyHard | HelloInterview: problem breakdown
Problem Statement
Design a scalable API service for serving large language model inference. Users send prompts and receive streamed completions. The system must handle high request volume, long-running inference, multi-tenant isolation, and token-based billing.
πReal-world: This is the most current system-design problem, and its constraints are unlike the classics. The bottleneck isn't a database β it's scarce, expensive GPUs: inference is slow and memory-bound, so the whole design orbits a request queue feeding a GPU worker pool, with tricks like continuous/in-flight batching (packing multiple users' tokens through the GPU together) to maximize utilization. Responses stream token-by-token over SSE because a full completion takes seconds, and users tolerate "typing" but not a long blank wait. ChatGPT's launch in late 2022 famously hit 1M users in 5 days and repeatedly showed "at capacity" β a vivid lesson in GPU scarcity as the scaling limit. For a security audience the rich angles are right here: multi-tenant isolation (one customer's prompt/data must never leak into another's context), prompt-injection and abuse filtering on inputs, token-based billing integrity (and the "denial-of-wallet" abuse where an attacker runs up your inference bill), and rate limiting as a cost-control and safety control β see
../ai-security/owasp-llm-2025.md.
Requirements
Functional
- Accept prompt + parameters β return completion (full or streamed)
- Maintain conversation context (multi-turn chat)
- Token counting for billing
- Model selection (different model sizes/capabilities)
- Streaming output (token by token as generated)
Non-Functional
- 10M requests/day β ~115 requests/s average; 1000 req/s peak
- Time to first token (TTFT): < 500ms
- Single inference: 10-60 seconds for long outputs
- GPU utilization: maximize (GPUs are expensive)
- Multi-tenant: one user's slow request cannot starve others
- Token limits enforced per API key
Core Design
Why This Is Hard
LLM inference is:
runs on expensive GPU hardware; each GPU handles 1-4 concurrent requests
10-60 seconds per request, not milliseconds
output tokens determined at inference time; can't pre-allocate
user expects to see tokens as generated, not wait for full completion
Standard request/response architecture doesn't scale for this.
Request Queuing with Priority Lanes
Client β API Gateway β Token Counter + Auth β Request Queue
β
Inference Scheduler
β
GPU Worker Pool
(batch requests)Priority lanes
- Real-time API (low latency, interactive): premium lane, preemptible
- Batch API (high throughput, non-interactive): best-effort lane, up to 50% cheaper
Continuous Batching
NaΓ―ve: one GPU processes one request at a time β 1/60th of GPU time utilized.
Continuous batching (key LLM optimization): multiple requests share the same GPU forward pass. As one request finishes, a new one joins the batch without waiting for all current requests to complete.
Time β
Req A: ββββββββ (done)
Req B: ββββββββββββββββ (long)
Req C: ββββ (done)
Req D: ββββββββ (started when A finished)GPU is never idle waiting for one long request. Throughput scales with batch size.
Key-Value (KV) Cache
During inference, transformer attention layers compute key-value tensors that can be reused across tokens in the same sequence. Caching these (KV cache) dramatically speeds up autoregressive generation.
Prefix cachingif two requests share the same system prompt, the KV cache for the prefix is computed once and reused. Important for:
- System prompt reuse (same system prompt across all requests in a session)
- RAG document context (same document used in many queries)
Streaming via SSE
HTTP Server-Sent Events (SSE) is the standard for streaming tokens:
HTTP/1.1 200 OK
Content-Type: text/event-stream
data: {"token": "The"}
data: {"token": " quick"}
data: {"token": " brown"}
data: [DONE]Client reads the stream token by token, rendering as it arrives.
Architecture
Client β API Gateway (auth, rate limit, token count)
β
Request Router (model selection, priority)
β
Queue (Redis Streams / SQS per model tier)
β
Inference Cluster (Kubernetes + GPU nodes)
|
ββββββββββββββββββββββββββ
β Inference Server β β vLLM / TensorRT-LLM
β (continuous batching) β
β KV cache manager β
ββββββββββββββββββββββββββ
β
Response Streamer β SSE β Client
β
Usage DB (token counts, latency, model used)Context Window Management
Each conversation has a history that grows with each turn. LLMs have a context window limit (e.g., 128k tokens). Strategy:
- Keep last N turns
- Summarize old context when approaching limit
- Always include system prompt + last M turns
Conversation history stored in Redis (TTL = session length); if user returns after TTL, start fresh.
Model Serving Infrastructure
open-source LLM serving library with continuous batching + PagedAttention (KV cache paging, like OS virtual memory)
(NVIDIA): model serving framework
A100 (80GB VRAM) for large models; H100 for fastest; A10G/L4 for smaller models / cost efficiency
Key Design Decisions
| Decision | Choice | Reason |
|---|---|---|
| GPU batching | Continuous batching (vLLM) | Maximize GPU utilization; handle variable-length outputs |
| Streaming | SSE | Standard; works through proxies and browsers |
| Priority | Separate queues per tier | Real-time users don't wait behind batch jobs |
| Context | Redis (TTL) | Fast per-session retrieval; no persistent storage needed |
| Billing | Token counting at gateway + inference | Accurate billing; enforce per-key limits |
Security Considerations
| Threat | Mitigation |
|---|---|
| Prompt injection | Input validation; system prompt hardening; output filtering; jailbreak detection |
| Data leakage between tenants | Strict KV cache isolation (don't share cached prefixes across tenants); no cross-tenant state |
| Insecure output (XSS, SQLi via LLM) | Treat LLM output as untrusted input; sanitize before rendering or executing |
| API key abuse | Rate limiting per key; anomaly detection on usage spikes |
| Model extraction / inversion | Rate limit; block systematic probing; watermarking outputs |
| PII in prompts | Log scrubbing; user consent; no prompt logging by default (or opt-in) |
| Excessive token consumption (DoS) | Max tokens per request; per-key monthly quota; hard spending limits |
| System prompt exfiltration | Respond consistently to "repeat your instructions" probing; don't include secrets in system prompt |
Interview Tips
is the key LLM-specific concept β GPU utilization is the bottleneck, and continuous batching maximizes it.
shows you understand how transformers work and how to exploit that for performance.
know the protocol, why WebSocket isn't needed here (unidirectional output).
mention prompt injection, tenant isolation, and output sanitization. Link to ai-security/owasp-llm-2025.md for full depth.