Security Notes
System Design

Design an LLM API Service (ChatGPT-scale)

5 min read 7 sections

DifficultyHard | HelloInterview: problem breakdown


Problem Statement

Design a scalable API service for serving large language model inference. Users send prompts and receive streamed completions. The system must handle high request volume, long-running inference, multi-tenant isolation, and token-based billing.

πŸ“–

Real-world: This is the most current system-design problem, and its constraints are unlike the classics. The bottleneck isn't a database β€” it's scarce, expensive GPUs: inference is slow and memory-bound, so the whole design orbits a request queue feeding a GPU worker pool, with tricks like continuous/in-flight batching (packing multiple users' tokens through the GPU together) to maximize utilization. Responses stream token-by-token over SSE because a full completion takes seconds, and users tolerate "typing" but not a long blank wait. ChatGPT's launch in late 2022 famously hit 1M users in 5 days and repeatedly showed "at capacity" β€” a vivid lesson in GPU scarcity as the scaling limit. For a security audience the rich angles are right here: multi-tenant isolation (one customer's prompt/data must never leak into another's context), prompt-injection and abuse filtering on inputs, token-based billing integrity (and the "denial-of-wallet" abuse where an attacker runs up your inference bill), and rate limiting as a cost-control and safety control β€” see ../ai-security/owasp-llm-2025.md.


Requirements

Functional

  • Accept prompt + parameters β†’ return completion (full or streamed)
  • Maintain conversation context (multi-turn chat)
  • Token counting for billing
  • Model selection (different model sizes/capabilities)
  • Streaming output (token by token as generated)

Non-Functional

  • 10M requests/day β†’ ~115 requests/s average; 1000 req/s peak
  • Time to first token (TTFT): < 500ms
  • Single inference: 10-60 seconds for long outputs
  • GPU utilization: maximize (GPUs are expensive)
  • Multi-tenant: one user's slow request cannot starve others
  • Token limits enforced per API key

Core Design

Why This Is Hard

LLM inference is:

GPU-bound

runs on expensive GPU hardware; each GPU handles 1-4 concurrent requests

Long-running

10-60 seconds per request, not milliseconds

Variable length

output tokens determined at inference time; can't pre-allocate

Streaming

user expects to see tokens as generated, not wait for full completion

Standard request/response architecture doesn't scale for this.

Request Queuing with Priority Lanes

Client β†’ API Gateway β†’ Token Counter + Auth β†’ Request Queue
                                                    ↓
                                          Inference Scheduler
                                                    ↓
                                        GPU Worker Pool
                                          (batch requests)

Priority lanes

  • Real-time API (low latency, interactive): premium lane, preemptible
  • Batch API (high throughput, non-interactive): best-effort lane, up to 50% cheaper

Continuous Batching

NaΓ―ve: one GPU processes one request at a time β†’ 1/60th of GPU time utilized.

Continuous batching (key LLM optimization): multiple requests share the same GPU forward pass. As one request finishes, a new one joins the batch without waiting for all current requests to complete.

Time β†’
Req A: β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ (done)
Req B:   β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ (long)
Req C:         β–ˆβ–ˆβ–ˆβ–ˆ (done)
Req D:             β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ (started when A finished)

GPU is never idle waiting for one long request. Throughput scales with batch size.

Key-Value (KV) Cache

During inference, transformer attention layers compute key-value tensors that can be reused across tokens in the same sequence. Caching these (KV cache) dramatically speeds up autoregressive generation.

Prefix cachingif two requests share the same system prompt, the KV cache for the prefix is computed once and reused. Important for:

  • System prompt reuse (same system prompt across all requests in a session)
  • RAG document context (same document used in many queries)

Streaming via SSE

HTTP Server-Sent Events (SSE) is the standard for streaming tokens:

HTTP/1.1 200 OK
Content-Type: text/event-stream

data: {"token": "The"}
data: {"token": " quick"}
data: {"token": " brown"}
data: [DONE]

Client reads the stream token by token, rendering as it arrives.


Architecture

Client β†’ API Gateway (auth, rate limit, token count)
              ↓
        Request Router (model selection, priority)
              ↓
        Queue (Redis Streams / SQS per model tier)
              ↓
        Inference Cluster (Kubernetes + GPU nodes)
              |
     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     β”‚ Inference Server        β”‚ ← vLLM / TensorRT-LLM
     β”‚ (continuous batching)   β”‚
     β”‚ KV cache manager        β”‚
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              ↓
        Response Streamer β†’ SSE β†’ Client
              ↓
        Usage DB (token counts, latency, model used)

Context Window Management

Each conversation has a history that grows with each turn. LLMs have a context window limit (e.g., 128k tokens). Strategy:

  • Keep last N turns
  • Summarize old context when approaching limit
  • Always include system prompt + last M turns

Conversation history stored in Redis (TTL = session length); if user returns after TTL, start fresh.

Model Serving Infrastructure

vLLM

open-source LLM serving library with continuous batching + PagedAttention (KV cache paging, like OS virtual memory)

Triton Inference Server

(NVIDIA): model serving framework

GPU types

A100 (80GB VRAM) for large models; H100 for fastest; A10G/L4 for smaller models / cost efficiency


Key Design Decisions

DecisionChoiceReason
GPU batchingContinuous batching (vLLM)Maximize GPU utilization; handle variable-length outputs
StreamingSSEStandard; works through proxies and browsers
PrioritySeparate queues per tierReal-time users don't wait behind batch jobs
ContextRedis (TTL)Fast per-session retrieval; no persistent storage needed
BillingToken counting at gateway + inferenceAccurate billing; enforce per-key limits

Security Considerations

ThreatMitigation
Prompt injectionInput validation; system prompt hardening; output filtering; jailbreak detection
Data leakage between tenantsStrict KV cache isolation (don't share cached prefixes across tenants); no cross-tenant state
Insecure output (XSS, SQLi via LLM)Treat LLM output as untrusted input; sanitize before rendering or executing
API key abuseRate limiting per key; anomaly detection on usage spikes
Model extraction / inversionRate limit; block systematic probing; watermarking outputs
PII in promptsLog scrubbing; user consent; no prompt logging by default (or opt-in)
Excessive token consumption (DoS)Max tokens per request; per-key monthly quota; hard spending limits
System prompt exfiltrationRespond consistently to "repeat your instructions" probing; don't include secrets in system prompt

Interview Tips

Continuous batching

is the key LLM-specific concept β€” GPU utilization is the bottleneck, and continuous batching maximizes it.

KV cache and prefix caching

shows you understand how transformers work and how to exploit that for performance.

Streaming with SSE

know the protocol, why WebSocket isn't needed here (unidirectional output).

Security = OWASP LLM Top 10

mention prompt injection, tenant isolation, and output sanitization. Link to ai-security/owasp-llm-2025.md for full depth.