Security Notes
System Design

Design an Online Code Judge (LeetCode)

4 min read 5 sections

DifficultyMedium | HelloInterview: problem breakdown


Problem Statement

Design an online judge where users submit code solutions to algorithmic problems. The system compiles and runs the code against test cases, and returns whether the solution is correct (pass/fail) along with performance metrics.

๐Ÿ“–

Real-world & security angle: This is the purest "run untrusted code safely" problem in the catalog โ€” it's literally arbitrary code from strangers, so it's a security design first and a queue design second. The whole game is sandboxing: you cannot run submissions on a normal host, so you isolate each run in a throwaway container or, better, a stronger boundary like gVisor or a Firecracker microVM (the same untrusted-workload tooling from ../kubernetes/fundamentals.md), with no network, dropped capabilities, a read-only filesystem, and hard CPU/memory/time limits to stop fork bombs, crypto-miners, and attempts to read other users' submissions or the test answers. Real judges have been escaped โ€” contestants have read the expected outputs or broken out of weak sandboxes โ€” which is exactly why defense-in-depth matters. Architecturally it's a submission queue โ†’ isolated worker pool โ†’ results via WebSocket, but if you don't lead with isolation, you've failed the question.


Requirements

Functional

  • Users submit code in multiple languages (Python, Java, C++, Go, etc.)
  • Run code against hidden test cases
  • Return: pass/fail per test case, runtime, memory usage
  • Problem library with description, constraints, examples
  • Submission history

Non-Functional

  • 500k submissions/day โ†’ ~6 submissions/s average; 100/s peak during contests
  • Execution timeout: configurable per problem (typically 1-5 seconds)
  • Sandboxed execution โ€” user code must not access the host filesystem, network, or other processes
  • Consistent results: same code โ†’ same result

Core Design

The Critical Requirement: Sandboxed Execution

Running arbitrary user-submitted code is one of the most dangerous operations in systems design. The execution environment must be completely isolated.

Approaches (security level low โ†’ high):

  1. Docker container per submission โ€” run code in a minimal container. Container limits CPU, memory, network. Cheap but not foolproof (container escapes possible with kernel bugs).

  2. gVisor (Google's userspace kernel) โ€” intercepts syscalls before they reach the real kernel. Better isolation than Docker alone. Used by Google's judge (Codeforces uses it).

  3. Firecracker microVM โ€” lightweight VM (100ms startup). True hypervisor isolation. Used by AWS Lambda. Each submission = fresh microVM.

  4. WASM sandbox โ€” compile to WebAssembly; runs in a wasm runtime with no OS access. Very safe but language support is limited.

For the interview, describe Docker + seccomp + resource limits:

Docker run options:
  --cpus=1
  --memory=512m
  --network=none          # no internet access
  --read-only             # read-only filesystem
  --security-opt seccomp=restricted.json  # block dangerous syscalls
  --pids-limit=100        # no fork bombs
  --no-new-privileges

Queue-Based Execution

User submits code โ†’ API โ†’ Submission DB (PENDING) โ†’ Job Queue (Kafka/SQS)
                                                           โ†“
                                                   Execution Workers
                                                   (auto-scaling pool)
                                                           โ†“
                                          Spin up container โ†’ compile โ†’ run test cases
                                                           โ†“
                                          Update Submission DB (ACCEPTED/WA/TLE/MLE/RE)
                                                           โ†“
                                          Notify user (WebSocket)

Execution Result Codes

CodeMeaning
ACCEPTEDAll test cases passed
WRONG_ANSWERIncorrect output
TIME_LIMIT_EXCEEDEDExecution > time limit
MEMORY_LIMIT_EXCEEDEDMemory > memory limit
RUNTIME_ERRORException / crash
COMPILE_ERRORCode doesn't compile

Contest Mode (High Burst)

During contests, thousands of submissions arrive simultaneously. Queue absorbs the burst; workers scale out (Kubernetes HPA). Worker pool should be pre-warmed โ€” starting a cold container takes 200ms+.

Worker warm poolmaintain N pre-started containers ready to receive code. After execution, re-initialize the container (or discard and start fresh from image).


Security Considerations

ThreatMitigation
Fork bomb--pids-limit=100; cgroup limit
Infinite loopHard wall-clock timeout via timeout command + SIGKILL
Network exfiltration--network=none; no outbound connectivity
Filesystem writeRead-only root filesystem; tmpfs for allowed write paths
Container escapeseccomp profile blocking dangerous syscalls; no privileged containers
Compiler exploitUse sandboxed compilers; restrict compiler filesystem access
Resource DoSPer-user submission rate limiting; memory hard limit
Code injection via test caseTest cases are server-controlled; never user-controlled

Security engineering notethis is an excellent real-world example of sandboxed execution that comes up in security interviews. The design maps directly to how CI/CD runners, Lambda functions, and browser DevTools sandboxes work.


Interview Tips

The sandbox is the entire interview

spend most time here. Docker + seccomp is the minimum; mention gVisor or Firecracker as production-grade alternatives.

Queue-based architecture

submission execution is async; never block the HTTP response for compilation+execution.

Warm worker pool

cold start latency is a UX problem; pre-warmed containers solve it.

This is a security design problem

frame it that way. The code execution sandbox has exactly the same requirements as a malware analysis sandbox.