Design Strava (Activity Tracking & Social Fitness)
DifficultyMedium | HelloInterview: problem breakdown
Problem Statement
Design a fitness tracking platform where users can record running/cycling activities (GPS tracks), analyze performance metrics, compete on leaderboard segments, and follow friends' activities.
๐Real-world: Strava is the cautionary privacy story in this whole catalog. In 2018, Strava's public "global heatmap" of aggregated activity revealed the locations and internal layouts of secret military bases โ soldiers jogging with fitness trackers traced the perimeters of classified sites in Afghanistan and Syria, visible because the surrounding desert was dark and the base was lit up with activity. It forced policy changes across multiple militaries. The lesson is exactly the one a security person should carry: aggregate/"anonymized" location data is rarely anonymous โ patterns re-identify people and places, so privacy must be designed in (opt-out zones, "privacy zones" around home/work that Strava added, and careful aggregation). Technically the design is a time-series GPS ingestion + leaderboard problem (Redis sorted sets for segment rankings, precomputed segment efforts), but the headline interview move is to raise the location-privacy risk unprompted.
Requirements
Functional
- Upload GPS track (lat/lng + timestamp sequence)
- Compute stats: distance, pace, elevation gain, splits
- Segment matching: detect if activity passes through a defined segment; rank on leaderboard
- Social feed: see friends' activities
- Route map visualization
Non-Functional
- 1M activities uploaded/day
- GPS track: 1 point/second ร 1h run = 3600 points ร 16 bytes = ~57 KB per activity
- Segment matching: fast enough that results appear within 30 seconds of upload
- Leaderboard: top 10 per segment updated in near real-time
Core Design
Activity Upload and Stats
Upload: GPS track (array of {lat, lng, alt, timestamp})
โ Object store (S3) for raw track
โ Stats worker: compute distance (Haversine), pace, elevation profile
โ Store summary in Activity DBHaversine formulacomputes great-circle distance between two lat/lng points. Applied to each consecutive pair โ sum = total distance.
Segment Matching
A segment is a defined route between two GPS points (start/end). Given an activity's track, determine if the athlete traveled through this segment.
Algorithm:
1. For each segment:
- Check if activity's bounding box overlaps segment's bounding box (fast pre-filter)
- If yes: check if start of segment matches any point in activity (within 15m)
- If yes: follow the activity track from that point, check if end of segment is reached
2. If matched: extract elapsed time between start and end โ leaderboard entryScale100k segments ร 1M activities/day = 100B potential matches/day. Can't check all pairs.
Optimization
- Spatial index on segments (R-tree/geohash) โ only check segments near the activity's geographic area
- Pre-filter: bounding box overlap eliminates most segments quickly
- Only run matching on activities in the same city/region as the segment
Leaderboard
Redis sorted set per segment:
ZADD segment:{segment_id}:leaderboard <elapsed_seconds> <user_id>
Query top 10:
ZRANGEBYSCORE segment:{segment_id}:leaderboard -inf +inf LIMIT 10On new best time: update sorted set, notify user if new record.
Social Feed
Same pattern as news feed: activities as posts. Fan-out on write for regular users; fan-out on read for very popular athletes (pros with 100k followers).
Architecture
Upload API โ S3 (raw track) โ Processing Queue
โ
Stats Worker (distance, pace, elevation)
โ
Segment Matcher (spatial index)
โ
Activity DB (Postgres) + Segment Leaderboard (Redis)
โ
Fan-out Worker โ Feed Cache (Redis)Security Considerations
| Threat | Mitigation |
|---|---|
| Location stalking | Allow private activities; flybys detection (who else was near at same time) is optional and controversial |
| Segment cheating | GPS validation (speed sanity, altitude consistency); flag/appeal mechanism |
| Data exfiltration | Rate limit Strava API; no bulk download of others' tracks |
| Home address exposure | Segment / activity start/end can reveal home location; "hide start/end" feature for X meters |
Interview Tips
spatial pre-filtering (bounding box + geohash) is the key optimization to make it tractable.
1M activities ร 57KB = 57 GB/day of raw tracks. Object store (S3), not a database.
same pattern as rate limiting and Top K.