Security Notes
System Design

Design Strava (Activity Tracking & Social Fitness)

3 min read 6 sections

DifficultyMedium | HelloInterview: problem breakdown


Problem Statement

Design a fitness tracking platform where users can record running/cycling activities (GPS tracks), analyze performance metrics, compete on leaderboard segments, and follow friends' activities.

๐Ÿ“–

Real-world: Strava is the cautionary privacy story in this whole catalog. In 2018, Strava's public "global heatmap" of aggregated activity revealed the locations and internal layouts of secret military bases โ€” soldiers jogging with fitness trackers traced the perimeters of classified sites in Afghanistan and Syria, visible because the surrounding desert was dark and the base was lit up with activity. It forced policy changes across multiple militaries. The lesson is exactly the one a security person should carry: aggregate/"anonymized" location data is rarely anonymous โ€” patterns re-identify people and places, so privacy must be designed in (opt-out zones, "privacy zones" around home/work that Strava added, and careful aggregation). Technically the design is a time-series GPS ingestion + leaderboard problem (Redis sorted sets for segment rankings, precomputed segment efforts), but the headline interview move is to raise the location-privacy risk unprompted.


Requirements

Functional

  • Upload GPS track (lat/lng + timestamp sequence)
  • Compute stats: distance, pace, elevation gain, splits
  • Segment matching: detect if activity passes through a defined segment; rank on leaderboard
  • Social feed: see friends' activities
  • Route map visualization

Non-Functional

  • 1M activities uploaded/day
  • GPS track: 1 point/second ร— 1h run = 3600 points ร— 16 bytes = ~57 KB per activity
  • Segment matching: fast enough that results appear within 30 seconds of upload
  • Leaderboard: top 10 per segment updated in near real-time

Core Design

Activity Upload and Stats

Upload: GPS track (array of {lat, lng, alt, timestamp})
         โ†’ Object store (S3) for raw track
         โ†’ Stats worker: compute distance (Haversine), pace, elevation profile
         โ†’ Store summary in Activity DB

Haversine formulacomputes great-circle distance between two lat/lng points. Applied to each consecutive pair โ†’ sum = total distance.

Segment Matching

A segment is a defined route between two GPS points (start/end). Given an activity's track, determine if the athlete traveled through this segment.

Algorithm:
1. For each segment:
   - Check if activity's bounding box overlaps segment's bounding box (fast pre-filter)
   - If yes: check if start of segment matches any point in activity (within 15m)
   - If yes: follow the activity track from that point, check if end of segment is reached
2. If matched: extract elapsed time between start and end โ†’ leaderboard entry

Scale100k segments ร— 1M activities/day = 100B potential matches/day. Can't check all pairs.

Optimization

  • Spatial index on segments (R-tree/geohash) โ€” only check segments near the activity's geographic area
  • Pre-filter: bounding box overlap eliminates most segments quickly
  • Only run matching on activities in the same city/region as the segment

Leaderboard

Redis sorted set per segment:
  ZADD segment:{segment_id}:leaderboard <elapsed_seconds> <user_id>
  
Query top 10:
  ZRANGEBYSCORE segment:{segment_id}:leaderboard -inf +inf LIMIT 10

On new best time: update sorted set, notify user if new record.

Social Feed

Same pattern as news feed: activities as posts. Fan-out on write for regular users; fan-out on read for very popular athletes (pros with 100k followers).


Architecture

Upload API โ†’ S3 (raw track) โ†’ Processing Queue
                                    โ†“
                           Stats Worker (distance, pace, elevation)
                                    โ†“
                           Segment Matcher (spatial index)
                                    โ†“
                      Activity DB (Postgres) + Segment Leaderboard (Redis)
                                    โ†“
                           Fan-out Worker โ†’ Feed Cache (Redis)

Security Considerations

ThreatMitigation
Location stalkingAllow private activities; flybys detection (who else was near at same time) is optional and controversial
Segment cheatingGPS validation (speed sanity, altitude consistency); flag/appeal mechanism
Data exfiltrationRate limit Strava API; no bulk download of others' tracks
Home address exposureSegment / activity start/end can reveal home location; "hide start/end" feature for X meters

Interview Tips

Segment matching is the interesting problem

spatial pre-filtering (bounding box + geohash) is the key optimization to make it tractable.

GPS data at scale

1M activities ร— 57KB = 57 GB/day of raw tracks. Object store (S3), not a database.

Leaderboard with Redis sorted set

same pattern as rate limiting and Top K.