MUHAMMAD
FARI
MADYAN
[ Press ESC or Click to Skip ]

System Design Roadmap

1

Part 1: Networking Basics – Packets, TCP/UDP, TLS & REST

2

Part 2: Core Building Blocks – Load Balancers, CDN, Caching, API Gateway, DBs & Kafka

3

Part 3: Distributed Patterns & Concepts – CAP Theorem, Hashing, Sharding & Sagas

4

Part 4: Real-World Interview Systems – URL Shortener, Chat, Food Delivery & Notification

This article is available in Indonesian

🇮🇩 Baca dalam Bahasa Indonesia
🇬🇧 English📚 System Design Roadmap

System Design Roadmap Part 4: Real-World Interview Systems – URL Shortener, Chat, Food Delivery & Notification

Ace system design interviews with complete end-to-end architectures: TinyURL URL shortener, WhatsApp-scale real-time chat, Zomato/Gojek geospatial food delivery, distributed notification engines, and Redis rate limiters.

Muhammad Fari MadyanAuthor

11 min read

·

Oct 7, 2026


Introduction: The Crucible of the System Design Interview

In Parts 1 through 3, we mastered the science: networking mechanics, scalability building blocks, and distributed patterns.

Now comes the art: The System Design Interview.

In a 45-minute FAANG or Tier-1 tech interview, you are given an intentionally ambiguous prompt like:

"Design WhatsApp" or "Design a Food Delivery System like DoorDash or Gojek."

Candidates who fail immediately begin drawing boxes. Candidates who succeed follow a structured, disciplined framework.

┌─────────────────────────────────────────────────────────────────────────────┐ │ THE 4-STEP SYSTEM DESIGN INTERVIEW BLUEPRINT │ ├─────────────────────────────────────────────────────────────────────────────┤ │ Step 1 (05 mins): Scope & Requirements (Functional, Non-Functional, Scale) │ │ Step 2 (10 mins): High-Level Architecture (Core API contracts & Flow) │ │ Step 3 (20 mins): Deep-Dive Bottlenecks (Data models, Scaling, Edge Cases) │ │ Step 4 (05 mins): Wrap-Up (Metrics, SPOFs, Monitoring, Future Extensibility)│ └─────────────────────────────────────────────────────────────────────────────┘

In this final chapter, we apply our roadmap to design 5 classic production systems step-by-step.


Case 1: Design a URL Shortener (e.g., TinyURL)

A URL shortener converts a long URL (https://example.com/products/deals/item-984218?ref=campaign) into an alias (https://tiny.url/aZ9k2m).

┌─────────────────────────────────────────────────────────────────────────────┐ │ TINYURL ARCHITECTURE │ └─────────────────────────────────────────────────────────────────────────────┘ [ Client ] ──POST /api/v1/urls──► [ Load Balancer ] │ ▼ [ URL Service ] / \ 1. Fetch Pre-Generated/ \ 2. Cache Write Unique ID Token / \ ▼ ▼ [ Key Generation ] [ Redis Cache ] [ Service (KGS) ] (Top 20% URLs) │ ▼ [ PostgreSQL / DynamoDB ] (short_key PK -> original_url)

1. Requirements & Math

  • Traffic: 100M new URLs created per month. Read-to-write ratio = $10:1$ (1 Billion reads/month).
  • Write RPS: $100\text{M} / (30 \times 86,400\text{s}) \approx 40\text{ writes/sec}$.
  • Read RPS: $\approx 400\text{ reads/sec}$ (peak: 2,000 RPS).
  • Storage: Over 5 years = $100\text{M} \times 12 \times 5 = 6\text{ Billion records}$. At $500\text{ bytes/record} \approx 3\text{ TB}$ (easily managed by a single distributed table or sharded database).

2. The Core Technical Challenge: Generating the Short Key

A short key must be compact and collision-free. Using Base62 ([a-zA-Z0-9]), a 7-character string yields: $$62^7 \approx 3.52\text{ Trillion unique combinations}$$ More than enough for decades!

Two approaches to generate the 7-character string:

  1. Hash of Original URL (MD5 / SHA-256): Taking the first 7 characters of MD5 causes hash collisions when different users shorten the same URL or during truncation.
  2. Key Generation Service (KGS) with Pre-allocated Ranges (Industry Best Practice):
    • A standalone service uses a distributed sequencer (or Apache ZooKeeper / Twitter Snowflake) to generate unique 64-bit integer IDs.
    • Workers claim blocks of 10,000 IDs in memory.
    • The integer is converted directly into Base62.
    • Result: Guaranteed $O(1)$ collision-free tokens with zero database coordination overhead!

3. Redirection: HTTP 301 vs. HTTP 302

When a user visits https://tiny.url/aZ9k2m:

  • HTTP 301 Moved Permanently: The browser caches the redirect locally. Subsequent clicks bypass your server entirely.
    • Pros: Minimum server load and fastest user redirect.
    • Cons: You cannot track click analytics or revoke URLs immediately.
  • HTTP 302 Found (Temporary Redirect): The browser queries the URL shortener on every single click.
    • Pros: Precise click telemetry (geographic location, referrers, device stats).
    • Cons: Increases server traffic.
    • Interview Recommendation: State this exact trade-off! Choose 302 if analytics are required, or 301 if raw cost efficiency is prioritized.

Case 2: Design a Real-Time Chat System (e.g., WhatsApp)

WhatsApp serves billions of users exchanging encrypted text, media, and presence statuses with sub-second latency.

┌─────────────────────────────────────────────────────────────────────────────┐ │ WHATSAPP-SCALE CHAT ENGINE │ └─────────────────────────────────────────────────────────────────────────────┘ [ Sender App ] [ Receiver App ] │ ▲ │ WebSocket │ WebSocket ▼ │ ┌───────────────────────────┐ ┌───────────────────────────┐ │ Chat Gateway Server 1 │ │ Chat Gateway Server 2 │ └─────────────┬─────────────┘ └─────────────▲─────────────┘ │ │ ▼ │ ┌───────────────────────────┐ ┌─────────────┴─────────────┐ │ Message Ingestion Engine │ ──► [ Kafka Topic ] ───► │ Push / Routing Engine │ └─────────────┬─────────────┘ └───────────────────────────┘ │ ▲ ▼ │ ┌───────────────────────────┐ ┌─────────────┴─────────────┐ │ Cassandra Message Store │ │ User Session Cache │ │ (Partition: chat_id) │ │ (Redis: user_id -> Host) │ └───────────────────────────┘ └───────────────────────────┘

1. Protocols: HTTP vs. WebSockets

  • HTTP Polling / Long Polling wastes battery and server resources on mobile devices due to redundant HTTP headers and TCP handshakes.
  • WebSockets establish a single, long-lived, bidirectional, lightweight TCP connection between client and chat gateway servers.

2. Message Flow & Delivery Semantics

  1. Sender writes a message over its active WebSocket to Chat Gateway 1.
  2. Chat Gateway 1 validates the message and returns a transient Message Ack (giving the single grey tick: Sent).
  3. The message is written to Apache Cassandra (Wide-Column NoSQL).
    • Why Cassandra? Chat messages are strictly sequential, time-series data indexed by chat_id and timestamp. Cassandra's LSM-Tree engine delivers blistering write throughput and linear scaling across commodity disks.
  4. The service queries Redis Session Store (user_id -> gateway_server_ip) to locate the recipient's active connection.
    • If Online: Forward message to Chat Gateway 2, which pushes it over the recipient's WebSocket (two grey ticks: Delivered). When the recipient opens the conversation, an event fires the blue ticks (Read).
    • If Offline: Enqueue a message to the Push Notification Service (Apple APNs / Google FCM).

3. Presence & Group Messaging Optimization

  • Presence (Online/Offline status): Clients send a heartbeat ping every 5 seconds over the WebSocket. Redis stores the status with a 15-second TTL (SET presence:user_42 "online" EX 15). If a heartbeat is missed, the key expires and the user is marked offline.
  • Group Chats: For a 500-person group, naive fanout (copying the message 500 times) overloads the network. Instead, the message is written once to a shared group mailbox partition in Cassandra. Each member's gateway simply reads the message using the group's chronological offset pointer.

Case 3: Design a Food Delivery App (e.g., Zomato / Gojek / DoorDash)

Food delivery requires real-time coordination across three distinct personas: Customers, Restaurants, and Delivery Drivers.

┌─────────────────────────────────────────────────────────────────────────────┐ │ FOOD DELIVERY GEOSPATIAL ARCHITECTURE │ └─────────────────────────────────────────────────────────────────────────────┘ [ Driver App ] ──Location Ping (every 4s)──► [ Location Tracking Service ] │ ▼ [ Redis Geospatial ] (GEOADD drivers:city lng lat id) │ [ Customer ] ──Order Placed──► [ Order State Machine ]│ │ │ ▼ ▼ [ Driver Dispatch Matching Engine ] (GEORADIUS: Find 10 drivers within 3km)

1. Geospatial Indexing: Finding Drivers in Real Time

Standard SQL queries (WHERE lat BETWEEN x AND y) perform slow table scans on dynamic data. We must partition 2D space:

┌──────────────────────────────────────┬──────────────────────────────────────┐ │ Geohash (Base32) │ Google S2 / Uber H3 │ ├──────────────────────────────────────┼──────────────────────────────────────┤ │ • Divides earth into rectangular │ • Divides earth into hierarchical │ │ hierarchical grid cells. │ hexagons (H3) or Hilbert curves(S2)│ │ • Strings share common prefixes: │ • Hexagonal cells have equidistant │ │ "qqgu1" is adjacent to "qqgu2". │ neighbors (perfect for radius math)│ │ • Easily stored in B-Trees / Redis. │ • Industry standard for ride-hailing.│ └──────────────────────────────────────┴──────────────────────────────────────┘

Production Choice: Use Redis Geospatial Data Structures (GEOADD, GEORADIUS, GEOSEARCH). Under the hood, Redis encodes coordinates into a 52-bit integer Geohash stored inside a Sorted Set (ZSET), allowing $O(\log N + M)$ radius queries in under 2 milliseconds!

2. Driver Location Ingestion Pipeline

  • 500,000 active drivers ping GPS coordinates every 4 seconds = 125,000 writes/second.
  • Pings are ingested via UDP or lightweight gRPC into a streaming cluster (Kafka topic: driver-locations).
  • Flink / Kafka Streams deduplicates and flushes the latest coordinate into Redis with a 30-second TTL.

3. The Order State Machine & Dispatch Engine

An order is a strict, finite state machine: $$\text{Created} \longrightarrow \text{Paid} \longrightarrow \text{Accepted by Restaurant} \longrightarrow \text{Driver Assigned} \longrightarrow \text{Picked Up} \longrightarrow \text{Delivered}$$

Dispatch Algorithm:

  1. When the restaurant confirms the meal is 10 minutes away from preparation, the Matching Engine executes a geospatial query for drivers within a 3km radius.
  2. Drivers are scored based on proximity, battery level, current acceptance rate, and vehicle type.
  3. A notification is sent to the top-ranked driver with an accept window of 15 seconds. If declined or timed out, a Distributed Lock (Redlock) releases the order and notifies Candidate #2.

Case 4: Design a Scalable Notification System

Modern platforms send billions of notifications daily across multiple communication channels (Mobile Push, SMS, Email, In-App Feeds).

┌─────────────────────────────────────────────────────────────────────────────┐ │ SCALABLE NOTIFICATION ENGINE PIPELINE │ └─────────────────────────────────────────────────────────────────────────────┘ [ Internal Microservices ] │ ▼ [ Notification API Gateway ] ──► Validates Rate Limits & User Preferences │ ▼ [ Priority Kafka Topics ] ├── topic: notifications-high (OTPs, Account Security, Payments) └── topic: notifications-low (Promotions, Weekly Summaries) │ ▼ [ Worker Fleet / Template Engine ] ──► Renders HTML / Push payloads │ ┌─────────┼─────────┬─────────┐ ▼ ▼ ▼ ▼ [APNs] [FCM] [Twilio] [SendGrid] (iOS) (Android) (SMS) (Email)

1. Requirements & Reliability Guarantees

  • Never drop critical messages: OTP verification codes must arrive within 5 seconds.
  • User opt-out preferences: Users can disable marketing SMS without disabling security emails.
  • Deduplication: Never spam a user with identical push alerts due to network retries.

2. Multi-Queue Prioritization with Kafka

Do not mix transactional and promotional messages in the same queue:

  • High-Priority Topic: OTPs, password resets, payment notifications. Workers are heavily provisioned with zero backlog tolerance.
  • Low-Priority Topic: Marketing campaigns, newsletter digests. Workers process batches during low-traffic windows.

3. Per-User Rate Limiting & Deduplication Window

  • A distributed rate limiter prevents spamming a single user with more than 3 marketing pushes per hour.
  • Deduplication Window: When generating a notification, compute a hash of (user_id, channel, event_type, date_hour). Store the hash in Redis with an EXPIRE of 1 hour. If an identical hash is submitted, silently drop the duplicate.

4. Third-Party Vendor Failover & DLQ

Third-party providers (Sendgrid, Twilio) experience outages.

  • Wrap external vendor API calls with Circuit Breakers.
  • If Twilio fails, automatically failover to a secondary SMS provider (e.g., MessageBird or AWS SNS).
  • If all retries fail with exponential backoff, route the message to a Dead Letter Queue (DLQ) for engineering inspection and replay.

Case 5: Design a Distributed Rate Limiter

APIs must enforce quotas (e.g., 100 requests per minute per IP or API Key) to prevent abuse and ensure fairness.

┌─────────────────────────────────────────────────────────────────────────────┐ │ DISTRIBUTED RATE LIMITER ARCHITECTURE │ └─────────────────────────────────────────────────────────────────────────────┘ [ Inbound HTTP Request ] │ ▼ [ API Gateway Filter Layer ] │ ▼ [ Local Memory L1 Cache (Optional) ] (Instant pass for known low-rate IPs) │ ▼ [ Redis Cluster (Atomic Lua Engine) ] - Key: "rl:{tenant_id}:{minute_window}" - Algorithm: Sliding Window Counter │ ┌───────────────┴───────────────┐ ▼ ▼ [ Count <= Quota ] [ Count > Quota ] Allow & Forward Reject with HTTP 429 to Microservice "Retry-After: 34"

1. Sliding Window Counter in Redis

A sliding window counter blends the request counts from the previous time window and the current time window:

$$\text{Estimated Count} = \text{Current Window Count} + \left(\text{Previous Window Count} \times \left(1 - \frac{\text{Current Time Offset}}{\text{Window Size}}\right)\right)$$

This delivers sub-millisecond calculation speed, consumes negligible memory ($<50\text{ bytes per key}$), and eliminates edge-of-window traffic spikes.

2. Standard HTTP Response Headers

Always communicate rate limiting quotas back to the client via standard RFC headers:

HTTP

HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 28
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1775568000

{
  "error": "rate_limit_exceeded",
  "message": "You have exceeded your quota of 100 requests per minute. Try again in 28 seconds."
}

The System Design Interview Master Cheatsheet

┌─────────────────────────────────────────────────────────────────────────────┐ │ FAANG INTERVIEW SURVIVAL MATRIX │ ├──────────────────────┬──────────────────────────────────────────────────────┤ │ Scenario │ Go-To Architecture Move │ ├──────────────────────┼──────────────────────────────────────────────────────┤ │ Read-Heavy Workload │ Redis Cache-Aside + PostgreSQL Read Replicas + CDN │ │ Write-Heavy Stream │ Kafka Ingestion + Cassandra / DynamoDB (LSM-Tree) │ │ Complex Financials │ Relational DB + ACID + Saga Pattern + Outbox │ │ Real-Time Latency │ WebSockets + Redis Session Registry + In-Memory ZSET │ │ Geospatial Search │ Redis Geospatial / Uber H3 / Google S2 Hexagons │ │ Unreliable Network │ Idempotency-Key + Exponential Backoff with Jitter │ │ Spiky Traffic Spurt │ API Gateway Rate Limiter (Token Bucket) + Kafka Buff │ └──────────────────────┴──────────────────────────────────────────────────────┘

Conclusion: Your Journey to System Design Mastery

Congratulations! You have completed the entire System Design Roadmap:

  1. Part 1: Networking Basics gave you the protocol foundations (Packets, TCP, UDP, TLS 1.3, DNS, REST).
  2. Part 2: Core Building Blocks provided the physical components (LBs, CDNs, Redis, API Gateways, SQL vs NoSQL, Kafka).
  3. Part 3: Distributed Patterns & Concepts equipped you with theory and resilient patterns (CAP, Consistent Hashing, Sharding, Idempotency, Outbox, Saga, CQRS).
  4. Part 4: Real-World Interview Systems synthesized everything into production designs for TinyURL, WhatsApp, Food Delivery, Notification Systems, and Rate Limiters.

System design is not about memorizing static diagrams. It is the discipline of weighing engineering trade-offs under constraints. When you walk into your next interview, do not seek the "perfect" architecture—show the interviewer how you navigate latency, cost, consistency, and resilience to engineer the right architecture.

Continue Reading

Previous article

← Previous Article

System Design Roadmap Part 3: Distributed Patterns & Concepts – CAP Theorem, Hashing, Sharding & Sagas

Next Article →

Rebuilding Laravel E-Learning to Java Spring Boot & React: Enterprise Architecture & Cloud-Ready Migration

Next article