Introduction: The Crucible of the System Design Interview
In Parts 1 through 3, we mastered the science: networking mechanics, scalability building blocks, and distributed patterns.
Now comes the art: The System Design Interview.
In a 45-minute FAANG or Tier-1 tech interview, you are given an intentionally ambiguous prompt like:
"Design WhatsApp" or "Design a Food Delivery System like DoorDash or Gojek."
Candidates who fail immediately begin drawing boxes. Candidates who succeed follow a structured, disciplined framework.
In this final chapter, we apply our roadmap to design 5 classic production systems step-by-step.
Case 1: Design a URL Shortener (e.g., TinyURL)
A URL shortener converts a long URL (https://example.com/products/deals/item-984218?ref=campaign) into an alias (https://tiny.url/aZ9k2m).
1. Requirements & Math
- Traffic: 100M new URLs created per month. Read-to-write ratio = $10:1$ (1 Billion reads/month).
- Write RPS: $100\text{M} / (30 \times 86,400\text{s}) \approx 40\text{ writes/sec}$.
- Read RPS: $\approx 400\text{ reads/sec}$ (peak: 2,000 RPS).
- Storage: Over 5 years = $100\text{M} \times 12 \times 5 = 6\text{ Billion records}$. At $500\text{ bytes/record} \approx 3\text{ TB}$ (easily managed by a single distributed table or sharded database).
2. The Core Technical Challenge: Generating the Short Key
A short key must be compact and collision-free. Using Base62 ([a-zA-Z0-9]), a 7-character string yields:
$$62^7 \approx 3.52\text{ Trillion unique combinations}$$
More than enough for decades!
Two approaches to generate the 7-character string:
- Hash of Original URL (
MD5/SHA-256): Taking the first 7 characters of MD5 causes hash collisions when different users shorten the same URL or during truncation. - Key Generation Service (KGS) with Pre-allocated Ranges (Industry Best Practice):
- A standalone service uses a distributed sequencer (or Apache ZooKeeper / Twitter Snowflake) to generate unique 64-bit integer IDs.
- Workers claim blocks of 10,000 IDs in memory.
- The integer is converted directly into Base62.
- Result: Guaranteed $O(1)$ collision-free tokens with zero database coordination overhead!
3. Redirection: HTTP 301 vs. HTTP 302
When a user visits https://tiny.url/aZ9k2m:
- HTTP 301 Moved Permanently: The browser caches the redirect locally. Subsequent clicks bypass your server entirely.
- Pros: Minimum server load and fastest user redirect.
- Cons: You cannot track click analytics or revoke URLs immediately.
- HTTP 302 Found (Temporary Redirect): The browser queries the URL shortener on every single click.
- Pros: Precise click telemetry (geographic location, referrers, device stats).
- Cons: Increases server traffic.
- Interview Recommendation: State this exact trade-off! Choose 302 if analytics are required, or 301 if raw cost efficiency is prioritized.
Case 2: Design a Real-Time Chat System (e.g., WhatsApp)
WhatsApp serves billions of users exchanging encrypted text, media, and presence statuses with sub-second latency.
1. Protocols: HTTP vs. WebSockets
- HTTP Polling / Long Polling wastes battery and server resources on mobile devices due to redundant HTTP headers and TCP handshakes.
- WebSockets establish a single, long-lived, bidirectional, lightweight TCP connection between client and chat gateway servers.
2. Message Flow & Delivery Semantics
- Sender writes a message over its active WebSocket to Chat Gateway 1.
- Chat Gateway 1 validates the message and returns a transient
Message Ack(giving the single grey tick: Sent). - The message is written to Apache Cassandra (Wide-Column NoSQL).
- Why Cassandra? Chat messages are strictly sequential, time-series data indexed by
chat_idandtimestamp. Cassandra's LSM-Tree engine delivers blistering write throughput and linear scaling across commodity disks.
- Why Cassandra? Chat messages are strictly sequential, time-series data indexed by
- The service queries Redis Session Store (
user_id -> gateway_server_ip) to locate the recipient's active connection.- If Online: Forward message to Chat Gateway 2, which pushes it over the recipient's WebSocket (two grey ticks: Delivered). When the recipient opens the conversation, an event fires the blue ticks (Read).
- If Offline: Enqueue a message to the Push Notification Service (Apple APNs / Google FCM).
3. Presence & Group Messaging Optimization
- Presence (Online/Offline status): Clients send a heartbeat ping every 5 seconds over the WebSocket. Redis stores the status with a 15-second TTL (
SET presence:user_42 "online" EX 15). If a heartbeat is missed, the key expires and the user is marked offline. - Group Chats: For a 500-person group, naive fanout (copying the message 500 times) overloads the network. Instead, the message is written once to a shared group mailbox partition in Cassandra. Each member's gateway simply reads the message using the group's chronological offset pointer.
Case 3: Design a Food Delivery App (e.g., Zomato / Gojek / DoorDash)
Food delivery requires real-time coordination across three distinct personas: Customers, Restaurants, and Delivery Drivers.
1. Geospatial Indexing: Finding Drivers in Real Time
Standard SQL queries (WHERE lat BETWEEN x AND y) perform slow table scans on dynamic data. We must partition 2D space:
Production Choice: Use Redis Geospatial Data Structures (GEOADD, GEORADIUS, GEOSEARCH). Under the hood, Redis encodes coordinates into a 52-bit integer Geohash stored inside a Sorted Set (ZSET), allowing $O(\log N + M)$ radius queries in under 2 milliseconds!
2. Driver Location Ingestion Pipeline
- 500,000 active drivers ping GPS coordinates every 4 seconds = 125,000 writes/second.
- Pings are ingested via UDP or lightweight gRPC into a streaming cluster (Kafka topic:
driver-locations). - Flink / Kafka Streams deduplicates and flushes the latest coordinate into Redis with a 30-second TTL.
3. The Order State Machine & Dispatch Engine
An order is a strict, finite state machine: $$\text{Created} \longrightarrow \text{Paid} \longrightarrow \text{Accepted by Restaurant} \longrightarrow \text{Driver Assigned} \longrightarrow \text{Picked Up} \longrightarrow \text{Delivered}$$
Dispatch Algorithm:
- When the restaurant confirms the meal is 10 minutes away from preparation, the Matching Engine executes a geospatial query for drivers within a 3km radius.
- Drivers are scored based on proximity, battery level, current acceptance rate, and vehicle type.
- A notification is sent to the top-ranked driver with an accept window of 15 seconds. If declined or timed out, a Distributed Lock (Redlock) releases the order and notifies Candidate #2.
Case 4: Design a Scalable Notification System
Modern platforms send billions of notifications daily across multiple communication channels (Mobile Push, SMS, Email, In-App Feeds).
1. Requirements & Reliability Guarantees
- Never drop critical messages: OTP verification codes must arrive within 5 seconds.
- User opt-out preferences: Users can disable marketing SMS without disabling security emails.
- Deduplication: Never spam a user with identical push alerts due to network retries.
2. Multi-Queue Prioritization with Kafka
Do not mix transactional and promotional messages in the same queue:
- High-Priority Topic: OTPs, password resets, payment notifications. Workers are heavily provisioned with zero backlog tolerance.
- Low-Priority Topic: Marketing campaigns, newsletter digests. Workers process batches during low-traffic windows.
3. Per-User Rate Limiting & Deduplication Window
- A distributed rate limiter prevents spamming a single user with more than 3 marketing pushes per hour.
- Deduplication Window: When generating a notification, compute a hash of
(user_id, channel, event_type, date_hour). Store the hash in Redis with anEXPIREof 1 hour. If an identical hash is submitted, silently drop the duplicate.
4. Third-Party Vendor Failover & DLQ
Third-party providers (Sendgrid, Twilio) experience outages.
- Wrap external vendor API calls with Circuit Breakers.
- If Twilio fails, automatically failover to a secondary SMS provider (e.g., MessageBird or AWS SNS).
- If all retries fail with exponential backoff, route the message to a Dead Letter Queue (DLQ) for engineering inspection and replay.
Case 5: Design a Distributed Rate Limiter
APIs must enforce quotas (e.g., 100 requests per minute per IP or API Key) to prevent abuse and ensure fairness.
1. Sliding Window Counter in Redis
A sliding window counter blends the request counts from the previous time window and the current time window:
$$\text{Estimated Count} = \text{Current Window Count} + \left(\text{Previous Window Count} \times \left(1 - \frac{\text{Current Time Offset}}{\text{Window Size}}\right)\right)$$
This delivers sub-millisecond calculation speed, consumes negligible memory ($<50\text{ bytes per key}$), and eliminates edge-of-window traffic spikes.
2. Standard HTTP Response Headers
Always communicate rate limiting quotas back to the client via standard RFC headers:
HTTP
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 28
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1775568000
{
"error": "rate_limit_exceeded",
"message": "You have exceeded your quota of 100 requests per minute. Try again in 28 seconds."
}
The System Design Interview Master Cheatsheet
Conclusion: Your Journey to System Design Mastery
Congratulations! You have completed the entire System Design Roadmap:
- Part 1: Networking Basics gave you the protocol foundations (Packets, TCP, UDP, TLS 1.3, DNS, REST).
- Part 2: Core Building Blocks provided the physical components (LBs, CDNs, Redis, API Gateways, SQL vs NoSQL, Kafka).
- Part 3: Distributed Patterns & Concepts equipped you with theory and resilient patterns (CAP, Consistent Hashing, Sharding, Idempotency, Outbox, Saga, CQRS).
- Part 4: Real-World Interview Systems synthesized everything into production designs for TinyURL, WhatsApp, Food Delivery, Notification Systems, and Rate Limiters.
System design is not about memorizing static diagrams. It is the discipline of weighing engineering trade-offs under constraints. When you walk into your next interview, do not seek the "perfect" architecture—show the interviewer how you navigate latency, cost, consistency, and resilience to engineer the right architecture.

