MUHAMMAD
FARI
MADYAN
[ Press ESC or Click to Skip ]

System Design Roadmap

1

Part 1: Networking Basics – Packets, TCP/UDP, TLS & REST

2

Part 2: Core Building Blocks – Load Balancers, CDN, Caching, API Gateway, DBs & Kafka

3

Part 3: Distributed Patterns & Concepts – CAP Theorem, Hashing, Sharding & Sagas

4

Part 4: Real-World Interview Systems – URL Shortener, Chat, Food Delivery & Notification

This article is available in Indonesian

🇮🇩 Baca dalam Bahasa Indonesia
🇬🇧 English📚 System Design Roadmap

System Design Roadmap Part 1: Networking Basics – Packets, TCP/UDP, TLS & REST

Master foundational networking concepts for system design: how the internet routes packets, TCP vs UDP trade-offs, HTTP/HTTPS security with TLS 1.3, DNS resolution deep-dive, and stateless REST API architectural principles.

Muhammad Fari MadyanAuthor

13 min read

·

Oct 7, 2026


Introduction: Why Networking Dictates Distributed Systems

When preparing for Senior or Staff Software Engineer System Design interviews, engineers often jump immediately to high-level blocks: "Let's slap Redis here, Kafka there, and put 10 microservices behind an API Gateway."

However, every distributed system is fundamentally a collection of independent computers exchanging bytes across an unreliable physical network. If you cannot reason about packet loss, connection handshakes, round-trip times (RTT), encryption overhead, and transport protocols, your high-level architecture is merely theoretical.

When an interviewer asks:

"What happens when you type https://api.myapp.com/v1/orders in your browser and press Enter?"

They are not just checking trivia. They are evaluating whether you understand latency budgets, DNS failover, TLS negotiation penalties, and REST idempotency under network partitions.

In this first part of our System Design Roadmap, we build the bedrock foundation: Networking Basics.

┌─────────────────────────────────────────────────────────────────────────────┐ │ THE CLIENT-TO-SERVER JOURNEY │ └─────────────────────────────────────────────────────────────────────────────┘ [ Client Browser ] │ │ 1. DNS Resolution (Browser -> OS -> Recursive -> Root -> TLD -> Auth) ▼ [ Resolved IP: 198.51.100.42 ] │ │ 2. TCP 3-Way Handshake (SYN -> SYN-ACK -> ACK) ▼ [ TCP Connection Established ] │ │ 3. TLS 1.3 Cryptographic Handshake (Key Exchange + Cipher) ▼ [ Secure HTTPS Pipeline ] │ │ 4. HTTP/2 or HTTP/3 REST Request (GET /v1/orders) ▼ [ Edge Network / Load Balancer / API Gateway ]

1. How the Internet Works: Packets, IP Addresses, and Routing

The internet is a packet-switched network. Unlike legacy telephone circuits that reserved dedicated copper wires for the duration of a call, the internet slices data into independent chunks called packets (datagrams).

Packet Anatomy and Encapsulation

Every layer of the network model prepends its own metadata headers:

┌─────────────────────────────────────────────────────────────────────────────┐ │ PACKET ENCAPSULATION STACK │ ├─────────────────────────────────────────────────────────────────────────────┤ │ Layer 2 (Data Link): [ Ethernet Header: MAC Dest | MAC Src | EtherType ] │ │ Layer 3 (Network): [ IP Header: Source IP | Dest IP | TTL | Protocol ] │ │ Layer 4 (Transport): [ TCP/UDP Header: Port Src | Port Dest | Seq / Ack ]│ │ Layer 7 (Application): [ HTTP Payload: Headers + Body Data ] │ └─────────────────────────────────────────────────────────────────────────────┘
  1. MTU (Maximum Transmission Unit): The maximum frame size that can traverse an Ethernet link (typically 1,500 bytes).
  2. MSS (Maximum Segment Size): The maximum TCP payload size excluding IP (20 bytes) and TCP (20 bytes) headers, usually 1,460 bytes.
  3. Fragmentation: If a packet exceeds the MTU of any intermediate router, it must either be fragmented (which degrades throughput) or dropped with an ICMP "Fragmentation Needed" signal (Path MTU Discovery).

IP Addresses and Subnetting

  • IPv4: 32-bit addresses formatted in four octets (192.168.1.1). Total space: ~4.3 billion addresses (exhausted, leading to NAT and CIDR).
  • IPv6: 128-bit hexadecimal addresses (2001:0db8:85a3::8a2e:0370:7334). Solves address exhaustion and eliminates the need for NAT.
  • CIDR (Classless Inter-Domain Routing): E.g., 10.0.0.0/16. The /16 signifies that the first 16 bits define the network mask, leaving $32 - 16 = 16$ bits ($2^{16} - 2 = 65,534$) for assignable host IPs inside a VPC (Virtual Private Cloud).

Routing: How Packets Traverse the Globe

Routers do not know the complete path from source to destination in advance. They consult internal Routing Tables and forward packets to the best Next Hop:

  • BGP (Border Gateway Protocol): The routing backbone of the global internet. The world is split into tens of thousands of Autonomous Systems (AS) managed by ISPs, cloud giants (AWS, Google, Cloudflare), and telecom operators. BGP negotiates peering paths between these AS networks.
  • Anycast Routing: In an Anycast configuration, multiple geographically distributed data centers advertise the exact same IP address via BGP. Routers along the path automatically steer incoming client traffic to the topologically closest data center. This is how CDNs and DNS resolvers (e.g., Cloudflare's 1.1.1.1 and Google's 8.8.8.8) achieve single-digit millisecond response times globally.

2. TCP vs. UDP: Reliability vs. Speed Trade-Offs

At Layer 4 (Transport), almost all internet applications choose between two fundamental protocols: TCP or UDP.

┌──────────────────────────────────────┬──────────────────────────────────────┐ │ TCP (Transmission Control) │ UDP (User Datagram Protocol) │ ├──────────────────────────────────────┼──────────────────────────────────────┤ │ • Connection-oriented │ • Connectionless │ │ • Guaranteed delivery (Retransmit) │ • Fire-and-forget (Loss-tolerant) │ │ • Strict in-order byte stream │ • No ordering guarantee │ │ • Flow control & Congestion control │ • No congestion throttles │ │ • Higher latency (3-way handshake) │ • Ultra-low latency (0-RTT delivery) │ │ • Heavy state per connection │ • Lightweight memory footprint │ └──────────────────────────────────────┴──────────────────────────────────────┘

Deep Dive: TCP Mechanisms

TCP provides the illusion of a clean, infallible stream of bytes over an inherently lossy network through three core mechanisms:

1. The 3-Way Handshake & 4-Way Teardown

Before transmitting a single byte of HTTP payload, TCP must synchronize state:

Client Server │ │ │ ─── SYN (Seq=X) ───────────────────────> │ 1. Client initiates │ <── SYN-ACK (Seq=Y, Ack=X+1) ─────────── │ 2. Server acknowledges & syncs │ ─── ACK (Seq=X+1, Ack=Y+1) ────────────> │ 3. Connection ESTABLISHED! │ │

This handshake costs 1 full Round-Trip Time (RTT). If the client is in Jakarta and the server is in Virginia (~220ms RTT), the handshake alone burns nearly a quarter of a second before application data moves!

2. Flow Control (Sliding Window)

Prevents a fast sender from overwhelming a slow receiver's buffer. The receiver continuously advertises a Receive Window (rwnd) in its TCP header. The sender cannot transmit more unacknowledged bytes than rwnd.

3. Congestion Control

Prevents senders from overwhelming the shared routers on the internet backbone:

  • Slow Start: Begins with a small Congestion Window (cwnd = 10 MSS). On every ACK received, cwnd doubles exponentially until reaching ssthresh (slow-start threshold).
  • Congestion Avoidance (AIMD): Additive Increase / Multiplicative Decrease. Once ssthresh is crossed, cwnd grows linearly. When packet drop is detected, cwnd is cut in half.

4. The Achilles' Heel: Head-of-Line (HoL) Blocking

Because TCP guarantees strict in-order delivery, if Packet #2 is dropped while Packets #3, #4, and #5 arrive safely, the operating system holds #3-#5 in its buffer. The application cannot consume #3-#5 until #2 is retransmitted and acknowledged.

When to Choose UDP

UDP simply wraps data with an 8-byte header (Source Port, Destination Port, Length, Checksum) and transmits immediately. If a packet drops, UDP does not care.

Ideal Use Cases:

  • Live Video Conferencing (Zoom, Google Meet, WebRTC): Dropping frame #45 is preferable to stalling the entire live conversation waiting for a retransmission.
  • Online Multiplayer Gaming: Real-time player coordinates must represent the present moment; stale positions are useless.
  • DNS Lookups: Fast request-response cycle that can easily be retried if unanswered.
  • QUIC / HTTP/3: Re-implements custom reliability, encryption, and congestion control in user space over UDP to bypass TCP's Head-of-Line blocking!

3. HTTP/HTTPS & TLS 1.3: Securing the Web Layer

┌─────────────────────────────────────────────────────────────────────────────┐ │ HTTP GENERATIONAL EVOLUTION │ ├─────────────────────────────────────────────────────────────────────────────┤ │ HTTP/1.1 (1997): Text-based, persistent keep-alive, 1 request per conn HoL │ │ HTTP/2 (2015): Binary framing, multiplexing streams over 1 TCP conn │ │ HTTP/3 (2022): QUIC over UDP, zero transport HoL, 0-RTT connection reuse │ └─────────────────────────────────────────────────────────────────────────────┘

HTTP Evolution

  1. HTTP/1.1: Human-readable text format (GET / HTTP/1.1\r\n). Introduced Connection: keep-alive to reuse TCP connections. However, it suffered from HTTP Head-of-Line blocking: only one request could be in flight per TCP connection at a time. Browsers worked around this by opening 6 simultaneous TCP connections per domain.
  2. HTTP/2: Introduced Binary Framing Layer. A single TCP connection is split into multiple concurrent bidirectional streams. Requests and responses interleave frames concurrently without blocking each other. Also added HPACK header compression and Server Push.
  3. HTTP/3 (QUIC): While HTTP/2 solved application-layer HoL blocking, a single packet drop at the TCP layer still froze all concurrent streams. HTTP/3 replaces TCP with QUIC (running over UDP). Each stream is truly isolated; a packet drop on stream A does not stall stream B!

TLS 1.3 Handshake: Speed Meets Cryptography

Plain HTTP transmits text in cleartext, vulnerable to eavesdropping and Man-In-The-Middle (MITM) attacks. HTTPS wraps HTTP inside Transport Layer Security (TLS).

While TLS 1.2 required 2 RTTs to complete its handshake, TLS 1.3 optimized it to just 1 RTT (and supports 0-RTT resumption for returning visitors):

Client Server │ │ │ ── ClientHello ─────────────────────────────> │ │ + Supported Ciphers │ │ + Key Share (ECDH Public Key) │ │ │ │ <── ServerHello ──────────────────────────── │ │ + Selected Cipher │ │ + Key Share (ECDH Public Key) │ │ + Server Certificate (X.509) │ │ + EncryptedExtensions + Finished │ │ │ │ ── [Application Data: GET /orders] ─────────> │ │ │

Cryptographic Pillars:

  • Asymmetric Encryption (RSA, ECC / ECDHE): Used only during the handshake to authenticate the server's certificate and securely derive a shared secret.
  • Symmetric Encryption (AES-GCM, ChaCha20-Poly1305): Once the shared secret is established, all ongoing payload bytes are encrypted symmetrically, which is computationally fast and hardware-accelerated by modern CPU instructions (AES-NI).
  • Forward Secrecy: If an attacker steals the server's private RSA certificate five years from now, they still cannot decrypt previously recorded network traffic because each session generated an ephemeral, single-use Diffie-Hellman key.

4. DNS Resolution: The Global Phonebook

Domain names (mfarim.com) exist because human beings are terrible at remembering numeric IP addresses (104.21.72.184).

The process of converting a human-friendly domain into a machine-routable IP address is called DNS Resolution.

┌─────────────────────────────────────────────────────────────────────────────┐ │ DNS RESOLUTION HIERARCHY │ └─────────────────────────────────────────────────────────────────────────────┘ 1. Client queries Local Recursive Resolver (ISP or 1.1.1.1 / 8.8.8.8) 2. Recursive Resolver queries Root Nameserver (".") └── Returns TLD Nameserver for ".com" 3. Recursive Resolver queries TLD Nameserver (".com") └── Returns Authoritative Nameserver for "mfarim.com" (e.g., Cloudflare DNS) 4. Recursive Resolver queries Authoritative Nameserver └── Returns A Record: 104.21.72.184 5. Recursive Resolver caches answer (honoring TTL) and returns to Client

Essential DNS Record Types for System Design

RecordTypePurpose in System Design
AAddressMaps domain directly to an IPv4 address (api.foo.com -> 198.51.100.1).
AAAAIPv6 AddressMaps domain to a 128-bit IPv6 address.
CNAMECanonical NameAlias from one domain to another domain (static.foo.com -> d123.cloudfront.net). Cannot live at root apex (@).
ALIAS / ANAMEVirtual AliasVendor-specific (Route53, Cloudflare) pseudo-record allowing alias at domain apex (foo.com -> alb-123.amazonaws.com).
MXMail ExchangeDirects email to mail servers (e.g., Google Workspace).
TXTArbitrary TextUsed for domain verification, SPF, DKIM, and SSL challenge validation.

DNS in High-Availability & Traffic Routing:

  1. TTL (Time to Live): The duration (in seconds) resolvers are permitted to cache the DNS record.
    • High TTL (e.g., 86400s / 24h): Low latency, reduces resolver load, but slow to propagate emergency IP cutovers.
    • Low TTL (e.g., 60s): Enables rapid blue/green failover, but increases DNS query latency and cost.
  2. GeoDNS Routing: Resolvers in Singapore receive Singapore IP addresses; resolvers in Frankfurt receive Frankfurt IP addresses.
  3. Weighted DNS Routing: Route 10% of global traffic to a new cluster for canary releases.

5. REST APIs: Designing Stateless, Resilient Endpoints

In modern web architecture, systems communicate primarily via REST (Representational State Transfer) APIs over HTTP.

The Golden Rule: Statelessness

A stateless server does not store client session context in its local memory between requests. Every request must contain all necessary data to authenticate and execute the transaction (e.g., in a Bearer JWT header).

STATEFUL (Bottleneck): STATELESS (Scalable): ┌──────────┐ Session in RAM ┌──────────┐ │ Client A │ ─────────────► [Server 1] │ Client A │ (Carries JWT Token) └──────────┘ [Server 2] └──────────┘ If Server 1 crashes, user │ Can route to ANY node! is logged out! ├───► [Server 1] └───► [Server 2]

HTTP Verbs, Idempotency, and Safety

Understanding the mathematical distinction between Safe and Idempotent methods is critical during system design interviews:

HTTP VerbPurposeSafe?Idempotent?Retry Safely on Network Timeout?
GETRetrieve resourceYesYesYes
HEADRetrieve headers onlyYesYesYes
PUTReplace resource completelyNoYesYes
DELETERemove resourceNoYesYes
POSTCreate resource / execute commandNoNoNo (Needs Idempotency Key)
PATCHPartial resource updateNoGenerally NoConditionally

Idempotent Definition: An operation is idempotent if executing it once yields the exact same server state as executing it $N$ times.

  • PUT /users/42 { "name": "Fari" } is idempotent. Repeated calls leave the name as "Fari".
  • POST /orders is not idempotent. If the network drops the response and the client naively retries, the customer gets charged twice!

Production-Grade Pagination: Offset vs. Keyset/Cursor

When an API returns lists (e.g., GET /v1/products), never return unbounded collections. You must choose an appropriate pagination strategy:

1. Offset-based Pagination

HTTP

GET /v1/products?limit=20&offset=100000
  • Pros: Easy to implement; supports jumping directly to "Page 50".
  • Cons: Severe database performance penalty ($O(N)$ scanning in PostgreSQL/MySQL). As the offset climbs, the database must scan 100,020 rows and discard the first 100,000. Additionally suffers from data drift (duplicated or skipped items when rows are inserted concurrently).

2. Keyset / Cursor-based Pagination (Industry Standard)

HTTP

GET /v1/products?limit=20&after_cursor=prod_98a72b

SQL

SELECT id, title, price 
FROM products 
WHERE id > 'prod_98a72b' 
ORDER BY id ASC 
LIMIT 20;
  • Pros: Consistent $O(1)$ B-Tree index seek time regardless of whether you are reading the 1st page or the 10,000,000th page. Immune to real-time insertion shifts.
  • Cons: Cannot skip ahead to page 42; only supports sequential navigation.

Interview Cheatsheet: Networking Traps & Trade-offs

When an interviewer tests your networking depth, keep these battle-tested principles in mind:

  1. "The Network is Reliable" Fallacy: The network is always slow, always lossy, and will eventually partition. Never assume a remote call succeeds. Every outbound call requires timeouts, circuit breakers, and retry policies.
  2. TCP Handshake Tax: In cross-region calls (e.g., EU client to US server), establishing a new TCP connection costs ~150-200ms before data transmits. Always maintain connection pools and HTTP/2 persistent connections.
  3. BGP Hijacking & Anycast: Mention Anycast when designing global edge networks (Cloudflare, AWS CloudFront, Route53) to absorb DDoS attacks and minimize latency.
  4. Idempotency Keys: Whenever designing payment or order creation APIs (POST), always mandate an Idempotency-Key: uuid header stored in Redis to guarantee exactly-once processing semantics over an unreliable network.

What's Next in Part 2?

Now that we understand how packets flow across the web, how connections are established, and how APIs are structured, we are ready to assemble high-scale architectures.

In Part 2: Core Building Blocks, we will break down:

  • Load Balancers (L4 vs. L7, algorithms, and high-availability setups)
  • Content Delivery Networks (CDNs) (Push vs. Pull, cache invalidation, Origin Shielding)
  • Distributed Caching with Redis (Cache-Aside, Write-Through, Write-Behind, and eviction strategies)
  • API Gateways (Rate limiting, auth offloading, SSL termination)
  • Databases (SQL vs. NoSQL, indexing, and storage engine internals)
  • Message Queues & Event Streaming (RabbitMQ vs. Apache Kafka)

Continue Reading

Next Article →

System Design Roadmap Part 2: Core Building Blocks – Load Balancers, CDN, Caching, API Gateway, DBs & Kafka

Next article