← Interview Preparation

Distributed Systems for Interviews

Multi-node correctness for senior+ loops — partial failure, consistency models, replication and quorums, partitioning, logical clocks, Raft, distributed transactions, and log-based messaging. Assumes Networking, OS, and Concurrency are done; no REST API design or single-database SQL. Primary textbook: Designing Data-Intensive Applications (Kleppmann). Topic 9 (Deep Cuts) is optional infra-oriented material — CRDTs, Spanner, HLC, resilience patterns, and advanced Raft.

Time budget: ≈43h

How to use this reference

  • Work through topics top to bottom — Raft assumes you understand replication; sagas assume you understand at-least-once delivery from messaging.
  • DDIA is the anchor book — each subtopic points to specific chapters, not cover-to-cover reading. Fowler's Patterns of Distributed Systems complements for concise pattern names.
  • We deliberately do not repeat Networking (TCP, DNS, LB), OS (processes, clocks), or Concurrency (mutexes on one machine) — those are prerequisites.
  • Optional tasks favor paper drills and browser visualizers (Raft); one optional Docker Kafka exercise if you already run containers.
  • Topic 9 — Deep Cuts (Infra) is optional. Skip it for standard backend system design; do it for storage, platform, or consistency-heavy roles (Spanner-style loops, etcd/Kafka platform teams). Adds ≈10h on top of the core path.
  • Topic 10 — Optional Coding Practice maps concepts to in-memory coding exercises and a few LeetCode problems — no cluster required. Skip if your loop is design-only.

The Reference

  1. 1

    Partial failure, CAP/PACELC, and how to estimate scale and assemble building blocks in a design round.

    1. 1.1Partial Failures & the Fallacies of Distributed Computing2/520m
    2. 1.2CAP & PACELC3/520m
    3. 1.3Capacity Estimation & Design Assembly3/535m
  2. 2

    What "consistent" means to clients — from linearizability to eventual.

    1. 2.1From Linearizability to Eventual Consistency4/530m
    2. 2.2Session Guarantees & Staleness Bounds3/520m
  3. 3

    Multiple copies of data — leaders, followers, lag, and quorum math.

    1. 3.1Leader-Based Replication3/525m
    2. 3.2Quorum Reads & Writes4/525m
  4. 4

    Splitting data across nodes — partition keys, hot spots, and resharding.

    1. 4.1Partition Keys, Skew & Resharding3/525m
  5. 5

    Clocks lie, events need order — Lamport, vectors, and conflict detection.

    1. 5.1Lamport & Vector Clocks4/525m
    2. 5.2Conflicts, LWW & Version Vectors3/520m
  6. 6

    Agreeing on one answer — Raft, leader election, locks, and fencing.

    1. 6.1Consensus & Raft4/530m
    2. 6.2Distributed Locks & Fencing Tokens4/520m
  7. 7

    Atomicity across services and databases — 2PC limits, Sagas, and the outbox.

    1. 7.1Two-Phase Commit & Its Limits4/525m
    2. 7.2Sagas & Transactional Outbox4/530m
  8. 8

    Commit logs, Kafka semantics, and delivery guarantees across distributed components.

    1. 8.1Commit Logs, Partitions & Consumer Groups3/525m
    2. 8.2Delivery Semantics & Exactly-Once4/525m
  9. 9

    Optional hard-core material for storage, platform, and consistency-heavy loops — skip unless your target role needs it.

    1. 9.1CRDTs & Conflict-Free Replication5/530m
    2. 9.2Spanner & Global Consistency5/530m
    3. 9.3Hybrid Logical Clocks4/525m
    4. 9.4Resilience Patterns4/525m
    5. 9.5Multi-Paxos & Zab5/525m
    6. 9.6Raft Advanced5/530m
  10. 10

    In-memory exercises and autograded problems — no cluster, Docker, or cloud setup required.

    1. 10.1Placement & Consistent Hashing3/515m
    2. 10.2Versioning & Staleness3/510m
    3. 10.3Rate Limits & Idempotency3/510m
    4. 10.4Replication & Delivery Semantics3/510m