System design explainer

Design distributed storage

The interview question behind S3: you upload a photo and it survives disks dying, racks failing, data centers flooding. How do you store exabytes that durably — and what's the cheapest redundancy that still survives?

The takeaway, up front

Durability is bought with redundancy, and there are exactly two ways to buy it: 3× replication — every chunk copied to 3 nodes; simple, fast to read, 200% storage overhead, survives any 2 failures — or Reed-Solomon erasure coding (6+3) — 6 data chunks plus 3 parity chunks spread over 9 nodes; only 50% overhead and survives any 3 failures, but every read touches 6 nodes and repairs cost CPU. Hot data gets replicated, cold data gets erasure-coded. The interview answer is knowing which wins when.

1 · Requirements

What are we actually building?

The analogy: you have an important document and two safes strategies. Replication is photocopying it 3 times and putting each copy in a different building — grab any copy to read. Erasure coding is shredding it into 9 strips where any 6 reassemble the original — far less paper, but reading means collecting 6 strips and doing a puzzle. Same safety goal, completely different cost profile.

Functional

  • PUT / GET / DELETE objects (blobs, GBs to TBs)
  • 11 nines of durability (99.999999999%)
  • Survive disk, node, and rack failures without data loss
  • Automatic repair: detect missing chunks, rebuild them
  • Range reads and streaming for large objects

Non-functional

  • Durability: 11 nines — the headline SLA
  • Availability: 99.9%+ for reads even during failures
  • Throughput: GB/s aggregate per cluster
  • Efficiency: minimize raw bytes per usable byte
  • Repair speed: rebuild lost redundancy before the next failure hits

🚫 Common misconception

"Erasure coding is strictly better than replication because it uses less storage." Cheaper, yes — better, no. EC reads need k nodes instead of 1 (higher tail latency), repairs read k blocks to rebuild 1 (the "repair tax"), and encoding burns CPU. That's why real systems tier: replicate hot data for speed, erasure-code cold data for cost. Anyone who answers "just use EC everywhere" hasn't thought about repair.

2 · Back-of-the-envelope

Capacity math

Assumption3× replicationRS(6+3)
Usable data100 PB
Raw storage needed300 PB (200% overhead)150 PB (50% overhead)
Nodes (12×16TB = 192TB each)~1,563 nodes~782 nodes
Failures tolerated (any)23
Read cost per object1 node6 nodes + decode CPU
Repair cost for 1 lost 192TB nodecopy 192TB from replicasread ~6×192TB ≈ 1.1PB, recompute, rewrite
Rebuild time at 50Gbps/node~8.5 hourslonger — repair traffic is the bottleneck

The line to say out loud: "EC saves 150 petabytes of disks — roughly 780 fewer servers — but every repair reads six times the lost data. The savings are real and the tax is real."

Go deeper: where do "11 nines" actually come from?

Not from the redundancy scheme alone — from repair speed. Data loss needs k+1 overlapping failures before repair finishes. With 3× replication, losing data means 3 specific nodes dying within the ~8-hour rebuild window: astronomically unlikely, hence the nines. EC(6+3) tolerates more failures but repairs slower, so the math roughly balances. The durability SLA is really a statement about your repair pipeline, not your encoding.

3 · Architecture

The system, end to end

flowchart TB
    CL["Client"]
    MD["Metadata service
placement via consistent hashing
Raft-replicated"] N1["Chunk server 1
dumb disks + checksums"] N2["Chunk server 2"] N3["Chunk server N"] AUD["Auditor / repair worker
scrubs + rebuilds missing chunks"] CL -->|"PUT: where do chunks go?"| MD MD -->|"placement map"| CL CL -->|"write chunks"| N1 CL -->|"write chunks"| N2 CL -->|"write chunks"| N3 N1 -. "checksums, heartbeats" .-> AUD N2 -. "checksums, heartbeats" .-> AUD N3 -. "checksums, heartbeats" .-> AUD AUD -->|"rebuild missing chunk"| N1 MD -. "chunk map updates" .-> AUD

Two kinds of smarts: the metadata service knows where everything is (small, precious, Raft-replicated), and the chunk servers are deliberately dumb (big, cheap, replaceable). The auditor closes the loop — durability without repair is just delayed data loss.

4 · Component deep-dives

Writing an object

sequenceDiagram
    autonumber
    participant C as Client
    participant M as Metadata service
    participant S as Chunk servers
    C->>M: PUT /objects/photo.jpg — placement?
    M-->>C: chunk map (9 nodes, rack-aware)
    C->>C: split into 6 data chunks,
encode 3 parity (EC) — or 3 copies C->>S: write chunks in parallel S-->>C: acks + checksums C->>M: commit (quorum of chunks durable) M-->>C: 201 Created Note over M,S: later: node dies → auditor
detects missing chunks → rebuilds

Read path: replication vs EC

flowchart TB
    A["GET /objects/photo.jpg"] --> B{"storage scheme?"}
    B -->|"3x replication"| C["Read chunk from
nearest healthy replica"] B -->|"RS 6+3"| D["Read any 6 of 9 chunks
in parallel"] C --> E["200 OK — 1 node touched"] D --> F["Decode: reconstruct
original 6 data chunks"] F --> G["200 OK — 6 nodes touched
+ CPU + tail latency risk"]
Go deeper: rack-aware placement

Copies on 3 nodes in the same rack don't survive a rack power failure. Placement must spread chunks across failure domains: different racks, ideally different power and network. Consistent hashing picks the nodes; a rack-awareness constraint vetoes bad picks. In the interview, saying "replicas across racks" unprompted is a quiet senior signal.

Go deeper: why metadata is the scary part

Lose a chunk server and the auditor rebuilds it. Lose the metadata service and you have a warehouse full of unlabeled boxes — the bytes exist but nothing is findable. That's why metadata is tiny, Raft-replicated across 3–5 nodes, snapshotted, and backed up, while chunk servers are cattle. Durability of data means nothing without durability of the map.

5 · API + data model

The contract

PUT /v1/objects/{key}              # body: bytes, optional ?scheme=repl|ec
→ 201 Created {"key": "photo.jpg", "size": 48321152, "scheme": "rs6+3"}

GET /v1/objects/{key}              # Range: bytes=0-1048575 supported
→ 200 OK (streamed)
HEAD /v1/objects/{key}             # size, checksum, scheme
DELETE /v1/objects/{key}           # tombstone; chunks GC'd lazily

Data model: Object {key, size, scheme, chunk_map: [{chunk_id, node_ids[], checksum}]} in the metadata service. Chunk {id, bytes, checksum} on chunk servers — content-addressed, immutable. Nodes are just {id, rack, status, free_bytes}; the system never trusts them, it verifies them.

6 · Trade-offs

What you give up, on purpose

3× replicationRS(6+3)
Storage overhead200% — expensive50% — cheap
Read latency1 node, lowest tail latency6 nodes + decode — higher p99
Repair cost1:1 copy — cheap6:1 read amplification — the repair tax
CPU~zeroEncode/decode on every write/read
Small objectsFineWasteful — parity overhead dominates; pad or replicate
Best forHot data, small objectsCold/archival data, large objects
7 · Failure modes

What breaks, and what saves you

8 · What I'd actually build

Opinionated, concrete, shippable

API: S3-compatible (don't invent a protocol). Metadata: Raft-replicated service, or Postgres with Patroni if you want boring. Chunk servers: dumb daemons on commodity disks, checksums on everything. Encoding: Intel ISA-L for Reed-Solomon — never hand-roll the math. Policy: replicate objects under 1MB and anything hot; EC(6+3) everything else; lifecycle rule migrates cold objects automatically. That's the architecture every major object store converges on.

9 · Interview tips

How to run the room

  1. Open with the two schemes and the overhead math. "3× = 200% overhead, RS(6+3) = 50%" in minute two frames everything.
  2. Say "rack-aware placement" before you're asked. Free senior points.
  3. The repair-tax argument is your differentiator — most candidates forget repairs entirely.
  4. Draw the read-path comparison. "1 node vs 6 nodes" is visceral.
  5. Have the metadata answer ready: "the map is more precious than the territory."
  6. End with tiering: hot → replicated, cold → EC. It shows you've seen a real system.
10 · Interactive widget

Replication vs erasure coding — kill nodes, see who survives

Pick a scheme, place the blocks across 9 nodes, then click nodes to kill them (or use the button). Watch which scheme keeps your 6 data chunks recoverable — and what each one costs in raw storage.

Scheme: 3× replication Blocks placed: 0 Nodes down: 0 Overhead: 200% (3.0× raw)
Place the blocks, then kill nodes to test durability.
Storage cost per usable byte
3× replication
3.0×
RS(6+3)
1.5×

Replication reads from 1 node; RS(6+3) must read 6 and decode. Repairing one lost node copies 1 block under replication, but reads 6 blocks under EC — that's the repair tax from the article, and it's why hot data stays replicated.

v2026.10.03-01