System design explainer
The interview question behind S3: you upload a photo and it survives disks dying, racks failing, data centers flooding. How do you store exabytes that durably — and what's the cheapest redundancy that still survives?
Durability is bought with redundancy, and there are exactly two ways to buy it: 3× replication — every chunk copied to 3 nodes; simple, fast to read, 200% storage overhead, survives any 2 failures — or Reed-Solomon erasure coding (6+3) — 6 data chunks plus 3 parity chunks spread over 9 nodes; only 50% overhead and survives any 3 failures, but every read touches 6 nodes and repairs cost CPU. Hot data gets replicated, cold data gets erasure-coded. The interview answer is knowing which wins when.
The analogy: you have an important document and two safes strategies. Replication is photocopying it 3 times and putting each copy in a different building — grab any copy to read. Erasure coding is shredding it into 9 strips where any 6 reassemble the original — far less paper, but reading means collecting 6 strips and doing a puzzle. Same safety goal, completely different cost profile.
"Erasure coding is strictly better than replication because it uses less storage." Cheaper, yes — better, no. EC reads need k nodes instead of 1 (higher tail latency), repairs read k blocks to rebuild 1 (the "repair tax"), and encoding burns CPU. That's why real systems tier: replicate hot data for speed, erasure-code cold data for cost. Anyone who answers "just use EC everywhere" hasn't thought about repair.
| Assumption | 3× replication | RS(6+3) |
|---|---|---|
| Usable data | 100 PB | |
| Raw storage needed | 300 PB (200% overhead) | 150 PB (50% overhead) |
| Nodes (12×16TB = 192TB each) | ~1,563 nodes | ~782 nodes |
| Failures tolerated (any) | 2 | 3 |
| Read cost per object | 1 node | 6 nodes + decode CPU |
| Repair cost for 1 lost 192TB node | copy 192TB from replicas | read ~6×192TB ≈ 1.1PB, recompute, rewrite |
| Rebuild time at 50Gbps/node | ~8.5 hours | longer — repair traffic is the bottleneck |
The line to say out loud: "EC saves 150 petabytes of disks — roughly 780 fewer servers — but every repair reads six times the lost data. The savings are real and the tax is real."
Not from the redundancy scheme alone — from repair speed. Data loss needs k+1 overlapping failures before repair finishes. With 3× replication, losing data means 3 specific nodes dying within the ~8-hour rebuild window: astronomically unlikely, hence the nines. EC(6+3) tolerates more failures but repairs slower, so the math roughly balances. The durability SLA is really a statement about your repair pipeline, not your encoding.
flowchart TB
CL["Client"]
MD["Metadata service
placement via consistent hashing
Raft-replicated"]
N1["Chunk server 1
dumb disks + checksums"]
N2["Chunk server 2"]
N3["Chunk server N"]
AUD["Auditor / repair worker
scrubs + rebuilds missing chunks"]
CL -->|"PUT: where do chunks go?"| MD
MD -->|"placement map"| CL
CL -->|"write chunks"| N1
CL -->|"write chunks"| N2
CL -->|"write chunks"| N3
N1 -. "checksums, heartbeats" .-> AUD
N2 -. "checksums, heartbeats" .-> AUD
N3 -. "checksums, heartbeats" .-> AUD
AUD -->|"rebuild missing chunk"| N1
MD -. "chunk map updates" .-> AUD
Two kinds of smarts: the metadata service knows where everything is (small, precious, Raft-replicated), and the chunk servers are deliberately dumb (big, cheap, replaceable). The auditor closes the loop — durability without repair is just delayed data loss.
sequenceDiagram
autonumber
participant C as Client
participant M as Metadata service
participant S as Chunk servers
C->>M: PUT /objects/photo.jpg — placement?
M-->>C: chunk map (9 nodes, rack-aware)
C->>C: split into 6 data chunks,
encode 3 parity (EC) — or 3 copies
C->>S: write chunks in parallel
S-->>C: acks + checksums
C->>M: commit (quorum of chunks durable)
M-->>C: 201 Created
Note over M,S: later: node dies → auditor
detects missing chunks → rebuilds
flowchart TB
A["GET /objects/photo.jpg"] --> B{"storage scheme?"}
B -->|"3x replication"| C["Read chunk from
nearest healthy replica"]
B -->|"RS 6+3"| D["Read any 6 of 9 chunks
in parallel"]
C --> E["200 OK — 1 node touched"]
D --> F["Decode: reconstruct
original 6 data chunks"]
F --> G["200 OK — 6 nodes touched
+ CPU + tail latency risk"]
Copies on 3 nodes in the same rack don't survive a rack power failure. Placement must spread chunks across failure domains: different racks, ideally different power and network. Consistent hashing picks the nodes; a rack-awareness constraint vetoes bad picks. In the interview, saying "replicas across racks" unprompted is a quiet senior signal.
Lose a chunk server and the auditor rebuilds it. Lose the metadata service and you have a warehouse full of unlabeled boxes — the bytes exist but nothing is findable. That's why metadata is tiny, Raft-replicated across 3–5 nodes, snapshotted, and backed up, while chunk servers are cattle. Durability of data means nothing without durability of the map.
PUT /v1/objects/{key} # body: bytes, optional ?scheme=repl|ec
→ 201 Created {"key": "photo.jpg", "size": 48321152, "scheme": "rs6+3"}
GET /v1/objects/{key} # Range: bytes=0-1048575 supported
→ 200 OK (streamed)
HEAD /v1/objects/{key} # size, checksum, scheme
DELETE /v1/objects/{key} # tombstone; chunks GC'd lazily
Data model: Object {key, size, scheme, chunk_map: [{chunk_id, node_ids[], checksum}]} in the metadata service. Chunk {id, bytes, checksum} on chunk servers — content-addressed, immutable. Nodes are just {id, rack, status, free_bytes}; the system never trusts them, it verifies them.
| 3× replication | RS(6+3) | |
|---|---|---|
| Storage overhead | 200% — expensive | 50% — cheap |
| Read latency | 1 node, lowest tail latency | 6 nodes + decode — higher p99 |
| Repair cost | 1:1 copy — cheap | 6:1 read amplification — the repair tax |
| CPU | ~zero | Encode/decode on every write/read |
| Small objects | Fine | Wasteful — parity overhead dominates; pad or replicate |
| Best for | Hot data, small objects | Cold/archival data, large objects |
API: S3-compatible (don't invent a protocol). Metadata: Raft-replicated service, or Postgres with Patroni if you want boring. Chunk servers: dumb daemons on commodity disks, checksums on everything. Encoding: Intel ISA-L for Reed-Solomon — never hand-roll the math. Policy: replicate objects under 1MB and anything hot; EC(6+3) everything else; lifecycle rule migrates cold objects automatically. That's the architecture every major object store converges on.
Pick a scheme, place the blocks across 9 nodes, then click nodes to kill them (or use the button). Watch which scheme keeps your 6 data chunks recoverable — and what each one costs in raw storage.