Back-of-the-Envelope Estimation

Naseebullah Ahmadi  Senior Software Engineer, London

The estimation step of a system design interview isn't a maths test. It's how you earn the right to say "one database is enough" or "this needs sharding." The handful of numbers worth memorising, a repeatable method for turning traffic into QPS, storage, and memory, and a worked URL shortener sized twice, at 1x and 100x, to show the numbers changing the design.

6 min read
#interview
In one line

Round everything to a power of ten, divide per-day numbers by 10⁵ to get per-second, multiply by a peak factor, and never stop at the number: every estimate should end in a design decision ("fits on one node", "needs a cache", "needs sharding"). The interviewer is grading whether your arithmetic drives your architecture, not whether you got 1,157 or 1,000.

Somewhere in the first ten minutes of a system design interview, the interviewer says "how much traffic are we talking about?" and a lot of candidates freeze, or worse, spend five minutes doing long division on the whiteboard to three significant figures.

Neither is what's being tested. The estimate exists to answer one question: what does this system actually need? One Postgres box or a sharded cluster. A cache or no cache. A CDN or not. You can't answer any of those honestly without a rough number, and a rough number is all you need.

The numbers worth carrying

You need three small tables in your head. Everything else you derive.

Time. The one that does the most work:

PeriodSecondsRound to
Day86,40010⁵
Month~2,600,0002.5 × 10⁶
Year~31,500,0003 × 10⁷

So 1M requests per day is about 10 per second (12, really). 100M per day is about 1,000 per second. That single conversion answers half the questions an interviewer will ask.

Size. Powers of two line up with powers of ten closely enough (2¹⁰ = 1,024 ≈ 10³) that you can treat them as the same:

PowerApprox.Unit
2¹⁰thousandKB
2²⁰millionMB
2³⁰billionGB
2⁴⁰trillionTB
2⁵⁰quadrillionPB

That means "a billion rows of 1 KB each" is a terabyte, without writing a single digit.

Latency. The classic latency numbers table has moved a lot since it was first published (SSDs especially), and Colin Scott's interactive version tracks how. The exact values don't matter. The gaps between the rows do:

OperationOrder of magnitude
Memory reference~100 ns
SSD random read~10 to 100 µs
Round trip in the same datacenter~0.5 ms
Disk seek (spinning)~5 ms
Round trip across an ocean~150 ms

Memory is hundreds to a thousand times faster than an SSD read, which is roughly ten times faster than a network hop, which is hundreds of times faster than crossing the planet. That's the whole argument for caches and CDNs in one line.

Finally, a few capacity ballparks so a number means something once you have it. These vary wildly with hardware and workload, so say them out loud as assumptions, not facts:

  • One app instance doing light work: ~1,000 req/s.
  • One #postgres primary: a few thousand simple writes per second, far more reads.
  • One #redis node: ~100,000 simple ops per second.
  • One commodity server: 64 to 256 GB of RAM, a few TB of SSD.

The method

Every estimate follows the same shape. Talking through it in this order is most of the signal the interviewer is looking for.

  1. 1

    State the assumptions out loud. Daily active users, actions per user, read:write ratio, record size, retention. Write them in a corner of the board. If the interviewer disagrees with one, they change one line, not your whole calculation.

  2. 2

    Traffic. Actions per day ÷ 10⁵ = average QPS. Multiply by 2 to 3 for peak. Do reads and writes separately; they almost never scale the same way.

  3. 3

    Storage. Writes per day × record size × retention. Then multiply by the replication factor if the question is about raw disk, not per-node disk.

  4. 4

    Memory. What's hot? The 80/20 rule is a fine default: cache the ~20% of data that serves ~80% of reads.

  5. 5

    Bandwidth. Peak QPS × response size. Usually boring, which is itself a useful finding.

  6. 6

    Close the loop. For every number, say what it decides. This is the step people skip, and it's the only one that matters.

Worked example: a URL shortener

Assumptions, written in the corner of the board:

  • 100M new short links per month.
  • Reads to writes at 100:1 (people click links far more than they make them).
  • ~500 bytes per record: a 7-character code, a long URL averaging a couple of hundred bytes, an owner, a timestamp, plus row and index overhead. Rounded up generously.
  • Keep everything for 5 years.

Traffic. 100M a month ÷ 2.5 × 10⁶ seconds is ~40 writes/s. At 100:1, that's ~4,000 reads/s on average, call it ~10,000 at peak.

Storage. 100M × 500 B = 50 GB a month, 600 GB a year, ~3 TB over five years.

Key space. Five years is 60 months × 100M = 6 × 10⁹ links. Base62 with 6 characters gives 62⁶ ≈ 5.7 × 10¹⁰, about 10x headroom. With 7 characters, 62⁷ ≈ 3.5 × 10¹², nearly 600x.

Memory. If 20% of last month's links take most of the clicks, that's 20M × 500 B = ~10 GB hot.

Bandwidth. 10,000 peak reads × ~500 B for a redirect response is ~5 MB/s. Nothing.

Concurrency. Little's law says requests in flight = arrival rate × time each one takes. At 10,000 req/s and ~10 ms per request, that's ~100 in flight: a small connection pool and a handful of app instances.

Now close the loop, which is where the interview is actually won:

What the numbers rule in

  • One Postgres primary: 40 writes/s and 3 TB is comfortably single-node territory. - Read replicas or a cache for the 10,000 peak reads. - One Redis node: 10 GB of hot data fits in memory with room to spare.

What the numbers rule out

  • Sharding. Nothing here needs it, and proposing it now adds cost with no benefit. - A CDN for bandwidth. 5 MB/s doesn't justify one (latency might, but that's a separate argument). - A distributed ID generator. A single database sequence handles 40 writes/s.

That second card is the part most candidates never say. Showing that you don't need a piece of infrastructure is as strong a signal as knowing how to build it.

Same system, 100x the traffic

Interviewers love to follow up with "now what if it's 100 times bigger?" Change one assumption, 10 billion new links a month, and rerun the same arithmetic:

Estimate1x100xWhat changes
Writes~40/s~4,000/sNow at the edge of one primary
Reads (peak)~10,000/s~1,000,000/sFar past one cache node; needs a cluster
Storage (5 yrs)~3 TB~300 TBNo single node holds this: shard
Hot set~10 GB~1 TBRedis cluster, partitioned by code
Bandwidth (peak)~5 MB/s~500 MB/sEdge caching of redirects starts to pay
Key space6 chars is enough7 is tight, use 86 × 10¹¹ links fill 17% of 62⁷

Every row in the right column is a design decision the left column didn't need. That's the point of the exercise: the architecture follows from the arithmetic, and the same arithmetic tells you when it stops being true. If you want to see what actually gives as traffic climbs through those tiers, What Breaks From 1k to 1M Requests Per Second walks one endpoint through exactly that.

  1. 1

    False precision

    Writing 1,157.4 QPS signals that you don't know which digits matter. Say "about a thousand" and move on; the interviewer will stop you if they need more.

  2. 2

    Silent assumptions

    Doing the maths in your head and announcing a result gives the interviewer nothing to check or steer. Every number on the board should trace back to an assumption written next to it.

  3. 3

    Averages without a peak

    Systems fall over at peak, not at the daily average. A service sized for 4,000 req/s will have a bad time at the 10,000 it sees every evening.

  4. 4

    Stopping at the number

    "3 TB of storage" is trivia until it's followed by "so one node holds it." An estimate that doesn't change or confirm a design choice was wasted time on the clock.


End of entry · Keep exploring

What's next in the notebook?

Keep reading — more from where that came from.

Featured next
11 min read
0%

How Does Instagram Instantly Know If Your Username Is Available?

Type a handle into a signup form and a green tick or a red cross appears almost before you stop typing, checked against a billion names. It isn't one clever database query. It's a debounced client, a Bloom filter that answers "definitely free" from memory, an index for everything else, and a unique constraint that has the final say.

12 min read
#algorithms, #math

Two Eggs, 100 Floors

Two eggs, a 100-floor building, and the fewest drops that are guaranteed to find the floor where eggs start breaking. Binary search loses. The answer is 14, and the reason why turns into a one-line recurrence that solves the problem for any number of eggs.

0%
10 min read
#engineering

Stop Calling, Start Announcing: An Intro to Event-Driven Systems

A checkout handler that calls the ledger, the warehouse, the email service and analytics in a row is as slow as all of them added together and as fragile as all of them multiplied. Event-driven design flips it: announce what happened and let whoever cares react. What that buys, what it costs, and when a plain call is still right.

0%
13 min read
#engineering

You Can't Commit to Two Systems at Once

A payment handler saves to the database, then publishes an event. Crash between the two lines and the payment exists but nothing downstream ever hears about it. Swap the order and it lies the other way. The transactional outbox stops asking two systems to agree and makes the event a row instead.

0%
11 min read
#engineering

Your Logging Sucks!

A checkout endpoint with a log line at every step looks like good observability, right up until a customer says "my payment failed" and you have thirteen unrelated lines from thirteen unrelated requests to sort through. The fix isn't more logs, it's one wide event per request instead.

0%
16 min read
#engineering

What Breaks From 1k to 1M Requests Per Second

The same endpoint, run through four traffic tiers. At 1k req/s almost any design survives. At 10k the database and the single instance give first. At 100k the cache and the load balancer become the systems under test. At 1M the architecture itself has to change, because the failure mode is no longer capacity, it's correlated behavior across clients you don't control.

0%
12 min read
#engineering

Migrating Schema-Per-Tenant Databases at Scale

Choosing physical tenant isolation over a shared, RLS-scoped schema buys two new problems: knowing where a tenant's data actually lives, and running one migration correctly hundreds of times instead of once. Neither has an app-code fix, both need their own infrastructure.

0%
21 min read
#engineering

Designing Multi-Tenant APIs That Scale

A missing tenant filter is a data leak, not a crash. Row-level security fixes that structurally, but rate limits, connection pools, and error codes built for one instance break the same quiet way once the API runs as several.

0%
8 min read
#algorithms

Valid Triangle Number

Sort the side lengths, fix the largest one at a time, and let the gap between two pointers count every valid pair against it in one step instead of testing them one by one.

0%
12 min read
#algorithms

3Sum

Sort the array, fix one value as a moving target, and reuse the exact two-pointer proof from Two Sum II to sweep the rest, with a couple of extra duplicate-skipping rules layered on top.

0%
8 min read
#algorithms

Two Sum II

A sorted array, two pointers closing in from both ends, and a proof that whichever side is off-target can be ruled out entirely rather than retried against a smaller search space.

0%
7 min read
#algorithms

Container With Most Water

A row of walls, two pointers starting at the widest container, and a proof that moving the taller wall can never beat what you already have, so only the shorter side is ever worth moving.

0%
7 min read
#algorithms

Two Pointers

Two indices walking through one ordered structure, discarding the side that cannot improve the answer at every step and replacing a nested loop with a single pass.

0%