Stop Calling, Start Announcing: An Intro to Event-Driven Systems

Naseebullah Ahmadi  Senior Software Engineer, London

A checkout handler that calls the ledger, the warehouse, the email service and analytics in a row is as slow as all of them added together and as fragile as all of them multiplied. Event-driven design flips it: announce what happened and let whoever cares react. What that buys, what it costs, and when a plain call is still right.

10 min read
#engineering
In one line

In a request-driven system the service that did something calls everyone who needs to know, so it waits on all of them and fails when any of them fails. In an event-driven system it publishes a fact (payment.captured) and moves on; the ledger, the warehouse and the email service each subscribe and react on their own time. You trade that coupling for eventual consistency, duplicate deliveries and harder debugging, so keep anything the caller needs before it can answer as a direct call.

A payment gets captured. Five things need to happen because of it: the ledger books it, the warehouse reserves the stock, the customer gets a receipt, analytics counts it, and later someone adds a loyalty points service. The question this post is about is simple: whose job is it to make all of that happen?

Start with the version that looks fine

The obvious answer is the checkout service's, since it's the one that knows the payment happened:

@itsnas TypeScript
// POST /payments
async function capturePayment(req: Request) {
  const payment = await payments.capture(req.body)
 
  await ledger.book(payment.id, payment.amount)
  await warehouse.reserve(payment.orderId)
  await email.sendReceipt(payment)
  await analytics.track('payment_captured', payment)
 
  return json(201, payment)
}
main
Nas (@itsnas)
Checkout knows about, and waits on, every consumer

It reads well and it works. Now look at what it costs.

Latency is the sum. The ledger takes 40ms, the warehouse 60ms, the email provider 300ms on a good day, analytics 50ms. The customer waits 450ms for a receipt email they'll read later.

Availability is the product. If each of those four services is up 99.9% of the time, checkout only succeeds when all four are: 0.999⁴ ≈ 99.6%. That's the difference between about 9 hours of downtime a year and about 35. Analytics being down now means customers can't pay.

Every new consumer is a change to checkout. The loyalty team can't ship until the payments team adds a fifth call, reviews it, and deploys it. Checkout has become the place where every other team's reaction to a payment lives.

The root of all three is the same: the service that did the thing is also responsible for telling everyone about it, one call at a time.

Announce what happened

Flip the responsibility. Checkout's job is to capture the payment and say so. Whoever cares listens:

@itsnas TypeScript
async function capturePayment(req: Request) {
  const payment = await payments.capture(req.body)
 
  await events.publish('payment.captured', {
    paymentId: payment.id,
    orderId: payment.orderId,
    amount: payment.amount,
  })
 
  return json(201, payment)
}
main
Nas (@itsnas)
Checkout states a fact and moves on

Each consumer subscribes on its own:

@itsnas TypeScript
events.subscribe('payment.captured', async event => {
  await ledger.book(event.paymentId, event.amount)
})
main
Nas (@itsnas)
The ledger owns its reaction; checkout never hears about it

The message goes through a broker, something like #kafka or #rabbitmq, which stores it so each subscriber can read it on its own schedule. (RabbitMQ keeps a message until it's taken; Kafka keeps everything for a retention window, days by default, whether anyone has read it or not.)

  1. Checkout to Broker: payment.captured
  2. Broker to Ledger
  3. Broker to Warehouse
  4. Broker to Email
  5. Broker to Analytics
Checkout publishes once. Each consumer subscribes independently, and a new one plugs into the broker without touching checkout.

Re-run the three costs. Checkout now waits on one write to the broker, not four services, and depends on one thing being up (the broker) instead of four. Analytics going down means analytics falls behind, and catches up from the broker when it comes back, as long as it's back within the retention window; payments keep working. And the loyalty team subscribes to payment.captured and ships without asking the payments team for anything.

Events are not commands

The word "message" covers two different things, and mixing them up is the most common way an event-driven design goes wrong.

CommandEvent
ExampleReserveStockPaymentCaptured
TenseImperative: do thisPast: this happened
HandlersExactly oneZero, one, or many
Sender expectsIt gets done, or an error backNothing; it's already done
Can it be refusedYes (out of stock)No; you can only react to it
Who owns the nameThe receiver (it's the receiver's API)The sender (it's the sender's fact)

The tell for a command wearing an event's name is an event named after what should happen next: SendReceiptEmail. That's checkout telling the email service what to do, with the coupling still there, just routed through a broker. Name the fact (PaymentCaptured) and let the email service decide that a receipt is its reaction.

Queues and topics

Brokers deliver in two shapes, and you usually need both.

A queue hands each message to exactly one of the workers reading it. Run three copies of the email service on one queue and each receipt goes to one copy, whichever is free, not to all three. That's how a consumer scales.

A topic hands each message to every subscriber. The ledger, the warehouse and the email service all get their own copy of payment.captured. That's how producers stay unaware of consumers.

Most brokers combine them. In Kafka, a topic delivers to every consumer group, and within a group each message goes to one member. So each service is a group: every service sees every payment, and within a service the work is split across its instances.

The bill

Nothing above is free. Here's what you sign up for.

Nothing is up to date at the same moment

The API returns 201 before the ledger has booked anything. If the customer's next click is "view my balance" and that page reads the ledger, it may not show the payment yet. This is eventual consistency: every consumer gets there, just not at the moment the producer answers. Design screens around it, for example by showing the payment from checkout's own record rather than the ledger's.

Every event may arrive twice

Assume every broker delivers at least once: it's the usual default, and the only setting that doesn't risk losing messages. A consumer that crashes after booking the ledger entry but before acknowledging the message gets it again. Every consumer has to be idempotent: record the event's id alongside its work, in the same transaction, and skip ids it has seen. It's the same idea as an idempotency key on an API.

Order is only guaranteed within a key

Kafka keeps order within a partition, not across a topic, and never across two topics. If payment.refunded can overtake payment.captured, the ledger will try to reverse something it hasn't booked. So the samples above, with one topic per event type, can't promise that order. When order matters, publish every payment event to one payments topic with the payment id as the partition key: all events for one payment then land on one partition, in order.

A request no longer has one stack trace

When a receipt doesn't arrive, there's no single call chain to read. The request ended at checkout; the failure happened seconds later in a different service. Put a trace id in every event's metadata and have every consumer log it, so one search pulls up the whole story. It's what wide events are for.

Some messages will never succeed

An event with a field the consumer can't parse fails every time. Retry it a few times with backoff, then move it to a dead-letter queue: a side queue someone inspects and replays, so one bad message doesn't block every message behind it.

Choreography or orchestration

Fan-out works when the reactions are independent. The ledger doesn't care whether the email went out. But some flows are a sequence with consequences. A refund has to reverse the ledger entry, then return the money through the provider, then restock, and if the provider refuses, the ledger reversal has to be undone.

You can build that as choreography: each service reacts to the previous one's event (LedgerReversed → provider refunds → RefundIssued → warehouse restocks). No one owns the whole flow, which is flexible and also the problem: to answer "where is refund 123 stuck?" you have to reconstruct it from three services' logs.

Or as orchestration: one refund service sends commands in order and tracks the state of each refund, including what to undo when a step fails. More central, easier to see. Both are forms of a saga: a sequence of local steps, each with a compensating step to undo it. They differ in whether one service owns the sequence or it emerges from the reactions.

The rule of thumb: independent reactions to a fact are choreography. A multi-step business process that can fail halfway belongs to an orchestrator, which still talks to the rest of the system through events and commands.

When a plain call is still right

One call was left out of the handler on purpose: the fraud check. It runs before the capture, and it stays a direct call:

@itsnas TypeScript
const verdict = await fraud.check(req.body)
if (verdict.block) return json(402, { error: 'Payment declined' })
main
Nas (@itsnas)
Checkout needs this answer before it can reply

Checkout can't answer the customer until it knows the answer, so an event doesn't help. There's nothing to do while waiting. The split that decides most cases:

The work...Make it
Must finish before you can answer the callerA direct call
Happens because of what you did, and can lagAn event
Is one of many steps that can fail and undoAn orchestrator
Lives in the same codebase and databaseA function call

That last row matters. If checkout, the ledger and the email sender are one app with one database, owned by one team, a broker adds latency, infrastructure and every item on the bill above, to decouple code that a function call already separates well enough. Events earn their place when separate services, owned by separate teams, need to react to the same thing without waiting on each other.

Back to the payment. Checkout captures it, says so, and answers the customer in the time it takes to write one message. The ledger, the warehouse, the receipt and the loyalty points all still happen, each in its own service, on its own schedule, and none of them can take payments down.


End of entry · Keep exploring

What's next in the notebook?

Keep reading — more from where that came from.

Featured next
13 min read
0%

You Can't Commit to Two Systems at Once

A payment handler saves to the database, then publishes an event. Crash between the two lines and the payment exists but nothing downstream ever hears about it. Swap the order and it lies the other way. The transactional outbox stops asking two systems to agree and makes the event a row instead.

15 min read
#engineering

Two Writes, One Row: Who Wins?

Two requests read the same row, both do their maths, both write back. One of them silently disappears. How the system should resolve that isn't one answer: it depends on whether the write is a delta, a quick piece of logic, or a human edit made minutes after the read.

0%
11 min read
#engineering

Your Logging Sucks!

A checkout endpoint with a log line at every step looks like good observability, right up until a customer says "my payment failed" and you have thirteen unrelated lines from thirteen unrelated requests to sort through. The fix isn't more logs, it's one wide event per request instead.

0%
16 min read
#engineering

What Breaks From 1k to 1M Requests Per Second

The same endpoint, run through four traffic tiers. At 1k req/s almost any design survives. At 10k the database and the single instance give first. At 100k the cache and the load balancer become the systems under test. At 1M the architecture itself has to change, because the failure mode is no longer capacity, it's correlated behavior across clients you don't control.

0%
12 min read
#engineering

Migrating Schema-Per-Tenant Databases at Scale

Choosing physical tenant isolation over a shared, RLS-scoped schema buys two new problems: knowing where a tenant's data actually lives, and running one migration correctly hundreds of times instead of once. Neither has an app-code fix, both need their own infrastructure.

0%
21 min read
#engineering

Designing Multi-Tenant APIs That Scale

A missing tenant filter is a data leak, not a crash. Row-level security fixes that structurally, but rate limits, connection pools, and error codes built for one instance break the same quiet way once the API runs as several.

0%
17 min read
#engineering

Why Payment Retries Need Idempotency

A plain payment endpoint looks correct until you trace what a double-click, a timed-out request, or a redelivered webhook actually does to it. Each one turns one payment into two. Idempotency keys are the fix, at two layers most write-ups skip.

0%
8 min read
#engineering, #frontend

Building a Typed Fetch Factory

How a single createFetcher factory infers request/response types from an OpenAPI schema and layers in caching, retries, and cancellation, and why each piece is built the way it is.

0%
7 min read
#algorithms

Two Pointers

Two indices walking through one ordered structure, discarding the side that cannot improve the answer at every step and replacing a nested loop with a single pass.

0%
2 min read
#engineering

AI Without Losing Judgment

AI can speed up delivery, but engineers still own architecture, quality, and decisions. A simple workflow to ship faster without outsourcing judgment.

0%