In a request-driven system the service that did something calls
everyone who needs to know, so it waits on all of them and fails
when any of them fails. In an event-driven system it publishes a
fact (payment.captured) and moves on; the ledger, the warehouse
and the email service each subscribe and react on their own time.
You trade that coupling for eventual consistency, duplicate
deliveries and harder debugging, so keep anything the caller needs
before it can answer as a direct call.
A payment gets captured. Five things need to happen because of it: the ledger books it, the warehouse reserves the stock, the customer gets a receipt, analytics counts it, and later someone adds a loyalty points service. The question this post is about is simple: whose job is it to make all of that happen?
Start with the version that looks fine
The obvious answer is the checkout service's, since it's the one that knows the payment happened:

It reads well and it works. Now look at what it costs.
Latency is the sum. The ledger takes 40ms, the warehouse 60ms, the email provider 300ms on a good day, analytics 50ms. The customer waits 450ms for a receipt email they'll read later.
Availability is the product. If each of those four services is up 99.9% of the time, checkout only succeeds when all four are: 0.999⁴ ≈ 99.6%. That's the difference between about 9 hours of downtime a year and about 35. Analytics being down now means customers can't pay.
Every new consumer is a change to checkout. The loyalty team can't ship until the payments team adds a fifth call, reviews it, and deploys it. Checkout has become the place where every other team's reaction to a payment lives.
The root of all three is the same: the service that did the thing is also responsible for telling everyone about it, one call at a time.
Announce what happened
Flip the responsibility. Checkout's job is to capture the payment and say so. Whoever cares listens:

Each consumer subscribes on its own:

The message goes through a broker, something like #kafka or #rabbitmq, which stores it so each subscriber can read it on its own schedule. (RabbitMQ keeps a message until it's taken; Kafka keeps everything for a retention window, days by default, whether anyone has read it or not.)
- Checkout to Broker: payment.captured
- Broker to Ledger
- Broker to Warehouse
- Broker to Email
- Broker to Analytics
Re-run the three costs. Checkout now waits on one write to the broker,
not four services, and depends on one thing being up (the broker)
instead of four. Analytics going down means analytics falls behind,
and catches up from the broker when it comes back, as long as it's
back within the retention window; payments keep working. And the
loyalty team subscribes to payment.captured and ships without asking
the payments team for anything.
Events are not commands
The word "message" covers two different things, and mixing them up is the most common way an event-driven design goes wrong.
| Command | Event | |
|---|---|---|
| Example | ReserveStock | PaymentCaptured |
| Tense | Imperative: do this | Past: this happened |
| Handlers | Exactly one | Zero, one, or many |
| Sender expects | It gets done, or an error back | Nothing; it's already done |
| Can it be refused | Yes (out of stock) | No; you can only react to it |
| Who owns the name | The receiver (it's the receiver's API) | The sender (it's the sender's fact) |
The tell for a command wearing an event's name is an event named after
what should happen next: SendReceiptEmail. That's checkout telling
the email service what to do, with the coupling still there, just
routed through a broker. Name the fact (PaymentCaptured) and let the
email service decide that a receipt is its reaction.
Queues and topics
Brokers deliver in two shapes, and you usually need both.
A queue hands each message to exactly one of the workers reading it. Run three copies of the email service on one queue and each receipt goes to one copy, whichever is free, not to all three. That's how a consumer scales.
A topic hands each message to every subscriber. The ledger, the
warehouse and the email service all get their own copy of
payment.captured. That's how producers stay unaware of consumers.
Most brokers combine them. In Kafka, a topic delivers to every consumer group, and within a group each message goes to one member. So each service is a group: every service sees every payment, and within a service the work is split across its instances.
The bill
Nothing above is free. Here's what you sign up for.
Nothing is up to date at the same moment
The API returns 201 before the ledger has booked anything. If the
customer's next click is "view my balance" and that page reads the
ledger, it may not show the payment yet. This is eventual consistency:
every consumer gets there, just not at the moment the producer
answers. Design screens around it, for example by showing the payment
from checkout's own record rather than the ledger's.
Every event may arrive twice
Assume every broker delivers at least once: it's the usual default, and the only setting that doesn't risk losing messages. A consumer that crashes after booking the ledger entry but before acknowledging the message gets it again. Every consumer has to be idempotent: record the event's id alongside its work, in the same transaction, and skip ids it has seen. It's the same idea as an idempotency key on an API.
Order is only guaranteed within a key
Kafka keeps order within a partition, not across a topic, and never
across two topics. If payment.refunded can overtake
payment.captured, the ledger will try to reverse something it hasn't
booked. So the samples above, with one topic per event type, can't
promise that order. When order matters, publish every payment event to
one payments topic with the payment id as the partition key: all
events for one payment then land on one partition, in order.
A request no longer has one stack trace
When a receipt doesn't arrive, there's no single call chain to read. The request ended at checkout; the failure happened seconds later in a different service. Put a trace id in every event's metadata and have every consumer log it, so one search pulls up the whole story. It's what wide events are for.
Some messages will never succeed
An event with a field the consumer can't parse fails every time. Retry it a few times with backoff, then move it to a dead-letter queue: a side queue someone inspects and replays, so one bad message doesn't block every message behind it.
Choreography or orchestration
Fan-out works when the reactions are independent. The ledger doesn't care whether the email went out. But some flows are a sequence with consequences. A refund has to reverse the ledger entry, then return the money through the provider, then restock, and if the provider refuses, the ledger reversal has to be undone.
You can build that as choreography: each service reacts to the
previous one's event (LedgerReversed → provider refunds →
RefundIssued → warehouse restocks). No one owns the whole flow,
which is flexible and also the problem: to answer "where is refund 123
stuck?" you have to reconstruct it from three services' logs.
Or as orchestration: one refund service sends commands in order and tracks the state of each refund, including what to undo when a step fails. More central, easier to see. Both are forms of a saga: a sequence of local steps, each with a compensating step to undo it. They differ in whether one service owns the sequence or it emerges from the reactions.
The rule of thumb: independent reactions to a fact are choreography. A multi-step business process that can fail halfway belongs to an orchestrator, which still talks to the rest of the system through events and commands.
When a plain call is still right
One call was left out of the handler on purpose: the fraud check. It runs before the capture, and it stays a direct call:

Checkout can't answer the customer until it knows the answer, so an event doesn't help. There's nothing to do while waiting. The split that decides most cases:
| The work... | Make it |
|---|---|
| Must finish before you can answer the caller | A direct call |
| Happens because of what you did, and can lag | An event |
| Is one of many steps that can fail and undo | An orchestrator |
| Lives in the same codebase and database | A function call |
That last row matters. If checkout, the ledger and the email sender are one app with one database, owned by one team, a broker adds latency, infrastructure and every item on the bill above, to decouple code that a function call already separates well enough. Events earn their place when separate services, owned by separate teams, need to react to the same thing without waiting on each other.
Back to the payment. Checkout captures it, says so, and answers the customer in the time it takes to write one message. The ledger, the warehouse, the receipt and the loyalty points all still happen, each in its own service, on its own schedule, and none of them can take payments down.

