Three services own checkout: inventory, payments, orders. Product wants all three durable or none. The intern wires three HTTP calls in order. Inventory decrements. Payments declines the card. The order row never writes. Stock is already gone. A compensating restock is a second protocol nobody shipped.

Two-phase commit: a coordinator asks every participant to prepare (vote), then all commit or all abort. Prepared resources stay held until that decision. If the coordinator dies in the gap, the cluster waits.

This post is that protocol. Families and the catalog live on the Algorithms Roadmap. It is not consensus — Raft is the awareness follow-up (a leader and a log; majority progress). It is not XA product config, not a JTA tutorial, and not Spring @Transactional on one JVM.

The job is one decision, not three lucky HTTP calls

A coordinator starts the transaction and owns the global outcome. Participants each own a resource — a stock reservation, a card capture, an order row. Nobody makes the result durable in phase one.

Two phases:

  1. Prepare (vote). The coordinator sends PREPARE. A participant that can finish writes a prepared record, holds its locks, and votes YES. A participant that cannot votes NO and may abort locally at once.
  2. Commit or abort. All YES → the coordinator decides COMMIT and broadcasts it. Any NO or a missing vote → ABORT. Every YES voter applies that same decision.

A YES is a promise: this participant can commit and will not unlock until the coordinator says so. Sequential HTTP never made that promise. Each call committed or failed on its own clock.

Note: Timeout during prepare is a no. The coordinator treats silence as NO and aborts. Timeout after a YES is not a local abort. That is the blocking contract below.

A walk: three participants

Coordinator C. Participants Inventory, Payments, Orders. Same checkout.

-- abort: Payments votes no --
C → all:  PREPARE
Inventory: YES   (stock reserved, lock held)
Payments:  NO    (card declined)
Orders:    YES   (row prepared)

C: any NO → ABORT
C → all:  ABORT
Inventory / Orders release. Nothing durable.

-- commit: all yes --
C → all:  PREPARE
Inventory: YES
Payments:  YES
Orders:    YES

C: all YES → persist COMMIT → broadcast COMMIT
all three durable. locks released.

-- block: C dies after prepare, before commit --
C → all:  PREPARE
all three: YES   (resources held)
C crashes. COMMIT never sent.

Inventory, Payments, Orders stay PREPARED.
They cannot commit — no decision arrived.
They cannot abort — recovered C may still say COMMIT.
Checkout is atomic. The cluster is stuck until C returns.

Three outcomes, one rule: no participant finishes until every vote is in, and no YES voter finishes until the decision arrives. The abort path is live. The commit path is live. The third path waits.

Why it blocks

The dangerous window is after every YES and before the decision is delivered. Prepared records and row locks are already taken. The coordinator is the only process that knows — or is about to know — COMMIT versus ABORT.

If C crashes in that window:

  • Participants that voted YES must not unilaterally abort. Another participant might already have received COMMIT if the crash happened mid-broadcast.
  • They must not unilaterally commit. C might never have decided, or might have decided ABORT for a vote they did not see.
  • So they hold. Locks stay. Checkout sits in PREPARED until C recovers and replays the decision (or a human breaks the protocol).

Coordinator crash after prepare, before commit, is unbounded wait — not a retry. That is the title. Atomic commit is the feature. Blocking is the bill.

A participant crash after YES is the same promise from the other side: on restart it re-enters PREPARED and asks C what was decided. It does not abort because it rebooted.

Note: Persisting the decision before the commit broadcast is how a recovered coordinator finishes the protocol instead of inventing a second outcome. The sketch below does not implement that log. The walk is the point.

Not consensus, not a transaction-manager product

2PC is atomic commit: one coordinator, every participant, all-or-nothing. Consensus is a different job — replicas agreeing on a sequence of values so the group can proceed when some members die.

Raft elects a leader and replicates a log. A majority can make progress after a leader crash. 2PC has no majority shortcut: Inventory and Payments do not elect a replacement for a missing C. Do not add a term number to this loop and call it Raft. Do not implement Paxos here.

Do not ship this coordinator and call the cluster Raft. Do not treat a 2PC lab as XA configuration, JTA, or a Spring @Transactional lecture. Those products may use two-phase commit inside a database or a manager. This page is the prepare / vote / decide procedure so you can see why it blocks.

Out of scope: three-phase commit, Paxos, writing a transaction manager, wiring XA datasources.

Java: votes, a decision, a coordinator loop

No java.util.TwoPhaseCommit. Enums for the two payloads, an interface for a participant, a loop that is honest about the shape and silent about the network.

enum Vote { YES, NO }
enum Decision { COMMIT, ABORT }

interface Participant {
    Vote prepare(); // YES ⇒ resources held until commit() or abort()
    void commit();
    void abort();
}

static Decision coordinate(List<Participant> parts) {
    boolean allYes = true;
    for (Participant p : parts) {
        if (p.prepare() != Vote.YES) {
            allYes = false;
        }
    }
    Decision d = allYes ? Decision.COMMIT : Decision.ABORT;
    // persist d here — crash before the next loop is the blocking window
    for (Participant p : parts) {
        if (d == Decision.COMMIT) {
            p.commit();
        } else {
            p.abort();
        }
    }
    return d;
}

In-process, this loop cannot crash between phases. Production 2PC is RPC plus a write-ahead log on C and on every participant that voted YES. Treat silence during prepare() as NO (abort). Treat silence after YES as wait.

Note: A production coordinator may skip remaining prepares after the first NO. It still sends ABORT to every participant that already voted YES. Stopping the first loop and then forgetting those voters leaves locks with no decision.

What you pay

Let n be participants. The interesting bill is the wait, not the big-O of a fan-out.

WhatCostWhy
Prepare + votes1 round, O(n) messagesCoordinator to each and back
Commit or abort1 round, O(n) messagesDecision fan-out
Extra spaceprepared log + held locksUntil the decision arrives
Crash in the gapunbounded waitYES voters cannot finish alone

Two rounds when everyone is alive. Unbounded hold when the coordinator is not. Quoting O(n) messages as “cheap distributed transactions” hides that wait.

When not to use two-phase commit

Skip this protocol when the job is not “every participant commits the same outcome, and we will accept blocking.”

  • One database, one JVM. Local ACID or Spring @Transactional is a session, not 2PC across services. Do not add a coordinator to a single DataSource.
  • You needed progress after a leader dies. That is consensus. Raft is the awareness post. 2PC will wait.
  • Compensating / eventual is enough. A saga undoes with a reverse action. It is not atomic commit. It also does not hold prepare-locks across a dead coordinator. Pick it when “stuck” is worse than “briefly visible and then reversed.”
  • You wanted XA knobs. Connection flags, recovery scans, application-server datasources — product manuals, not this loop.
  • You wanted 3PC or Paxos. Different procedures. Not this page.

All prepare, then all commit — and it blocks. Consistent hashing and token bucket are other systems procedures on the roadmap. Raft is the next post, not a 2PC variant. Do not open those labs here.

Cheat sheet

Job:         atomic commit across participants (all durable or none)
Roles:       one coordinator; n participants that own resources
Phase 1:     PREPARE; vote YES (hold locks) or NO
Phase 2:     all YES → COMMIT; any NO / missing vote → ABORT
YES means:   promise to finish; do not unlock until the decision
Blocks:      coordinator crash after prepare, before commit
Time/space:  two O(n) rounds when healthy; unbounded wait in the gap
JDK:         no TwoPhaseCommit type; this is a protocol sketch
Not this:    Raft/Paxos, XA/JTA config, Spring @Transactional, sagas

Do:

  • Prepare everyone, then broadcast one decision.
  • Treat a missing prepare vote as abort. Persist the decision before the commit fan-out.
  • Keep YES resources held until commit() or abort().

Don’t:

  • Let a YES voter abort because the coordinator is late.
  • Call this consensus or Raft. Majority leader election is a different post.
  • Sequential HTTP and a hope that the third call lands.
  • Lecture JTA, XA datasources, or @Transactional here.

Wrap-up

Two-phase commit replaces three lucky RPC calls with a coordinator: vote, then the same decision on every participant. YES holds the resource. NO aborts everyone. Atomic is the feature. Coordinator death after prepare is the wait. It is not a replicated log, not a majority vote, and not a Spring transaction annotation. The intern’s layout was three HTTP calls. The procedure is prepare, then one decision. When the job is a leader and a log that can move on after a crash, that is Raft.

Next optional step in the series Leader election and a replicated log — Raft at awareness depth, not a from-scratch cluster. Raft: Elect a Leader and Replicate a Log