Five nodes hold the same config log. The leader dies. Writes stall. The intern votes with wall clocks — “whoever has the newest timestamp is leader” — and two nodes both take client traffic. A second intern says run Two-Phase Commit on every SET. Neither job is this one. You needed one leader and a replicated log that does not forget a committed write.

Raft elects one leader per term with randomized timeouts and a majority vote, then that leader appends the log and followers match. A committed entry is never overwritten. That is the invariant. This post is awareness of what etcd (and Kafka-adjacent metadata quorums) are doing. It is not a cluster you ship.

Families and the catalog live on the Algorithms Roadmap. The hub already named the cargo-cult: do not hand-roll Raft. Read the terms, walk one election, leave the protocol in the library.

Three roles, one term, a majority

Every node is one of three roles:

RoleJob
FollowerAccept heartbeats and log appends. Vote at most once per term.
CandidateIncrement the term, vote for self, ask the others.
LeaderAppend client writes, ship AppendEntries, decide commit.

A term is a monotonically increasing ballot. At most one leader per term. A majority is ⌊n/2⌋ + 1. Five nodes: three. You can lose two and still elect, still commit. Two is not enough. That is why the split vote below stalls instead of forking the log.

Followers grant a vote only if they have not already voted in that term and the candidate’s log is at least as complete as theirs (last index and last term). Stale candidates lose. You do not elect on wall-clock “newest.”

Note: Odd cluster sizes are the usual default so majority is unambiguous. Four nodes still need three. You paid a node and did not buy extra failure budget.

Randomized timeouts, then a vote

The leader sends heartbeats (AppendEntries with no new rows, or with them). Each follower draws an election timeout from a randomized window — not a shared 150 ms. If the heartbeat is late, the follower becomes a candidate: bump currentTerm, vote for itself, send RequestVote.

The random window is the split-vote brake. Identical timeouts make two candidates start together, split the majority, and sit leaderless until the next timeout. Jitter makes one campaign usually finish first. The retry after a split is the same rule, not a different algorithm.

A candidate that hears a higher term, or a valid leader heartbeat for this term, steps down to follower. There is no “keep campaigning because I started first.”

The leader appends; followers match

Clients write to the leader. The leader appends a log entry (index, term, command), then ships that entry to followers. A follower matches: it accepts only if the previous index and previous term already agree with its log. A hole or a conflicting term is a reject. The leader retries from an earlier index until the prefixes line up.

An entry is committed when it sits on a majority. The leader then applies it and tells followers the new commit index. Reads that must see committed state go to the leader (or through a linearizable path the library already owns). Followers are the copy, not a second writer.

leader log:     [1:SET a=1 | 2:SET b=2 | 3:SET c=3]
follower match: [1:SET a=1 | 2:SET b=2 |          ]   ← still catching 3

commit index = 2   (majority already has 1 and 2)
entry 3 is appended, not committed, until a third node stores it

That is replication. It is not 2PC’s prepare-then-commit handshake, and it is not “gossip the map until it looks equal.”

Safety: committed stays committed

Once an entry is committed, no later leader overwrites that index with a different command. State the invariant. Do not prove the paper. The election rule (majority already has the committed prefix; votes require an up-to-date log) plus log matching are why a new term cannot bury a committed SET. Uncommitted tail entries can be replaced after a crash — that is the unfinished write, not the committed one.

If you needed a proof, you needed the Raft paper and a library’s test suite, not this page.

Not two-phase commit

Two-Phase Commit is blocking atomic commit: a coordinator asks every participant to prepare, then commit (or abort). A coordinator crash in the middle can leave participants blocked. 2PC is not leader election. It does not choose who owns the log. It does not keep a five-node replica set available when two nodes die.

Raft’s job is: keep one replicated log moving while a minority of nodes fail. 2PC’s job is: all named participants accept the same commit-or-abort. Mixing them is how you get a “consensus” service that cannot elect and a transaction protocol that cannot lose the coordinator. Different procedures. The 2PC post owns the blocking handshake. This page does not relitigate it.

A walk: five nodes, a death, a split, a retry

Nodes n1–n5. Majority = 3. n1 is leader in term 4. Then it dies.

term 4, n1 Leader. Committed log:
  idx 1  SET a=1   term 4
  idx 2  SET b=2   term 4
n2 n3 n4 n5 Followers. Logs match. Heartbeats stop.

t ≈ 150 ms   n2 timeout → Candidate, term 5, vote self
t ≈ 160 ms   n3 timeout → Candidate, term 5, vote self

term-5 votes:
  n2: self + n4                 = 2
  n3: self + n5                 = 2
  n1 dead; nobody has 3

split vote. No leader. Writes still stall.
Each candidate draws a fresh random timeout.

t ≈ 310 ms   n3 fires first → Candidate, term 6, vote self
             RequestVote → n2, n4, n5
             n4 and n5 grant (idle in term 6)
             n2 steps down, grants (higher term)
             n3 has 4 ≥ 3  → Leader, term 6

n3 ships AppendEntries / heartbeats.
Followers match idx 1–2. Those committed rows stay.
Next client SET: n3 appends idx 3, waits for a majority match, then commits.

The first campaign failed because two timers fired together. The second campaign succeeded because jitter let one candidate collect three votes before the other restarted. That is the whole trick. A wall-clock “newest node” rule would have let n2 and n3 both serve writes in term 5.

Note: A node that is down does not vote. Majority is of the cluster size, not of who answered. Two live nodes in a five-node set cannot elect.

Java: roles as a sketch, not a cluster

No java.util.Raft. The JDK will not elect this for you. etcd, Consul, and Kafka’s metadata quorum (KRaft) already run a Raft-shaped protocol. The enum is vocabulary so the walk above has types. It is not a server.

enum Role {
    FOLLOWER, CANDIDATE, LEADER
}

final class NodeState {
    Role role = Role.FOLLOWER;
    int currentTerm = 0;
    Integer votedFor; // node id; null = none this term

    void becomeCandidate(int selfId) {
        role = Role.CANDIDATE;
        currentTerm++;
        votedFor = selfId;
    }

    void becomeLeader() {
        if (role != Role.CANDIDATE) {
            throw new IllegalStateException("only a candidate becomes leader");
        }
        role = Role.LEADER;
    }

    void stepDown(int discoveredTerm) {
        currentTerm = discoveredTerm;
        role = Role.FOLLOWER;
        votedFor = null;
    }
}

becomeCandidate is the timeout firing. becomeLeader is majority in this term. stepDown is a higher term or a real leader heartbeat. Missing from this sketch, on purpose: RPC, the log, matching, snapshots, membership changes, disk, clocks. Those are why you do not paste this into production.

Do not grow this enum into a cluster. The hub’s cargo-cult line is the same one as Timsort and gzip: call the library. Awareness is knowing what the library is doing when the leader dies. Implementation is etcd’s job.

Who actually runs this

etcd uses Raft for the Kubernetes control-plane store. Consul uses Raft for its catalog. Kafka’s KRaft metadata quorum is Raft-shaped — this post is not a Kafka course. You operate those systems: odd replica counts, disk, peer list, backup. You do not reimplement RequestVote.

Paxos is a different paper and a different vocabulary. Out of scope. If a design doc says “we will write Raft in a weekend,” the answer is the hub line, not a starter repo.

When not to reach for Raft

Skip this protocol when the job is not “replicated log plus leader election under minority failure.”

  • You wanted atomic commit across two databases. That is Two-Phase Commit. Blocking, coordinator-shaped, not an election.
  • You wanted a single-process lock. A leader enum in one JVM is not consensus. synchronized and a database row are cheaper and honest.
  • You wanted to ship a cluster. Use etcd, Consul, or the metadata quorum your product already runs. Do not hand-roll Raft.
  • You wanted Paxos trivia. Different paper. This page will not recap it.
  • You wanted Kafka internals. KRaft is a pointer that the same family shows up in the broker’s metadata. The product course is elsewhere.

Awareness of election and log matching — that is the job. Huffman, interval scheduling, and backtracking are other procedures. They do not elect a leader.

Cheat sheet

Job:         one leader, one replicated log, minority failures tolerated
Roles:       Follower / Candidate / Leader
Term:        monotonic ballot; at most one leader per term
Majority:    ⌊n/2⌋ + 1   (5 nodes → 3)
Election:    randomized timeout → candidate → RequestVote → majority
Replication: leader appends; followers match previous index/term
Commit:      entry on a majority, then apply
Safety:      a committed entry is never overwritten
Not this:    2PC, Paxos recap, a from-scratch Java cluster, Kafka-as-product
Who runs it: etcd, Consul, Kafka KRaft metadata — call those, do not paste this

Do:

  • Learn the roles, the term, and why majority is three of five.
  • Expect split votes; expect a randomized retry to break them.
  • Treat committed as durable across leader change. Uncommitted tail may vanish.
  • Point at etcd (or the quorum you already run) when someone needs a cluster.

Don’t:

  • Hand-roll Raft because the enum fit in a gist.
  • Elect on wall clocks or “newest node.” That is two leaders.
  • Call 2PC an election. Atomic commit is a different handshake.
  • Prove the paper on a whiteboard and ship the whiteboard.

Wrap-up

Raft is leader election plus log replication: random timeouts so campaigns do not tie, a majority so two nodes cannot fork the truth, matching so followers share a prefix, and a commit rule so a committed entry survives the next term. Five nodes, leader dies, split vote, retry — that walk is the awareness. etcd already runs it. 2PC does not. The Java roles are labels, not a cluster. Do not hand-roll this.

When the next systems post needs an approximate distinct count in tiny memory instead of a replicated log, that is HyperLogLog.

Next optional step in the series Approximate how many distinct keys you saw when an exact set will not fit. HyperLogLog: Approximate Distinct Count in Tiny Memory