Five nodes hold the same config log. The leader dies. Writes stall. The intern votes with wall clocks — “whoever has the newest timestamp is leader” — and two nodes both take client traffic. A second intern says run Two-Phase Commit on every SET. Neither job is this one. You needed one leader and a replicated log that does not forget a committed write.
Raft elects one leader per term with randomized timeouts and a majority vote, then that leader appends the log and followers match. A committed entry is never overwritten. That is the invariant. This post is awareness of what etcd (and Kafka-adjacent metadata quorums) are doing. It is not a cluster you ship.
Families and the catalog live on the Algorithms Roadmap. The hub already named the cargo-cult: do not hand-roll Raft. Read the terms, walk one election, leave the protocol in the library.
Three roles, one term, a majority
Every node is one of three roles:
| Role | Job |
|---|---|
| Follower | Accept heartbeats and log appends. Vote at most once per term. |
| Candidate | Increment the term, vote for self, ask the others. |
| Leader | Append client writes, ship AppendEntries, decide commit. |
A term is a monotonically increasing ballot. At most one leader per term. A majority is ⌊n/2⌋ + 1. Five nodes: three. You can lose two and still elect, still commit. Two is not enough. That is why the split vote below stalls instead of forking the log.
Followers grant a vote only if they have not already voted in that term and the candidate’s log is at least as complete as theirs (last index and last term). Stale candidates lose. You do not elect on wall-clock “newest.”
Note: Odd cluster sizes are the usual default so majority is unambiguous. Four nodes still need three. You paid a node and did not buy extra failure budget.
Randomized timeouts, then a vote
The leader sends heartbeats (AppendEntries with no new rows, or with them). Each follower draws an election timeout from a randomized window — not a shared 150 ms. If the heartbeat is late, the follower becomes a candidate: bump currentTerm, vote for itself, send RequestVote.
The random window is the split-vote brake. Identical timeouts make two candidates start together, split the majority, and sit leaderless until the next timeout. Jitter makes one campaign usually finish first. The retry after a split is the same rule, not a different algorithm.
A candidate that hears a higher term, or a valid leader heartbeat for this term, steps down to follower. There is no “keep campaigning because I started first.”
The leader appends; followers match
Clients write to the leader. The leader appends a log entry (index, term, command), then ships that entry to followers. A follower matches: it accepts only if the previous index and previous term already agree with its log. A hole or a conflicting term is a reject. The leader retries from an earlier index until the prefixes line up.
An entry is committed when it sits on a majority. The leader then applies it and tells followers the new commit index. Reads that must see committed state go to the leader (or through a linearizable path the library already owns). Followers are the copy, not a second writer.
leader log: [1:SET a=1 | 2:SET b=2 | 3:SET c=3]
follower match: [1:SET a=1 | 2:SET b=2 | ] ← still catching 3
commit index = 2 (majority already has 1 and 2)
entry 3 is appended, not committed, until a third node stores it
That is replication. It is not 2PC’s prepare-then-commit handshake, and it is not “gossip the map until it looks equal.”
Safety: committed stays committed
Once an entry is committed, no later leader overwrites that index with a different command. State the invariant. Do not prove the paper. The election rule (majority already has the committed prefix; votes require an up-to-date log) plus log matching are why a new term cannot bury a committed SET. Uncommitted tail entries can be replaced after a crash — that is the unfinished write, not the committed one.
If you needed a proof, you needed the Raft paper and a library’s test suite, not this page.
Not two-phase commit
Two-Phase Commit is blocking atomic commit: a coordinator asks every participant to prepare, then commit (or abort). A coordinator crash in the middle can leave participants blocked. 2PC is not leader election. It does not choose who owns the log. It does not keep a five-node replica set available when two nodes die.
Raft’s job is: keep one replicated log moving while a minority of nodes fail. 2PC’s job is: all named participants accept the same commit-or-abort. Mixing them is how you get a “consensus” service that cannot elect and a transaction protocol that cannot lose the coordinator. Different procedures. The 2PC post owns the blocking handshake. This page does not relitigate it.
A walk: five nodes, a death, a split, a retry
Nodes n1–n5. Majority = 3. n1 is leader in term 4. Then it dies.
term 4, n1 Leader. Committed log:
idx 1 SET a=1 term 4
idx 2 SET b=2 term 4
n2 n3 n4 n5 Followers. Logs match. Heartbeats stop.
t ≈ 150 ms n2 timeout → Candidate, term 5, vote self
t ≈ 160 ms n3 timeout → Candidate, term 5, vote self
term-5 votes:
n2: self + n4 = 2
n3: self + n5 = 2
n1 dead; nobody has 3
split vote. No leader. Writes still stall.
Each candidate draws a fresh random timeout.
t ≈ 310 ms n3 fires first → Candidate, term 6, vote self
RequestVote → n2, n4, n5
n4 and n5 grant (idle in term 6)
n2 steps down, grants (higher term)
n3 has 4 ≥ 3 → Leader, term 6
n3 ships AppendEntries / heartbeats.
Followers match idx 1–2. Those committed rows stay.
Next client SET: n3 appends idx 3, waits for a majority match, then commits.
The first campaign failed because two timers fired together. The second campaign succeeded because jitter let one candidate collect three votes before the other restarted. That is the whole trick. A wall-clock “newest node” rule would have let n2 and n3 both serve writes in term 5.
Note: A node that is down does not vote. Majority is of the cluster size, not of who answered. Two live nodes in a five-node set cannot elect.
Java: roles as a sketch, not a cluster
No java.util.Raft. The JDK will not elect this for you. etcd, Consul, and Kafka’s metadata quorum (KRaft) already run a Raft-shaped protocol. The enum is vocabulary so the walk above has types. It is not a server.
enum Role {
FOLLOWER, CANDIDATE, LEADER
}
final class NodeState {
Role role = Role.FOLLOWER;
int currentTerm = 0;
Integer votedFor; // node id; null = none this term
void becomeCandidate(int selfId) {
role = Role.CANDIDATE;
currentTerm++;
votedFor = selfId;
}
void becomeLeader() {
if (role != Role.CANDIDATE) {
throw new IllegalStateException("only a candidate becomes leader");
}
role = Role.LEADER;
}
void stepDown(int discoveredTerm) {
currentTerm = discoveredTerm;
role = Role.FOLLOWER;
votedFor = null;
}
}
becomeCandidate is the timeout firing. becomeLeader is majority in this term. stepDown is a higher term or a real leader heartbeat. Missing from this sketch, on purpose: RPC, the log, matching, snapshots, membership changes, disk, clocks. Those are why you do not paste this into production.
Do not grow this enum into a cluster. The hub’s cargo-cult line is the same one as Timsort and gzip: call the library. Awareness is knowing what the library is doing when the leader dies. Implementation is etcd’s job.
Who actually runs this
etcd uses Raft for the Kubernetes control-plane store. Consul uses Raft for its catalog. Kafka’s KRaft metadata quorum is Raft-shaped — this post is not a Kafka course. You operate those systems: odd replica counts, disk, peer list, backup. You do not reimplement RequestVote.
Paxos is a different paper and a different vocabulary. Out of scope. If a design doc says “we will write Raft in a weekend,” the answer is the hub line, not a starter repo.
When not to reach for Raft
Skip this protocol when the job is not “replicated log plus leader election under minority failure.”
- You wanted atomic commit across two databases. That is Two-Phase Commit. Blocking, coordinator-shaped, not an election.
- You wanted a single-process lock. A leader enum in one JVM is not consensus.
synchronizedand a database row are cheaper and honest. - You wanted to ship a cluster. Use etcd, Consul, or the metadata quorum your product already runs. Do not hand-roll Raft.
- You wanted Paxos trivia. Different paper. This page will not recap it.
- You wanted Kafka internals. KRaft is a pointer that the same family shows up in the broker’s metadata. The product course is elsewhere.
Awareness of election and log matching — that is the job. Huffman, interval scheduling, and backtracking are other procedures. They do not elect a leader.
Cheat sheet
Job: one leader, one replicated log, minority failures tolerated
Roles: Follower / Candidate / Leader
Term: monotonic ballot; at most one leader per term
Majority: ⌊n/2⌋ + 1 (5 nodes → 3)
Election: randomized timeout → candidate → RequestVote → majority
Replication: leader appends; followers match previous index/term
Commit: entry on a majority, then apply
Safety: a committed entry is never overwritten
Not this: 2PC, Paxos recap, a from-scratch Java cluster, Kafka-as-product
Who runs it: etcd, Consul, Kafka KRaft metadata — call those, do not paste this
Do:
- Learn the roles, the term, and why majority is three of five.
- Expect split votes; expect a randomized retry to break them.
- Treat committed as durable across leader change. Uncommitted tail may vanish.
- Point at etcd (or the quorum you already run) when someone needs a cluster.
Don’t:
- Hand-roll Raft because the enum fit in a gist.
- Elect on wall clocks or “newest node.” That is two leaders.
- Call 2PC an election. Atomic commit is a different handshake.
- Prove the paper on a whiteboard and ship the whiteboard.
Wrap-up
Raft is leader election plus log replication: random timeouts so campaigns do not tie, a majority so two nodes cannot fork the truth, matching so followers share a prefix, and a commit rule so a committed entry survives the next term. Five nodes, leader dies, split vote, retry — that walk is the awareness. etcd already runs it. 2PC does not. The Java roles are labels, not a cluster. Do not hand-roll this.
When the next systems post needs an approximate distinct count in tiny memory instead of a replicated log, that is HyperLogLog.