43 KiB
m11p5 — Membership, Discovery, Elasticity (COMPLETE — 2026-06-12)
Phase spec and exit gate: docs/roadmap-to-cluster.md §4/m11p5. Predecessors: phase-2.md (StreamSegments — the catch-up stream the snapshot resumes into), phase-3.md (quorum commit index — the arithmetic conf-changes mutate), phase-4.md (terms, fencing, the durable-state disciplines, and three carried hazards this phase closes: quarantine-clearing rides reseed, reseeded nodes lack pre-baseline history, pre-baseline data is unreachable through a term's stream).
Goal: nodes are cattle; topology is data, not files. A new node joins by
contacting any seed, catches up via snapshot + WAL stream in minutes, and
serves quorum; a dead node's replacement is kubectl delete pod; addresses
are DNS names; membership changes ride the replicated log under the m11p4
term machinery.
Design review: a four-lens adversarial review (safety / crash-recovery / liveness / compatibility) of the first draft produced 33 findings (9 BLOCKER, 13 MAJOR, 11 MINOR), all folded in below before implementation. The blockers that reshaped the design: the even-voter-count
majority()arithmetic error (§3.0), the vacuous cross-term conf-change gate and the baseline-jump record skip (both closed by the activation membership re-append, §3.2), learner marks poisoning the quorum (§3.3), and two crash windows in the reseed swap (§2.3).
Design (as adopted)
1. DNS peer addresses: the bind/advertise split
grpc_addr becomes an advertised address — hostname or IP, validated by
the same syntactic validate_host_port that http_addr already passes
(hostnames accepted since m8p10). A new optional per-region grpc_bind
(SocketAddr) controls the local bind:
grpc_bindpresent → bind it.- Absent,
grpc_addrparses as a literal SocketAddr → bind that (today's behavior, byte-for-byte: every existing topology keeps working). - Absent,
grpc_addris a hostname → bind0.0.0.0:<port from grpc_addr>.
resolve_grpc_addr's dual use (own bind + every peer dial) splits: the bind
path stays SocketAddr; the peer path stops parsing entirely. tidal-net's
GrpcTransportConfig.peers retypes from HashMap<ShardId, SocketAddr> to
HashMap<ShardId, String> (host:port); PeerPool builds the tonic URI from
the string — Channel::from_shared with a hostname makes hyper re-resolve
DNS on every reconnect, fixing the pod-rescheduled-onto-a-new-IP case the
SocketAddr type made structurally impossible. grpc_server_ready probes the
bind address, substituting 127.0.0.1 for an unspecified bind IP. TLS note:
SNI follows the URI host, so DNS peer names require DNS-SAN certs —
documented in the grpc_tls topology block; no code change.
This alone kills the per-pod ConfigMap variant: one shared topology file can
name every region by its per-pod DNS name while each pod binds 0.0.0.0.
Verification (review finding): the re-resolution property is the entire
point of the retype and localhost-based tests cannot regress it. A tier-2
test exercises changed-address reconvergence where the platform permits
(dual-family localhost rebind), and the k8s pod-reschedule drill (delete
pod, new pod IP, ship/stream traffic resumes with no peer restarts) is a
required line item of the exit-gate evidence, not an incidental.
2. Snapshot transfer (FetchSnapshot): the reseed and joiner path
Artifact = TidalDb::create_backup (tidal/src/db/backup.rs:211) with
three review-driven hardenings:
- the WAL-checkpoint-marker write into the copy becomes fatal (a staged artifact with a stale marker double-counts signals on the joiner);
create_backupalso refreshes the replication-HWM row in the same step as the ledger checkpoint (the copied row otherwise understates the artifact by up to one 30 s cycle → re-folds on the joiner);- the snapshot-staging variant skips the non-fatal text-index flushes (derived state, rebuilt on open) to shorten the quiesce window.
The leader stages the artifact under <data_dir>/snapshots/ and serves it
as a new server-streaming RPC in the WalShipping service:
rpc FetchSnapshot(SnapshotRequest) returns (stream SnapshotChunk)
SnapshotRequest { shard_id, from_seqno, term }—from_seqnois the puller's frontier+1. The handler answers a no-snapshot-needed header chunk when the WAL can still servefrom_seqno(the joiner just streams), else streams: manifest chunk (snapshot seq, file list with size + BLAKE3, term, leader_region) then file chunks (path, offset, bytes), term-stamped and term-fenced exactly likeStreamSegments(same-term puller; stale source steps down; puller gates every chunk).- Excluded from the manifest (node-identity / node-local files):
election_state,stream_baseline,membership(the §3.6 cache),tidaldb.lock,snapshots/, the reseed marker. Everything else ships byte-identical (TSEG segments must arrive bit-exact or the joiner's open refuses — identify-or-refuse end to end). - Caching + single-flight: one staged artifact at a time; concurrent
joiners share it. A second
FetchSnapshotarriving while staging is mid-flight gets a retryableUNAVAILABLE(neverFAILED_PRECONDITION, which would mis-route it into the reseed class). An artifact is reusable while the live WAL still coversartifact_seq + 1(else refresh). - Quiesce honesty (review finding): every
create_backup429s all signal writes for checkpoint + flush + copy (seconds at 100k items). Gate writers retry 429 honoringretry_after_ms— a retried-then-acked write is neither lost nor an SLO violation (the established busy-writer harness posture). The measured window is recorded in the gate evidence.
2.1 Retention pin (review-corrected). The pin is taken at backup
start, at the captured seq S, before any copying — a pin installed
after staging cannot protect the copy window from the 30 s checkpoint
thread. Pin semantics: online compaction must preserve the segment chain
covering S + 1 onward (the pin clamps the floor such that the segment
containing S + 1 survives — an off-by-one that deletes the straddling
segment is data loss inside the artifact; unit-tested). The pin holds
while: the artifact is staged, any FetchSnapshot stream is active, and
for a per-consumer grace until that puller's first successful
StreamSegments pull from S + 1 (bounded by a hard cap + loud metric so
a dead joiner cannot pin retention forever). The reuse TTL
(snapshot_artifact_ttl_ms) governs artifact refresh only — never
the pin of an active consumer; consecutive reseed loops per peer are
counted and alerted at ≥ 2.
2.2 Install is a boot-time operation — no live data-dir swap under a
running engine, ever. The pre-open step runs only on (a) seed-join
boots and (b) reseed_required-marker boots — plain topology boots never
contact a peer before serving (boot-order independence at cold start is
load-bearing; a behind-a-compacted-leader restart is detected at runtime
per §2.4 and reseeds on the next boot). A marker boot that cannot
complete the handshake within a bounded window falls back to opening the
existing data dir — degraded, marker latched, quarantine fence intact,
voting enabled — because the reseed needs a leader and the leader may
need this node's vote (review finding: reseed-blocks-voting deadlocks the
cluster). Fresh joiners have nothing to fall back to and nothing to vote
with; they retry the seed loop indefinitely with bounded backoff.
2.3 The swap protocol (review-corrected ordering). Fetch into a sibling staging dir, then:
- Verify every file's BLAKE3 + manifest completeness.
- Copy the node's own identity files INTO staging:
election_state(vote safety, §2.5) andmembershipcache — before the sentinel. - Write the
COMPLETEsentinel into staging; fsync. rename(data_dir → data_dir.discard),rename(staging → data_dir), fsync parent, delete.discard.
Filesystem constraint: the swap renames data_dir itself, so data_dir
must not be a filesystem/volume mount root — staging and .discard live in
its parent, same filesystem (EXDEV/permission failures surface loudly at
reseed time, never silently). The k8s cluster manifests therefore mount the
PVC at /data and set --data-dir /data/db. Leader-side staging has the
inverse problem: it lives inside the data dir (<data_dir>/snapshots/) so
it stays on the PVC, which means create_backup's recursive copy must
exclude it (along with the identity files) or it would copy itself.
Boot-time recovery (runs before ElectionStore classification — the
review found that classifying first re-opens restart amnesia):
staging-without-COMPLETE → delete staging, retry/fall back;
staging-with-COMPLETE + data_dir present → redo step 4 from the top;
staging-with-COMPLETE + data_dir absent (the mid-rename crash — the
window where a naive boot would classify Fresh at term 0 and re-vote) →
complete the rename, then classify; stray .discard → delete. Every
window is idempotent redo; the identity files travel inside staging so
there is no instant at which the canonical path lacks them post-swap.
2.4 Runtime detection → durable marker, not live surgery. A running
follower that hits a typed snapshot-required refusal (see below) — or
the m11p4 divergence quarantine — durably latches reseed_required
(ElectionStore file discipline: magic, version, checksum, tmp + fsync +
rename + dir fsync), surfaces reseed_required: true in
/cluster/status/local plus a tidaldb_cluster_reseed_required gauge,
and keeps serving degraded (the 30 s retry timer keeps its standing
wake-up). The reseed executes on the next boot. POST /cluster/reseed
sets the marker on demand; replication.reseed_self_restart: true
(default false; set in the k8s manifests) drains and exits cleanly
once the marker latches — refused (loudly, in status + gauge) when the
node's local view shows the remaining voters cannot sustain quorum without
it (review finding: self-restart during a 2-voter window is total write
unavailability). Quarantine clearing (the p4 carried hazard): quarantine
latches the marker; a successful reseed boot clears the quarantine latch
and tidaldb_cluster_divergence.
The three-way term-join rule (found by the tier-3 gate, not the review).
The first reseed drill exposed that a follower behind at a leadership change
never asks for its missing range: the term-join baseline jump advances its
frontier past pre-baseline history with lag=0 — silently missing data,
exactly p4's carried hazard, with nothing to fire the snapshot-required
refusal. Detection cannot use the baseline (cross-numbering; p4 as-built
#2); it uses the comparison the join check already makes. The heartbeat
carries the leader's election-time position in the previous stream's
numbering (prev_log); the join check becomes three-way:
own > prev_log→ divergent suffix → quarantine (existing, p4);own < prev_log→ the node lacks(own, prev_log]of the previous stream, which is pre-baseline in the new stream and will never ship → latchreseed_required(new); the vote restriction makes this exact: an elected leader's position ≥ every granter's, so a strictly-smaller position is genuinely missing committed-era history;own == prev_log(ortail_term == T, the within-term rejoin) → clean.
A snapshot-installed node joins clean by construction: its WAL is the
leader's copy, so its tail_term equals the leader's term. Two corollaries
from the same drill: the install boot must issue its post-install catch-up
pull toward the discovered leader (the boot self-heal targets the boot
topology leader — which can be the node itself), and the FetchSnapshot
needed decision must be baseline-aware (needed = from_seqno ≤ stream_baseline OR from_seqno < earliest WAL seq) — a "WAL still covers
it" answer is a lie when the stream clamp will never serve at or below the
baseline, and it loops the marker boot forever.
Typed refusals (review finding — the conflation hazard): today all
three StreamSegments refusal classes are FAILED_PRECONDITION, and a
term-mismatch during an ordinary election must NOT latch reseed markers
fleet-wide. The refusal gains structured trailer metadata
(x-tidal-catchup: snapshot-required | rejoin | stepping-down); the
marker latches only on snapshot-required. Pre-p5 sources emit no
trailer → no latch (conservative).
2.5 Vote safety across reseed. The node's own election_state (term,
voted_for) survives the swap inside staging (§2.3 step 2) — it is the
node's promise "I voted for X in term T"; wiping it would allow a re-vote
in an already-voted term (two leaders in one term). A fresh joiner has
none — it never voted. A corrupt election_state refuses boot even
with a reseed marker latched (a reseed does not restore term knowledge;
the m11p4 rule stands).
2.6 Post-open seeding (review-corrected durability). The install boot,
before starting the receiver, transport serving, or any pull: seed
applied_seqno(leader_shard) = the installed artifact's recovered WAL tail (valid at the install instant only — own-WAL numbering diverges from
stream numbering as soon as applies begin; p4 as-built #1), then
synchronously persist the replication-state checkpoint, then delete
the install sentinel, then start replication. A crash anywhere before the
sentinel delete finds the WAL unchanged (no receiver ran) and re-derives
the same seed idempotently. The copied HWM row's other-shard entries are
kept (true for any copy-holder once §2's create_backup refreshes the row
in-step). wal_term_mark recovers from the copied log's kind-3 markers by
construction, so the node's LogPosition is consistent in the stream's
numbering with no new bookkeeping.
2.7 Fresh-joiner term sequencing. A seed-joiner learns (term, leader)
from the join response, persists election_state {term} before
pulling (persist-before-act), and stamps the snapshot fetch and subsequent
stream pulls with that term. The boot-time loop iterates
(seeds ∪ cached membership) → poll /cluster/status for the current
leader → handshake → fetch, with bounded backoff and indefinite
persistence, re-resolving leadership after every aborted stream (an
election mid-fetch aborts on the term fence; the loop re-discovers and
re-stamps — review finding: without re-discovery the joiner retries a
deposed source forever). "Backup already in progress" maps into the same
retry loop. Process exit is reserved for unrecoverable config errors —
crash-loop is never the retry mechanism.
3. Membership as data: kind-4 records on the one log
3.0 Prerequisite arithmetic fix (review BLOCKER, exists today).
ElectionConfig::majority() is peers.len()/2 + 1 — correct at n=3,
wrong for every even voter count: n=4 → 2-of-4 (disjoint vote quorums
{A,B} and {C,D} elect two leaders in one term), n=2 → 1-of-2 (a follower
self-elects with zero RPCs while check-quorum keeps the old leader alive).
CommitIndex already uses the correct div_ceil form — the two formulas
in one subsystem disagree. p5 makes even sizes mandatory transit states
(3→4→5→4→3), so this lands first, with even-n property tests, before
any reconfigure work: majority() = (peers.len() + 1)/2 + 1 (true
floor(n/2)+1 over the full voter set).
A MembershipRecord is a new WAL blob kind (kind-4, beside kind-0
signals, kind-1/2 item/embedding blobs, kind-3 term markers), journaled by
the leader, replicated through the normal stream, folded by followers into a
ClusterMembership cell (the WalTermMark pattern: recovered from the WAL
at boot, advanced on apply, single-lock value):
MembershipRecord {
version: u64, // monotonic conf version; latest record wins
term: u64, // the appending leadership
members: Vec<MemberEntry {
id: u16, // PERMANENT — never renumbered, never reused
name: String,
grpc_addr: String, // advertised, DNS-capable
http_addr: String,
role: Voter | Learner | Removed, // tombstones keep ids burned
}>
}
Records carry the FULL membership (not deltas): recovery folds the latest record at-or-below the durable frontier, no merge logic. Membership epoch 0 = the topology file — a cluster with no kind-4 record behaves byte-for-byte as today. The first record snapshots the topology-derived membership (ids = the existing positional ids, preserved) plus the change.
3.1 Capability gating (review finding — the brick hazard). A kind-4
record shipped to a pre-p5 follower is an unknown batch kind →
WalError::Corruption → the receiver's torn-state halt latch, permanent
across restarts (boot self-heal re-pulls the same record). The m11p4
kind-3 precedent does NOT transfer: term markers were self-gated by the
vote quorum needing upgraded binaries; kind-4 needs only one upgraded
leader plus one verb — and auto-promotion appends records autonomously.
Therefore: HeartbeatResponse and ReportApplied gain a capabilities
bit-field (proto3 zero-default = pre-p5 = incapable), and the leader
refuses JoinCluster and every conf-change until all current voters
have reported kind-4 capability. The failure mode is stated honestly in
runbook §8: a kind-4 record delivered to a pre-p5 follower halts its
receiver until binary upgrade; both-followers-halted = quorum-write
outage. Downgrade rule (kind-3 precedent verbatim): once any kind-4 record
is in a node's WAL, downgrade below p5 requires a reseed.
3.2 The activation membership record (review BLOCKER ×2 — this is the linchpin). On every leadership activation, immediately after the kind-3 term marker, the new leader re-appends its full current membership as a fresh kind-4 record (the Raft no-op-entry analogue, adapted to the baseline-jump world). This single rule closes two blockers at once:
- The vacuous gate:
CommitIndex::activateresetscommittedto the new leader's baseline, so "commit index ≥ record seq" is trivially true for every pre-baseline record — and pre-baseline records are structurally unshippable (the stream clamp). The one-at-a-time gate therefore requires same-term quorum commitment of the activation record (and of each subsequent change), evaluated fromupdate_peer_for_termfolds — never from the activation-resetcommitted(). - The baseline-jump skip: a follower behind at a transfer jumps its frontier over any kind-4 record the old term committed, with no gap, no signal, while staying electable under a stale roster. The activation record sits above the new baseline, so every joiner of the term receives the current roster in-stream before any same-term traffic.
3.3 Conf-change discipline (Raft single-server changes):
- One at a time: the leader refuses a new change until the prior record's seq is same-term quorum-committed (§3.2 predicate). Majorities of consecutive single-change configs overlap; the gate's integrity across leadership changes is exactly what §3.2 restores.
- Join → Learner → Voter:
POST /cluster/join {name, grpc_addr, http_addr}on any node forwards to the leader, which assignsid = max(all ids ever) + 1, appends a Learner record, and answers only after that record is quorum-committed (bounded wait → retryable 503; review finding: answering pre-commit strands the joiner invisibly when the leader dies). Joins are idempotent by name: a re-join from a known member returns its existing id and current role, appending nothing. The response carries membership, current term, leader addresses, and the assigned id. - Learner marks (review BLOCKER): the role-blind
CommitIndexmust not see learner marks — two learners "committing" a write no voter holds is acked-write loss.CommitIndexbecomes role-aware: voter marks feed the k-th-largest selection; learner marks are tracked in a side map (promotion input, transfer-wait input) and never count towardneeded. - Auto-promotion is a standing leader duty derived solely from the
applied
ClusterMembership— re-armed on every leadership activation and on every membership apply, evaluated on commit-index publishes (review finding: a duty anchored to the join handler dies with the joining-era leader, leaving a permanent learner). Promotion fires when the learner's durable mark is withinlearner_promote_lagofflushed_seqor has stayed within one ship-round for K consecutive evaluations (the Raft "rounds stop shrinking" criterion — a fixed distance alone starves under sustained load). Apromotion_pending (lag=N)status field + metric make a stuck scale-up diagnosable. - Remove: an operator verb appends a
Removedrecord (one at a time, quorum-commit-gated like every change). The removed peer's ship cell is retired only after the record is quorum-committed AND delivered to (or acked by) the removed peer, with a bounded give-up (review finding: retiring on append means the removed node never learns, never stops campaigning, and zombie-serves as Ready). Heartbeat/vote/pull refusals to a removed member carry a typedremovedsignal that flips its readiness to 503, suppresses campaigning, and is exempt from the reseed marker. Its id is burned. - Effective-set switching: each node's effective peer sets derive
from its latest APPLIED record (leader: on append). The four
construction-frozen peer-set copies become mutable behind one fenced
reconfigure path executed inside a single membership-apply critical
section — the sets can never disagree about the roster:
CommitIndex::reconfigure(voters, learners)— resize marks +neededunder the index's own lock. In-flightack=quorumwaits are re-evaluated against the new config, not failed (review finding: epoch-bump-as-Demoted on every conf-change is an availability dip with wrong error semantics); a shrink may satisfy waiters instantly. Unknown-peer reports stop being silently dropped: a report from a known learner records its side-map mark; a report from an unknown id logs at WARN with a counter metric.ShipQueue::{add_peer, remove_peer}— spawn/retire per-peer sender cells and threads at runtime (removal per the delivery rule above).PeerPool::{add_peer, remove_peer}— peers map behind a RwLock; channels are lazy; DNS strings mean no resolution at insert.ElectionState::reconfigure(voters)— a new pure-machine input: recomputemajority()from the post-change voter set, resizepeer_ack, and clear in-flight prevote/vote grant sets (review finding: grants from just-removed voters must not count against the resized majority). The driver'speer_shardsand the node's HTTP tables update in the same apply step.
- Election interaction: voters =
role == Voterin the latest applied record. Learners never campaign — enforced as a gate in the state machine (auto-election suppressed while not a voter), not a comment (the review found p4's quarantine election-suppression is comment-only; that gap is verified and fixed in-phase with a unit test). A node that hasn't applied the newest record may transiently refuse a legitimate candidate — the next timeout retries; safety is the overlap argument, liveness is the retry.
3.4 Seed-join boot (--seed http://host:port, repeatable): the joiner
skips validate_multiproc's "every region declared" gate — its identity is
--region <name> + advertised addresses from new flags/env
(--advertise-grpc, --advertise-http, --metrics). It runs the §2.7
loop, persists the returned membership to the durable local cache, persists
the term, then proceeds through the normal boot with the fetched roster
(snapshot install per §2 when needed). A restart boots from the cache
without the seed.
3.5 Knob source for seed boots (review finding). A bare --seed boot
with no local config would silently fall back to the compiled-in
default-cluster.yaml — wrong ack default, wrong quorum timeout, wrong
election timing, no metrics listener, no TLS material. Rule: a --seed
boot still requires the local topology/config file for the behavioral knob
blocks (replication:, wal:, election:, timeouts:, grpc_tls) —
the k8s manifests already mount the shared bootstrap ConfigMap on every
pod including N≥3, so this costs nothing there — and a bare --seed with
neither --topology nor TIDAL_CONFIG refuses to boot naming the
rule. The regions: list in that file is ignored for the roster (the
join response is the roster); the C2 election-timing validation runs on
the knob blocks as always. The joiner's metrics listener comes from
--metrics (new flag, mirroring standalone's).
3.6 Durable membership cache (data_dir/membership, ElectionStore file
discipline: magic, version, checksum, tmp+fsync+rename+dir-fsync).
Precedence: the WAL-recovered ClusterMembership cell wins over the
cache at open (the cache exists for the pre-open boot loop — seed lists,
leader discovery — not as a second source of roster truth); the cache is
rewritten from the cell after every applied record. It is excluded from
snapshot manifests and survives the reseed swap inside staging (§2.3).
3.7 Mixed-version rule: §3.1's capability gate makes "complete the
binary upgrade before the first conf-change" structurally enforced, not
operator discipline. Pre-p5 peers answer Unimplemented to
JoinCluster/FetchSnapshot; the joiner reports it loudly and retries the
next seed.
4. Kubernetes reference: one StatefulSet
New k8s/cluster/ kustomize set beside the standalone one. Naming
(review finding — the lineages must not collide): the cluster set lives
in its own namespace tidaldb-cluster with StatefulSet tidaldb, Services
tidaldb-peers (headless, publishNotReadyAddresses: true) and tidaldb
(client-facing, readiness-gated); the two kustomize sets are mutually
exclusive per namespace. Secret shape: tidaldb-credentials /
TIDAL_API_KEY (the stress/Ref-A lineage — the in-repo stress Jobs are the
exit-gate harness and already use it); the stress Jobs' literal-ClusterIP
--targets move to the new DNS names in this phase.
- One StatefulSet,
replicas: 3, podManagementPolicyParallel,TIDAL_REGIONfromPOD_NAME(fieldRef), argscluster --listen 0.0.0.0:9500 --data-dir /data ..., gRPC containerPort, topologySpreadConstraints. - Bootstrap topology ConfigMap — ONE file shared by all pods (per-pod
DNS advertised addresses +
0.0.0.0binds), naming the initial 3 regions astidaldb-{0,1,2}.tidaldb-peers.tidaldb-cluster.svc.cluster.local. Scaling past 3 does NOT edit it: pod N ≥ 3 boots with--seed http://tidaldb-peers...+ the same mounted file for knob blocks (§3.5) and joins as a learner. - Readiness (review-pinned predicate):
converged :=the boot catch-up pull has completed at least once ANDlag_events ≤ learner_promote_lag(hysteresis — neverlag == 0, which an open-loop load keeps perpetually false);joiner :=this boot installed a snapshot or seed-joined with a fresh data dir (derivable at boot, no new state). 503 while quarantined, removed, or a joiner-boot has not yet first-converged (sticky-ready after). A restarted existing voter is Ready on today's terms — no regression for PVC-retained restarts. The predicate is recorded in the runbook so probe behavior is diagnosable. - PDB
maxUnavailable: 1;reseed_self_restart: true; rolling node replace iskubectl delete pod(PVC retained → boot catch-up) or PVC delete + pod delete (fresh reseed via snapshot). - Scale-down drill: remove verb first, then
kubectl scale --replicas, lowest-ordinal-last.
5. What goes where
| Crate | Work |
|---|---|
tidal |
majority() even-n fix + property tests (FIRST); kind-4 MembershipRecord encode/decode/recover/apply + ClusterMembership cell; role-aware CommitIndex (reconfigure, learner side-map, waiter re-evaluation, WARN-counter for unknown reporters); ShipQueue::{add_peer,remove_peer}; ElectionState::reconfigure + learner campaign gate (+ verify/fix the p4 quarantine campaign gate); compaction retention pin (straddling-segment-safe, unit-tested); create_backup hardenings (fatal marker, HWM refresh in-step, staging variant skipping text flush) |
tidal-net |
proto: FetchSnapshot/SnapshotChunk, JoinCluster, capabilities on HeartbeatResponse/AppliedReport; typed catch-up refusal trailer (x-tidal-catchup); SnapshotSource trait (late-bound) + server handler (StreamSegments-shaped); standalone snapshot-fetch client + PeerPool verbs; peers retype to String + dynamic add/remove; changed-address reconvergence test |
tidal-server |
bind/advertise split; boot-time install (§2.2 scope, §2.3 swap + recovery, §2.5/§2.6 ordering); seed-join boot loop (§2.7) + flags (--seed, --advertise-*, --metrics) + §3.5 knob-source rule; membership runtime (one fenced apply path; activation record; capability gate; auto-promotion duty; removal delivery); /cluster/join, /cluster/members (+remove), /cluster/reseed verbs; reseed marker + quarantine wiring + self-restart quorum refusal; readiness predicate; metrics; status fields |
k8s/ |
k8s/cluster/ per §4; stress-Job retargeting |
| tests | unit: even-n majority properties, record/store round-trips, reconfigure paths (incl. grant-set clearing, learner marks, waiter re-evaluation), retention-pin straddle, swap-recovery windows, topology DNS derivation table; tidal-net sockets: snapshot stream (fake-source), join RPC, capability gate, typed refusal, changed-address; tier-3: mp_seed_join_snapshot_catchup (vs a compacted leader), reseed drill (quarantine → marker → restart → converged, gauges cleared, no wipe_data_dir), DNS-hostname topology + chaos, mp_scale_3_5_3_under_load_zero_loss (ledger invariants, 429-retrying writers, p99 probes) |
6. Configuration
| Knob | Default | Meaning |
|---|---|---|
regions[].grpc_bind |
derived (§1) | local gRPC bind when grpc_addr is a DNS name |
replication.reseed_self_restart |
false | drain + clean exit once reseed_required latches (k8s: true); refused when remaining voters can't sustain quorum |
replication.snapshot_artifact_ttl_ms |
600000 | staged-artifact reuse window (never the active-consumer pin, §2.1) |
replication.learner_promote_lag |
1024 | promotion distance; also the readiness hysteresis threshold (§4) |
CLI --seed <url> (repeatable), --advertise-grpc, --advertise-http, --metrics |
— | seed-join boot (requires the local knob file, §3.5) |
Exit gate (from the roadmap, restated as tests)
- Seed join + snapshot catch-up: a fresh node with an empty data dir
joins via
--seedagainst a leader whose WAL has rotated AND compacted past seq 1 → converges via snapshot + stream with zero operator verbs and no full-log replay; tier-3mp_seed_join_snapshot_catchup. - Scale 3→5→3 online under load: background
ack=quorumwriters (WriterTally, lost==0, retrying 429s per §2) through both gateways while two nodes join, promote to voters, then one is removed and decommissioned → m11p3 ledger invariants (frontier + content) hold throughout; the joiner serves quorum (its reports advance the commit index); p99 impact <2× for <60 s; 100k-item joiner catch-up ≤5 min (env-scaled gate run recorded here). - Reseed self-healing: a quarantined divergent node (p4 drill) latches
reseed_required, restarts, reseeds via snapshot, rejoins clean, and the divergence gauge clears — nowipe_data_dirin the test. - DNS: a tier-3 cluster whose topology names peers by hostname boots, replicates, and survives the existing chaos drills; unit tests pin the bind/advertise derivation table; the k8s pod-reschedule drill is recorded as exit-gate evidence (§1).
Status
- Design adopted — 33 adversarial-review findings (9 BLOCKER) folded in pre-implementation
- §3.0 majority() fix + DNS/bind-advertise split —
tidal,tidal-net,tidal-server - Snapshot transfer + boot-time install + reseed marker — all crates
- Membership records + conf-changes + seed join — all crates
- Kubernetes reference (
k8s/cluster/) - Exit-gate suites green; gate evidence recorded
- Docs (runbook §§1b/3/6/8/9/11, kubernetes.md, CHANGELOG, roadmap, spec 01 §2.2, monitoring.md)
As built (deltas vs the adopted design)
-
majority()even-n fix landeddiv_ceil-spelled, not the(n+1)/2+1form the design wrote.(peers.len() + 1).div_ceil(2) + 1is the exact samefloor(n/2)+1over the full voter set (n = peers.len()+1), but spelled withdiv_ceilso the two arithmetic sites in the subsystem —ElectionConfig::majority()andCommitIndex— now read identically (the design's whole point: the two formulas can no longer drift). Even-n property tests landed first, as planned. -
The three-way term-join rule is implemented, AND
FetchSnapshot'sneededdecision is baseline-aware (§2.4 as adopted). Both were found by the first tier-3 reseed drill, not the pre-implementation review, and both are in the shipped design above — recorded here too because they were the single largest correctness delta from the first draft: a follower behind at a leadership change advances its frontier past pre-baseline history withlag=0and never asks for the gap, so detection rides the heartbeat'sprev_logcomparison (own < prev_log→ latchreseed_required), and the snapshot handler answersneededfromfrom_seqno ≤ stream_baseline OR from_seqno < earliest WAL seqrather than a naive "the WAL still covers it" (which loops a marker boot forever at or below the baseline). -
Three product bugs surfaced and were fixed during the C3 build (the membership-runtime + verb stage), all in the as-built behavior above: (a) the membership-apply path had to enter the async reactor context correctly before driving the fenced reconfigure — a stray non-reactor enter wedged the first conf-change; (b) a freshly-joined Learner did not receive leader heartbeats until its first applied record, so its promotion clock never started — heartbeats now flow to learners (they are in the ship set, just not the quorum); (c) boot-time reconcile of the durable membership cache against the WAL-recovered
ClusterMembershipcell had to prefer the cell unconditionally (§3.6) — an early version let a stale cache shadow a newer in-stream roster on restart. -
Auto-promotion landed in the C3 stage, alongside the membership runtime — not as a separate increment. It is the standing leader duty of §3.3 (re-armed on every leadership activation and every membership apply, evaluated on commit-index publishes), driven solely from the applied
ClusterMembership, so it survives the joining-era leader's death. -
Removal grace mechanics (D1). The
Removedrecord retires the peer's ship cell only after the record is quorum-committed AND delivered-to / acked-by the removed peer, with a bounded give-up that incrementstidaldb_cluster_remove_delivery_giveups_totaland a typedremovedrefusal that flips the removed node's readiness to 503 and suppresses its campaigning (§3.3 as adopted; D1 is where the delivery-then-retire ordering and the give-up counter were wired end to end). -
needed_peersmoved from a construction-frozen constant to a function of the latest applied record. The four peer-set copies (CommitIndex, ShipQueue, PeerPool, ElectionState) wereconst-derived at boot in p4; the fenced reconfigure path (§3.3) made them a function of the appliedClusterMembership, evaluated inside the single membership-apply critical section so they can never disagree about the roster. -
Re-appending the activation membership record is non-fatal on failure. The §3.2 activation re-append (the linchpin that closes the vacuous gate and the baseline-jump skip) logs and continues if the append itself fails rather than aborting the activation — a leader that cannot journal its roster is still a leader for fencing purposes, and the next apply/activation retries the record. (The conf-change gate still requires same-term quorum commitment of the record before the NEXT change, so safety is unaffected; only the cosmetic "every activation always has a record" property degrades gracefully under a write fault.)
-
The faster election machinery changed
/cluster/promote's topology-era behavior — fixed by gating the transfer leg on the era. m11p4's promote handler relied on a behavioral accident: in the term-0 topology era,transfer_to(target)never took (the target couldn't acknowledge aTimeoutNowat a term it hadn't joined), so the handler always fell through to the legacy fan-out and its documented{ok, leader, baseline, acked, failed}response shape. p5's election work made that same transfer succeed, returning the elected-era{ok, leader, term, transfer}shape early and breaking the runbook's deserialization plus the non-leader forward leg (cluster_runbook×2 +cluster_lifecycle×1, surfaced by the D2 regression sweep). Fix: a topology-era leader does not fence-transfer — it falls through to the legacy fan-out as documented — while the self-campaign leg (the reseed drill's re-leadership path) and the genuinely-elected era keep the fenced transfer.term/transferremain additive fields on the elected path.
Exit-gate evidence
All gate suites are tier-3 (--features cluster-e2e), real OS processes over
localhost loopback, release test binary + debug-spawned servers. The k3s
(Ref-A) re-run is pending infra access (the standing M11 caveat, see
Carried hazards) — the StatefulSet reference (k8s/cluster/) is the
deliverable, but the under-load gate numbers below are localhost, stated
honestly as such.
| Gate (roadmap, restated as test) | Suite / test | Result |
|---|---|---|
| 1. Seed join + snapshot catch-up (fresh node, leader compacted past seq 1, zero verbs, no full-log replay) | cluster_reseed.rs + cluster_membership.rs::mp_seed_join_snapshot_catchup |
ok — converges via snapshot + stream; env-scaled: 320 items join→converged 5.3s / 44 MB, 2000 → 13.0s / 269 MB, 5000 → 26.4s / 655 MB, every probed item searchable. The catch-up metric is ~26s at 5000 heavy items, large headroom under the 5-min budget. |
| 2. Scale 3→5→3 online under load (ack=quorum writers retrying 429s, ledger invariants hold, joiner serves quorum, p99 <2× for <60s) | cluster_membership.rs::mp_scale_3_5_3_under_load_zero_loss |
ok — lost=0, max_acked_seq=1467; INVARIANT A (max acked seq ≤ leader frontier) + INVARIANT B (sampled acked content searchable) held at every 3→4→5→4 transition; per-window write p99 before=54.6ms, join-A=105.5ms, join-B=68.6ms, after=102.6ms (all <2× steady, within the <2×-for-<60s Ref-A figure; recorded, not hard-asserted per the non-flaky posture). |
3. Reseed self-healing (quarantined divergent node latches reseed_required, restarts, reseeds via snapshot, gauge clears, no wipe_data_dir) |
cluster_reseed.rs |
ok — quarantine → marker → restart → converged; tidaldb_cluster_divergence_quarantined + tidaldb_cluster_reseed_required clear after a genuine reseed; no wipe_data_dir in the test. |
| 4. DNS (hostname-topology cluster boots, replicates, survives chaos; bind/advertise derivation table pinned) | cluster_membership.rs::mp_dns_hostname_topology_replicates + unit derivation tests |
ok — a 3-node hostname topology (which a pre-m11p5 SocketAddr::parse would have refused) boots, replicates a leader write to a follower, and survives a SIGKILL+restart drill (the resolver re-dials on reconnect — the point of the String peer retype). The k8s pod-reschedule drill (changed pod IP) is the deliberate Ref-A line item, pending infra access. |
Full suite tail (2026-06-12): cluster_membership 5 passed / 0 failed (86.78s,
incl. "[scale] exit gate 2 met" and "[dns] exit gate 4 met"); cluster_multiproc
5 passed, cluster_chaos 4 passed, cluster_quorum 2 passed,
cluster_election 3 passed, cluster_reseed 2 passed. cargo test -p tidaldb --lib 1896 passed / 0 failed; cargo clippy --workspace --all-targets --all-features -- -D warnings clean; cargo fmt --check -p tidal-server exit 0.
The D2 sweep surfaced a promote regression in
cluster_lifecycle/cluster_runbook— fixed post-D (see As-built #8): all eight tier-3 suites are green. Final post-fix sweep (2026-06-12): multiproc 5, chaos 4, lifecycle 2, runbook 9, quorum 2, election 3, reseed 2, membership 5 — all 0 failed when run per-suite;tidaldb --lib1896 passed; tidal-net 10 + tidal-server 19 ok-binaries; workspace clippy-D warningsclean; fmt clean;scripts/check-docs.shOK. The suites are load-marginal when run back-to-back on one saturated host (a colocated sweep flaked one runbook drill that passes 9/9 isolated) — the same documented sensitivity class as p4'smp_rolling_upgradenote; widenTIDAL_TEST_CONVERGENCE_BUDGET_SECSon slow runners.
Carried hazards (tracked, not regressions)
-
Ref-A (k3s) run pending infra access. Every under-load and catch-up number above is localhost loopback on a debug-spawned server, not Ref-A. The k3s deployment has been unreachable from the local environment since m11p1 (the standing caveat in p2/p3); p5's k8s-centric gate cannot fall back to localhost the way p1–p3 did, so the StatefulSet reference (
k8s/cluster/) is the deliverable and the k8s pod-reschedule drill (gate 4, §1) + 100k-item Ref-A catch-up remain the named Ref-A line items. The 5000-item localhost run is the largest that completed in reasonable wall time — the ceiling was the ~13 items/sec HTTP seed rate against a debug server (the seed cost, not the catch-up path the gate measures), not a catch-up limit. -
Tier-3 suites are load-marginal under colocation. Each suite is green in isolation (the way gate runs are recorded), but running all eight back-to-back on one saturated host can flake a FAST_ELECTION-timed drill (one runbook failure in a full-sweep run that passes 9/9 isolated, three times over). Same class as p4's
mp_rolling_upgradenote; the reseed suite's content probes were widened from a fixed 15 s to the harnessconvergence_budget()in-phase. m11p9's nightly CI should run suites serially or on isolated runners. -
The quiesce window is recorded, not yet measured at 100k on Ref-A. §2's
create_backup429s signal writes for checkpoint + flush + copy; gate writers retry 429 honoringretry_after_ms(neither lost nor an SLO violation), and the localhost windows are within budget — but the "seconds at 100k items" figure the design names is a Ref-A measurement still owed once infra access returns. -
DNS changed-IP coverage is the k8s drill, not a localhost test. The re-resolution property (a pod rescheduled onto a new IP, ship/stream resumes with no peer restarts) is the entire point of the
String-peer retype and cannot be regressed by a localhost test that never changes addresses. Themp_dns_hostname_topology_replicatestest proves the hostname-topology boot- reconnect re-dial; the changed-address reconvergence is the k8s pod-reschedule line item under hazard 1.