tidaldb/docs/planning/milestone-11/phase-9.md
jx12n 580142df49 feat(m12): election-divergence-fix + soak-eval streak + release tooling
Durable `leader_acked` frontier in `ShardReplica` tracks the highest seqno
acked under `ack=leader` (journal-only, un-replicated); `decide_join` now
quarantines on THIS node's own frontier rather than comparing stream numbers
across stream boundaries — eliminates false-quarantine churn on rolling
restarts. `SHUTDOWN_HANDOFF_WAIT` (3s) drains the leader's tail to quorum
before step-down so the next leader inherits a clean prefix. New
`load_leader_acked`/`persist_leader_acked` helpers; `cluster_reseed.rs` gains
the divergence-fix regression suite; `replication_ops.rs` threads the signal.

Soak-eval: `tidal_stress::soak_eval` + `soak-eval` binary implement the
30-night streak (ledger.tsv × restarts.tsv → streak.tsv); monitor and nightly
CronJob k8s YAMLs updated; phase-9 doc clarifies the dual-stream streak
definition (ledger PASS AND zero pod restarts in window). `run-reliability.sh`
gates the election-divergence suite before any k8s push.

Release tooling: `docker/release/` multi-stage Dockerfile + DR image;
`scripts/build-release.sh` single repeatable cross-compile+buildx path.
2026-06-18 13:08:53 -06:00

16 KiB
Raw Blame History

m11p9 — Continuous Correctness (COMPLETE — 2026-06-13)

Phase spec and exit gate: docs/roadmap-to-cluster.md §4/m11p9. Closes the roadmap's G-C (Continuous correctness) — trust is a pipeline, not a milestone — and wires the named owner-test for every other §2 guarantee (guarantee-traceability.md). Predecessors: every prior m11 phase (this phase industrializes their tier-3 suites and adds the fault classes they lacked).

Goal: make correctness continuous — new fault classes (disk-full, slow-fsync, asymmetric partition) as REAL faults; the invariant checks the durability and election gates carried inline promoted to first-class reusable checkers; a nightly chaos + soak pipeline that fails on a regression; and a guarantee→test matrix so the GA bar is auditable.

Design (as adopted)

1. Fault injection compiled out of production (tidal/src/fault.rs)

Two new fault classes need REAL faults at the storage boundary, not engine flags:

  • TIDAL_FAULT_FSYNC_DELAY_MS — sleep before every durable WAL fsync (wal::sync_file_durable). Models a slow disk: data still reaches stable storage, each group-commit fsync costs delay + real.
  • TIDAL_FAULT_DISK_FULL_AFTER_BYTES — after N cumulative segment bytes this process lifetime, every segment write returns a real ENOSPC (std::io::Error::from_raw_os_error(28) — the errno on Linux AND macOS, wrapped in the same WalError::Io a genuine full disk produces) at segment::write_batch_bytes.

Both are behind the non-default fault-injection cargo feature (tidaldb + a tidal-server passthrough). Production builds (docker/standalone/Dockerfile: cargo build -p tidal-server --release --locked) never pass it, so the hooks are compiled out entirely — a disk-write or fsync fault a stray env var could trip in production is exactly the 3am footgun we refuse to ship, so the safety is structural, not a default. Even compiled in, faults are inert until an env var arms them (the posture of TIDAL_HLC_SKEW_MS / the kill-point knobs). The MultiProcCluster harness's tidal_server_bin() builds the spawned binary with the feature; every non-fault suite spawns the same binary and is unaffected.

The injection returns ENOSPC before the partial write_all (the standard clean-fault model). The follower apply path already handles a WAL write error correctly (proven by the suite below): the receiver halts (degraded, alive), NEVER advances its applied frontier past the fault (no gap-swallowing), and a restart re-pulls the suffix via the catch-up stream.

2. First-class invariant checkers (tidal-server/tests/support/invariants.rs)

The roadmap's named invariants were inline copies in three suites; m11p9 promotes them to one audited module every suite consumes:

  • AckLedger — no acknowledged-write loss: records every write the client saw a 2xx + x-tidal-seq for, then proves (A) no acked seqno exceeds the max-applied survivor's contiguous frontier, and (B) every acked item is present on the promoted/recovered leader via /search. Extracted verbatim from the m11p3 ledger gate, which now CONSUMES it (re-verified: 3/3 kill-points zero acked loss after the refactor).
  • assert_feed_parity / feed_pairs — cross-replica decay/score parity to 1e-6 between continuously-running replicas.
  • feed_item_ids — cross-replica DATA parity (the materialized item set), robust across a node RESTART where a velocity/rate score legitimately differs (its time-bucketed windowed counts are rebuilt from a burst WAL-replay rather than continuous accumulation, even with identical durable data — see the disk-full test).
  • assert_single_leader_per_term — membership safety: at most one leader per (shard, term) across a set of status snapshots (reads the m11p6 shards[] rows or the flat fields). A second leader at the same term is a split brain.
  • MonotonicCounters — per-entity monotonic counters: a node's applied frontier and the leader's commit index never move backward across observations (legitimate epoch resets — a reseeding node, a leader change — are forgotten, not flagged), so a regression means durably-acked state was lost.

3. New fault-class suite (tidal-server/tests/cluster_faults.rs)

Four tier-3 tests over real OS processes, asserting through the checkers:

  • mp_disk_full_follower_degrades_no_acked_loss — a follower hits ENOSPC mid-replication, its receiver halts (degraded, alive), the healthy majority keeps acking quorum writes, NO acked write is lost, and a restart with space recovers it to full content + item-set parity.
  • mp_slow_fsync_follower_lags_but_quorum_holds — a slow disk lags one follower; the fast follower supplies quorum so every write commits; the slow node converges once the burst ends; no loss.
  • mp_slow_fsync_both_followers_force_honest_quorum_timeout — neither follower can confirm inside the budget ⇒ ack=quorum returns a retryable 503 naming the laggards (never a false 2xx) while ack=leader is unaffected; the followers recover.
  • mp_asymmetric_partition_no_split_brain_no_loss — a follower that loses INBOUND links (can still send) cannot disrupt the cluster: pre-vote + check-quorum keep the standing leader, there is never a second leader at the same term, no acked write is lost, and it rejoins cleanly on heal.

The asymmetric partition reuses the harness's directed-edge proxies (proxied_rewrite(&["ap-south"]) + region("ap-south").sever_all() severs only inbound to ap-south; its outbound dials are unproxied, so it keeps sending the disruptive RequestVotes pre-vote neutralizes) — no new mechanism, the existing PartitionProxy was already directional.

4. Soak + regression gates (tidal-stress)

tidal-stress gained, for the nightly soak:

  • --json-summary <path> — a flat, machine-readable per-stage roll-up (p99 / throughput / error-rate / client-shed + the gate verdict), hand-rolled (no serializer dep, matching the metrics module's posture) for trend-line archival.
  • --max-p99-ms, --max-error-pct, --fail-on-knee — PASS/FAIL regression gates. A breach returns StressError::Gate → non-zero exit (the tool always exited 0 before, so a regression scrolled past in green). Unarmed = the original informational behavior, byte-for-byte.

5. Nightly pipeline (.woodpecker.yaml) — Woodpecker, never GitHub Actions

The pipeline now carries two event-routed flows (workflow-level when: event: [push, cron], per-step when):

  • push — the m11p8 rolling-upgrade release gate, then the Kaniko image build (unchanged).
  • cron nightlynightly-chaos runs every tier-3 suite serially with elevated kill-points (TIDAL_QUORUM_KILLPOINTS=25, TIDAL_ELECTION_KILLPOINTS=15) and the fault-injection feature (so cluster_faults runs); nightly-soak boots a target and runs tidal-stress with the regression gates, archiving the JSON summary. Point TIDAL_SOAK_TARGET at the live Ref-A cluster and raise TIDAL_SOAK_SECS=3600 for the GA-bar 1-hour 100k-DAU soak.

6. Guarantee traceability

guarantee-traceability.md maps every §2 guarantee (G-D, G-A, G-S, G-E, G-Sec, G-O, G-Op, G-C) to its named automated test(s), the fault each injects, and the checker each asserts through.

Exit gate (from the roadmap)

  • Every guarantee in §2 maps to a named automated test.
  • Nightly suite green 30 consecutive days before GA.

How the streak is computed (and what resets it)

The bar is NOT "30 green soak Job exits". It is 30 consecutive nights where the soak passed its SLO gates AND no cluster pod restarted under load. The two halves come from two streams on the shared result PVC:

  • ledger.tsv — one row per night from the nightly CronJob: PASS/FAIL on the armed gates (--fail-on-knee, --max-p99-ms 150, --max-error-pct 1).
  • restarts.tsv — the monitor's 5-minute snapshot of each tidaldb-{0,1,2} pod's cumulative restartCount.

soak-eval (the proven core in tidal_stress::soak_eval, table-driven-tested; the binary runs in the monitor's evaluator container and once per night in the CronJob) joins them into streak.tsv:

  • a night is GREEN iff its ledger row is PASS and no pod's cumulative restart count rose within that night's window;
  • the streak is the trailing run of consecutive green nights;
  • any non-green night resets the streak to 0. A night resets it when EITHER the soak breached an SLO gate (Job FAIL) OR a pod restarted under load (the zero-unrecovered-restart half — a green ledger row with a restart in-window is a FAIL, not a PASS: that was the false-confidence hole).

The evaluator emits exactly one alert per non-green transition (a durable ALERT-<date>.txt the http surface serves, plus an optional NOTIFY_WEBHOOK POST). Dependency: the streak is only reachable once the graceful-restart reseed-on-rejoin churn is fixed — before that, every rollout/upgrade/node reboot increments a restart count and would reset the streak forever. The production- readiness divergence fix (tasks 01/02) is what makes a restart-in-window a real, rare anomaly worth failing on.

Exit-gate evidence (local; debug builds, real OS processes)

Gate Target Measured
New fault classes pass disk-full / slow-fsync / asymmetric cluster_faults 4/4 green (73s): disk-full follower froze at applied 71 while node1 reached 240, recovered after restart to 120-item parity, zero acked loss; slow follower lagged 15 vs 0 then converged; both-slow → honest 503 naming ["eu-west","ap-south"] while ack=leader 204'd; asymmetric → leader held us-east, single-leader-per-term, ap-south term bounded (0→0), rejoined.
First-class checkers faithful extraction = no regression cluster_quorum::mp_quorum_ledger_zero_acked_loss_across_killpoints refactored onto AckLedger: 3/3 kill-points zero acked loss (29s) — extraction verified against the live gate.
Soak regression gate works PASS exits 0, breach exits ≠0 Against a real standalone server: a within-threshold run exited 0 with a valid JSON summary (p99 39.7ms, 0% error); an impossible --max-p99-ms 0.01 run exited 1 with the breach in stderr and "passed": false in the JSON.
Every guarantee → named test §2 coverage guarantee-traceability.md: G-D/G-A/G-E/G-Sec/G-Op each map to a green named test; G-S to the m11p6 owner-test (in progress); G-O to the m11p8 metric/dashboard/alert artifacts; G-C to this apparatus.
Nightly green 30 days calendar Pipeline established, not yet accrued — this is a wall-clock criterion the cron pipeline produces over 30 nights; it cannot be completed in one session. Honestly out-of-session (like the Ref-A k3s runs).

Status

  • fault-injection feature (slow-fsync + disk-full WAL hooks), compiled out of production
  • First-class invariant checkers (support/invariants.rs): ledger, decay parity, single-leader-per-term, monotonic counters
  • cluster_quorum refactored onto AckLedger (extraction re-verified)
  • cluster_faults.rs: disk-full, slow-fsync (×2), asymmetric partition — 4/4 green
  • tidal-stress --json-summary + --max-p99-ms / --max-error-pct / --fail-on-knee + non-zero exit (verified vs a real server)
  • Woodpecker nightly cron pipeline (chaos suites + gated soak) beside the push release gate
  • Guarantee → automated-test traceability matrix
  • Docs (this phase, runbook chaos/soak section, monitoring soak note) + CHANGELOG + ROADMAP + roadmap-to-cluster as-built
  • Calendar: nightly green for 30 consecutive days before GA (the pipeline accrues this; not an in-session deliverable)

Adversarial review (6-dimension, read-only)

A read-only review fanned out over six dimensions (production-safety, checkers, fault-tests, soak-gates, Woodpecker, doc-honesty), each finding adversarially verified before adjudication. 10 confirmed findings, all fixed:

  • Woodpecker (blocker): the soak step referenced environment:-block vars as ${VAR}, which Woodpecker's preprocessor blanks BEFORE the shell — the gate would have run with an empty ramp/thresholds (never executing). Fixed with the $${VAR} escape so the container shell expands them at runtime.
  • Gate threshold (blocker): a --max-p99-ms nan/inf reported the gate ARMED while every > NaN comparison is false — silently passing a regression. Now rejected (finite, non-negative) at the CLI boundary.
  • JSON validity (blocker): a non-finite target_rps (--ramp "inf:30") emitted bare inf/NaN, invalid JSON. Fixed at the parse boundary + a to_json null-sanitizer (defense in depth).
  • Checkers (warnings): assert_single_leader_per_term false-failed on a duplicate same-region snapshot (assert_ne! → gated if prev != region); MonotonicCounters now re-baselines a region's role-relative applied frontier on a leader↔follower flip (not just reseed); assert_items_present distinguishes a transport/non-2xx probe failure from true absence (no false ACKNOWLEDGED LOSS).
  • Coverage (warning): added a summary.rs test module (gate breach/pass, fail-on-knee, JSON round-trip through a real parser, null-sanitize) — 6 tests.
  • Doc honesty (warnings): the G-Sec/G-Op owner-tests (mTLS, cluster_security, tidalctl backup/restore, WAL-archival) were NOT in the nightly — ADDED a nightly-security-ops cron step so the "runs nightly" claim is true; the G-S row is marked "(NOT YET WRITTEN — m11p6 L4)" so it doesn't read as delivered.
  • Cron name (nit, external): the nightly is silently-green until a cron named nightly is created in Woodpecker repo settings — documented in the pipeline header; a one-time repo-settings step.

After the fixes the fault suite re-ran 4/4 green and the ledger gate 3/3 (the checker changes are faithful); the slow-fsync and asymmetric tests moved to item-SET parity (a burst catch-up reconstructs velocity differently — the same restart-robustness rationale as disk-full).

Formal seven-dimension review (/review-codefix-all)

A second pass ran the code-reviewer seven-dimension protocol (each dimension a distinct frame, each finding adversarially verified): 0 blockers, 0 critical, 7 warnings, 12 suggestions (APPROVE-with-fixes). All 7 warnings + 10 of 12 suggestions fixed; final score 90/100. Notable fixes: a real diagnostic bug (assert_items_present's two guards were swapped — a true acknowledged-loss printed "PROBE UNREACHABLE" and vice-versa); StressError::Gate was overloaded across breach / invalid-threshold / IO-write and split into Gate/BadGate/ Summary so the operator-facing message names the real cause; the soak port moved off iknowyou's 59521 to a TIDAL_SOAK_PORT (59526); the soak step now fails fast on a boot failure and reaps its background server; doc-truthing the "every suite consumes the shared checkers" claim (only cluster_quorum + cluster_faults do — the legacy suites' migration is a tracked follow-up); S=1-scope notes on the flat-field checkers; a real from_stage unit test; and the observe_cluster_frontiers double-status-fetch removed. Two suggestions deferred with justification: consolidating the role-vs-is_leader leadership read into one helper (steady-state-equal, >5min, regression risk on green gates) and Woodpecker YAML anchors for the repeated apt-get/cargo lines (untestable CI refactor, benign duplication).

Verification status

  • Workspace cargo fmt --all -- --check: clean.
  • cargo clippy -D warnings: clean on tidaldb (fault-injection on + off), tidal-server (cluster-e2e fault-injection), and tidal-stress (all targets).
  • Production safety: cargo check -p tidaldb with the feature OFF compiles the fault module and both WAL hooks out entirely (the shipped binary is unchanged).
  • The tree is UNCOMMITTED (continues the m11p1p8 uncommitted tree; the user commits).