# m12p5 — Idle-readiness fix + elasticity under load (T4) Status as of 2026-06-14. Closes the **idle-readiness stall** and the **cert SAN scale gap**; the 1M/1536-dim T4 run on k3s remains the project's standing Ref-A/k3s dependency (the machinery and the fix that unblocks it are proven locally over real OS processes). ## 1. The idle-readiness stall (WORKLOG 2026-06-13: an 11.5h hang) A snapshot-installed joiner (`install_boot` / `seed_joiner`) serves `503` on its readiness probe (`/health` → `region_health` → `is_ready`) until its sticky `converged` latch flips. Pre-m12p5 the ONLY thing that flips that latch is `note_lag_for_readiness`, and it is called from exactly one place: `local_status` — i.e. when something hits the node's `/cluster/status/local` (an operator/monitoring poll, or the leader's `/cluster/status` aggregator querying it). On a cluster with write traffic this is masked: ship traffic keeps the lag gauge fresh and monitoring keeps poking status. On a **fully idle** cluster (no writes, no status polls) neither happens, so a freshly caught-up joiner sits `503` forever and never joins the Service VIP. The k8s `readinessProbe` is `/health` (`statefulset.yaml`), so the pod never receives traffic — an 11.5h observed stall. Root-cause specifics (read on `44b768b`): - `node.rs::note_lag_for_readiness` sets `converged` only when a lag value ≤ `learner_promote_lag` is *recomputed* — there is no periodic self-driven check. - `region_health` (the readiness handler) reads `is_ready()` but never recomputes lag, so polling readiness does not advance convergence. - The lag gauge (`leader_seqno_for(shard)`) is seeded only by **received ship segments**; on idle nothing ships, so the gauge would even read a stale/0 value — converging on it directly is unsafe (a behind-but-unshipped joiner would read lag 0 from an uninitialized gauge and false-converge). ## 2. The fix — converge from the heartbeat The leader **heartbeat** flows every heartbeat interval regardless of write traffic and already proves "the leader is alive at term T". m12p5 makes it also carry the leader's **live flushed frontier**: - **Proto** (`wal_shipping.proto`): new `HeartbeatRequest.leader_last_seq = 14` — the leader's `ship_feed.flushed_seq()` at heartbeat time (same stream numbering as a follower's per-shard `applied_seqno`). proto3 zero-default `0` = a pre-m12p5 leader → the follower falls back to the status-poll path (no behavior change). - **Leader** (`election_heartbeat`): stamps `leader_last_seq` on every heartbeat. - **Follower** (`election_driver::on_heartbeat` → `node::note_leader_frontier_for_readiness`): on every ACCEPTED heartbeat, folds `leader_last_seq` into the lag gauge (monotonic — keeps the `lag_segments` metric/`local_status` truthful on idle for every follower) and, for a joiner that is not yet converged, computes `lag = leader_last_seq − applied_seqno(leader_shard)` and drives the existing `note_lag_for_readiness`. The convergence uses a **real** leader frontier (never the uninitialized-0 gauge), so a still-behind joiner stays `503` until it actually catches up. - **Leader-elect edge** (`become_leader_for_term`): a joiner that WINS leadership receives no heartbeats, so it converges its latch on activation — by construction it is caught up to its own log. Without this a promoted-then-elected joiner would stay `503` forever on an idle cluster. Net: a caught-up joiner converges within a heartbeat interval (~100ms here) of catching up, with no write traffic and no status poll. ## 3. The cert SAN scale gap `k8s/cluster/certs.yaml` (and `scripts/gen-cluster-certs.sh`) enumerated SANs for `tidaldb-0/1/2` only. Scaling the StatefulSet to 5 (T4) gives `tidaldb-3/4` no matching SAN, so the inter-node mTLS handshake to the new pods fails. Fix: a **wildcard pod SAN** `*.tidaldb-peers.tidaldb-cluster.svc.cluster.local` covers every ordinal, so scale-up/down needs no cert re-issue. The explicit `tidaldb-0/1/2` names and the headless/client Service names are retained (belt-and-suspenders for any strict verifier). tidalDB dials peers over tonic → rustls → webpki, which matches a wildcard against the single leftmost DNS label per RFC 6125 — exactly the pod-ordinal label. Verified for real (openssl leaf generated by the updated script): ``` $ openssl verify -CAfile ca.crt \ -verify_hostname tidaldb-7.tidaldb-peers.tidaldb-cluster.svc.cluster.local tls.crt tls.crt: OK # tidaldb-7 is NOT explicitly listed — matched by the wildcard ``` ## 4. Verification ### Idle-readiness regression test (real OS processes, tier-3) `tidal-server/tests/cluster_membership.rs::mp_idle_cluster_snapshot_joiner_flips_ready_without_traffic`: 1. 3-node elected cluster; heavy seed → graceful leader restart → WAL compaction past seq 1, so the later joiner takes the **snapshot-install** path (`install_boot`/`seed_joiner` true → readiness IS gated on `converged`). A small `needed=false` join boots a voter and never engages the gate, so it cannot reproduce the stall — this setup is load-bearing. 2. Go **fully idle**, then `add_node` (which returns on `/health/startup`, an unconditional 200 — NOT on cluster-readiness). 3. **Gate**: poll ONLY the joiner's `/health` (never `/cluster/status/local`, which would drive the old latch and mask the bug). With zero writes the joiner flips ready in **<101ms** (one measured run: 257µs; another: 101ms). 4. **Honesty**: head/tail items are then searchable on the joiner — `converged ⟹ caught up`. **Negative control (proves it is a real gate, not wiring).** With the heartbeat-convergence call (`note_leader_frontier_for_readiness`) commented out, the identical test 503s for the full 30s convergence budget and fails — exactly the pre-m12p5 stall. Re-enabling the call makes it pass in <101ms. ### Suites green (2026-06-14, local) - `cluster_membership` (tier-3, real processes): **6/6** — includes the new idle test, `mp_seed_join_snapshot_catchup`, and the T4 `mp_scale_3_5_3_under_load_zero_loss`. - `tidal-net` unit/integration (proto roundtrip + election RPC incl. the new `leader_last_seq` on the wire): green. - `tidaldb` lib: 1903 passed. `clippy -D warnings` clean across `tidaldb`, `tidal-net`, `tidal-server` (incl. `--features cluster-e2e --all-targets`). ## 5. What remains — T4 at 1M/1536-dim on k3s The exit gate's full form ("joiner reaches lag=0 ≤ 5 min on a 1M/1536D corpus; p99 impact < 2× baseline for < 60s; zero loss") requires the production-shape corpus, which needs multi-node k3s/Ref-A. This is the SAME standing dependency called out for m12p1-p4: the local kubeconfig cannot reach Ref-A. What is proven locally: - the 3→5→3 scale-up machinery, zero acked loss, quorum on the grown set (`mp_scale_3_5_3_under_load_zero_loss`); - the idle-readiness fix that lets an idle scale-up actually join the VIP — the specific blocker that would have stalled a real k3s scale-up. When Ref-A is reachable: deploy `k8s/cluster/` (now wildcard-SAN), seed 1M×1536-D, `kubectl scale statefulset tidaldb --replicas=5` under the tidal-stress load, record joiner-lag-to-0 and the p99 envelope, then scale back to 3.