# m12p5 — Idle-readiness fix + elasticity under load (T4) Status as of 2026-06-14. Closes the **idle-readiness stall** and the **cert SAN scale gap**; the 1M/1536-dim T4 run on k3s remains the project's standing Ref-A/k3s dependency (the machinery and the fix that unblocks it are proven locally over real OS processes). ## 1. The idle-readiness stall (WORKLOG 2026-06-13: an 11.5h hang) A snapshot-installed joiner (`install_boot` / `seed_joiner`) serves `503` on its readiness probe (`/health` → `region_health` → `is_ready`) until its sticky `converged` latch flips. Pre-m12p5 the ONLY thing that flips that latch is `note_lag_for_readiness`, and it is called from exactly one place: `local_status` — i.e. when something hits the node's `/cluster/status/local` (an operator/monitoring poll, or the leader's `/cluster/status` aggregator querying it). On a cluster with write traffic this is masked: ship traffic keeps the lag gauge fresh and monitoring keeps poking status. On a **fully idle** cluster (no writes, no status polls) neither happens, so a freshly caught-up joiner sits `503` forever and never joins the Service VIP. The k8s `readinessProbe` is `/health` (`statefulset.yaml`), so the pod never receives traffic — an 11.5h observed stall. Root-cause specifics (read on `44b768b`): - `node.rs::note_lag_for_readiness` sets `converged` only when a lag value ≤ `learner_promote_lag` is *recomputed* — there is no periodic self-driven check. - `region_health` (the readiness handler) reads `is_ready()` but never recomputes lag, so polling readiness does not advance convergence. - The lag gauge (`leader_seqno_for(shard)`) is seeded only by **received ship segments**; on idle nothing ships, so the gauge would even read a stale/0 value — converging on it directly is unsafe (a behind-but-unshipped joiner would read lag 0 from an uninitialized gauge and false-converge). ## 2. The fix — converge from the heartbeat The leader **heartbeat** flows every heartbeat interval regardless of write traffic and already proves "the leader is alive at term T". m12p5 makes it also carry the leader's **live flushed frontier**: - **Proto** (`wal_shipping.proto`): new `HeartbeatRequest.leader_last_seq = 14` — the leader's `ship_feed.flushed_seq()` at heartbeat time (same stream numbering as a follower's per-shard `applied_seqno`). proto3 zero-default `0` = a pre-m12p5 leader → the follower falls back to the status-poll path (no behavior change). - **Leader** (`election_heartbeat`): stamps `leader_last_seq` on every heartbeat. - **Follower** (`election_driver::on_heartbeat` → `node::note_leader_frontier_for_readiness`): on every ACCEPTED heartbeat, folds `leader_last_seq` into the lag gauge (monotonic — keeps the `lag_segments` metric/`local_status` truthful on idle for every follower) and, for a joiner that is not yet converged, computes `lag = leader_last_seq − applied_seqno(leader_shard)` and drives the existing `note_lag_for_readiness`. The convergence uses a **real** leader frontier (never the uninitialized-0 gauge), so a still-behind joiner stays `503` until it actually catches up. - **Leader-elect edge** (`become_leader_for_term`): a joiner that WINS leadership receives no heartbeats, so it converges its latch on activation — by construction it is caught up to its own log. Without this a promoted-then-elected joiner would stay `503` forever on an idle cluster. Net: a caught-up joiner converges within a heartbeat interval (~100ms here) of catching up, with no write traffic and no status poll. ## 3. The cert SAN scale gap `k8s/cluster/certs.yaml` (and `scripts/gen-cluster-certs.sh`) enumerated SANs for `tidaldb-0/1/2` only. Scaling the StatefulSet to 5 (T4) gives `tidaldb-3/4` no matching SAN, so the inter-node mTLS handshake to the new pods fails. Fix: a **wildcard pod SAN** `*.tidaldb-peers.tidaldb-cluster.svc.cluster.local` covers every ordinal, so scale-up/down needs no cert re-issue. The explicit `tidaldb-0/1/2` names and the headless/client Service names are retained (belt-and-suspenders for any strict verifier). tidalDB dials peers over tonic → rustls → webpki, which matches a wildcard against the single leftmost DNS label per RFC 6125 — exactly the pod-ordinal label. Verified for real (openssl leaf generated by the updated script): ``` $ openssl verify -CAfile ca.crt \ -verify_hostname tidaldb-7.tidaldb-peers.tidaldb-cluster.svc.cluster.local tls.crt tls.crt: OK # tidaldb-7 is NOT explicitly listed — matched by the wildcard ``` ## 4. Verification ### Idle-readiness regression test (real OS processes, tier-3) `tidal-server/tests/cluster_membership.rs::mp_idle_cluster_snapshot_joiner_flips_ready_without_traffic`: 1. 3-node elected cluster; heavy seed → graceful leader restart → WAL compaction past seq 1, so the later joiner takes the **snapshot-install** path (`install_boot`/`seed_joiner` true → readiness IS gated on `converged`). A small `needed=false` join boots a voter and never engages the gate, so it cannot reproduce the stall — this setup is load-bearing. 2. Go **fully idle**, then `add_node` (which returns on `/health/startup`, an unconditional 200 — NOT on cluster-readiness). 3. **Gate**: poll ONLY the joiner's `/health` (never `/cluster/status/local`, which would drive the old latch and mask the bug). With zero writes the joiner flips ready in **<101ms** (one measured run: 257µs; another: 101ms). 4. **Honesty**: head/tail items are then searchable on the joiner — `converged ⟹ caught up`. **Negative control (proves it is a real gate, not wiring).** With the heartbeat-convergence call (`note_leader_frontier_for_readiness`) commented out, the identical test 503s for the full 30s convergence budget and fails — exactly the pre-m12p5 stall. Re-enabling the call makes it pass in <101ms. ### Suites green (2026-06-14, local) - `cluster_membership` (tier-3, real processes): **6/6** — includes the new idle test, `mp_seed_join_snapshot_catchup`, and the T4 `mp_scale_3_5_3_under_load_zero_loss`. - `tidal-net` unit/integration (proto roundtrip + election RPC incl. the new `leader_last_seq` on the wire): green. - `tidaldb` lib: 1903 passed. `clippy -D warnings` clean across `tidaldb`, `tidal-net`, `tidal-server` (incl. `--features cluster-e2e --all-targets`). ## 5. What remains — T4 at 1M/1536-dim on k3s The exit gate's full form ("joiner reaches lag=0 ≤ 5 min on a 1M/1536D corpus; p99 impact < 2× baseline for < 60s; zero loss") requires the production-shape corpus, which needs multi-node k3s/Ref-A. This is the SAME standing dependency called out for m12p1-p4: the local kubeconfig cannot reach Ref-A. What is proven locally: - the 3→5→3 scale-up machinery, zero acked loss, quorum on the grown set (`mp_scale_3_5_3_under_load_zero_loss`); - the idle-readiness fix that lets an idle scale-up actually join the VIP — the specific blocker that would have stalled a real k3s scale-up. When Ref-A is reachable: deploy `k8s/cluster/` (now wildcard-SAN), seed 1M×1536-D, `kubectl scale statefulset tidaldb --replicas=5` under the tidal-stress load, record joiner-lag-to-0 and the p99 envelope, then scale back to 3. ## 6. Real-k8s T4 run on `kind` — and the seed-join-over-TLS bug chain it exposed The idle-readiness fix was first proven over real OS processes (§4) but those run **plaintext** inter-node. Running an actual `kubectl scale` on a real (mTLS) k8s cluster — `kind-canopy`, the cluster m12p4 used — exercised the seed-join path over TLS **for the first time**, and surfaced a chain of real bugs that no in-process test could have caught. All are fixed; the run is GREEN. Overlay: `k8s/cluster-t4-kind/` (single replication group — T4 is replica elasticity, not sharding; the proper-CA `certs.yaml`; kind `standard` StorageClass; 10-min startup budget for the 1536-dim index rebuild). Seed/load Jobs: `tidal-stress/k8s/t4-{seed,load}-job.yaml`. ### The bug chain (every one TLS-only; the plaintext in-process harness is blind to them) 1. **Seed dial scheme** (`statefulset.yaml`) — the scale-up pod dialed `--seed "http://…:9500"`, but the m11p7 `:9500` plane serves TLS and `forward::peer_url` honors an explicit scheme verbatim, so the joiner spoke plaintext to a TLS port → "could not join via any seed within 120s". Fixed to `https://`. 2. **rustls `CryptoProvider` install order** (`tidal-server/main.rs` + `tidal-net` `ensure_crypto_provider` made `pub`) — the seed-join / reseed boot thread builds a blocking `reqwest` (rustls) client *before* the gRPC transport installs the process default provider, so the client **panicked** ("could not automatically determine the process-level CryptoProvider"). Now installed at the top of `main()`. 3. **Discovery target** (`statefulset.yaml`) — `--seed` pointed at the **headless** `tidaldb-peers` Service (`publishNotReadyAddresses: true`), which resolves to *every* pod including the still-joining joiner itself → discovery round-robined onto not-ready pods. Re-pointed at the **ready-only** client Service `tidaldb`. 4. **Discovery poll timeout** (`join_boot.rs`) — `STATUS_POLL_TIMEOUT` was 500 ms, too tight for a cold TLS handshake under CPU contention; every poll timed out silently. Raised to 5 s (env-overridable `TIDAL_SEED_STATUS_TIMEOUT_MS`) and the discovery loop now logs each failure at WARN with the full error source chain (it was previously swallowed — a silent loop is what made bugs 1–6 each present as the same opaque 120-s timeout). 5. **cert-manager PKI shape** (`certs.yaml`) — the leaf was issued DIRECTLY from a `selfSigned` Issuer, producing a self-signed end-entity cert (`CA:FALSE`) whose `ca.crt` is a copy of the leaf. Lenient stacks tolerated it, but the joiner's strict webpki verifier rejected the peer cert as **`UnknownIssuer`**. Replaced with a real two-tier PKI: a `selfSigned` Issuer → a CA **certificate** (`CA:TRUE`) → a `ca:` Issuer that signs the leaf. (`scripts/gen-cluster-certs.sh` already did this correctly — the two were merely inconsistent.) 6. **`grpc_tls` for a node not in the topology** (`join_boot.rs`) — **the actual `UnknownIssuer` root cause.** `own_grpc_tls`/`self_tls_spec` looked up the joiner's OWN region in the knob file to find its TLS material — but a seed-joiner is NEVER in the shared-ConfigMap `regions:` list, so it got `None` → the seed client was built with **no CA** and the synthesized topology was plaintext. Fixed with `grpc_tls_for`: fall back to ANY region's block (every pod mounts the same cert files at the same paths). ### Result (GREEN) `kubectl scale statefulset tidaldb --replicas=5` on an **idle**, 1536-dim-seeded cluster: both new pods seed-join over mTLS (wildcard SAN covers `tidaldb-3/4`), catch up, and flip `/health` **Ready in 13 s** — driven by the idle-readiness heartbeat convergence (no writes, no status poll). Both joiners auto-promote to Voter and show **full content parity** (honest convergence); `/cluster/status` reports all five regions `lag=0, reachable`. Pre-m12p5 this would have stalled at 503 indefinitely (the 11.5-h WORKLOG hang). Zero acked loss held across a subsequent under-load scale-down (every probed seeded item present). The clean continuous-quorum decommission (roster shrinks via the remove verb so quorum follows the smaller set) is the in-process T4 (`mp_scale_3_5_3_under_load_zero_loss`); a StatefulSet scale-down *without* the remove verb keeps a 3-of-5 roster and is momentarily quorum-fragile by design. ### One bug found and root-caused but NOT fixed (deliberately) At a corpus large enough to **compact the leader's WAL below the snapshot point** (≥ ~10k×1536 here), a seed-joiner installs a snapshot (`snapshot_seq` correct, fix #2/#6 made the gRPC snapshot fetch work) but then `node.rs:734` seeds the post-install catch-up frontier from `db.last_wal_seq()` — which is **0** for a STATE-ONLY artifact (a compacted leader ships no WAL). The joiner then requests catch-up `from_seqno=1`, the compacted source refuses ("WAL compacted below seqno 1"), and it latches a reseed marker and stalls degraded. The sentinel already carries the correct `snapshot_seq`, but a naïve "use `snapshot_seq`" fix **broke `mp_seed_join_snapshot_catchup`** — the frontier is own-WAL numbering with a `stream_baseline` translation, not stream numbering, so the real fix lives in the own-WAL↔stream mapping for a state-only install. That is durable replication-frontier machinery: shipping it blind on a remote cluster risks data loss, so it is left reverted and tracked as a follow-up that must FIRST extend the in-process snapshot test to the deep-compaction case. The GREEN run above uses a small corpus (stream catch-up, no snapshot install) to keep that path out of scope. The full 1M/1536-dim gate remains the standing Ref-A/k3s dependency.