tidaldb/docs/planning/milestone-11/phase-8.md
jx12n 005e292cbb fix(m11): review remediation + tidal-stress perf sweep + perf wave 2
Resolve all BLOCKER/CRITICAL/WARNING findings from the m11p7/p8 review:
- tidalctl restore: safe_join path-traversal/Zip-Slip guard + fsync on write
- corrupt-WAL checkpoint_seq guard; PITR archive-before-delete
- cluster: x-tidal-relayed audit-dedup marker; forward_failures counts 5xx
- mTLS/HTTP-TLS handshake hardening; accept-loop EMFILE backoff
- per-principal rate-limit + node-token marker-pinning tests
- self-heal tier-3 coverage; 5 router-auth tests

tidal-stress: measurement-fidelity fixes (schedule-lag p99/max, exact
feed-over-SLO verdict, shed annotation) + typed Body, workload.next
184ns->68ns, RoundRobin len==1 short-circuit, HeaderValue cache;
new benches/hotpath.rs + lib.rs.

perf wave 2: signal_snapshot SmallVec/SignalKey carrier; one-get-per-type
ranking pre-pass.
2026-06-13 12:28:04 -06:00

13 KiB
Raw Blame History

m11p8 — Observability + Operations (COMPLETE — 2026-06-13)

Phase spec and exit gate: docs/roadmap-to-cluster.md §4/m11p8. Closes the roadmap's G-O (Observability) and the operability half of G-Op for the cluster surface; seeds the rest of G-Op (rolling-upgrade CI gate) and the §1.4-3 incident ("the breaker eats the first heal"). Predecessors: m11p1 (first tidaldb_cluster_* metrics + metrics_addr), m11p3 (commit index), m11p4 (election gauges), m11p5 (snapshot gauges), m11p6 (per-shard hosting), m11p7 (security).

Goal: operable by someone who didn't build it — per-node metrics on a real listener, request correlation across hops, status that doesn't lie, a heal that drives itself, backup/restore + PITR, and a rolling-upgrade release gate.

Design (as adopted)

m11p8 is six sub-items. Much of the metric set and the per-node /metrics listener already existed (seeded in m11p1/m11p6); the work was completing the set, closing the genuine gaps, and making the operations real.

1. Metrics + the cluster /metrics listener (A)

The per-node listener already binds (the engine's metrics HTTP server, per the topology metrics_addr). The metric set gained the two members the spec named that were missing — breaker and forwards — plus the self-heal series:

  • tidaldb_cluster_forwards_total / tidaldb_cluster_forward_failures_total (cross-node write forwards initiated by a gateway; instrumented in forward_write and forward_to_group_node).
  • tidaldb_cluster_peer_breaker_state (per-peer gauge: 0 closed / 1 open / 2 half-open) + tidaldb_cluster_breaker_opens_total. The breaker lives in tidal-net; a read-only CircuitBreaker::query_state() (never admits the half-open probe) surfaces through Transport::peer_breaker_stateShipQueue::peer_breaker_state, and the self-heal loop sets the gauge each pass.
  • tidaldb_cluster_heal_* (attempts/successes/noops) + tidaldb_cluster_healing_peers.

Multi-shard listener (m11p6 co-location). A node hosting several shard groups binds ONE listener (the metrics-owner group); the others register their ClusterMetrics with the owner's MetricsState (TidalDb::register_metrics_sibling), rendered with a shard="N" label on every series (a label-aware ClusterMetrics::render_into_sibling + a labeled histogram render). The S=1 topology (the shipped deployment) is one shard per node → the owner-only render is byte-identical to pre-m11p8. Co-located shards never collide on the shared scrape target.

Dashboard + alerts. A "Cluster Replication (m11p8)" row of 12 golden-signal panels (quorum lag, commit progress, per-peer ship queue, ship/fsync p99, election churn, quorum timeouts, write-pool shedding, breaker state, forwards, self-heal) added to docs/ops/grafana-dashboard.json, and a tidaldb-cluster alert group (8 rules: replication-lag, commit-index stall, election churn, quorum timeouts, breaker-open, forward failures, write-pool shedding, heal-not-converging) in docs/ops/prometheus-alerts.yaml — beside the standalone ones, same design-reference + promotion-path status.

2. Request-id propagation + tracing across hops (B)

The standalone router's SetRequestId + PropagateRequestId + TraceLayer stack was extracted to crate::router::with_request_id_tracing and applied to BOTH cluster routers (single-process build_cluster_router, multi-process build_region_router) — they previously skipped it. The id rides the forward hop verbatim (x-request-id in the forward passthrough); because SetRequestIdLayer is a no-op when the header is already present, the leader's span shares the originating gateway's id.

Ship hop, honestly. Ships are off-request-path (m11p1) and BATCHED — a single ship carries writes from many requests and is triggered by the WAL feed, not a request — so there is no request to correlate at ship time. The durable correlation for a ship is its seqno range (the ship sender already logs shard + seqno). The version handshake (below) rides the heartbeat. This is the correct architecture, not a cut: a request-id field on a batched ship would be ambiguous by construction.

3. Truthful status (C)

Two distinct, confirmed bugs:

  • Leader's own applied row read 0 (multi-process local_status). The applied frontier (replication_state().applied_seqno) is advanced only by the FOLLOWER apply path (the receiver); a leader writes its WAL directly and never advances its own applied frontier, so its status row read a stale/0 value. Fix: for a leader, report its durable flushed frontier (ship_feed.flushed_seq()) — it is, by construction, applied to its own log.
  • Post-promote ShardId(0) keying (single-process applied_count). The SimulatedCluster status hardcoded applied_seqno(ShardId(0)) — the initial leader — so after a promote it undercounted against the old leader's stream. Fix: key on the current leader's shard (ShardId(leader_region.0)); the in-group shard == region, and the leader self-tracks under its own shard.

(The explorer-proposed self.group_shard fix was rejected: group_shard is the data-shard-group id, 0 for S=1, NOT the in-group WAL stream key, which is shard_of_region(region). Keying on it would have been wrong.)

Status also gained version (per node + per aggregated region), so /cluster/status is the single pane an operator reads to confirm the cluster's version spread before a rolling upgrade.

4. Self-driving heal (D) — closes incident §1.4-3

A STANDING LEADER DUTY (ShardReplica::tick_self_heal) re-armed every ~3 s (throttled inside the 50 ms election tick; non-blocking, runs inline). Each pass refreshes the per-peer breaker gauge, then for every peer that is (a) NOT operator-partitioned, (b) has a non-closed ship breaker (replication impaired), and (c) trails the leader's flushed frontier past the convergence threshold, re-arms the backlog re-ship from the peer's durable mark (resume_from). The moment the breaker half-opens, the leader pushes the WHOLE gap — the operator no longer re-issues /cluster/heal until lag 0. Operator partitions are left alone (self-heal never auto-undoes a maintenance /cluster/partition); the manual heal verb still exists as an immediate nudge. Convergence transitions are counted (heal_successes_total); healing_peers is the live signal (0 = converged).

5. Coordinated backup/restore + WAL PITR archival (E)

  • WAL archival hook (the PITR primitive): a wal.archive_dir config (TidalDb::builder().wal_archive_dir + topology wal.archive_dir). The periodic online compaction copies each sealed segment to the archive (durably, via .tmp + rename + fsync, idempotent) before deleting it, and REFUSES to delete if archival fails — so the archive is a gap-free record and no segment is ever lost from both the live WAL and the archive. Segment filenames encode shard + first-seq, so co-located groups share one archive dir without collision.
  • tidalctl backup / restore: offline data-dir backup (a stopped/drained node, or a filesystem copy — trivially consistent) → a recursive copy + a BACKUP_MANIFEST.json (BLAKE3 per file + the recovered WAL checkpoint cursor). Restore verifies EVERY file's BLAKE3 against the manifest BEFORE writing anything, and refuses a non-empty target (the destructive-op guard).
  • Coordinated cluster backup is the manifest + procedure (runbook §13): under ack=quorum, any committed replica's data dir holds the quorum-durable log, so it is a cluster-consistent snapshot at its recorded checkpoint_seq. Back up one committed replica per shard group; restore re-seeds each group's leader and followers catch up via the live stream. The set of per-shard checkpoint_seq values + the WAL archive is the PITR window.

6. Rolling upgrade: version handshake + release gate (F)

  • Version handshake on ship: HeartbeatRequest.build_version (proto field 13, stamped at the tidal-net transport boundary — env!("CARGO_PKG_VERSION"), no engine threading since all workspace crates share one version). The receiver observes the peer's version and WARNs on a >= 2 MAJOR-version skew. Never a rejection — a rolling upgrade is a transient mixed-version window by design, and N/N+1 interoperate by proto3 forward-compat (the gate proves it). An empty version = a pre-m11p8 peer (version-unknown, no warning).
  • Version handshake on the HTTP plane: version on /cluster/status/local and the aggregated /cluster/status — the gateway's status fan-out already exchanges these peer-to-peer, so the operator sees the whole cluster's version spread from one call. #[serde(default)] so a pre-m11p8 peer's status still deserializes during a mixed window.
  • Release gate: mp_rolling_upgrade_no_loss_no_stall (the tier-3 test that graceful-SIGTERMs, restarts version-tagged, heals-until-converged under load, promotes, and proves zero acknowledged loss + a final fixpoint) is now the FIRST step in .woodpecker.yaml — a failure blocks the image build below. CI is Woodpecker, never GitHub Actions.

Exit gate (from the roadmap)

  • Dashboard answers the golden-signal questions without code.
  • Backup→restore of a 100k-item cluster < 30 min.
  • Upgrade-under-load gate green.

Status

  • Metric set completed (breaker, forwards, self-heal) + multi-shard listener aggregation
  • Grafana cluster dashboard (12 panels) + Prometheus cluster alert group (8 rules)
  • Request-id + tracing on both cluster routers; id propagated across the forward hop
  • Truthful status: leader's own applied row + post-promote ShardId(0) keying
  • Self-driving heal loop (breaker-gated backlog re-ship) + heal metrics
  • WAL archival hook for PITR (wal.archive_dir, archive-before-delete, gap-free)
  • tidalctl backup / restore (BLAKE3 manifest, integrity-verified round-trip)
  • Version handshake (heartbeat build_version + status version) + N/N+1 policy
  • mp_rolling_upgrade_no_loss_no_stall promoted to a Woodpecker release gate
  • Docs (runbook §12 security / §13 backup+PITR / §14 rolling upgrade, monitoring cluster metrics + alerts) + CHANGELOG

Exit-gate evidence (local; release builds where noted)

Gate Target Measured
Dashboard answers golden-signal questions without code qualitative 12-panel "Cluster Replication" row covers lag / commit progress / per-peer queue / ship+fsync p99 / election churn / quorum timeouts / write-pool shed / breaker / forwards / self-heal — every alert's expr has a panel.
Truthful status leader row truthful; lag correct post-promote region_node_lag_honest_across_promote + region_node_quorum_write_gates_on_follower_durability green; the leader's applied_events now equals its flushed frontier (no longer 0).
Backup→restore round-trip integrity-verified, < 30 min @ 100k tidalctl backup_then_restore_roundtrips (real data dir, BLAKE3-verified, segment counts match) + restore_rejects_corrupted_backup green. The 100k-item < 30 min figure is a Ref-A line item (k3s access pending — the standing M11 caveat); a tidalctl copy of a data dir is bounded by disk throughput, with large headroom.
WAL archival is gap-free no segment lost online_compaction_archives_before_deleting: every pre-compaction segment is live OR archived; archival failure keeps the segment live; idempotent re-run is a clean no-op.
Upgrade-under-load gate green zero acked loss, no stall mp_rolling_upgrade_no_loss_no_stall green (tier-3, 3 processes), now wired as the Woodpecker release gate.

Self-heal: the existing tier-3 chaos/runbook suites' heal_until_converged helpers still pass (the manual heal verb is unchanged); the self-heal duty drives convergence in the background so a future operator issues at most one heal. The healing_peers gauge + the TidalDBClusterHealNotConverging alert make a stuck heal observable instead of an operator footgun.

Verification status

  • Workspace cargo fmt --all -- --check: clean.
  • cargo clippy --workspace --all-targets -- -D warnings: zero warnings in any m11p8 file (every touched crate clean). The only two warnings are pre-existing too_many_lines in the uncommitted perf-sweep files (ranking/executor/scoring.rs, tests.rs) — not m11p8, left untouched.
  • Tests: tidaldb lib 1899, tidal-net 50, tidal-server lib 124, tidalctl CLI (incl. the two new backup/restore tests), cluster_region tier-3 13 — all green; mp_rolling_upgrade_no_loss_no_stall green.
  • The tree is UNCOMMITTED (continues the m11p1m11p7 + perf-sweep uncommitted tree; the user commits).