Resolve all BLOCKER/CRITICAL/WARNING findings from the m11p7/p8 review: - tidalctl restore: safe_join path-traversal/Zip-Slip guard + fsync on write - corrupt-WAL checkpoint_seq guard; PITR archive-before-delete - cluster: x-tidal-relayed audit-dedup marker; forward_failures counts 5xx - mTLS/HTTP-TLS handshake hardening; accept-loop EMFILE backoff - per-principal rate-limit + node-token marker-pinning tests - self-heal tier-3 coverage; 5 router-auth tests tidal-stress: measurement-fidelity fixes (schedule-lag p99/max, exact feed-over-SLO verdict, shed annotation) + typed Body, workload.next 184ns->68ns, RoundRobin len==1 short-circuit, HeaderValue cache; new benches/hotpath.rs + lib.rs. perf wave 2: signal_snapshot SmallVec/SignalKey carrier; one-get-per-type ranking pre-pass.
13 KiB
m11p8 — Observability + Operations (COMPLETE — 2026-06-13)
Phase spec and exit gate: docs/roadmap-to-cluster.md §4/m11p8.
Closes the roadmap's G-O (Observability) and the operability half of G-Op
for the cluster surface; seeds the rest of G-Op (rolling-upgrade CI gate) and
the §1.4-3 incident ("the breaker eats the first heal").
Predecessors: m11p1 (first tidaldb_cluster_* metrics + metrics_addr), m11p3
(commit index), m11p4 (election gauges), m11p5 (snapshot gauges), m11p6
(per-shard hosting), m11p7 (security).
Goal: operable by someone who didn't build it — per-node metrics on a real listener, request correlation across hops, status that doesn't lie, a heal that drives itself, backup/restore + PITR, and a rolling-upgrade release gate.
Design (as adopted)
m11p8 is six sub-items. Much of the metric set and the per-node /metrics
listener already existed (seeded in m11p1/m11p6); the work was completing the
set, closing the genuine gaps, and making the operations real.
1. Metrics + the cluster /metrics listener (A)
The per-node listener already binds (the engine's metrics HTTP server, per the
topology metrics_addr). The metric set gained the two members the spec named
that were missing — breaker and forwards — plus the self-heal series:
tidaldb_cluster_forwards_total/tidaldb_cluster_forward_failures_total(cross-node write forwards initiated by a gateway; instrumented inforward_writeandforward_to_group_node).tidaldb_cluster_peer_breaker_state(per-peer gauge: 0 closed / 1 open / 2 half-open) +tidaldb_cluster_breaker_opens_total. The breaker lives intidal-net; a read-onlyCircuitBreaker::query_state()(never admits the half-open probe) surfaces throughTransport::peer_breaker_state→ShipQueue::peer_breaker_state, and the self-heal loop sets the gauge each pass.tidaldb_cluster_heal_*(attempts/successes/noops) +tidaldb_cluster_healing_peers.
Multi-shard listener (m11p6 co-location). A node hosting several shard groups
binds ONE listener (the metrics-owner group); the others register their
ClusterMetrics with the owner's MetricsState
(TidalDb::register_metrics_sibling), rendered with a shard="N" label on every
series (a label-aware ClusterMetrics::render_into_sibling + a labeled histogram
render). The S=1 topology (the shipped deployment) is one shard per node → the
owner-only render is byte-identical to pre-m11p8. Co-located shards never
collide on the shared scrape target.
Dashboard + alerts. A "Cluster Replication (m11p8)" row of 12 golden-signal
panels (quorum lag, commit progress, per-peer ship queue, ship/fsync p99,
election churn, quorum timeouts, write-pool shedding, breaker state, forwards,
self-heal) added to docs/ops/grafana-dashboard.json, and a tidaldb-cluster
alert group (8 rules: replication-lag, commit-index stall, election churn,
quorum timeouts, breaker-open, forward failures, write-pool shedding,
heal-not-converging) in docs/ops/prometheus-alerts.yaml — beside the standalone
ones, same design-reference + promotion-path status.
2. Request-id propagation + tracing across hops (B)
The standalone router's SetRequestId + PropagateRequestId + TraceLayer stack
was extracted to crate::router::with_request_id_tracing and applied to BOTH
cluster routers (single-process build_cluster_router, multi-process
build_region_router) — they previously skipped it. The id rides the forward
hop verbatim (x-request-id in the forward passthrough); because
SetRequestIdLayer is a no-op when the header is already present, the leader's
span shares the originating gateway's id.
Ship hop, honestly. Ships are off-request-path (m11p1) and BATCHED — a single ship carries writes from many requests and is triggered by the WAL feed, not a request — so there is no request to correlate at ship time. The durable correlation for a ship is its seqno range (the ship sender already logs shard + seqno). The version handshake (below) rides the heartbeat. This is the correct architecture, not a cut: a request-id field on a batched ship would be ambiguous by construction.
3. Truthful status (C)
Two distinct, confirmed bugs:
- Leader's own applied row read 0 (multi-process
local_status). The applied frontier (replication_state().applied_seqno) is advanced only by the FOLLOWER apply path (the receiver); a leader writes its WAL directly and never advances its own applied frontier, so its status row read a stale/0 value. Fix: for a leader, report its durable flushed frontier (ship_feed.flushed_seq()) — it is, by construction, applied to its own log. - Post-promote
ShardId(0)keying (single-processapplied_count). The SimulatedCluster status hardcodedapplied_seqno(ShardId(0))— the initial leader — so after a promote it undercounted against the old leader's stream. Fix: key on the current leader's shard (ShardId(leader_region.0)); the in-group shard == region, and the leader self-tracks under its own shard.
(The explorer-proposed self.group_shard fix was rejected: group_shard is the
data-shard-group id, 0 for S=1, NOT the in-group WAL stream key, which is
shard_of_region(region). Keying on it would have been wrong.)
Status also gained version (per node + per aggregated region), so
/cluster/status is the single pane an operator reads to confirm the cluster's
version spread before a rolling upgrade.
4. Self-driving heal (D) — closes incident §1.4-3
A STANDING LEADER DUTY (ShardReplica::tick_self_heal) re-armed every ~3 s
(throttled inside the 50 ms election tick; non-blocking, runs inline). Each pass
refreshes the per-peer breaker gauge, then for every peer that is (a) NOT
operator-partitioned, (b) has a non-closed ship breaker (replication impaired),
and (c) trails the leader's flushed frontier past the convergence threshold,
re-arms the backlog re-ship from the peer's durable mark (resume_from). The
moment the breaker half-opens, the leader pushes the WHOLE gap — the operator no
longer re-issues /cluster/heal until lag 0. Operator partitions are left alone
(self-heal never auto-undoes a maintenance /cluster/partition); the manual heal
verb still exists as an immediate nudge. Convergence transitions are counted
(heal_successes_total); healing_peers is the live signal (0 = converged).
5. Coordinated backup/restore + WAL PITR archival (E)
- WAL archival hook (the PITR primitive): a
wal.archive_dirconfig (TidalDb::builder().wal_archive_dir+ topologywal.archive_dir). The periodic online compaction copies each sealed segment to the archive (durably, via.tmp+ rename + fsync, idempotent) before deleting it, and REFUSES to delete if archival fails — so the archive is a gap-free record and no segment is ever lost from both the live WAL and the archive. Segment filenames encode shard + first-seq, so co-located groups share one archive dir without collision. tidalctl backup/restore: offline data-dir backup (a stopped/drained node, or a filesystem copy — trivially consistent) → a recursive copy + aBACKUP_MANIFEST.json(BLAKE3 per file + the recovered WAL checkpoint cursor). Restore verifies EVERY file's BLAKE3 against the manifest BEFORE writing anything, and refuses a non-empty target (the destructive-op guard).- Coordinated cluster backup is the manifest + procedure (runbook §13): under
ack=quorum, any committed replica's data dir holds the quorum-durable log, so it is a cluster-consistent snapshot at its recordedcheckpoint_seq. Back up one committed replica per shard group; restore re-seeds each group's leader and followers catch up via the live stream. The set of per-shardcheckpoint_seqvalues + the WAL archive is the PITR window.
6. Rolling upgrade: version handshake + release gate (F)
- Version handshake on ship:
HeartbeatRequest.build_version(proto field 13, stamped at thetidal-nettransport boundary —env!("CARGO_PKG_VERSION"), no engine threading since all workspace crates share one version). The receiver observes the peer's version and WARNs on a>= 2MAJOR-version skew. Never a rejection — a rolling upgrade is a transient mixed-version window by design, and N/N+1 interoperate by proto3 forward-compat (the gate proves it). An empty version = a pre-m11p8 peer (version-unknown, no warning). - Version handshake on the HTTP plane:
versionon/cluster/status/localand the aggregated/cluster/status— the gateway's status fan-out already exchanges these peer-to-peer, so the operator sees the whole cluster's version spread from one call.#[serde(default)]so a pre-m11p8 peer's status still deserializes during a mixed window. - Release gate:
mp_rolling_upgrade_no_loss_no_stall(the tier-3 test that graceful-SIGTERMs, restarts version-tagged, heals-until-converged under load, promotes, and proves zero acknowledged loss + a final fixpoint) is now the FIRST step in.woodpecker.yaml— a failure blocks the image build below. CI is Woodpecker, never GitHub Actions.
Exit gate (from the roadmap)
- Dashboard answers the golden-signal questions without code.
- Backup→restore of a 100k-item cluster < 30 min.
- Upgrade-under-load gate green.
Status
- Metric set completed (breaker, forwards, self-heal) + multi-shard listener aggregation
- Grafana cluster dashboard (12 panels) + Prometheus cluster alert group (8 rules)
- Request-id + tracing on both cluster routers; id propagated across the forward hop
- Truthful status: leader's own applied row + post-promote
ShardId(0)keying - Self-driving heal loop (breaker-gated backlog re-ship) + heal metrics
- WAL archival hook for PITR (
wal.archive_dir, archive-before-delete, gap-free) tidalctl backup/restore(BLAKE3 manifest, integrity-verified round-trip)- Version handshake (heartbeat
build_version+ statusversion) + N/N+1 policy mp_rolling_upgrade_no_loss_no_stallpromoted to a Woodpecker release gate- Docs (runbook §12 security / §13 backup+PITR / §14 rolling upgrade, monitoring cluster metrics + alerts) + CHANGELOG
Exit-gate evidence (local; release builds where noted)
| Gate | Target | Measured |
|---|---|---|
| Dashboard answers golden-signal questions without code | qualitative | 12-panel "Cluster Replication" row covers lag / commit progress / per-peer queue / ship+fsync p99 / election churn / quorum timeouts / write-pool shed / breaker / forwards / self-heal — every alert's expr has a panel. |
| Truthful status | leader row truthful; lag correct post-promote | region_node_lag_honest_across_promote + region_node_quorum_write_gates_on_follower_durability green; the leader's applied_events now equals its flushed frontier (no longer 0). |
| Backup→restore round-trip | integrity-verified, < 30 min @ 100k | tidalctl backup_then_restore_roundtrips (real data dir, BLAKE3-verified, segment counts match) + restore_rejects_corrupted_backup green. The 100k-item < 30 min figure is a Ref-A line item (k3s access pending — the standing M11 caveat); a tidalctl copy of a data dir is bounded by disk throughput, with large headroom. |
| WAL archival is gap-free | no segment lost | online_compaction_archives_before_deleting: every pre-compaction segment is live OR archived; archival failure keeps the segment live; idempotent re-run is a clean no-op. |
| Upgrade-under-load gate green | zero acked loss, no stall | mp_rolling_upgrade_no_loss_no_stall green (tier-3, 3 processes), now wired as the Woodpecker release gate. |
Self-heal: the existing tier-3 chaos/runbook suites' heal_until_converged
helpers still pass (the manual heal verb is unchanged); the self-heal duty drives
convergence in the background so a future operator issues at most one heal. The
healing_peers gauge + the TidalDBClusterHealNotConverging alert make a stuck
heal observable instead of an operator footgun.
Verification status
- Workspace
cargo fmt --all -- --check: clean. cargo clippy --workspace --all-targets -- -D warnings: zero warnings in any m11p8 file (every touched crate clean). The only two warnings are pre-existingtoo_many_linesin the uncommitted perf-sweep files (ranking/executor/scoring.rs,tests.rs) — not m11p8, left untouched.- Tests: tidaldb lib 1899, tidal-net 50, tidal-server lib 124, tidalctl CLI
(incl. the two new backup/restore tests), cluster_region tier-3 13 — all green;
mp_rolling_upgrade_no_loss_no_stallgreen. - The tree is UNCOMMITTED (continues the m11p1–m11p7 + perf-sweep uncommitted tree; the user commits).