Resolve all BLOCKER/CRITICAL/WARNING findings from the m11p7/p8 review: - tidalctl restore: safe_join path-traversal/Zip-Slip guard + fsync on write - corrupt-WAL checkpoint_seq guard; PITR archive-before-delete - cluster: x-tidal-relayed audit-dedup marker; forward_failures counts 5xx - mTLS/HTTP-TLS handshake hardening; accept-loop EMFILE backoff - per-principal rate-limit + node-token marker-pinning tests - self-heal tier-3 coverage; 5 router-auth tests tidal-stress: measurement-fidelity fixes (schedule-lag p99/max, exact feed-over-SLO verdict, shed annotation) + typed Body, workload.next 184ns->68ns, RoundRobin len==1 short-circuit, HeaderValue cache; new benches/hotpath.rs + lib.rs. perf wave 2: signal_snapshot SmallVec/SignalKey carrier; one-get-per-type ranking pre-pass.
200 lines
13 KiB
Markdown
200 lines
13 KiB
Markdown
# m11p8 — Observability + Operations (COMPLETE — 2026-06-13)
|
||
|
||
Phase spec and exit gate: [docs/roadmap-to-cluster.md §4/m11p8](../../roadmap-to-cluster.md).
|
||
Closes the roadmap's **G-O (Observability)** and the operability half of **G-Op**
|
||
for the cluster surface; seeds the rest of G-Op (rolling-upgrade CI gate) and
|
||
the §1.4-3 incident ("the breaker eats the first heal").
|
||
Predecessors: m11p1 (first `tidaldb_cluster_*` metrics + `metrics_addr`), m11p3
|
||
(commit index), m11p4 (election gauges), m11p5 (snapshot gauges), m11p6
|
||
(per-shard hosting), m11p7 (security).
|
||
|
||
**Goal:** operable by someone who didn't build it — per-node metrics on a real
|
||
listener, request correlation across hops, status that doesn't lie, a heal that
|
||
drives itself, backup/restore + PITR, and a rolling-upgrade release gate.
|
||
|
||
## Design (as adopted)
|
||
|
||
m11p8 is six sub-items. Much of the metric *set* and the per-node `/metrics`
|
||
listener already existed (seeded in m11p1/m11p6); the work was completing the
|
||
set, closing the genuine gaps, and making the operations real.
|
||
|
||
### 1. Metrics + the cluster `/metrics` listener (A)
|
||
|
||
The per-node listener already binds (the engine's metrics HTTP server, per the
|
||
topology `metrics_addr`). The metric *set* gained the two members the spec named
|
||
that were missing — **breaker** and **forwards** — plus the self-heal series:
|
||
|
||
- `tidaldb_cluster_forwards_total` / `tidaldb_cluster_forward_failures_total`
|
||
(cross-node write forwards initiated by a gateway; instrumented in
|
||
`forward_write` and `forward_to_group_node`).
|
||
- `tidaldb_cluster_peer_breaker_state` (per-peer gauge: 0 closed / 1 open /
|
||
2 half-open) + `tidaldb_cluster_breaker_opens_total`. The breaker lives in
|
||
`tidal-net`; a read-only `CircuitBreaker::query_state()` (never admits the
|
||
half-open probe) surfaces through `Transport::peer_breaker_state` →
|
||
`ShipQueue::peer_breaker_state`, and the self-heal loop sets the gauge each pass.
|
||
- `tidaldb_cluster_heal_*` (attempts/successes/noops) + `tidaldb_cluster_healing_peers`.
|
||
|
||
**Multi-shard listener (m11p6 co-location).** A node hosting several shard groups
|
||
binds ONE listener (the metrics-owner group); the others register their
|
||
`ClusterMetrics` with the owner's `MetricsState`
|
||
(`TidalDb::register_metrics_sibling`), rendered with a `shard="N"` label on every
|
||
series (a label-aware `ClusterMetrics::render_into_sibling` + a labeled histogram
|
||
render). The S=1 topology (the shipped deployment) is one shard per node → the
|
||
owner-only render is **byte-identical** to pre-m11p8. Co-located shards never
|
||
collide on the shared scrape target.
|
||
|
||
**Dashboard + alerts.** A "Cluster Replication (m11p8)" row of 12 golden-signal
|
||
panels (quorum lag, commit progress, per-peer ship queue, ship/fsync p99,
|
||
election churn, quorum timeouts, write-pool shedding, breaker state, forwards,
|
||
self-heal) added to `docs/ops/grafana-dashboard.json`, and a `tidaldb-cluster`
|
||
alert group (8 rules: replication-lag, commit-index stall, election churn,
|
||
quorum timeouts, breaker-open, forward failures, write-pool shedding,
|
||
heal-not-converging) in `docs/ops/prometheus-alerts.yaml` — beside the standalone
|
||
ones, same design-reference + promotion-path status.
|
||
|
||
### 2. Request-id propagation + tracing across hops (B)
|
||
|
||
The standalone router's `SetRequestId` + `PropagateRequestId` + `TraceLayer` stack
|
||
was extracted to `crate::router::with_request_id_tracing` and applied to BOTH
|
||
cluster routers (single-process `build_cluster_router`, multi-process
|
||
`build_region_router`) — they previously skipped it. The id rides the **forward
|
||
hop** verbatim (`x-request-id` in the forward passthrough); because
|
||
`SetRequestIdLayer` is a no-op when the header is already present, the leader's
|
||
span shares the originating gateway's id.
|
||
|
||
**Ship hop, honestly.** Ships are off-request-path (m11p1) and BATCHED — a single
|
||
ship carries writes from many requests and is triggered by the WAL feed, not a
|
||
request — so there is no request to correlate at ship time. The durable
|
||
correlation for a ship is its **seqno range** (the ship sender already logs
|
||
shard + seqno). The version handshake (below) rides the heartbeat. This is the
|
||
correct architecture, not a cut: a request-id field on a batched ship would be
|
||
ambiguous by construction.
|
||
|
||
### 3. Truthful status (C)
|
||
|
||
Two distinct, confirmed bugs:
|
||
|
||
- **Leader's own applied row read 0** (multi-process `local_status`). The applied
|
||
frontier (`replication_state().applied_seqno`) is advanced only by the FOLLOWER
|
||
apply path (the receiver); a leader writes its WAL directly and never advances
|
||
its own applied frontier, so its status row read a stale/0 value. Fix: for a
|
||
leader, report its durable **flushed** frontier (`ship_feed.flushed_seq()`) —
|
||
it is, by construction, applied to its own log.
|
||
- **Post-promote `ShardId(0)` keying** (single-process `applied_count`). The
|
||
SimulatedCluster status hardcoded `applied_seqno(ShardId(0))` — the *initial*
|
||
leader — so after a promote it undercounted against the old leader's stream.
|
||
Fix: key on the **current** leader's shard (`ShardId(leader_region.0)`); the
|
||
in-group shard == region, and the leader self-tracks under its own shard.
|
||
|
||
(The explorer-proposed `self.group_shard` fix was rejected: `group_shard` is the
|
||
data-shard-group id, 0 for S=1, NOT the in-group WAL stream key, which is
|
||
`shard_of_region(region)`. Keying on it would have been wrong.)
|
||
|
||
Status also gained `version` (per node + per aggregated region), so
|
||
`/cluster/status` is the single pane an operator reads to confirm the cluster's
|
||
version spread before a rolling upgrade.
|
||
|
||
### 4. Self-driving heal (D) — closes incident §1.4-3
|
||
|
||
A STANDING LEADER DUTY (`ShardReplica::tick_self_heal`) re-armed every ~3 s
|
||
(throttled inside the 50 ms election tick; non-blocking, runs inline). Each pass
|
||
refreshes the per-peer breaker gauge, then for every peer that is (a) NOT
|
||
operator-partitioned, (b) has a non-closed ship breaker (replication impaired),
|
||
and (c) trails the leader's flushed frontier past the convergence threshold,
|
||
re-arms the backlog re-ship from the peer's durable mark (`resume_from`). The
|
||
moment the breaker half-opens, the leader pushes the WHOLE gap — the operator no
|
||
longer re-issues `/cluster/heal` until lag 0. Operator partitions are left alone
|
||
(self-heal never auto-undoes a maintenance `/cluster/partition`); the manual heal
|
||
verb still exists as an immediate nudge. Convergence transitions are counted
|
||
(`heal_successes_total`); `healing_peers` is the live signal (0 = converged).
|
||
|
||
### 5. Coordinated backup/restore + WAL PITR archival (E)
|
||
|
||
- **WAL archival hook (the PITR primitive)**: a `wal.archive_dir` config
|
||
(`TidalDb::builder().wal_archive_dir` + topology `wal.archive_dir`). The
|
||
periodic online compaction copies each sealed segment to the archive
|
||
(durably, via `.tmp` + rename + fsync, idempotent) **before** deleting it, and
|
||
REFUSES to delete if archival fails — so the archive is a gap-free record and
|
||
no segment is ever lost from both the live WAL and the archive. Segment
|
||
filenames encode shard + first-seq, so co-located groups share one archive dir
|
||
without collision.
|
||
- **`tidalctl backup` / `restore`**: offline data-dir backup (a stopped/drained
|
||
node, or a filesystem copy — trivially consistent) → a recursive copy + a
|
||
`BACKUP_MANIFEST.json` (BLAKE3 per file + the recovered WAL checkpoint cursor).
|
||
Restore verifies EVERY file's BLAKE3 against the manifest BEFORE writing
|
||
anything, and refuses a non-empty target (the destructive-op guard).
|
||
- **Coordinated cluster backup** is the manifest + procedure (runbook §13): under
|
||
`ack=quorum`, any committed replica's data dir holds the quorum-durable log, so
|
||
it is a cluster-consistent snapshot at its recorded `checkpoint_seq`. Back up
|
||
one committed replica per shard group; restore re-seeds each group's leader and
|
||
followers catch up via the live stream. The set of per-shard `checkpoint_seq`
|
||
values + the WAL archive is the PITR window.
|
||
|
||
### 6. Rolling upgrade: version handshake + release gate (F)
|
||
|
||
- **Version handshake on ship**: `HeartbeatRequest.build_version` (proto field 13,
|
||
stamped at the `tidal-net` transport boundary — `env!("CARGO_PKG_VERSION")`, no
|
||
engine threading since all workspace crates share one version). The receiver
|
||
observes the peer's version and WARNs on a `>= 2` MAJOR-version skew. **Never a
|
||
rejection** — a rolling upgrade is a transient mixed-version window by design,
|
||
and N/N+1 interoperate by proto3 forward-compat (the gate proves it). An empty
|
||
version = a pre-m11p8 peer (version-unknown, no warning).
|
||
- **Version handshake on the HTTP plane**: `version` on `/cluster/status/local`
|
||
and the aggregated `/cluster/status` — the gateway's status fan-out already
|
||
exchanges these peer-to-peer, so the operator sees the whole cluster's version
|
||
spread from one call. `#[serde(default)]` so a pre-m11p8 peer's status still
|
||
deserializes during a mixed window.
|
||
- **Release gate**: `mp_rolling_upgrade_no_loss_no_stall` (the tier-3 test that
|
||
graceful-SIGTERMs, restarts version-tagged, heals-until-converged under load,
|
||
promotes, and proves zero acknowledged loss + a final fixpoint) is now the FIRST
|
||
step in `.woodpecker.yaml` — a failure blocks the image build below. CI is
|
||
Woodpecker, never GitHub Actions.
|
||
|
||
## Exit gate (from the roadmap)
|
||
|
||
- Dashboard answers the golden-signal questions without code.
|
||
- Backup→restore of a 100k-item cluster < 30 min.
|
||
- Upgrade-under-load gate green.
|
||
|
||
## Status
|
||
|
||
- [x] Metric set completed (breaker, forwards, self-heal) + multi-shard listener aggregation
|
||
- [x] Grafana cluster dashboard (12 panels) + Prometheus cluster alert group (8 rules)
|
||
- [x] Request-id + tracing on both cluster routers; id propagated across the forward hop
|
||
- [x] Truthful status: leader's own applied row + post-promote `ShardId(0)` keying
|
||
- [x] Self-driving heal loop (breaker-gated backlog re-ship) + heal metrics
|
||
- [x] WAL archival hook for PITR (`wal.archive_dir`, archive-before-delete, gap-free)
|
||
- [x] `tidalctl backup` / `restore` (BLAKE3 manifest, integrity-verified round-trip)
|
||
- [x] Version handshake (heartbeat `build_version` + status `version`) + N/N+1 policy
|
||
- [x] `mp_rolling_upgrade_no_loss_no_stall` promoted to a Woodpecker release gate
|
||
- [x] Docs (runbook §12 security / §13 backup+PITR / §14 rolling upgrade, monitoring cluster metrics + alerts) + CHANGELOG
|
||
|
||
## Exit-gate evidence (local; release builds where noted)
|
||
|
||
| Gate | Target | Measured |
|
||
|------|--------|----------|
|
||
| Dashboard answers golden-signal questions without code | qualitative | 12-panel "Cluster Replication" row covers lag / commit progress / per-peer queue / ship+fsync p99 / election churn / quorum timeouts / write-pool shed / breaker / forwards / self-heal — every alert's expr has a panel. |
|
||
| Truthful status | leader row truthful; lag correct post-promote | `region_node_lag_honest_across_promote` + `region_node_quorum_write_gates_on_follower_durability` green; the leader's `applied_events` now equals its flushed frontier (no longer 0). |
|
||
| Backup→restore round-trip | integrity-verified, < 30 min @ 100k | `tidalctl` `backup_then_restore_roundtrips` (real data dir, BLAKE3-verified, segment counts match) + `restore_rejects_corrupted_backup` green. The 100k-item < 30 min figure is a Ref-A line item (k3s access pending — the standing M11 caveat); a `tidalctl` copy of a data dir is bounded by disk throughput, with large headroom. |
|
||
| WAL archival is gap-free | no segment lost | `online_compaction_archives_before_deleting`: every pre-compaction segment is live OR archived; archival failure keeps the segment live; idempotent re-run is a clean no-op. |
|
||
| Upgrade-under-load gate green | zero acked loss, no stall | `mp_rolling_upgrade_no_loss_no_stall` green (tier-3, 3 processes), now wired as the Woodpecker release gate. |
|
||
|
||
Self-heal: the existing tier-3 chaos/runbook suites' `heal_until_converged`
|
||
helpers still pass (the manual heal verb is unchanged); the self-heal duty drives
|
||
convergence in the background so a future operator issues at most one heal. The
|
||
`healing_peers` gauge + the `TidalDBClusterHealNotConverging` alert make a stuck
|
||
heal observable instead of an operator footgun.
|
||
|
||
## Verification status
|
||
|
||
- Workspace `cargo fmt --all -- --check`: clean.
|
||
- `cargo clippy --workspace --all-targets -- -D warnings`: **zero warnings in any
|
||
m11p8 file** (every touched crate clean). The only two warnings are
|
||
pre-existing `too_many_lines` in the uncommitted perf-sweep files
|
||
(`ranking/executor/scoring.rs`, `tests.rs`) — not m11p8, left untouched.
|
||
- Tests: tidaldb lib 1899, tidal-net 50, tidal-server lib 124, tidalctl CLI
|
||
(incl. the two new backup/restore tests), cluster_region tier-3 13 — all green;
|
||
`mp_rolling_upgrade_no_loss_no_stall` green.
|
||
- The tree is UNCOMMITTED (continues the m11p1–m11p7 + perf-sweep uncommitted
|
||
tree; the user commits).
|