docs: refresh ops runbooks to the live rc7 / full-placement reality

The runbooks had drifted to the retired m8/m11p5 design while all m12 production
reality (topology, perf, fixes, DR) sat only in a profiling doc no operator opens.
This promotes that reality into the runbooks and fixes the contradictions.

Contradictions fixed:
- runbooks/cluster.md: the "NEITHER IS QUORUM-ACKED HA YET" status banner was FALSE
  (quorum-ack + automatic election have been live since m11p3/p4). Rewritten to
  state the deployed reality (single-StatefulSet full-placement RF3, rc7).
- README.md: the cluster section called the HA cluster a "built-in simulated
  cluster / multi-region fabric" demo and showed promote-by-region as failover.
  Rewritten — real quorum HA, automatic failover, /cluster/promote is a maintenance
  verb. Kept the honest caveats (experimental gate, global-signals-only).

Reality promoted into the runbooks:
- Live topology (single STS, ns tidaldb-cluster, 3 voters, full-placement RF3,
  gRPC 9601/9602/9603, HTTPS+mTLS :9500), the five shipped fixes, and the real
  build+digest-pin procedure (cross-compile -> trixie -> amd64 PLATFORM manifest,
  not the index/attestation digest) in cluster.md + kubernetes.md.
- Stale constants: soak ramp 3900 -> 200 rps; cluster grace 60 -> 600s; the
  pre-m12 4.5k/s signal-write perf table annotated + the 1536-D read reality added.
- ops/capacity-planning.md: new "Ref-A 3-node fleet — measured capacity" section
  (read p99/ceiling, ~250 rps write knee, 1M needs >16GiB nodes, pod resources).
- ops/recovery.md: new cluster-recovery routing section + scoped the quiesce-and-
  copy note to standalone (the cluster uses tidalctl + the DR runbook).

New docs:
- runbooks/disaster-recovery.md: the proven S3/R2 backup -> restore -> byte-verify
  -> query-proof procedure, full-cluster rebuild, PITR posture (previously
  undocumented despite being proven against real S3).
- runbooks/on-call.md: incident response — symptom -> golden signal -> runbook,
  severity, escalation, and the open alert-wiring step.
- runbooks/README.md: the runbook index + current production facts.

Open follow-up (infra, not docs): ops/prometheus-alerts.yaml is accurate but
design-reference; promoting it to a live PrometheusRule is the one unwired step.
This commit is contained in:
jx12n 2026-06-19 19:53:29 -06:00
parent 0847c3d36f
commit 25296bcc5b
8 changed files with 645 additions and 51 deletions

View File

@ -214,33 +214,42 @@ The UI shows seeded users, exploration labels, and real-time adaptation; see `ap
### 5. Run the cluster server + Docker image
Need a single endpoint that fronts the built-in simulated cluster? Use
`tidal-server` in `cluster` mode. It spins up the multi-region fabric,
ships WAL batches between regions, and exposes `/signals`, `/feed`,
`/search` plus cluster-management routes.
Need a real high-availability endpoint? Run `tidal-server` in `cluster`
mode. This is a genuine HA cluster — quorum-acked writes, **automatic**
leader election + failover, elastic seed-join membership, inter-node mTLS,
and per-node Prometheus metrics — deployed in production on k3s as one
`StatefulSet` (3 pods = 3 regions = 3 voters, full-placement RF3 so every
pod hosts all shard groups, HTTPS + mTLS on `:9500`). It exposes `/signals`,
`/feed`, `/search` plus cluster-management routes. It stays behind the
`--experimental-cluster` gate.
```bash
cargo run -p tidal-server -- \
cluster \
--listen 0.0.0.0:9500 \
--schema tidal-server/config/default-schema.yaml \
--topology tidal-server/config/default-cluster.yaml
--topology tidal-server/config/default-cluster.yaml \
--experimental-cluster
```
Key endpoints:
```bash
curl http://127.0.0.1:9500/health
curl -X POST http://127.0.0.1:9500/signals -d '{ "entity_id": 1, "signal": "view", "weight": 1.0 }'
curl "http://127.0.0.1:9500/feed?profile=trending&region=eu-west"
curl http://127.0.0.1:9500/cluster/status
curl -X POST http://127.0.0.1:9500/cluster/promote -d '{ "region": "eu-west" }'
curl https://127.0.0.1:9500/health
curl -X POST https://127.0.0.1:9500/signals -d '{ "entity_id": 1, "signal": "view", "weight": 1.0 }'
curl "https://127.0.0.1:9500/feed?profile=trending&region=eu-west"
curl https://127.0.0.1:9500/cluster/status
# /cluster/promote is a fenced MAINTENANCE verb: a graceful, voluntary
# leadership handoff. It is NOT the failover path — kill the leader and the
# survivors elect a successor automatically, with zero operator action.
curl -X POST https://127.0.0.1:9500/cluster/promote -d '{ "region": "eu-west" }'
```
Cluster mode currently replicates global signals (no `user_id` /
`creator_id` contexts) so that followers can stay in sync with the leaders
WAL stream. See **[docs/runbooks/cluster.md](docs/runbooks/cluster.md)** for
operational steps, failure drills, and API references.
Cluster mode replicates **global signals only** (no `user_id` /
`creator_id` contexts) so that followers stay in sync with the leader's
replicated log. For Kubernetes deployment, scaling, failover drills, and the
operational API see **[docs/runbooks/kubernetes.md](docs/runbooks/kubernetes.md)**
and **[docs/runbooks/cluster.md](docs/runbooks/cluster.md)**.
Prefer containers? Build the provided image and run it anywhere:
@ -301,5 +310,6 @@ Milestones completed:
- Session and agent context (short-lived signals, preference decay)
- Crash recovery, graceful degradation, rate limiting, diagnostics
- Scale: tested to 1M items; scale benchmarks passing
- High-availability cluster (experimental): quorum-acked writes, automatic election + failover, elastic seed-join membership, inter-node mTLS, per-node Prometheus — running in production on k3s
The API surface is stable for the implemented features. Breaking changes are possible before 1.0.

View File

@ -151,3 +151,63 @@ Do not configure swap for production tidalDB instances. USearch HNSW traversal a
### Disk Type
SSD is strongly recommended for all deployments. NVMe is recommended at 10M+ items. The WAL uses synchronous `fsync` on every segment rotation, and fjall's journal uses `persist(SyncAll)` during checkpoint. Spinning disk latency on these operations directly impacts signal write throughput.
---
## Cluster (Ref-A 3-node fleet) — measured capacity
The tables above are single-node, analytic estimates. This section is the **measured operating envelope** of the live cluster: **Ref-A = 3 nodes × 4 vCPU / 16 GiB**, full-placement RF3 (every pod replicates every shard group), 100k × 1536-D corpus. These are real `tidal-stress` numbers, not formulas. They are **Ref-A figures** — the enterprise **Ref-B** target is **≥5 × 8 vCPU / ≥16 GiB** nodes with **partitioned** placement (each pod hosts a subset of groups). See `docs/profiling/m12-cluster-deploy-findings.md` and `docs/profiling/m12p4-t5-sharded-throughput.md`.
### Read throughput
| metric | value | gate | verdict |
|--------|-------|------|---------|
| read p99 @ 100500 rps | **7.9711.47 ms** | ≤ 10 ms (G1) | **MET** @ 100k/1536-D |
| recall@10 | **0.9989** | ≥ 0.95 | **MET** |
| read ceiling (clean) | **~1000 read-ops/s** | — | CPU-bound |
| read ceiling (saturated) | **~1500 read-ops/s** | — | shed/error past here |
- Reads are CPU-bound: each read scatters **3 parallel HNSW searches** (one per shard group, full placement).
- Read throughput scales **~linearly with node count** on full placement — every node serves reads from its local replica of every group.
- Spread reads **round-robin across all 3 pods**; do not pin to the leader.
### Write throughput
| metric | value | note |
|--------|-------|------|
| write knee | **~250 rps** | the peach mix is WRITE-heavy (view/like/skip ≈ 90% of ops) |
| 30-night soak | **200 rps** | sustained with margin |
| write scaling vs node count | **~1.0×** | does NOT scale at full-placement RF3 |
- Write tput does **NOT** scale with node count at full-placement RF3: every per-shard quorum spans all 3 nodes, so **every follower applies every 1536-D write**. Adding nodes adds replication work, not write capacity (~1.0×, not 2.5×).
- The old T5 **2.5× write-scaling** target needs **≥5 nodes + PARTITIONED placement**. T5 / G-S has been **re-scoped to read-throughput scaling**.
### Per-pod memory at 1536-D
| corpus | per-pod RSS | pod mem limit | verdict |
|--------|-------------|---------------|---------|
| 100k × 1536-D | **~1.9 GiB** (HNSW load peak) | 4 GiB | fits |
| 1M × 1536-D | **~78 GB** | 4 GiB | **OOMs the 16 GiB nodes** |
- Full placement → **each pod holds the WHOLE corpus** (no per-shard sharding of memory).
- 1M × 1536-D ≈ 78 GB/pod RSS overruns the ~13 GiB allocatable (already ~7 GiB of co-tenants) → the **1M production-read gate needs nodes >16 GiB** (a Ref-B requirement).
### Startup / boot
- HNSW rebuild/load at 1536-D is **CPU-bound** (~5 min single-core at 100k); 1M scales up from there.
- `startupProbe` budget is **~20 min** (`failureThreshold` 240 × 5 s).
- Graceful shutdown **saves the graphs** (grace **600 s**), so a clean restart **skips the rebuild** (load, not rebuild). A SIGKILL/crash skips the save → next boot rebuilds.
### Pod resources (live)
| resource | request | limit |
|----------|---------|-------|
| CPU | 500m | 3 |
| memory | 1 GiB | 4 GiB |
| PVC | — | 5 GiB/pod, `local-path` |
- PVC is **local NVMe** (`local-path`) — longhorn's fsync overhead was unacceptable for the WAL path.
### Ref-B target (enterprise)
These are Ref-A figures. The enterprise **Ref-B** target is **≥5 × 8 vCPU / ≥16 GiB** nodes with **partitioned** placement — the shape required to (a) scale writes past ~1.0×, and (b) hold a 1M × 1536-D corpus without OOM.

View File

@ -2,9 +2,29 @@
This document covers error scenarios, their causes, data at risk, and step-by-step recovery procedures for tidalDB.
> **Standalone vs cluster.** Everything below the cluster section (corrupt keyspace, WAL truncate, stale lock, schema mismatch, disk-full, quiesce-and-copy backup) is for the **standalone / single-node engine**. If you are operating the replicated cluster (namespace `tidaldb-cluster`), start with the next section — most cluster failure modes self-heal and have a dedicated runbook, and the cluster has a real backup/restore tool (`tidalctl`), not quiesce-and-copy.
---
## Error Scenarios
## Cluster recovery (ns `tidaldb-cluster`)
This is the routing table `kubernetes.md` points operators to for **cluster** incidents. The cluster is a quorum-acked, sharded × replicated deployment: a single node loss is survivable (2-of-3 quorum), and divergent/corrupt nodes self-quarantine and reseed from a clean snapshot rather than serving bad data. Recognize the mode, then follow the pointer. Deep procedures live in [`docs/runbooks/cluster.md`](../runbooks/cluster.md); object-store DR and PITR live in [`docs/runbooks/disaster-recovery.md`](../runbooks/disaster-recovery.md).
| Failure mode | Recognize | Respond → pointer |
|:-------------|:----------|:------------------|
| **Reseed loop / CrashLoop** | Pod restarts repeatedly with "reseed marker latched" in logs; `reseed_required` metric set. | Resolved in the live image (**rc5+**): a shard that heals via catch-up now **aborts** the self-restart, and a genuine reseed snapshot-installs and converges to lag=0. If seen, confirm the running image is rc5+. See [cluster.md — reseed / membership verbs](../runbooks/cluster.md#membership-verbs-m11p5--online-add--remove--inspect--reseed). |
| **No quorum / writes 503** | `ack=quorum` writes return a retryable 503 naming laggards; `/cluster/status` shows missing voters. | Single-node loss is survivable (2/3 quorum) — the cluster keeps serving. **Both** followers down = honest retryable 503 until one returns; check leader + voter liveness via `/cluster/status`. See [cluster.md — durability contract (§8)](../runbooks/cluster.md#8-write-durability-contract-the-ack-knob-and-its-cost-m11p3). |
| **Write-burst quorum-loss spiral / breaker stuck open** | `tidaldb_cluster_peer_breaker_state` shows a peer pinned at `1` (open) while that peer is actually alive and reachable. | This is the write-burst **false-partition**, fixed in **rc7**. Roll to **≥ rc7** if seen on an older image; the per-peer breaker no longer latches open against a live peer. See [cluster.md — circuit breaker / partition drill](../runbooks/cluster.md#10-partition-drill). |
| **Divergence / quarantine** | A node would serve a divergent prefix; it flips to 503 readiness instead. | By design: the node **auto-quarantines** (503 readiness) and reseeds **without wiping**, and never serves a divergent prefix. No manual wipe. See [cluster.md](../runbooks/cluster.md). |
| **Stuck leader / dual-leader view** | `/cluster/status` shows no progress or two nodes each claiming leadership. | Restart the leader to force re-election; auto-election (m11p4) converges on its own. See [cluster.md — failover (§9)](../runbooks/cluster.md#9-failover-multi-process). |
| **Read-SLA collapse** | p99 read latency blows past 10 ms; reads hang. | Was CPU oversubscription, fixed in **rc12** (SEARCH_GATE + parallel scatter + cpu limit 2→3). Confirm running image is rc12+ and that load is spread across **all 3** pods (region-pinned reads, not pinned to one). See [cluster.md](../runbooks/cluster.md). |
| **PVC loss / corruption on one pod** | One pod's data dir is corrupt or its PVC is gone. | Delete that pod's **PVC and pod**; it reseeds fresh from the quorum (snapshot install) and converges to lag=0. On the live image the corrupt-PVC case is recoverable this way — no full-cluster action needed. See [cluster.md — reseed](../runbooks/cluster.md#membership-verbs-m11p5--online-add--remove--inspect--reseed). |
| **Full-cluster loss / rebuild from backup** | Quorum cannot be restored from surviving pods (multi-node data loss). | Rebuild from object store with `tidalctl restore` (BLAKE3-manifested, per-shard `checkpoint_seq`, proven against real S3). See [disaster-recovery.md](../runbooks/disaster-recovery.md). |
| **Point-in-time recovery (PITR)** | Need to restore the corpus to a specific point in time. | Use the `tidalctl` + DR runbook PITR procedure. See [disaster-recovery.md](../runbooks/disaster-recovery.md). |
---
## Error Scenarios (standalone engine)
### 1. `StorageError::Corruption` on open
@ -149,7 +169,9 @@ This document covers error scenarios, their causes, data at risk, and step-by-st
## Backup and Restore
tidalDB's underlying storage engine (fjall 3.x) does not yet expose a native backup API ([fjall issue #52](https://github.com/fjall-rs/fjall/issues/52)). Until that ships, the recommended backup procedure is quiesce-and-copy.
> **Standalone vs cluster.** This section is the **standalone** backup method. The **cluster** does not use quiesce-and-copy: it has a real backup/restore tool, `tidalctl backup` / `tidalctl restore` (BLAKE3-manifested, per-shard `checkpoint_seq`, proven against real S3). For cluster backups, object-store DR, and PITR, use `tidalctl` and the [disaster-recovery runbook](../runbooks/disaster-recovery.md) — not the procedure below.
For the **standalone** engine, tidalDB's underlying storage engine (fjall 3.x) does not yet expose a native backup API ([fjall issue #52](https://github.com/fjall-rs/fjall/issues/52)). Until that ships, the recommended standalone backup procedure is quiesce-and-copy.
### Creating a Backup

30
docs/runbooks/README.md Normal file
View File

@ -0,0 +1,30 @@
# tidalDB Runbooks — index
Operational documentation for the **cluster** deployment (ns `tidaldb-cluster`,
single StatefulSet, full-placement RF3, quorum-ack + automatic election). Live state
and history live in the `orchard9-k3sf` repo (`cluster-state.yaml`,
`deployments/history/tidaldb.md`).
| When you need to… | Open |
|---|---|
| **Respond to an incident / page** — symptom → signal → fix | [`on-call.md`](on-call.md) |
| Operate the cluster — API, failover, partition, membership, scale, rolling upgrade | [`cluster.md`](cluster.md) |
| Run it on Kubernetes — deploy, digest-pin, probes, scale up/down, troubleshoot pods | [`kubernetes.md`](kubernetes.md) |
| Back up / restore / DR — object-store export, restore, byte-verify, query-proof, R2, rebuild | [`disaster-recovery.md`](disaster-recovery.md) |
| Recover a single node / standalone engine — corrupt keyspace, WAL, lock, schema | [`../ops/recovery.md`](../ops/recovery.md) |
| Size a deployment — single-node tables + the measured Ref-A cluster envelope | [`../ops/capacity-planning.md`](../ops/capacity-planning.md) |
| Read the metrics / wire dashboards & alerts | [`../ops/monitoring.md`](../ops/monitoring.md), [`../ops/grafana-dashboard.json`](../ops/grafana-dashboard.json), [`../ops/prometheus-alerts.yaml`](../ops/prometheus-alerts.yaml) |
| Understand the live perf/topology findings (dev handoff) | [`../profiling/m12-cluster-deploy-findings.md`](../profiling/m12-cluster-deploy-findings.md) |
## Current production facts (2026-06-19)
- **Image:** `registry.threesix.ai/tidal/server@sha256:171505745b…` (`m12-writeburst-rc7`).
- **Shape:** 1 StatefulSet `tidaldb`, 3 pods = 3 regions = 3 voters, **full-placement RF3** (every pod hosts all 3 shard groups; gRPC 9601/9602/9603), HTTPS+mTLS on `:9500`, metrics `:9091`.
- **Guarantees live:** quorum-ack writes, automatic election/failover, elastic seed-join (learner→voter), inter-node mTLS, per-node Prometheus.
- **Shipped fixes in this image:** reseed-loop resolved, seed-join promotion, election-divergence/quarantine, read-SLA (CPU oversubscription), and the write-burst false-partition fix.
- **Known limits:** read p99 ≤10 ms to ~1000 rps (ceiling ~11.5k/s, CPU-bound); write knee ~250 rps; soak at 200 rps; 1M corpus needs >16 GiB nodes; 2.5× write-scaling needs ≥5 nodes + partitioned placement.
- **Must-watch signal:** `tidaldb_cluster_peer_breaker_state` (0/1/2) — a live peer stuck at 1 is the write-burst/partition pattern.
> Alert rules in `../ops/prometheus-alerts.yaml` are accurate but **design-reference
> — not yet loaded by the live Prometheus.** Promoting them to a `PrometheusRule` is
> the one open observability step (see [`on-call.md`](on-call.md) §3).

View File

@ -5,11 +5,27 @@ dev fabric **and** the multi-process region nodes), the operational API,
replication transport facts, failover and partition drills, and the honest
write-durability contract.
> ## STATUS: EXPERIMENTAL — TWO MODES, NEITHER IS QUORUM-ACKED HA YET
> ## STATUS: QUORUM-ACKED HA IS LIVE — RUNNING IN PRODUCTION ON k3s
>
> Cluster mode has **two shapes**, both behind the same experimental opt-in:
> Quorum-ack writes (m11p3) and automatic election/failover (m11p4) are **LIVE and
> deployed**. The reference cluster runs in production on k3s as a **single
> StatefulSet, full-placement RF3** deployment: ns `tidaldb-cluster`, `replicas: 3`,
> every pod a region (`tidaldb-0/1/2`) hosting **all three shard groups**, image
> `m12-writeburst-rc7`. ack=quorum is the **cluster deployment default** (topology
> `replication.ack: quorum`) and leader election + failover are **automatic** — kill
> the leader and the survivors elect a successor with zero operator verbs and zero
> acknowledged-write loss. See the live production topology in
> [§1](#1-launch-the-cluster-locally) and the deployed shape in
> [§3a](#3a-sharding--replication-shards-m11p6).
>
> **1. Multi-process (`--region`) — real process isolation.** Each
> **Honest caveats that still hold:** cluster mode replicates **global** retrieval
> signals only (no per-user personalization on the `/signals` route — see the
> personalization note in [§3](#3-topology-yaml)); and both launch modes still
> **refuse to start** unless you explicitly opt in with `--experimental-cluster` or
> `TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1` (the k8s manifests set the env var). The two
> launch shapes below remain:
>
> **1. Multi-process (`--region`) — real process isolation (the production shape).** Each
> `tidal-server cluster --region <name>` process owns **exactly one region**: one
> `TidalDb`, one [`GrpcTransport`](#4-grpc-replication-transport-tidal-net) whose
> server binds *this* region's `grpc_addr` and whose peers are every **sibling
@ -29,13 +45,14 @@ write-durability contract.
> the default because it needs no per-region topology addresses and no process
> orchestration.
>
> **Honest remaining limits (both modes):**
> * **Quorum durability is opt-in.** The default `204` is leader-durable
> (storage + WAL fsync; follower ship off the request path). Since m11p3,
> `ack=quorum` — topology default or per-request `x-tidal-ack` header —
> gates success on a **majority of the replica set durably holding the
> write**, surviving permanent leader loss (see
> [§8](#8-write-durability-contract-the-ack-knob-and-its-cost-m11p3)).
> **Durability / leadership / membership facts (both modes):**
> * **Quorum durability is the cluster default, per-request overridable.**
> `replication.ack: quorum` is the topology default the k8s reference cluster
> ships, so a write succeeds only once a **majority of the replica set durably
> holds it** (m11p3), surviving permanent leader loss. A caller can downgrade a
> single write to leader-durable with `x-tidal-ack: leader` (storage + WAL fsync;
> follower ship off the request path). See
> [§8](#8-write-durability-contract-the-ack-knob-and-its-cost-m11p3).
> * **Leadership is automatic** (m11p4). Every node runs a failure detector and
> a Raft-style election (pre-vote + vote + check-quorum + fenced transfer):
> kill the leader and the survivors elect a successor — typically under a
@ -52,11 +69,13 @@ write-durability contract.
> [§3](#3-topology-yaml), [§6](#6-cluster-management-api),
> [§9.1](#91-automatic-failover-m11p4--the-default)).
>
> **For a production deployment today**, run a **single `tidal-server standalone`**
> node backed by host-level redundancy and disk durability (see
> [kubernetes.md](kubernetes.md) and [server-deployment.md](../guides/server-deployment.md)),
> and reach for multi-process cluster mode for read-scale / multi-region
> deployments whose writes need `ack=quorum`'s failover-survivable contract.
> **Production deployment today** is the multi-process cluster on k3s described
> above — the single-StatefulSet full-placement RF3 reference in [`k8s/cluster/`](../../k8s/cluster/)
> (see [§1](#1-launch-the-cluster-locally), [§3a](#3a-sharding--replication-shards-m11p6),
> and [kubernetes.md](kubernetes.md)). The single-process shape below remains the
> dev/demo fabric and replication-correctness harness, and a single
> `tidal-server standalone` node (see [server-deployment.md](../guides/server-deployment.md))
> stays valid for deployments that do not need multi-region / `ack=quorum`.
>
> **Both modes refuse to start** unless you explicitly opt in with either the
> `--experimental-cluster` flag or the `TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1`
@ -93,6 +112,26 @@ Cluster mode is gated. Pass `--experimental-cluster` (or set
`TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1`) or the server exits with a mode-specific
error explaining why.
> **Live production topology (the deployed shape on k3s — 2026-06-19).** The
> reference cluster is **ONE StatefulSet `tidaldb`** in namespace
> `tidaldb-cluster`, `replicas: 3` = **3 pods = 3 regions = 3 voters**
> (`tidaldb-0/1/2`), **full-placement RF3**: every pod hosts **all three shard
> groups** (the 3-group `shards:` block in [§3a](#3a-sharding--replication-shards-m11p6)
> is the DEPLOYED shape, not optional). Each group's data lives under
> `/data/db/shard-0000N` on the pod's one PVC. Each pod binds one gRPC port per
> group — **shard 0 → 9601, shard 1 → 9602, shard 2 → 9603** (derived
> `node base port + shard id`). The HTTP plane is **`:9500` over HTTPS with
> inter-node mTLS** (m11p7 — every probe/curl uses `https://`); Prometheus
> `/metrics` is `:9091`. `replication.ack: quorum` is the deployment default and
> election/failover is automatic (`election.auto_election: true`). Live image
> `m12-writeburst-rc7`
> (`@sha256:171505745b801dcf231b531de6167dbc309a7182957811cbc2228f0a302572b1`).
> Networking is a **headless peer Service `tidaldb-peers`** (per-pod DNS
> `tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local`) plus a **ready-only
> client Service `tidaldb`** (the seed-join discovery target). See
> [`k8s/cluster/statefulset.yaml`](../../k8s/cluster/statefulset.yaml) and
> [`k8s/cluster/topology-configmap.yaml`](../../k8s/cluster/topology-configmap.yaml).
### 1a. Single-process (dev/demo default)
```bash
@ -288,17 +327,17 @@ shared file names every region by its per-pod DNS name while each pod binds
```yaml
regions:
- name: us-east
grpc_addr: "tidaldb-0.tidaldb-peers.svc.cluster.local:9601" # ADVERTISED (siblings dial; DNS re-resolved on reconnect)
grpc_addr: "tidaldb-0.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9601" # ADVERTISED (siblings dial; DNS re-resolved on reconnect)
grpc_bind: "0.0.0.0:9601" # LOCAL bind (optional; see derivation rule below)
http_addr: "10.0.1.10:9501" # this region's public HTTP gateway (forwarding + status)
metrics_addr: "10.0.1.10:9091" # optional Prometheus /metrics listener (set it in production)
- name: eu-west
grpc_addr: "tidaldb-1.tidaldb-peers.svc.cluster.local:9602"
grpc_addr: "tidaldb-1.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9602"
grpc_bind: "0.0.0.0:9602"
http_addr: "10.0.2.10:9502"
metrics_addr: "10.0.2.10:9091"
- name: ap-south
grpc_addr: "tidaldb-2.tidaldb-peers.svc.cluster.local:9603"
grpc_addr: "tidaldb-2.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9603"
grpc_bind: "0.0.0.0:9603"
http_addr: "10.0.3.10:9503"
metrics_addr: "10.0.3.10:9091"
@ -335,9 +374,19 @@ DNS-SAN certs (see `grpc_tls` below); no code change.
### 3a. Sharding × replication (`shards:`, m11p6)
Absent `shards:`, the cluster is **1 shard × RF = all regions** — one replicated
log, one elected leader, byte-for-byte everything above. Add an optional
`shards:` block to split the entity space into **S groups, each a replication
group at RF with its own elected leader**, leaders balanced across nodes. Writes
log, one elected leader, byte-for-byte everything above. Add a `shards:` block to
split the entity space into **S groups, each a replication group at RF with its
own elected leader**, leaders balanced across nodes.
> **The k3s reference cluster runs this block enabled: S=3, full placement, RF3**
> ([`k8s/cluster/topology-configmap.yaml`](../../k8s/cluster/topology-configmap.yaml)).
> Every pod replicates all three groups; tidaldb-0 leads shard 0, tidaldb-1 shard 1,
> tidaldb-2 shard 2, binding gRPC 9601/9602/9603 respectively. Full placement is
> the production shape, not a sharding experiment — a single pod loss never loses a
> group's quorum (2 of 3 survive per group) and any pod can serve a corpus-wide
> read locally.
Writes
to `/items`//`/embeddings`//`/signals` hash-route (the engine's FNV-1a router) to
the owning group's leader and replicate at RF; any gateway accepts any write.
@ -749,8 +798,8 @@ by the leader and are **quorum-commit-gated**, one change at a time. The leader
curl -X POST "$BASE/cluster/join" \
-H 'Content-Type: application/json' \
-d '{ "name": "eu-2",
"grpc_addr": "eu-2.tidaldb-peers.svc.cluster.local:9600",
"http_addr": "eu-2.tidaldb-peers.svc.cluster.local:9504" }'
"grpc_addr": "eu-2.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9600",
"http_addr": "eu-2.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9504" }'
# → { "id": 4, "role": "learner", "term": 7, "leader": "us-east", "members": [ … ] }
```
@ -1078,6 +1127,30 @@ on the heartbeat):
A snapshot-installed node always joins clean by construction (its WAL is the
leader's copy, so its tail term equals the leader's).
**Failure mode: write-burst false-partition (the headline m12 fix — rc7).**
- **SYMPTOM.** Under a sustained 1536-D `ack=quorum` ingest burst, `ack=quorum`
writes start **503-storming and do not self-heal**. The breaker gauge
`tidaldb_cluster_peer_breaker_state` shows **BOTH followers stuck at `1` (Open)**
while they are actually alive — still heartbeating, still applying. The commit
index stalls because the leader believes it cannot ship to a majority.
- **CAUSE.** A follower's CPU-heavy HNSW apply (each apply an `ef_construction=400`
insert at 1536-D) momentarily **starves its transport runtime**, so a leader ship
RPC misses the 10s request deadline. Pre-rc7 that tonic `DeadlineExceeded` was
counted as a transport failure (`record_failure`) and **opened the breaker** even
though the peer was alive and heartbeating — both followers' breakers latched
Open, commit stalled, quorum writes 503-stormed with no self-heal.
- **FIX.** Shipped in **`m12-writeburst-rc7`** (tidal-net `record_timeout`): a ship
deadline opens the breaker **only when there is no recent proof of life**
(`last_contact` stale ⇒ a genuine blackhole still opens it; `DeadlineExceeded` /
`Cancelled` route through `record_timeout`, while a genuine `Unavailable` still
opens immediately). Heuristic-only change — the commit / election / vote paths are
untouched.
- **RESPONSE if seen on an older image.** Roll the StatefulSet to **≥ rc7**
(`@sha256:171505745b801dcf231b531de6167dbc309a7182957811cbc2228f0a302572b1`).
Confirm recovery by watching `tidaldb_cluster_peer_breaker_state` clear back to
`0` once load eases and `ack=quorum` writes stop 503-ing.
To run the pre-m11p4 posture (operator-driven failover, no automatic
elections, no check-quorum step-down), set in the topology:
@ -1288,6 +1361,10 @@ it all on.
## 13. Coordinated backup / restore + point-in-time recovery (m11p8)
> **See also the dedicated DR runbook:**
> [docs/runbooks/disaster-recovery.md](disaster-recovery.md) for full
> disaster-recovery procedures (region/cluster loss, restore drills, RPO/RTO).
The building blocks: the engine's crash-consistent `create_backup`, the WAL
**archive** (`wal.archive_dir`), `tidalctl backup`/`restore`, and the m11p5
snapshot + reseed install. Under `ack=quorum`, ANY committed replica's data dir
@ -1371,11 +1448,27 @@ mixed-version window by design.
> Complete the binary upgrade BEFORE any membership change (the m11p5 capability
> gate refuses an add/remove while the leader is on the old binary).
## Performance (measured over real localhost processes)
## Performance
| Operation | SLA | Measured (p99 / typical) |
> **Read this table by era.** The `/signals`-throughput row below is the **m11p1
> signal-write benchmark** (3-byte signal writes, small payloads, measured over
> real localhost processes) — it is **NOT** the m12 1536-D production shape and
> must not be cited as the live cluster's read/write ceiling. The m12 reality on
> the 3-node k3s fleet at the 1536-D production corpus (see
> [`docs/profiling/m12-cluster-deploy-findings.md`](../profiling/m12-cluster-deploy-findings.md)):
> * **Reads:** p99 **7.9711.47 ms** at 100500 rps; G1 (p99 ≤ 10 ms) **MET at
> 100k** with **recall@10 0.9989** (clean index). Read **ceiling ~1000 rps
> clean** (~1500 rps saturated), **CPU-bound** — beyond it a node sheds/errors.
> * **Writes:** the peach mix is write-heavy and the **write knee is ~250 rps** on
> this fleet (1536-D `ack=quorum` ingest; each apply is an HNSW insert). This is
> why the soak ([§15](#15-continuous-correctness-chaos-suites--soak-m11p9)) was
> re-scoped to **200 rps** on 2026-06-19.
The legacy m11p1 signal-write benchmark (measured over real localhost processes):
| Operation | SLA | Measured (p99 / typical) — **m11p1 signal-write benchmark, pre-m12, NOT the 1536-D shape** |
|-----------|-----|--------------------------|
| Replicated `/signals` throughput (m11p1, 3 nodes, release build) | ≥ 2,000/s | **4,534 signal-writes/s** within SLO on the ramp (knee ~5.5k/s); **2,739/s sustained 10 min** (1.65M writes, 0.35% errors); was ~90/s pre-m11p1 |
| Replicated `/signals` throughput (m11p1, 3 nodes, release build) | ≥ 2,000/s | **4,534 signal-writes/s** within SLO on the ramp (knee ~5.5k/s); **2,739/s sustained 10 min** (1.65M writes, 0.35% errors); was ~90/s pre-m11p1. **Superseded for production sizing by the m12 1536-D figures above.** |
| Replication lag under that load (m11p1) | < 2s | 103 events (~40ms) across the 10-min sustain; 377 events on the 5k/s ramp (follower group-commit coalescing) |
| Cross-region replication (write → follower applied) | < 2s | ~110133ms p99 (m8p10) |
| Failover (`/cluster/promote` → first successful write) | < 10s | ~3134ms |
@ -1420,14 +1513,21 @@ these on a production node.
**Soak with regression gates** (`tidal-stress`):
```
tidal-stress --target http://<gateway> --ramp "3900:3600" --mix peach \
tidal-stress --target https://<gateway> --ramp "200:3600" --mix peach \
--json-summary soak.json --max-error-pct 1 --max-p99-ms 250 --fail-on-knee
```
> **Soak runs at 200 rps (re-scoped 2026-06-19).** The peach mix is write-heavy
> and the write knee on this 3-node fleet is **~250 rps** (1536-D `ack=quorum`
> ingest), so the soak ramp is a single **measured-sustainable 200-rps** stage
> (`--ramp "200:3600"`) rather than the retired `3900:3600` constant (a pre-m12
> signal-write figure that does not hold at the 1536-D production shape). Point
> `--target` at an `https://` gateway — the `:9500` plane serves TLS.
`--fail-on-knee` (built-in SLO), `--max-p99-ms`, and `--max-error-pct` make the
run exit non-zero on a regression; `--json-summary` writes a machine-readable
per-stage roll-up for trend lines. A bounded version runs nightly; the GA-bar
1-hour 100k-DAU soak points `--target` at the live Ref-A cluster (`--ramp 3900:3600`).
100k-DAU soak points `--target` at the live Ref-A cluster at 200 rps.
**Nightly CI** (`.woodpecker.yaml`, cron `nightly` — Woodpecker, never GitHub
Actions): the chaos suites with elevated kill-points (`TIDAL_QUORUM_KILLPOINTS`,
@ -1442,6 +1542,9 @@ day's correctness or performance regression. The guarantee→test map is
multi-region cluster reference in [`k8s/cluster/`](../../k8s/cluster/) — one
StatefulSet + headless Service peer discovery + PDB, with `--seed`-based scale
and `kubectl delete pod` node-replace, shipped in m11p5).
- **Disaster recovery** — [docs/runbooks/disaster-recovery.md](disaster-recovery.md)
(backup/restore/PITR procedures, region/cluster loss, RPO/RTO; the operational
companion to [§13](#13-coordinated-backup--restore--point-in-time-recovery-m11p8)).
- **Server deployment guide** — [docs/guides/server-deployment.md](../guides/server-deployment.md)
(standalone and cluster launch, config, env, health probes).
- **Monitoring & alerts** — [docs/ops/monitoring.md](../ops/monitoring.md)

View File

@ -0,0 +1,183 @@
# tidalDB Disaster Recovery Runbook
Object-store backup, restore, byte-verification, and the boot-on-restore query-proof
for the **cluster** deployment (ns `tidaldb-cluster`). For single-node/standalone
recovery (corrupt keyspace, stale lock, quiesce-and-copy) see
[`../ops/recovery.md`](../ops/recovery.md). For the live operational API see
[`cluster.md`](cluster.md) and [`kubernetes.md`](kubernetes.md).
> **Status (2026-06-19): PROVEN against real S3.** The full
> export → restore → byte-verify → **query-proof** round-trip is green against a
> real S3 server (self-hosted MinIO on the fleet). Production Cloudflare R2 is the
> *same* `tidalctl` code path — only the endpoint URL and the token differ (see
> [R2](#production-r2)).
---
## What DR protects, and the recovery posture
- **Unit of backup = a per-shard data dir at rest.** The cluster is full-placement
RF3 (every pod hosts all 3 shard groups under `/data/db/shard-0000N`). A backup
captures one shard's data dir (fjall keyspaces + `wal/` + `vector/` HNSW graphs +
the `checkpoint.meta` marker). Back up each of the 3 shards.
- **Integrity = BLAKE3 manifest + per-shard `checkpoint_seq`.** `tidalctl backup`
writes `BACKUP_MANIFEST.json` (per-file BLAKE3 + `checkpoint_seq`) and uploads it
**last** as the atomicity marker. `tidalctl restore` verifies every file's BLAKE3
before writing a byte and refuses a non-empty target or a path-traversal manifest.
- **RPO/RTO posture (honest):** backup operates on a data dir **at rest** (a
drained/stopped node or a streamed point-in-time copy), so the snapshot is
trivially consistent. There is **no arbitrary point-in-time recovery yet** — the
WAL-archive primitive exists but `tidalctl replay --until <seq>` is not shipped
(see [PITR](#pitr-posture)). Restore is offline (boot a server on the restored
dir); RTO is dominated by HNSW load (~5 min/100k single-core, or instant if the
persisted graph restored cleanly).
---
## DR infrastructure (live)
Self-hosted MinIO is the real S3 backend (it is a real S3 server, **not** a mock).
Manifests live in the **orchard9-k3sf** repo:
| Artifact | Path (orchard9-k3sf) | Purpose |
|---|---|---|
| MinIO | `deployments/k8s/base/tidaldb-dr/minio.yaml` | S3 backend, ns `tidaldb-dr`, ClusterIP `minio.tidaldb-dr.svc:9000`, 5Gi PVC |
| MinIO secret | `scripts/dr-minio-secret.sh` | generates `minio-creds` (`MINIO_ROOT_USER`/`MINIO_ROOT_PASSWORD`) out-of-band; never committed |
| Byte-verify drill | `deployments/k8s/base/tidaldb-dr/dr-drill-job.yaml` | export → restore → `sha256` byte-equivalence |
| Query-proof drill | `deployments/k8s/base/tidaldb-dr/dr-queryproof-job.yaml` | seed → backup → restore → **boot a server on the restore → serve** |
The drill image `registry.threesix.ai/tidal/tidalctl:m12-dr-qp` bundles `tidalctl`
+ `tidal-server` + `mc` (built from `docker/release/dr.Dockerfile`).
---
## Procedure 1 — Back up a shard to object storage
`tidalctl` reads `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` from the env and
forces path-style addressing (region is the placeholder `"auto"`). Run it where it
can reach both the data dir and the endpoint (an in-cluster Job for the live shards;
the live RWO PVCs stay attached to the running StatefulSet, so back up a streamed
point-in-time copy, not the attached volume).
```bash
AWS_ACCESS_KEY_ID=$MINIO_ROOT_USER AWS_SECRET_ACCESS_KEY=$MINIO_ROOT_PASSWORD \
tidalctl backup \
--path /src/db/shard-00000 \
--out /work/backup \
--s3-endpoint http://minio.tidaldb-dr.svc:9000 \
--s3-bucket tidaldb-dr \
--s3-prefix shard0
```
Writes a local BLAKE3-manifested artifact **and** mirrors every object to
`s3://tidaldb-dr/shard0/`, uploading `BACKUP_MANIFEST.json` last. Verify the marker
landed: `mc ls --recursive dr/tidaldb-dr/shard0 | sort -k4` (manifest timestamp must
be after every data file). Repeat with `--s3-prefix shard1`/`shard2` for the other
two shards.
## Procedure 2 — Restore + byte-verify
```bash
AWS_ACCESS_KEY_ID=$MINIO_ROOT_USER AWS_SECRET_ACCESS_KEY=$MINIO_ROOT_PASSWORD \
tidalctl restore \
--path /target/restored \
--s3-endpoint http://minio.tidaldb-dr.svc:9000 \
--s3-bucket tidaldb-dr \
--s3-prefix shard0
```
Restore downloads the prefix, **BLAKE3-verifies every file**, `safe_join`s each path
(rejects `..`/absolute/backslash), and refuses a non-empty target. Confirm:
```bash
tidalctl status --path /target/restored # status:ok, checkpoint_seq preserved
# sha256 byte-equivalence vs source (the drill does this for all files)
```
The only expected `diff -r` deltas are empty scaffold dirs (`keyspaces/{2,3}/tables`)
— they carry no data and are correctly excluded from the file manifest.
## Procedure 3 — Query-proof (boot a server on the restore and SERVE)
Byte-equivalence proves the *bytes*; this proves the restore *serves*. A raw
point-in-time stream of a LIVE pod is **not** a standalone-bootable fjall root
(`fjall recovering … No such file or directory`); a consistency-clean source comes
from a **graceful shutdown** (SIGTERM → checkpoint + fsync + persist HNSW graphs —
the only supported way; there is no online checkpoint CLI/HTTP). The
`dr-queryproof-job.yaml` does this end-to-end:
```
seed (100 items via API) → SIGTERM (clean checkpoint) → tidalctl backup → minio
→ tidalctl restore → tidal-server standalone --data-dir <restore> → assert it serves:
/health items == N · /vector_search non-empty · /search?query=<tok> == entity · /feed non-empty
```
Run + read the verdict:
```bash
export KUBECONFIG=~/.kube/orchard9-k3sf.yaml
kubectl -n tidaldb-dr delete job dr-queryproof --ignore-not-found
kubectl -n tidaldb-dr apply -f deployments/k8s/base/tidaldb-dr/dr-queryproof-job.yaml
kubectl -n tidaldb-dr logs job/dr-queryproof # last line: "DR QUERY-PROOF: PASS"
```
---
## Production R2
Identical `tidalctl` code path — swap the endpoint and supply an R2 token:
```bash
AWS_ACCESS_KEY_ID=<r2-access-key> AWS_SECRET_ACCESS_KEY=<r2-secret> \
tidalctl backup --path <shard> --out /work/backup \
--s3-endpoint https://<account-id>.r2.cloudflarestorage.com \
--s3-bucket tidaldb-dr --s3-prefix shard0
```
> **Blocker:** this needs an **account-level** Cloudflare token with
> *Workers R2 Storage: Edit* (account `c75bf009a23e747ef34b776ae8095dc4`). The
> zone/DNS-scoped `THREESIX_CLOUDFLARE_API_TOKEN` in the shell **cannot** mint R2
> S3 keys (auth error 9109) — it is the one external dependency for live-R2 DR.
---
## Full-cluster rebuild from backup
1. Provision the cluster fresh: `kubectl apply -k k8s/cluster/` (creates ns,
ConfigMaps, certs, Services, StatefulSet). Do **not** seed yet.
2. For each of the 3 shards, `tidalctl restore` the latest backup into the
corresponding pod's `/data/db/shard-0000N` **before** the server opens it (use an
init step or restore onto the PVC offline). The restore must land a consistency-
clean dir (a graceful-shutdown backup, not a live stream).
3. Bring the StatefulSet up; each pod opens its restored shards, the WAL
`checkpoint.meta` skips replay of already-checkpointed events, and the cluster
elects a leader and converges. Verify `3/3 Ready`, single leader, and
`/cluster/status` reports the expected `checkpoint_seq` per shard.
4. If only one pod's data is lost, do **not** restore from object storage — delete
that pod's PVC + pod and let it reseed fresh from the live quorum (snapshot
install → converges lag=0). Object-store restore is for total loss.
---
## PITR posture
- **What exists:** the WAL is gap-free and archived; `tidalctl backup` fences a
consistent `checkpoint_seq`; restore replays the WAL from the checkpoint forward.
- **What does NOT exist yet:** arbitrary point-in-time recovery to a chosen seq
(`tidalctl replay --until <seq>` is not shipped). Recovery lands you at the
backup's `checkpoint_seq` + whatever WAL was captured, not an arbitrary instant.
- **Cadence:** run the byte-verify + query-proof drills on every release and on a
scheduled cadence; record the result in `orchard9-k3sf/deployments/history/tidaldb.md`.
---
## Quick reference
| Need | Command / pointer |
|---|---|
| Back up a shard | `tidalctl backup --path <shard> --out <dir> --s3-endpoint … --s3-bucket … --s3-prefix …` |
| Restore + verify | `tidalctl restore --path <fresh> --s3-endpoint … --s3-bucket … --s3-prefix …` then `tidalctl status --path <fresh>` |
| Prove it serves | apply `dr-queryproof-job.yaml`, read `DR QUERY-PROOF: PASS` |
| One pod lost | delete its PVC+pod → reseeds from quorum (no object-store restore) |
| Total loss | rebuild fresh + `tidalctl restore` each shard before boot |
| Live R2 | same path + account-level R2 token (the one external blocker) |

View File

@ -233,6 +233,42 @@ IP is reachable with no peer restart), while `grpc_bind` is the local socket
(`0.0.0.0:9601`). One topology ConfigMap names all three regions by their stable
pod DNS, and every pod mounts it unmodified.
### Shard layout: full placement (every pod hosts all three groups)
The deployed shape is **3 shard groups, RF3, full placement** — the `shards:`
block in the topology ConfigMap is enabled (an absent block would mean one group,
RF = all pods). Every pod replicates **all three** groups; leadership balances
one-per-pod (tidaldb-0 leads shard 0, tidaldb-1 shard 1, tidaldb-2 shard 2) and
the entity space is hash-partitioned across them (~even, ≈⅓ each).
- **Per-shard data dir:** each group's WAL + checkpoints + per-shard reseed marker
live under `/data/db/shard-00000`, `/data/db/shard-00001`, `/data/db/shard-00002`
in the one PVC. (Boot-install and divergent-suffix detection are per-shard — a
divergent group heals from *its own* leader.)
- **One gRPC port per group:** `replicas[].grpc_addr` is omitted in the `shards:`
block, so each bind is **derived as `node base port + shard id`** — shard 0 →
9601, shard 1 → 9602, shard 2 → 9603. The StatefulSet declares all three
containerPorts (`grpc`/`grpc-1`/`grpc-2`) for clarity and NetworkPolicy; the bind
itself is driven by the topology. Collapse back to a single `grpc` port only if
the `shards:` block is removed (legacy single group).
> The m12 findings flagged the **m12p4 3-shard catch-up + cross-shard read** layer
> as not yet production-ready on real k3s (rc7); those five bugs were root-caused
> and **fixed in rc8/rc9**, and the LIVE rc7 image below carries the reseed-loop,
> seed-join-promotion, election-divergence, read-SLA, and write-burst fixes.
### Resources at 1536-D
The 1536-dim production shape sets the pod resources, deliberately:
| Field | Value | Why |
|-------|-------|-----|
| `limits.cpu` | `"3"` (raised from 2) | The cgroup CPU quota is what the engine reads for `available_parallelism()` (SEARCH_GATE / worker-thread sizing). At `2` a cross-shard search burst **starved the async reactor + the election/heartbeat/apply control plane** — reads hung to the 30 s route timeout and the starved control plane churned elections into reseed self-exit. `3` leaves ~1 core for kubelet/system on the 4-core nodes. |
| `requests.cpu` | `500m` | Kept low so the pod still schedules (server nodes alloc ≈ 3). |
| `limits.memory` | `4Gi` | A 100k×1536-D HNSW load peaks **~1.9Gi**. 1M needs more headroom — plan **>16Gi nodes** for the 1M gate. |
| `startupProbe.failureThreshold` | `240` (~20 min) | HNSW rebuild/load at 1536-D is CPU-bound: **~5 min single-core at 100k**, so the startup budget must cover a cold rebuild (plus headroom for the 1M gate). |
| probe `scheme` | `HTTPS` (all three) | The `:9500` HTTP plane serves TLS (inter-node mTLS, m11p7). kubelet does not verify the server cert for httpGet probes, so the cert's DNS-only SANs are fine. |
### What's in `k8s/cluster/`
| File | Purpose |
@ -240,7 +276,7 @@ pod DNS, and every pod mounts it unmodified.
| `namespace.yaml` | The `tidaldb-cluster` namespace (mutually exclusive with `tidaldb`) |
| `topology-configmap.yaml` | The ONE bootstrap topology shared by all pods: 3 regions by per-pod DNS (`grpc_addr` advertised, `grpc_bind` `0.0.0.0`), `replication.ack: quorum`, `replication.reseed_self_restart: true`, the election block |
| `schema-configmap.yaml` | The schema YAML every region loads (`--schema`) |
| `statefulset.yaml` | `replicas: 3`, `podManagementPolicy: Parallel`, `TIDAL_REGION` from `POD_NAME`, durable PVC at `/data` with `--data-dir /data/db`, the three probes (readiness now cluster-aware), uid 10001, `terminationGracePeriodSeconds: 60`, topology spread, and the ordinal-branching scale-up wrapper |
| `statefulset.yaml` | `replicas: 3`, `podManagementPolicy: Parallel`, `TIDAL_REGION` from `POD_NAME`, durable PVC at `/data` with `--data-dir /data/db`, the three probes (readiness now cluster-aware, all `scheme: HTTPS`), uid 10001, `terminationGracePeriodSeconds: 600` with `TIDAL_SHUTDOWN_DRAIN_MS=3000`, the full-placement 3-shard layout (one gRPC port per group), 1536-D resources (`cpu: "3"`, `memory: 4Gi`), topology spread, and the ordinal-branching scale-up wrapper |
| `service-peers.yaml` | Headless peer Service (`publishNotReadyAddresses: true`) — stable per-pod DNS, keeps not-ready joiners resolvable for peers |
| `service-client.yaml` | Client Service — readiness-gated, drops not-ready/joining/quarantined pods from load balancing |
| `poddisruptionbudget.yaml` | `maxUnavailable: 1` — a 3-voter cluster keeps quorum across one disruption |
@ -285,11 +321,22 @@ gauges.
### Deploy the cluster
```bash
# 1. Build/publish an image whose runtime user is uid 10001 (matches the
# securityContext so the PVC is writable). One image serves every subcommand.
docker build -f docker/deploy/Dockerfile -t <registry>/tidaldb:<tag> .
docker push <registry>/tidaldb:<tag>
# Set image: in k8s/cluster/statefulset.yaml (pin by @sha256 in production).
# 1. Build/publish the release image (runtime user uid 10001 to match the
# securityContext so the PVC is writable; one image serves every subcommand).
# Use the release script — NOT a bare `docker build`:
./scripts/build-release.sh <tag> server
# It HOST cross-compiles macOS-arm64 -> x86_64-unknown-linux-gnu
# (GCC 15.2 / glibc 2.41), then packages a `debian:trixie-slim` runtime
# (the binary needs `libmvec.so.1` — ABSENT on bookworm — plus libstdc++6 /
# libgcc-s1) via the `amd64builder` buildx builder (QEMU), and pushes to
# registry.threesix.ai.
#
# CRITICAL — pin the linux/amd64 PLATFORM manifest digest, NOT the OCI index
# digest and NOT the `unknown/unknown` attestation manifest. Verify before pin:
docker buildx imagetools inspect registry.threesix.ai/tidal/server:<tag> --raw
# Expect mediaType application/vnd.oci.image.manifest.v1+json, ~6 layers,
# architecture amd64. (Pinning the index or attestation digest yields an
# ImagePullBackOff or a no-arch pull.)
# 2. Namespace + the credentials secret (stress/Ref-A shape: name
# tidaldb-credentials, key TIDAL_API_KEY). Same key on EVERY pod and client.
@ -310,6 +357,41 @@ for p in tidaldb-0 tidaldb-1 tidaldb-2; do
done
```
#### Upgrade by digest (RollingUpdate)
The LIVE image is
`registry.threesix.ai/tidal/server:m12-writeburst-rc7`
(`@sha256:171505745b801dcf231b531de6167dbc309a7182957811cbc2228f0a302572b1`).
It carries the **reseed-loop, seed-join-promotion, election-divergence, read-SLA,
and write-burst** fixes (all shipped). To roll a new build:
```bash
# Set the image by its amd64 platform-manifest DIGEST (verified above), then watch
# the RollingUpdate. The StatefulSet rolls highest-ordinal-first, one pod at a time.
kubectl set image statefulset/tidaldb \
tidaldb=registry.threesix.ai/tidal/server@sha256:<amd64-digest> \
-n tidaldb-cluster
kubectl -n tidaldb-cluster rollout status statefulset/tidaldb --timeout=600s
```
Then verify **3/3 Ready** and check each pod's **boot-reseed outcome**:
```bash
kubectl -n tidaldb-cluster get pods -l app.kubernetes.io/name=tidaldb
for p in tidaldb-0 tidaldb-1 tidaldb-2; do
kubectl -n tidaldb-cluster exec "$p" -- \
curl -s localhost:9500/cluster/status/local | jq '{role, term, lag_events, reseed_required}'
done
```
The ideal boot-reseed outcome is `[(0,NotNeeded),(1,NotNeeded),(2,NotNeeded)]`
every pod's PVC-retained WAL caught up via stream, no reseed. A **snapshot-install
fallback** (a pod rejoined behind WAL retention and reseeded via snapshot) is
**SAFE, not a failure**: it converges to `lag=0` on its own; readiness simply
stays 503 until the install first-converges. Only a pod that *loops*
(quarantine → self-restart → re-detect) is a real problem — see
[disaster-recovery.md](disaster-recovery.md).
### Scale up (3 → N): seed-join as a learner, auto-promote
`kubectl scale` is the whole story — **no topology edits**. Pods with ordinal
@ -388,7 +470,18 @@ marker on demand.
upgrade) take at most one pod at a time, preserving quorum (2 of 3). On SIGTERM a
pod flips readiness to 503 (leaving both Services), drains, lets the leader
lease/heartbeat windows lapse so a successor is elected cleanly, then checkpoints
+ fsyncs the WAL; `terminationGracePeriodSeconds: 60` covers the sequence.
+ fsyncs the WAL **and saves every hosted shard's HNSW graph** before exit.
`terminationGracePeriodSeconds: 600` (raised from 60 in m12p6) covers that
sequence. At 1536-D the long pole is the graph save: each pod hosts all three
shard groups, and a single group's USearch serialize+fsync (~32k vectors/slot) is
slow, so the three concurrent saves need ample budget or k8s SIGKILLs mid-save and
the next boot rebuilds from raw vectors (a ~5 min single-core stall at 100k).
`TIDAL_SHUTDOWN_DRAIN_MS=3000` shortens the post-SIGTERM in-flight drain (from the
15 s default) so the save *starts* promptly inside the grace window; a clean save
typically finishes in well under a minute, so 600 s is a ceiling, not the norm —
but it makes rolling restarts on a loaded cluster slow, since each pod saves
before it exits.
### The exit-gate harness
@ -403,4 +496,5 @@ run the capacity ramp, the quorum-throughput gate (`stress-job-t2a.yaml`,
- [Server deployment guide](../guides/server-deployment.md) — config, auth, the served OpenAPI spec
- [Build a feed app](../guides/build-a-feed-app.md) — what to run against this server
- [Cluster runbook](cluster.md) — the multi-region mode's operational API (launch, promote, heal, reseed)
- [Disaster recovery](disaster-recovery.md) — backup, restore, and DR for the cluster (snapshot/restore, behind-compaction reseed, divergence recovery)
- [Monitoring](../ops/monitoring.md) · [Capacity planning](../ops/capacity-planning.md) · [Recovery](../ops/recovery.md)

92
docs/runbooks/on-call.md Normal file
View File

@ -0,0 +1,92 @@
# tidalDB On-Call & Incident Response
First page to open during a tidalDB incident. Routes a symptom → the golden signal
that confirms it → the runbook that fixes it. For the operational API see
[`cluster.md`](cluster.md); for k8s mechanics [`kubernetes.md`](kubernetes.md); for
backup/restore [`disaster-recovery.md`](disaster-recovery.md).
**Ownership:** the tidalDB maintainers (source: `git.threesix.ai/jordan/tidaldb`).
The live cluster runs in `orchard9-k3sf` (ns `tidaldb-cluster`); deployment history
is `orchard9-k3sf/deployments/history/tidaldb.md`; live state is
`orchard9-k3sf/cluster-state.yaml`.
---
## 0. First 60 seconds — orient
```bash
export KUBECONFIG=~/.kube/orchard9-k3sf.yaml
kubectl get pods -n tidaldb-cluster -l app.kubernetes.io/name=tidaldb -o wide # 3/3 Ready? restarts?
# health + leader (HTTPS + mTLS plane), via port-forward:
kubectl port-forward -n tidaldb-cluster tidaldb-0 19500:9500 >/dev/null 2>&1 &
curl -fsSk https://127.0.0.1:19500/health # ok:true + agreed leader across pods
curl -fsSk https://127.0.0.1:19500/cluster/status # per-shard leaders, commit index, quarantine, reseed
```
Golden signals (Grafana "tidalDB Overview" → Cluster Replication row, or `/metrics`):
| Signal | Metric | Healthy |
|---|---|---|
| Per-peer circuit breaker | `tidaldb_cluster_peer_breaker_state` | 0 (closed). **1 (open) on a live peer = the write-burst/partition pattern.** |
| Commit progress / quorum lag | `relay_durable_seq` vs `relay_last_seq` | lag bounded, advancing under load |
| Reseed pending | `reseed_required` | 0 |
| Election churn | `leader_changes_total`, `election_term` | flat |
| Quorum timeouts | `quorum_timeouts_total` | flat (a rising rate = writes can't reach quorum) |
| Divergence | `divergence_quarantined` | 0 |
---
## 1. Severity
| SEV | Definition | Examples |
|---|---|---|
| **SEV1** | Writes unavailable or acked data at risk | quorum lost (≥2 voters down or both breakers stuck open), divergent prefix served, total cluster loss |
| **SEV2** | Degraded but serving | one voter down (still 2/3 quorum), read p99 > 10 ms sustained, a pod CrashLooping, write-burst 503s on an old image |
| **SEV3** | No user impact yet | single transient reseed, elevated election term, one breaker briefly open then clears, soak night failed |
---
## 2. Symptom → signal → runbook
| Symptom | Confirm with | Action |
|---|---|---|
| **Pod CrashLooping** ("reseed marker latched") | `kubectl logs`, `reseed_required` | Reseed-loop is resolved on the live image (rc5+); a genuine reseed snapshot-installs and converges lag=0. If looping on an OLD image → roll to the live digest. [`cluster.md` reseed] |
| **Writes 503 / quorum lost** | `/cluster/status` (voters down?), `quorum_timeouts_total` rising | One voter down is survivable (2/3). Both followers down → honest retryable 503 until one returns. Restart/replace the down voter; check node health. [`cluster.md` §9] |
| **Write-burst 503-storm, no self-heal; both breakers stuck at 1 while peers alive** | `tidaldb_cluster_peer_breaker_state == 1` on live (heartbeating) followers | Fixed in **rc7** (write-burst false-partition). If on an older image, **roll to ≥ rc7**. Confirm the gauge clears as load eases. [`cluster.md` §9.1] |
| **Stuck leader / dual-leader view** (leader sees peers partitioned, reads 408) | `/cluster/status` disagrees across pods | Restart the leader pod to force re-election; auto-election converges. [`cluster.md` §9] |
| **Divergence / node serving a stale prefix** | `divergence_quarantined > 0` | The node auto-quarantines (503 readiness) and reseeds without wiping — it never serves a divergent prefix. Let it reseed; verify it rejoins lag=0. [`cluster.md` §9.1] |
| **Read p99 > 10 ms, reads hang** | Grafana read-latency; check load spread | Read-SLA collapse was CPU oversubscription, fixed in rc12 (SEARCH_GATE + parallel scatter + cpu limit 3). Ensure reads are spread round-robin across all 3 pods; check the ~1000-rps read ceiling isn't exceeded. [`capacity-planning.md`] |
| **OOMKilled** | pod `lastState.terminated.reason=OOMKilled` | 1536-D full placement peaks ~1.9 GiB at 100k; **1M needs > 16 GiB nodes** (≈78 GB/pod). Do not raise the limit past node headroom — it's a node-size limit. [`capacity-planning.md`] |
| **One pod's PVC lost/corrupt** | pod won't open its data dir | Delete that pod's PVC + pod → it reseeds fresh from the live quorum (snapshot install), converges lag=0. Do NOT object-store-restore for a single-pod loss. [`ops/recovery.md`] |
| **Total cluster loss** | — | Rebuild fresh + `tidalctl restore` each shard before boot. [`disaster-recovery.md`] |
| **Soak night failed** | `kubectl get jobs -n tidaldb-cluster -l app.kubernetes.io/name=tidal-soak` | A FAIL resets the 30-night streak. Read the night's log for the SLO breach (p99 > 150 ms or error > 1%) or an under-load restart. [`cluster.md` §15] |
---
## 3. Alerting status (action required)
The alert rules in [`../ops/prometheus-alerts.yaml`](../ops/prometheus-alerts.yaml)
are **accurate but DESIGN-REFERENCE — they are not yet loaded by the live
Prometheus.** The rules already cover the must-watch signals
(`TidalDBClusterBreakerOpen` on `peer_breaker_state == 1`,
`TidalDBClusterCommitIndexStall`, `TidalDBClusterElectionChurn`,
`TidalDBClusterQuorumTimeouts`, `DivergenceQuarantine`, `ReseedPending`).
**Open infra step:** promote them into the observability stack as a `PrometheusRule`
(or vmalert rule file) so they actually page. Until then, on-call watches the
Grafana Cluster Replication row manually. This is the one item that is an
observability *deployment*, not a doc fix — track it in `orchard9-k3sf`.
---
## 4. Escalation
1. Page the on-call operator → run §0, classify per §1.
2. SEV1 (acked data at risk / writes down): stop further rollouts, capture
`/cluster/status` + breaker gauge + `kubectl logs` from all 3 pods, then engage
the tidalDB maintainers. Do **not** delete PVCs or force-promote under a
divergence/quorum-loss condition without confirming the data-safety path in
[`cluster.md`](cluster.md) — the system is designed to quarantine rather than
serve bad data; let it.
3. Record the incident + resolution in
`orchard9-k3sf/deployments/history/tidaldb.md` (newest first).