The runbooks had drifted to the retired m8/m11p5 design while all m12 production reality (topology, perf, fixes, DR) sat only in a profiling doc no operator opens. This promotes that reality into the runbooks and fixes the contradictions. Contradictions fixed: - runbooks/cluster.md: the "NEITHER IS QUORUM-ACKED HA YET" status banner was FALSE (quorum-ack + automatic election have been live since m11p3/p4). Rewritten to state the deployed reality (single-StatefulSet full-placement RF3, rc7). - README.md: the cluster section called the HA cluster a "built-in simulated cluster / multi-region fabric" demo and showed promote-by-region as failover. Rewritten — real quorum HA, automatic failover, /cluster/promote is a maintenance verb. Kept the honest caveats (experimental gate, global-signals-only). Reality promoted into the runbooks: - Live topology (single STS, ns tidaldb-cluster, 3 voters, full-placement RF3, gRPC 9601/9602/9603, HTTPS+mTLS :9500), the five shipped fixes, and the real build+digest-pin procedure (cross-compile -> trixie -> amd64 PLATFORM manifest, not the index/attestation digest) in cluster.md + kubernetes.md. - Stale constants: soak ramp 3900 -> 200 rps; cluster grace 60 -> 600s; the pre-m12 4.5k/s signal-write perf table annotated + the 1536-D read reality added. - ops/capacity-planning.md: new "Ref-A 3-node fleet — measured capacity" section (read p99/ceiling, ~250 rps write knee, 1M needs >16GiB nodes, pod resources). - ops/recovery.md: new cluster-recovery routing section + scoped the quiesce-and- copy note to standalone (the cluster uses tidalctl + the DR runbook). New docs: - runbooks/disaster-recovery.md: the proven S3/R2 backup -> restore -> byte-verify -> query-proof procedure, full-cluster rebuild, PITR posture (previously undocumented despite being proven against real S3). - runbooks/on-call.md: incident response — symptom -> golden signal -> runbook, severity, escalation, and the open alert-wiring step. - runbooks/README.md: the runbook index + current production facts. Open follow-up (infra, not docs): ops/prometheus-alerts.yaml is accurate but design-reference; promoting it to a live PrometheusRule is the one unwired step.
8.6 KiB
tidalDB Disaster Recovery Runbook
Object-store backup, restore, byte-verification, and the boot-on-restore query-proof
for the cluster deployment (ns tidaldb-cluster). For single-node/standalone
recovery (corrupt keyspace, stale lock, quiesce-and-copy) see
../ops/recovery.md. For the live operational API see
cluster.md and kubernetes.md.
Status (2026-06-19): PROVEN against real S3. The full export → restore → byte-verify → query-proof round-trip is green against a real S3 server (self-hosted MinIO on the fleet). Production Cloudflare R2 is the same
tidalctlcode path — only the endpoint URL and the token differ (see R2).
What DR protects, and the recovery posture
- Unit of backup = a per-shard data dir at rest. The cluster is full-placement
RF3 (every pod hosts all 3 shard groups under
/data/db/shard-0000N). A backup captures one shard's data dir (fjall keyspaces +wal/+vector/HNSW graphs + thecheckpoint.metamarker). Back up each of the 3 shards. - Integrity = BLAKE3 manifest + per-shard
checkpoint_seq.tidalctl backupwritesBACKUP_MANIFEST.json(per-file BLAKE3 +checkpoint_seq) and uploads it last as the atomicity marker.tidalctl restoreverifies every file's BLAKE3 before writing a byte and refuses a non-empty target or a path-traversal manifest. - RPO/RTO posture (honest): backup operates on a data dir at rest (a
drained/stopped node or a streamed point-in-time copy), so the snapshot is
trivially consistent. There is no arbitrary point-in-time recovery yet — the
WAL-archive primitive exists but
tidalctl replay --until <seq>is not shipped (see PITR). Restore is offline (boot a server on the restored dir); RTO is dominated by HNSW load (~5 min/100k single-core, or instant if the persisted graph restored cleanly).
DR infrastructure (live)
Self-hosted MinIO is the real S3 backend (it is a real S3 server, not a mock). Manifests live in the orchard9-k3sf repo:
| Artifact | Path (orchard9-k3sf) | Purpose |
|---|---|---|
| MinIO | deployments/k8s/base/tidaldb-dr/minio.yaml |
S3 backend, ns tidaldb-dr, ClusterIP minio.tidaldb-dr.svc:9000, 5Gi PVC |
| MinIO secret | scripts/dr-minio-secret.sh |
generates minio-creds (MINIO_ROOT_USER/MINIO_ROOT_PASSWORD) out-of-band; never committed |
| Byte-verify drill | deployments/k8s/base/tidaldb-dr/dr-drill-job.yaml |
export → restore → sha256 byte-equivalence |
| Query-proof drill | deployments/k8s/base/tidaldb-dr/dr-queryproof-job.yaml |
seed → backup → restore → boot a server on the restore → serve |
The drill image registry.threesix.ai/tidal/tidalctl:m12-dr-qp bundles tidalctl
tidal-server+mc(built fromdocker/release/dr.Dockerfile).
Procedure 1 — Back up a shard to object storage
tidalctl reads AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY from the env and
forces path-style addressing (region is the placeholder "auto"). Run it where it
can reach both the data dir and the endpoint (an in-cluster Job for the live shards;
the live RWO PVCs stay attached to the running StatefulSet, so back up a streamed
point-in-time copy, not the attached volume).
AWS_ACCESS_KEY_ID=$MINIO_ROOT_USER AWS_SECRET_ACCESS_KEY=$MINIO_ROOT_PASSWORD \
tidalctl backup \
--path /src/db/shard-00000 \
--out /work/backup \
--s3-endpoint http://minio.tidaldb-dr.svc:9000 \
--s3-bucket tidaldb-dr \
--s3-prefix shard0
Writes a local BLAKE3-manifested artifact and mirrors every object to
s3://tidaldb-dr/shard0/, uploading BACKUP_MANIFEST.json last. Verify the marker
landed: mc ls --recursive dr/tidaldb-dr/shard0 | sort -k4 (manifest timestamp must
be after every data file). Repeat with --s3-prefix shard1/shard2 for the other
two shards.
Procedure 2 — Restore + byte-verify
AWS_ACCESS_KEY_ID=$MINIO_ROOT_USER AWS_SECRET_ACCESS_KEY=$MINIO_ROOT_PASSWORD \
tidalctl restore \
--path /target/restored \
--s3-endpoint http://minio.tidaldb-dr.svc:9000 \
--s3-bucket tidaldb-dr \
--s3-prefix shard0
Restore downloads the prefix, BLAKE3-verifies every file, safe_joins each path
(rejects ../absolute/backslash), and refuses a non-empty target. Confirm:
tidalctl status --path /target/restored # status:ok, checkpoint_seq preserved
# sha256 byte-equivalence vs source (the drill does this for all files)
The only expected diff -r deltas are empty scaffold dirs (keyspaces/{2,3}/tables)
— they carry no data and are correctly excluded from the file manifest.
Procedure 3 — Query-proof (boot a server on the restore and SERVE)
Byte-equivalence proves the bytes; this proves the restore serves. A raw
point-in-time stream of a LIVE pod is not a standalone-bootable fjall root
(fjall recovering … No such file or directory); a consistency-clean source comes
from a graceful shutdown (SIGTERM → checkpoint + fsync + persist HNSW graphs —
the only supported way; there is no online checkpoint CLI/HTTP). The
dr-queryproof-job.yaml does this end-to-end:
seed (100 items via API) → SIGTERM (clean checkpoint) → tidalctl backup → minio
→ tidalctl restore → tidal-server standalone --data-dir <restore> → assert it serves:
/health items == N · /vector_search non-empty · /search?query=<tok> == entity · /feed non-empty
Run + read the verdict:
export KUBECONFIG=~/.kube/orchard9-k3sf.yaml
kubectl -n tidaldb-dr delete job dr-queryproof --ignore-not-found
kubectl -n tidaldb-dr apply -f deployments/k8s/base/tidaldb-dr/dr-queryproof-job.yaml
kubectl -n tidaldb-dr logs job/dr-queryproof # last line: "DR QUERY-PROOF: PASS"
Production R2
Identical tidalctl code path — swap the endpoint and supply an R2 token:
AWS_ACCESS_KEY_ID=<r2-access-key> AWS_SECRET_ACCESS_KEY=<r2-secret> \
tidalctl backup --path <shard> --out /work/backup \
--s3-endpoint https://<account-id>.r2.cloudflarestorage.com \
--s3-bucket tidaldb-dr --s3-prefix shard0
Blocker: this needs an account-level Cloudflare token with Workers R2 Storage: Edit (account
c75bf009a23e747ef34b776ae8095dc4). The zone/DNS-scopedTHREESIX_CLOUDFLARE_API_TOKENin the shell cannot mint R2 S3 keys (auth error 9109) — it is the one external dependency for live-R2 DR.
Full-cluster rebuild from backup
- Provision the cluster fresh:
kubectl apply -k k8s/cluster/(creates ns, ConfigMaps, certs, Services, StatefulSet). Do not seed yet. - For each of the 3 shards,
tidalctl restorethe latest backup into the corresponding pod's/data/db/shard-0000Nbefore the server opens it (use an init step or restore onto the PVC offline). The restore must land a consistency- clean dir (a graceful-shutdown backup, not a live stream). - Bring the StatefulSet up; each pod opens its restored shards, the WAL
checkpoint.metaskips replay of already-checkpointed events, and the cluster elects a leader and converges. Verify3/3 Ready, single leader, and/cluster/statusreports the expectedcheckpoint_seqper shard. - If only one pod's data is lost, do not restore from object storage — delete that pod's PVC + pod and let it reseed fresh from the live quorum (snapshot install → converges lag=0). Object-store restore is for total loss.
PITR posture
- What exists: the WAL is gap-free and archived;
tidalctl backupfences a consistentcheckpoint_seq; restore replays the WAL from the checkpoint forward. - What does NOT exist yet: arbitrary point-in-time recovery to a chosen seq
(
tidalctl replay --until <seq>is not shipped). Recovery lands you at the backup'scheckpoint_seq+ whatever WAL was captured, not an arbitrary instant. - Cadence: run the byte-verify + query-proof drills on every release and on a
scheduled cadence; record the result in
orchard9-k3sf/deployments/history/tidaldb.md.
Quick reference
| Need | Command / pointer |
|---|---|
| Back up a shard | tidalctl backup --path <shard> --out <dir> --s3-endpoint … --s3-bucket … --s3-prefix … |
| Restore + verify | tidalctl restore --path <fresh> --s3-endpoint … --s3-bucket … --s3-prefix … then tidalctl status --path <fresh> |
| Prove it serves | apply dr-queryproof-job.yaml, read DR QUERY-PROOF: PASS |
| One pod lost | delete its PVC+pod → reseeds from quorum (no object-store restore) |
| Total loss | rebuild fresh + tidalctl restore each shard before boot |
| Live R2 | same path + account-level R2 token (the one external blocker) |