tidaldb/docs/runbooks/kubernetes.md
jx12n 25296bcc5b docs: refresh ops runbooks to the live rc7 / full-placement reality
The runbooks had drifted to the retired m8/m11p5 design while all m12 production
reality (topology, perf, fixes, DR) sat only in a profiling doc no operator opens.
This promotes that reality into the runbooks and fixes the contradictions.

Contradictions fixed:
- runbooks/cluster.md: the "NEITHER IS QUORUM-ACKED HA YET" status banner was FALSE
  (quorum-ack + automatic election have been live since m11p3/p4). Rewritten to
  state the deployed reality (single-StatefulSet full-placement RF3, rc7).
- README.md: the cluster section called the HA cluster a "built-in simulated
  cluster / multi-region fabric" demo and showed promote-by-region as failover.
  Rewritten — real quorum HA, automatic failover, /cluster/promote is a maintenance
  verb. Kept the honest caveats (experimental gate, global-signals-only).

Reality promoted into the runbooks:
- Live topology (single STS, ns tidaldb-cluster, 3 voters, full-placement RF3,
  gRPC 9601/9602/9603, HTTPS+mTLS :9500), the five shipped fixes, and the real
  build+digest-pin procedure (cross-compile -> trixie -> amd64 PLATFORM manifest,
  not the index/attestation digest) in cluster.md + kubernetes.md.
- Stale constants: soak ramp 3900 -> 200 rps; cluster grace 60 -> 600s; the
  pre-m12 4.5k/s signal-write perf table annotated + the 1536-D read reality added.
- ops/capacity-planning.md: new "Ref-A 3-node fleet — measured capacity" section
  (read p99/ceiling, ~250 rps write knee, 1M needs >16GiB nodes, pod resources).
- ops/recovery.md: new cluster-recovery routing section + scoped the quiesce-and-
  copy note to standalone (the cluster uses tidalctl + the DR runbook).

New docs:
- runbooks/disaster-recovery.md: the proven S3/R2 backup -> restore -> byte-verify
  -> query-proof procedure, full-cluster rebuild, PITR posture (previously
  undocumented despite being proven against real S3).
- runbooks/on-call.md: incident response — symptom -> golden signal -> runbook,
  severity, escalation, and the open alert-wiring step.
- runbooks/README.md: the runbook index + current production facts.

Open follow-up (infra, not docs): ops/prometheus-alerts.yaml is accurate but
design-reference; promoting it to a live PrometheusRule is the one unwired step.
2026-06-19 19:53:29 -06:00

501 lines
27 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# tidalDB on Kubernetes
How to run tidalDB on Kubernetes. Two manifest sets ship in this repo, mutually
exclusive per namespace:
- **[`k8s/`](../../k8s/)** — the hardened single-node **standalone** deployment
(namespace `tidaldb`, `replicas: 1`). Apply with `kubectl apply -k k8s/`. This
is the recommended default: a `StatefulSet` with a durable volume, the three
health probes, metrics, secret-backed auth, and graceful rolling updates.
- **[`k8s/cluster/`](../../k8s/cluster/)** — the multi-region **cluster**
reference (namespace `tidaldb-cluster`, `replicas: 3`, automatic election,
quorum-ack writes, elastic membership). Apply with `kubectl apply -k
k8s/cluster/`. Still `--experimental-cluster`-gated. See
[Cluster mode on Kubernetes](#cluster-mode-on-kubernetes) below.
> ## Deploy the STANDALONE server, one replica (recommended default)
>
> tidalDB is single-node-first: the server wraps one embedded engine whose state
> (WAL + checkpoints + indexes) lives on a data dir. It scales **vertically**
> (a bigger pod), not by adding replicas — there is no shared-storage multi-writer
> mode, so `replicas: 1` in the standalone StatefulSet is load-bearing. Run one
> standalone pod, back it with a durable `PersistentVolume`, and recover from the
> WAL on restart (see [recovery](../ops/recovery.md)).
>
> **The standalone pod remains the recommended deployment.** The multi-region
> `cluster` mode is now genuinely HA — quorum-acked writes (m11p3), automatic
> election/failover (m11p4), and elastic membership (m11p5) all exist — but it is
> still `--experimental-cluster`-gated, so choose it deliberately when you need
> multi-node availability, and keep a single standalone pod when you do not. The
> cluster reference is [`k8s/cluster/`](../../k8s/cluster/); see
> [Cluster mode on Kubernetes](#cluster-mode-on-kubernetes).
## What's in `k8s/`
| File | Purpose |
|------|---------|
| `namespace.yaml` | The `tidaldb` namespace |
| `schema-configmap.yaml` | The schema YAML the server loads (`--schema`); edit for your signals |
| `statefulset.yaml` | The server: durable PVC, probes, security context, resources |
| `service.yaml` | Headless `Service` for stable DNS + in-cluster clients |
| `poddisruptionbudget.yaml` | `maxUnavailable: 0` — a drain can't silently kill the single node |
| `secret.example.yaml` | Template for the API-key secret (create the real one out-of-band) |
| `servicemonitor.yaml` | Optional Prometheus-Operator scrape config (apply separately) |
| `kustomization.yaml` | Ties the core resources together for `kubectl apply -k` |
## Prerequisites
- A Kubernetes cluster (1.25+) and `kubectl` pointed at it. For local testing,
[`kind`](https://kind.sigs.k8s.io/) is used in the walkthrough below.
- A container registry the cluster can pull from (for real clusters), or a local
image loaded into the node (for `kind`). The image is built from
[`docker/deploy/Dockerfile`](../../docker/deploy/Dockerfile).
- A default `StorageClass` (for dynamic `PersistentVolumeClaim` provisioning).
`kind`, GKE, EKS, and AKS all ship one.
## Deploy
### 1. Build and publish the image
```bash
# From the repo root — the build context must be the workspace root.
docker build -f docker/deploy/Dockerfile -t <registry>/tidaldb:<tag> .
docker push <registry>/tidaldb:<tag>
```
Set that reference in `k8s/statefulset.yaml` (`image:`), pinned by digest in
production (`@sha256:...`).
### 2. Create the namespace and the API-key secret
The secret is deliberately **not** in the kustomization so no key lands in git.
Create it directly:
```bash
kubectl create namespace tidaldb
kubectl -n tidaldb create secret generic tidaldb-api-key \
--from-literal=api-key="$(openssl rand -hex 32)"
```
In production, manage it with External Secrets Operator, Sealed Secrets, or
Vault Agent instead. The StatefulSet injects it as `TIDAL_API_KEY` — clients
then send `Authorization: Bearer <key>` on every data route. **If the secret is
empty the server runs unauthenticated and logs a WARN — never do that on a
shared network.**
### 3. Edit the schema (optional)
`k8s/schema-configmap.yaml` carries the schema the server loads. Edit it to model
your signals, text fields, embedding slots, and (optionally) ranking profiles —
the format is documented in [server-deployment.md](../guides/server-deployment.md).
The schema is read once at boot; roll the StatefulSet to apply changes.
### 4. Apply
```bash
kubectl apply -k k8s/
kubectl -n tidaldb rollout status statefulset/tidaldb --timeout=180s
```
### 5. Verify
```bash
kubectl -n tidaldb port-forward statefulset/tidaldb 9400:9400 &
KEY=$(kubectl -n tidaldb get secret tidaldb-api-key -o jsonpath='{.data.api-key}' | base64 -d)
curl -s localhost:9400/health # {"ok":true,...}
curl -s localhost:9400/openapi.json | jq .info # served API contract
curl -s -X POST localhost:9400/items \
-H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
-d '{"entity_id":1,"metadata":{"title":"hello","category":"demo","created_at":"1700000000"}}'
curl -s -H "Authorization: Bearer $KEY" "localhost:9400/feed?profile=trending&limit=5"
```
## How the health probes map
The server exposes three unauthenticated endpoints, wired to the three probe
types in `statefulset.yaml`:
| Probe | Endpoint | Behavior |
|-------|----------|----------|
| `startupProbe` | `GET /health/startup` | 200 once the HTTP listener is up; high `failureThreshold` covers slow WAL replay / index load on large data dirs (see [capacity-planning](../ops/capacity-planning.md)) |
| `livenessProbe` | `GET /health/live` | 200 while the process is alive; restart if it stops answering |
| `readinessProbe` | `GET /health` | 200 ready / **503 while draining** — on SIGTERM the pod leaves the Service endpoints before it stops accepting |
## Rolling updates and graceful shutdown
On `kubectl rollout restart` (or any pod delete), the kubelet sends SIGTERM. The
server flips readiness to 503 (so it leaves the Service), drains in-flight
requests, then checkpoints and fsyncs the WAL before exit.
`terminationGracePeriodSeconds: 60` gives that room — raise it if your data dir
is large. Because there is one replica, a restart is a brief planned outage
while the new pod replays the WAL; the `PodDisruptionBudget` (`maxUnavailable:
0`) prevents an *involuntary* drain from taking the node down without operator
intent.
## Persistence and backup
The `volumeClaimTemplate` provisions a `PersistentVolumeClaim` (`/data`, 10Gi by
default — size it from [capacity-planning](../ops/capacity-planning.md)). The WAL
+ checkpoints there are the source of truth and survive pod restarts. For backup
and disaster recovery (snapshotting the PVC, restoring a corrupt data dir), see
[recovery](../ops/recovery.md).
## Metrics
The pod exposes Prometheus metrics on `:9091/metrics` (**unauthenticated** — it is
not exposed by the headless Service externally; keep it cluster-internal). Scrape
it one of two ways:
- **Prometheus Operator:** `kubectl apply -f k8s/servicemonitor.yaml` (requires the
`monitoring.coreos.com` CRDs).
- **Plain Prometheus:** the pod carries `prometheus.io/scrape`, `prometheus.io/port`,
and `prometheus.io/path` annotations.
Alert rules and a dashboard ship in [`docs/ops/prometheus-alerts.yaml`](../ops/prometheus-alerts.yaml)
and [`docs/ops/grafana-dashboard.json`](../ops/grafana-dashboard.json); see
[monitoring](../ops/monitoring.md).
## Local walkthrough with `kind`
This is the exact flow used to verify the manifests end-to-end:
```bash
# 1. Create a local cluster.
kind create cluster --name tidaldb
# 2. Build the image and load it into the kind node (no registry needed).
docker build -f docker/deploy/Dockerfile -t tidaldb:deploy .
kind load docker-image tidaldb:deploy --name tidaldb
# 3. Namespace + API-key secret.
kubectl create namespace tidaldb
kubectl -n tidaldb create secret generic tidaldb-api-key \
--from-literal=api-key="$(openssl rand -hex 32)"
# 4. Apply and wait for ready.
kubectl apply -k k8s/
kubectl -n tidaldb rollout status statefulset/tidaldb --timeout=240s
# 5. Verify, then tear down.
kubectl -n tidaldb port-forward statefulset/tidaldb 9400:9400 &
curl -s localhost:9400/health
kind delete cluster --name tidaldb
```
The manifests set `image: tidaldb:deploy` with `imagePullPolicy: IfNotPresent`,
which is exactly what `kind load` + a local tag need. For a real cluster, swap in
your registry image.
> **`kind create cluster` fails with "could not find a log line that matches …
> Multi-User System"?** On Docker Desktop the node's `systemd` can die at boot
> with `Failed to create control group inotify object: Too many open files`
> (`docker logs <cluster>-control-plane` shows it). The Docker VM's inotify
> limits are too low; raise them in the VM kernel, then recreate:
>
> ```bash
> docker run --rm --privileged alpine sysctl -w fs.inotify.max_user_instances=8192
> kind delete cluster --name tidaldb && kind create cluster --name tidaldb
> ```
>
> This is a kind-on-Docker-Desktop prerequisite, unrelated to tidalDB.
## Troubleshooting
| Symptom | Likely cause |
|---------|--------------|
| Pod `Pending` | No default `StorageClass`, or the PVC can't bind — `kubectl -n tidaldb describe pvc data-tidaldb-0` |
| Pod `CrashLoopBackOff` at boot | Bad schema YAML in the ConfigMap, or a data dir from an incompatible schema — check logs; see [recovery § schema mismatch](../ops/recovery.md) |
| Pod never `Ready`, but `Running` | Readiness probe failing — `kubectl -n tidaldb logs statefulset/tidaldb`; a large data dir may need a longer `startupProbe` |
| `401 Unauthorized` on data routes | Wrong/empty `tidaldb-api-key` secret; clients must send `Authorization: Bearer <key>` |
| Writes lost after restart | Data dir not on the PVC — confirm `--data-dir /data` and the `data` volume mount |
## Cluster mode on Kubernetes
The cluster reference is [`k8s/cluster/`](../../k8s/cluster/): **ONE `StatefulSet`
named `tidaldb`, `replicas: 3`**, in its own namespace `tidaldb-cluster`. Each
pod is a region; the three pods (`tidaldb-0/1/2`) form the initial voter set.
This is real HA — automatic election/failover, quorum-acked writes, and
membership changes that ride the replicated log — but the mode is still
`--experimental-cluster`-gated (set via `TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1` in
the manifest). It is **mutually exclusive** with the standalone set per namespace:
they share the StatefulSet name `tidaldb`, and the standalone set's `replicas: 1`
is load-bearing. Deploy one or the other.
### Why one StatefulSet (not one per region)
Before m11p5, a pod bound its own `grpc_addr` literally, and a pod cannot bind a
Service ClusterIP — so each region needed a per-pod topology variant, defeating
the "every process parses the same file" contract. The m11p5 **bind/advertise
split** removes that: `grpc_addr` is the address peers *dial* (a per-pod headless
DNS name, re-resolved by tonic on every reconnect, so a rescheduled pod on a new
IP is reachable with no peer restart), while `grpc_bind` is the local socket
(`0.0.0.0:9601`). One topology ConfigMap names all three regions by their stable
pod DNS, and every pod mounts it unmodified.
### Shard layout: full placement (every pod hosts all three groups)
The deployed shape is **3 shard groups, RF3, full placement** — the `shards:`
block in the topology ConfigMap is enabled (an absent block would mean one group,
RF = all pods). Every pod replicates **all three** groups; leadership balances
one-per-pod (tidaldb-0 leads shard 0, tidaldb-1 shard 1, tidaldb-2 shard 2) and
the entity space is hash-partitioned across them (~even, ≈⅓ each).
- **Per-shard data dir:** each group's WAL + checkpoints + per-shard reseed marker
live under `/data/db/shard-00000`, `/data/db/shard-00001`, `/data/db/shard-00002`
in the one PVC. (Boot-install and divergent-suffix detection are per-shard — a
divergent group heals from *its own* leader.)
- **One gRPC port per group:** `replicas[].grpc_addr` is omitted in the `shards:`
block, so each bind is **derived as `node base port + shard id`** — shard 0 →
9601, shard 1 → 9602, shard 2 → 9603. The StatefulSet declares all three
containerPorts (`grpc`/`grpc-1`/`grpc-2`) for clarity and NetworkPolicy; the bind
itself is driven by the topology. Collapse back to a single `grpc` port only if
the `shards:` block is removed (legacy single group).
> The m12 findings flagged the **m12p4 3-shard catch-up + cross-shard read** layer
> as not yet production-ready on real k3s (rc7); those five bugs were root-caused
> and **fixed in rc8/rc9**, and the LIVE rc7 image below carries the reseed-loop,
> seed-join-promotion, election-divergence, read-SLA, and write-burst fixes.
### Resources at 1536-D
The 1536-dim production shape sets the pod resources, deliberately:
| Field | Value | Why |
|-------|-------|-----|
| `limits.cpu` | `"3"` (raised from 2) | The cgroup CPU quota is what the engine reads for `available_parallelism()` (SEARCH_GATE / worker-thread sizing). At `2` a cross-shard search burst **starved the async reactor + the election/heartbeat/apply control plane** — reads hung to the 30 s route timeout and the starved control plane churned elections into reseed self-exit. `3` leaves ~1 core for kubelet/system on the 4-core nodes. |
| `requests.cpu` | `500m` | Kept low so the pod still schedules (server nodes alloc ≈ 3). |
| `limits.memory` | `4Gi` | A 100k×1536-D HNSW load peaks **~1.9Gi**. 1M needs more headroom — plan **>16Gi nodes** for the 1M gate. |
| `startupProbe.failureThreshold` | `240` (~20 min) | HNSW rebuild/load at 1536-D is CPU-bound: **~5 min single-core at 100k**, so the startup budget must cover a cold rebuild (plus headroom for the 1M gate). |
| probe `scheme` | `HTTPS` (all three) | The `:9500` HTTP plane serves TLS (inter-node mTLS, m11p7). kubelet does not verify the server cert for httpGet probes, so the cert's DNS-only SANs are fine. |
### What's in `k8s/cluster/`
| File | Purpose |
|------|---------|
| `namespace.yaml` | The `tidaldb-cluster` namespace (mutually exclusive with `tidaldb`) |
| `topology-configmap.yaml` | The ONE bootstrap topology shared by all pods: 3 regions by per-pod DNS (`grpc_addr` advertised, `grpc_bind` `0.0.0.0`), `replication.ack: quorum`, `replication.reseed_self_restart: true`, the election block |
| `schema-configmap.yaml` | The schema YAML every region loads (`--schema`) |
| `statefulset.yaml` | `replicas: 3`, `podManagementPolicy: Parallel`, `TIDAL_REGION` from `POD_NAME`, durable PVC at `/data` with `--data-dir /data/db`, the three probes (readiness now cluster-aware, all `scheme: HTTPS`), uid 10001, `terminationGracePeriodSeconds: 600` with `TIDAL_SHUTDOWN_DRAIN_MS=3000`, the full-placement 3-shard layout (one gRPC port per group), 1536-D resources (`cpu: "3"`, `memory: 4Gi`), topology spread, and the ordinal-branching scale-up wrapper |
| `service-peers.yaml` | Headless peer Service (`publishNotReadyAddresses: true`) — stable per-pod DNS, keeps not-ready joiners resolvable for peers |
| `service-client.yaml` | Client Service — readiness-gated, drops not-ready/joining/quarantined pods from load balancing |
| `poddisruptionbudget.yaml` | `maxUnavailable: 1` — a 3-voter cluster keeps quorum across one disruption |
| `secret.example.yaml` | Template for the `tidaldb-credentials` / `TIDAL_API_KEY` secret (create the real one out-of-band) |
| `kustomization.yaml` | Ties it together for `kubectl apply -k k8s/cluster/` (secret excluded) |
### The two Services
- **`tidaldb-peers`** (headless, `clusterIP: None`, `publishNotReadyAddresses:
true`): gives each pod the stable DNS name
`tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local`, which the topology
advertises. `publishNotReadyAddresses: true` is load-bearing — a joiner is
*not ready* until it first converges, but peers must still resolve it to feed
it a snapshot + catch-up stream; without this the joiner could never reach a
seed to become ready (a deadlock).
- **`tidaldb`** (client-facing, VIP, readiness-gated): the address in-cluster
clients hit. Default readiness gating drops not-ready/joining/quarantined/
draining pods, so a client is never routed to a node that is still catching up.
### The readiness predicate (cluster-aware `/health`)
The `readinessProbe` stays `GET /health`, but in cluster mode it is now
predicate-driven (m11p5 §4). `/health` returns **503** when the node is:
- shutting down (SIGTERM drain — leaves both Services before it stops accepting);
- **quarantined** (an m11p4 divergent leader-acked suffix — serves status and
votes, refuses the data plane until reseeded);
- **removed / decommissioned** (a `Removed` membership record reached it, or a
voter's typed `removed` signal told it so);
- a **joiner** (seed-join learner) or **install boot** (snapshot-reseeded) that
has **not yet first-converged** — convergence means the boot catch-up pull
completed at least once AND lag fell to `≤ learner_promote_lag` (hysteresis;
**never** `lag == 0`, which an open-loop write load keeps perpetually false).
Sticky-ready after the first convergence.
A **restarted, PVC-retained voter is Ready on today's terms** — no regression for
ordinary pod restarts. `GET /cluster/status/local` surfaces the inputs
(`lag_events`, `quarantined`, `role`, `term`, `reseed_required`) for diagnosis,
plus the `tidaldb_cluster_reseed_required` and `tidaldb_cluster_divergence`
gauges.
### Deploy the cluster
```bash
# 1. Build/publish the release image (runtime user uid 10001 to match the
# securityContext so the PVC is writable; one image serves every subcommand).
# Use the release script — NOT a bare `docker build`:
./scripts/build-release.sh <tag> server
# It HOST cross-compiles macOS-arm64 -> x86_64-unknown-linux-gnu
# (GCC 15.2 / glibc 2.41), then packages a `debian:trixie-slim` runtime
# (the binary needs `libmvec.so.1` — ABSENT on bookworm — plus libstdc++6 /
# libgcc-s1) via the `amd64builder` buildx builder (QEMU), and pushes to
# registry.threesix.ai.
#
# CRITICAL — pin the linux/amd64 PLATFORM manifest digest, NOT the OCI index
# digest and NOT the `unknown/unknown` attestation manifest. Verify before pin:
docker buildx imagetools inspect registry.threesix.ai/tidal/server:<tag> --raw
# Expect mediaType application/vnd.oci.image.manifest.v1+json, ~6 layers,
# architecture amd64. (Pinning the index or attestation digest yields an
# ImagePullBackOff or a no-arch pull.)
# 2. Namespace + the credentials secret (stress/Ref-A shape: name
# tidaldb-credentials, key TIDAL_API_KEY). Same key on EVERY pod and client.
kubectl create namespace tidaldb-cluster
kubectl -n tidaldb-cluster create secret generic tidaldb-credentials \
--from-literal=TIDAL_API_KEY="$(openssl rand -hex 32)"
# 3. Apply and wait for the 3-pod voter set.
kubectl apply -k k8s/cluster/
kubectl -n tidaldb-cluster rollout status statefulset/tidaldb --timeout=300s
# 4. Confirm three reachable regions with low lag, and check each node's role.
kubectl -n tidaldb-cluster exec tidaldb-0 -- \
curl -s localhost:9500/cluster/status | jq '.regions[] | {name, lag_events, reachable}'
for p in tidaldb-0 tidaldb-1 tidaldb-2; do
kubectl -n tidaldb-cluster exec "$p" -- \
curl -s localhost:9500/cluster/status/local | jq '{role, term, membership_role, lag_events}'
done
```
#### Upgrade by digest (RollingUpdate)
The LIVE image is
`registry.threesix.ai/tidal/server:m12-writeburst-rc7`
(`@sha256:171505745b801dcf231b531de6167dbc309a7182957811cbc2228f0a302572b1`).
It carries the **reseed-loop, seed-join-promotion, election-divergence, read-SLA,
and write-burst** fixes (all shipped). To roll a new build:
```bash
# Set the image by its amd64 platform-manifest DIGEST (verified above), then watch
# the RollingUpdate. The StatefulSet rolls highest-ordinal-first, one pod at a time.
kubectl set image statefulset/tidaldb \
tidaldb=registry.threesix.ai/tidal/server@sha256:<amd64-digest> \
-n tidaldb-cluster
kubectl -n tidaldb-cluster rollout status statefulset/tidaldb --timeout=600s
```
Then verify **3/3 Ready** and check each pod's **boot-reseed outcome**:
```bash
kubectl -n tidaldb-cluster get pods -l app.kubernetes.io/name=tidaldb
for p in tidaldb-0 tidaldb-1 tidaldb-2; do
kubectl -n tidaldb-cluster exec "$p" -- \
curl -s localhost:9500/cluster/status/local | jq '{role, term, lag_events, reseed_required}'
done
```
The ideal boot-reseed outcome is `[(0,NotNeeded),(1,NotNeeded),(2,NotNeeded)]`
every pod's PVC-retained WAL caught up via stream, no reseed. A **snapshot-install
fallback** (a pod rejoined behind WAL retention and reseeded via snapshot) is
**SAFE, not a failure**: it converges to `lag=0` on its own; readiness simply
stays 503 until the install first-converges. Only a pod that *loops*
(quarantine → self-restart → re-detect) is a real problem — see
[disaster-recovery.md](disaster-recovery.md).
### Scale up (3 → N): seed-join as a learner, auto-promote
`kubectl scale` is the whole story — **no topology edits**. Pods with ordinal
`≥ 3` boot with `--seed` (the StatefulSet's ordinal-branching wrapper adds it
automatically) and learn their roster/id/term from a seed, joining as a
**learner**. The leader auto-promotes a learner to a voter once its durable mark
is within `learner_promote_lag` of the leader's frontier (or stops falling behind
for K rounds under sustained load).
```bash
# Grow to 5 voters. Pods tidaldb-3 and tidaldb-4 seed-join + auto-promote.
kubectl -n tidaldb-cluster scale statefulset/tidaldb --replicas=5
kubectl -n tidaldb-cluster rollout status statefulset/tidaldb --timeout=600s
# Watch the new members converge and promote (promotion_pending shows the lag):
kubectl -n tidaldb-cluster exec tidaldb-0 -- \
curl -s localhost:9500/cluster/members | jq '.members[] | {id, name, role}'
```
A scaled pod still mounts the shared topology ConfigMap — a `--seed` boot
**requires** the local config for the behavioral knob blocks (`replication`,
`wal`, `election`, `timeouts`, `grpc_tls`); its `regions:` list is ignored for
the seed joiner's roster (the join response is authoritative).
### Scale down (N → fewer): remove verb FIRST, then scale
Decommission a member **before** removing its pod, so the cluster stops counting
it toward quorum cleanly. Remove the highest-ordinal members (StatefulSet deletes
lowest-ordinal-last on scale-down).
```bash
# Going 5 -> 3: decommission tidaldb-4, then tidaldb-3, then scale.
# Remove verb: POST /cluster/members/remove {"region": "<name>"} (any node
# forwards to the leader; one-at-a-time, quorum-commit-gated).
kubectl -n tidaldb-cluster exec tidaldb-0 -- curl -s -X POST \
-H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
-d '{"region":"tidaldb-4"}' localhost:9500/cluster/members/remove
kubectl -n tidaldb-cluster exec tidaldb-0 -- curl -s -X POST \
-H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
-d '{"region":"tidaldb-3"}' localhost:9500/cluster/members/remove
# Each Remove is delivered to the removed peer (which flips to 503 and stops
# campaigning) before its ship cell retires. Only then is it safe to scale.
kubectl -n tidaldb-cluster scale statefulset/tidaldb --replicas=3
```
Removing pods without the remove verb leaves zombie members the cluster still
expects — never scale down before decommissioning.
### Node / pod replacement
- **`kubectl delete pod tidaldb-N` (PVC retained):** the StatefulSet recreates
the pod onto its existing PVC. It boots a follower, replays its WAL, and
catches up via the `StreamSegments` stream (boot-time + the 30 s catch-up
timer). No operator verb. This is the routine replace — already proven by the
p4 leader-kill regression (lag → 0 within ~10 s).
- **PVC + pod delete (fresh reseed):** delete the PVC *and* the pod. The new pod
comes up with an empty data dir, latches an install boot, and **reseeds via
snapshot** (`FetchSnapshot` + stream) from the leader, then rejoins clean. Use
this when the data dir is corrupt or the node fell behind a compacted leader.
Readiness stays 503 until the snapshot install first converges.
### Self-healing reseed (quarantine / behind-compaction)
`replication.reseed_self_restart: true` (set in the topology ConfigMap) makes a
node that durably latches `reseed_required` (an m11p4 divergence quarantine, or a
typed `snapshot-required` refusal) drain and exit(0); the StatefulSet restarts it
and the **boot-time install** reseeds it via snapshot, clearing the quarantine
and the divergence gauge with no operator verb and no PVC wipe. The self-restart
is **refused** (loudly, in `/cluster/status/local` + the gauge) when the
remaining voters cannot sustain quorum without this node — exiting during a
2-voter window would be a total write outage. `POST /cluster/reseed` latches the
marker on demand.
### PodDisruptionBudget and graceful shutdown
`maxUnavailable: 1` lets a voluntary disruption (node drain, autoscaler, rolling
upgrade) take at most one pod at a time, preserving quorum (2 of 3). On SIGTERM a
pod flips readiness to 503 (leaving both Services), drains, lets the leader
lease/heartbeat windows lapse so a successor is elected cleanly, then checkpoints
+ fsyncs the WAL **and saves every hosted shard's HNSW graph** before exit.
`terminationGracePeriodSeconds: 600` (raised from 60 in m12p6) covers that
sequence. At 1536-D the long pole is the graph save: each pod hosts all three
shard groups, and a single group's USearch serialize+fsync (~32k vectors/slot) is
slow, so the three concurrent saves need ample budget or k8s SIGKILLs mid-save and
the next boot rebuilds from raw vectors (a ~5 min single-core stall at 100k).
`TIDAL_SHUTDOWN_DRAIN_MS=3000` shortens the post-SIGTERM in-flight drain (from the
15 s default) so the save *starts* promptly inside the grace window; a clean save
typically finishes in well under a minute, so 600 s is a ceiling, not the norm —
but it makes rolling restarts on a loaded cluster slow, since each pod saves
before it exits.
### The exit-gate harness
The in-cluster load Jobs in [`tidal-stress/k8s/`](../../tidal-stress/k8s/) target
the cluster by per-pod DNS (`tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500`),
in namespace `tidaldb-cluster`, reading the `tidaldb-credentials` secret. They
run the capacity ramp, the quorum-throughput gate (`stress-job-t2a.yaml`,
`--ack quorum`), and the leader-kill chaos drill (`stress-job-t2b.yaml`).
## See also
- [Server deployment guide](../guides/server-deployment.md) — config, auth, the served OpenAPI spec
- [Build a feed app](../guides/build-a-feed-app.md) — what to run against this server
- [Cluster runbook](cluster.md) — the multi-region mode's operational API (launch, promote, heal, reseed)
- [Disaster recovery](disaster-recovery.md) — backup, restore, and DR for the cluster (snapshot/restore, behind-compaction reseed, divergence recovery)
- [Monitoring](../ops/monitoring.md) · [Capacity planning](../ops/capacity-planning.md) · [Recovery](../ops/recovery.md)