tidaldb/docs/runbooks/kubernetes.md
jx12n 25296bcc5b docs: refresh ops runbooks to the live rc7 / full-placement reality
The runbooks had drifted to the retired m8/m11p5 design while all m12 production
reality (topology, perf, fixes, DR) sat only in a profiling doc no operator opens.
This promotes that reality into the runbooks and fixes the contradictions.

Contradictions fixed:
- runbooks/cluster.md: the "NEITHER IS QUORUM-ACKED HA YET" status banner was FALSE
  (quorum-ack + automatic election have been live since m11p3/p4). Rewritten to
  state the deployed reality (single-StatefulSet full-placement RF3, rc7).
- README.md: the cluster section called the HA cluster a "built-in simulated
  cluster / multi-region fabric" demo and showed promote-by-region as failover.
  Rewritten — real quorum HA, automatic failover, /cluster/promote is a maintenance
  verb. Kept the honest caveats (experimental gate, global-signals-only).

Reality promoted into the runbooks:
- Live topology (single STS, ns tidaldb-cluster, 3 voters, full-placement RF3,
  gRPC 9601/9602/9603, HTTPS+mTLS :9500), the five shipped fixes, and the real
  build+digest-pin procedure (cross-compile -> trixie -> amd64 PLATFORM manifest,
  not the index/attestation digest) in cluster.md + kubernetes.md.
- Stale constants: soak ramp 3900 -> 200 rps; cluster grace 60 -> 600s; the
  pre-m12 4.5k/s signal-write perf table annotated + the 1536-D read reality added.
- ops/capacity-planning.md: new "Ref-A 3-node fleet — measured capacity" section
  (read p99/ceiling, ~250 rps write knee, 1M needs >16GiB nodes, pod resources).
- ops/recovery.md: new cluster-recovery routing section + scoped the quiesce-and-
  copy note to standalone (the cluster uses tidalctl + the DR runbook).

New docs:
- runbooks/disaster-recovery.md: the proven S3/R2 backup -> restore -> byte-verify
  -> query-proof procedure, full-cluster rebuild, PITR posture (previously
  undocumented despite being proven against real S3).
- runbooks/on-call.md: incident response — symptom -> golden signal -> runbook,
  severity, escalation, and the open alert-wiring step.
- runbooks/README.md: the runbook index + current production facts.

Open follow-up (infra, not docs): ops/prometheus-alerts.yaml is accurate but
design-reference; promoting it to a live PrometheusRule is the one unwired step.
2026-06-19 19:53:29 -06:00

27 KiB
Raw Blame History

tidalDB on Kubernetes

How to run tidalDB on Kubernetes. Two manifest sets ship in this repo, mutually exclusive per namespace:

  • k8s/ — the hardened single-node standalone deployment (namespace tidaldb, replicas: 1). Apply with kubectl apply -k k8s/. This is the recommended default: a StatefulSet with a durable volume, the three health probes, metrics, secret-backed auth, and graceful rolling updates.
  • k8s/cluster/ — the multi-region cluster reference (namespace tidaldb-cluster, replicas: 3, automatic election, quorum-ack writes, elastic membership). Apply with kubectl apply -k k8s/cluster/. Still --experimental-cluster-gated. See Cluster mode on Kubernetes below.

tidalDB is single-node-first: the server wraps one embedded engine whose state (WAL + checkpoints + indexes) lives on a data dir. It scales vertically (a bigger pod), not by adding replicas — there is no shared-storage multi-writer mode, so replicas: 1 in the standalone StatefulSet is load-bearing. Run one standalone pod, back it with a durable PersistentVolume, and recover from the WAL on restart (see recovery).

The standalone pod remains the recommended deployment. The multi-region cluster mode is now genuinely HA — quorum-acked writes (m11p3), automatic election/failover (m11p4), and elastic membership (m11p5) all exist — but it is still --experimental-cluster-gated, so choose it deliberately when you need multi-node availability, and keep a single standalone pod when you do not. The cluster reference is k8s/cluster/; see Cluster mode on Kubernetes.

What's in k8s/

File Purpose
namespace.yaml The tidaldb namespace
schema-configmap.yaml The schema YAML the server loads (--schema); edit for your signals
statefulset.yaml The server: durable PVC, probes, security context, resources
service.yaml Headless Service for stable DNS + in-cluster clients
poddisruptionbudget.yaml maxUnavailable: 0 — a drain can't silently kill the single node
secret.example.yaml Template for the API-key secret (create the real one out-of-band)
servicemonitor.yaml Optional Prometheus-Operator scrape config (apply separately)
kustomization.yaml Ties the core resources together for kubectl apply -k

Prerequisites

  • A Kubernetes cluster (1.25+) and kubectl pointed at it. For local testing, kind is used in the walkthrough below.
  • A container registry the cluster can pull from (for real clusters), or a local image loaded into the node (for kind). The image is built from docker/deploy/Dockerfile.
  • A default StorageClass (for dynamic PersistentVolumeClaim provisioning). kind, GKE, EKS, and AKS all ship one.

Deploy

1. Build and publish the image

# From the repo root — the build context must be the workspace root.
docker build -f docker/deploy/Dockerfile -t <registry>/tidaldb:<tag> .
docker push <registry>/tidaldb:<tag>

Set that reference in k8s/statefulset.yaml (image:), pinned by digest in production (@sha256:...).

2. Create the namespace and the API-key secret

The secret is deliberately not in the kustomization so no key lands in git. Create it directly:

kubectl create namespace tidaldb
kubectl -n tidaldb create secret generic tidaldb-api-key \
  --from-literal=api-key="$(openssl rand -hex 32)"

In production, manage it with External Secrets Operator, Sealed Secrets, or Vault Agent instead. The StatefulSet injects it as TIDAL_API_KEY — clients then send Authorization: Bearer <key> on every data route. If the secret is empty the server runs unauthenticated and logs a WARN — never do that on a shared network.

3. Edit the schema (optional)

k8s/schema-configmap.yaml carries the schema the server loads. Edit it to model your signals, text fields, embedding slots, and (optionally) ranking profiles — the format is documented in server-deployment.md. The schema is read once at boot; roll the StatefulSet to apply changes.

4. Apply

kubectl apply -k k8s/
kubectl -n tidaldb rollout status statefulset/tidaldb --timeout=180s

5. Verify

kubectl -n tidaldb port-forward statefulset/tidaldb 9400:9400 &
KEY=$(kubectl -n tidaldb get secret tidaldb-api-key -o jsonpath='{.data.api-key}' | base64 -d)

curl -s localhost:9400/health                       # {"ok":true,...}
curl -s localhost:9400/openapi.json | jq .info      # served API contract
curl -s -X POST localhost:9400/items \
  -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -d '{"entity_id":1,"metadata":{"title":"hello","category":"demo","created_at":"1700000000"}}'
curl -s -H "Authorization: Bearer $KEY" "localhost:9400/feed?profile=trending&limit=5"

How the health probes map

The server exposes three unauthenticated endpoints, wired to the three probe types in statefulset.yaml:

Probe Endpoint Behavior
startupProbe GET /health/startup 200 once the HTTP listener is up; high failureThreshold covers slow WAL replay / index load on large data dirs (see capacity-planning)
livenessProbe GET /health/live 200 while the process is alive; restart if it stops answering
readinessProbe GET /health 200 ready / 503 while draining — on SIGTERM the pod leaves the Service endpoints before it stops accepting

Rolling updates and graceful shutdown

On kubectl rollout restart (or any pod delete), the kubelet sends SIGTERM. The server flips readiness to 503 (so it leaves the Service), drains in-flight requests, then checkpoints and fsyncs the WAL before exit. terminationGracePeriodSeconds: 60 gives that room — raise it if your data dir is large. Because there is one replica, a restart is a brief planned outage while the new pod replays the WAL; the PodDisruptionBudget (maxUnavailable: 0) prevents an involuntary drain from taking the node down without operator intent.

Persistence and backup

The volumeClaimTemplate provisions a PersistentVolumeClaim (/data, 10Gi by default — size it from capacity-planning). The WAL

  • checkpoints there are the source of truth and survive pod restarts. For backup and disaster recovery (snapshotting the PVC, restoring a corrupt data dir), see recovery.

Metrics

The pod exposes Prometheus metrics on :9091/metrics (unauthenticated — it is not exposed by the headless Service externally; keep it cluster-internal). Scrape it one of two ways:

  • Prometheus Operator: kubectl apply -f k8s/servicemonitor.yaml (requires the monitoring.coreos.com CRDs).
  • Plain Prometheus: the pod carries prometheus.io/scrape, prometheus.io/port, and prometheus.io/path annotations.

Alert rules and a dashboard ship in docs/ops/prometheus-alerts.yaml and docs/ops/grafana-dashboard.json; see monitoring.

Local walkthrough with kind

This is the exact flow used to verify the manifests end-to-end:

# 1. Create a local cluster.
kind create cluster --name tidaldb

# 2. Build the image and load it into the kind node (no registry needed).
docker build -f docker/deploy/Dockerfile -t tidaldb:deploy .
kind load docker-image tidaldb:deploy --name tidaldb

# 3. Namespace + API-key secret.
kubectl create namespace tidaldb
kubectl -n tidaldb create secret generic tidaldb-api-key \
  --from-literal=api-key="$(openssl rand -hex 32)"

# 4. Apply and wait for ready.
kubectl apply -k k8s/
kubectl -n tidaldb rollout status statefulset/tidaldb --timeout=240s

# 5. Verify, then tear down.
kubectl -n tidaldb port-forward statefulset/tidaldb 9400:9400 &
curl -s localhost:9400/health
kind delete cluster --name tidaldb

The manifests set image: tidaldb:deploy with imagePullPolicy: IfNotPresent, which is exactly what kind load + a local tag need. For a real cluster, swap in your registry image.

kind create cluster fails with "could not find a log line that matches … Multi-User System"? On Docker Desktop the node's systemd can die at boot with Failed to create control group inotify object: Too many open files (docker logs <cluster>-control-plane shows it). The Docker VM's inotify limits are too low; raise them in the VM kernel, then recreate:

docker run --rm --privileged alpine sysctl -w fs.inotify.max_user_instances=8192
kind delete cluster --name tidaldb && kind create cluster --name tidaldb

This is a kind-on-Docker-Desktop prerequisite, unrelated to tidalDB.

Troubleshooting

Symptom Likely cause
Pod Pending No default StorageClass, or the PVC can't bind — kubectl -n tidaldb describe pvc data-tidaldb-0
Pod CrashLoopBackOff at boot Bad schema YAML in the ConfigMap, or a data dir from an incompatible schema — check logs; see recovery § schema mismatch
Pod never Ready, but Running Readiness probe failing — kubectl -n tidaldb logs statefulset/tidaldb; a large data dir may need a longer startupProbe
401 Unauthorized on data routes Wrong/empty tidaldb-api-key secret; clients must send Authorization: Bearer <key>
Writes lost after restart Data dir not on the PVC — confirm --data-dir /data and the data volume mount

Cluster mode on Kubernetes

The cluster reference is k8s/cluster/: ONE StatefulSet named tidaldb, replicas: 3, in its own namespace tidaldb-cluster. Each pod is a region; the three pods (tidaldb-0/1/2) form the initial voter set. This is real HA — automatic election/failover, quorum-acked writes, and membership changes that ride the replicated log — but the mode is still --experimental-cluster-gated (set via TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1 in the manifest). It is mutually exclusive with the standalone set per namespace: they share the StatefulSet name tidaldb, and the standalone set's replicas: 1 is load-bearing. Deploy one or the other.

Why one StatefulSet (not one per region)

Before m11p5, a pod bound its own grpc_addr literally, and a pod cannot bind a Service ClusterIP — so each region needed a per-pod topology variant, defeating the "every process parses the same file" contract. The m11p5 bind/advertise split removes that: grpc_addr is the address peers dial (a per-pod headless DNS name, re-resolved by tonic on every reconnect, so a rescheduled pod on a new IP is reachable with no peer restart), while grpc_bind is the local socket (0.0.0.0:9601). One topology ConfigMap names all three regions by their stable pod DNS, and every pod mounts it unmodified.

Shard layout: full placement (every pod hosts all three groups)

The deployed shape is 3 shard groups, RF3, full placement — the shards: block in the topology ConfigMap is enabled (an absent block would mean one group, RF = all pods). Every pod replicates all three groups; leadership balances one-per-pod (tidaldb-0 leads shard 0, tidaldb-1 shard 1, tidaldb-2 shard 2) and the entity space is hash-partitioned across them (~even, ≈⅓ each).

  • Per-shard data dir: each group's WAL + checkpoints + per-shard reseed marker live under /data/db/shard-00000, /data/db/shard-00001, /data/db/shard-00002 in the one PVC. (Boot-install and divergent-suffix detection are per-shard — a divergent group heals from its own leader.)
  • One gRPC port per group: replicas[].grpc_addr is omitted in the shards: block, so each bind is derived as node base port + shard id — shard 0 → 9601, shard 1 → 9602, shard 2 → 9603. The StatefulSet declares all three containerPorts (grpc/grpc-1/grpc-2) for clarity and NetworkPolicy; the bind itself is driven by the topology. Collapse back to a single grpc port only if the shards: block is removed (legacy single group).

The m12 findings flagged the m12p4 3-shard catch-up + cross-shard read layer as not yet production-ready on real k3s (rc7); those five bugs were root-caused and fixed in rc8/rc9, and the LIVE rc7 image below carries the reseed-loop, seed-join-promotion, election-divergence, read-SLA, and write-burst fixes.

Resources at 1536-D

The 1536-dim production shape sets the pod resources, deliberately:

Field Value Why
limits.cpu "3" (raised from 2) The cgroup CPU quota is what the engine reads for available_parallelism() (SEARCH_GATE / worker-thread sizing). At 2 a cross-shard search burst starved the async reactor + the election/heartbeat/apply control plane — reads hung to the 30 s route timeout and the starved control plane churned elections into reseed self-exit. 3 leaves ~1 core for kubelet/system on the 4-core nodes.
requests.cpu 500m Kept low so the pod still schedules (server nodes alloc ≈ 3).
limits.memory 4Gi A 100k×1536-D HNSW load peaks ~1.9Gi. 1M needs more headroom — plan >16Gi nodes for the 1M gate.
startupProbe.failureThreshold 240 (~20 min) HNSW rebuild/load at 1536-D is CPU-bound: ~5 min single-core at 100k, so the startup budget must cover a cold rebuild (plus headroom for the 1M gate).
probe scheme HTTPS (all three) The :9500 HTTP plane serves TLS (inter-node mTLS, m11p7). kubelet does not verify the server cert for httpGet probes, so the cert's DNS-only SANs are fine.

What's in k8s/cluster/

File Purpose
namespace.yaml The tidaldb-cluster namespace (mutually exclusive with tidaldb)
topology-configmap.yaml The ONE bootstrap topology shared by all pods: 3 regions by per-pod DNS (grpc_addr advertised, grpc_bind 0.0.0.0), replication.ack: quorum, replication.reseed_self_restart: true, the election block
schema-configmap.yaml The schema YAML every region loads (--schema)
statefulset.yaml replicas: 3, podManagementPolicy: Parallel, TIDAL_REGION from POD_NAME, durable PVC at /data with --data-dir /data/db, the three probes (readiness now cluster-aware, all scheme: HTTPS), uid 10001, terminationGracePeriodSeconds: 600 with TIDAL_SHUTDOWN_DRAIN_MS=3000, the full-placement 3-shard layout (one gRPC port per group), 1536-D resources (cpu: "3", memory: 4Gi), topology spread, and the ordinal-branching scale-up wrapper
service-peers.yaml Headless peer Service (publishNotReadyAddresses: true) — stable per-pod DNS, keeps not-ready joiners resolvable for peers
service-client.yaml Client Service — readiness-gated, drops not-ready/joining/quarantined pods from load balancing
poddisruptionbudget.yaml maxUnavailable: 1 — a 3-voter cluster keeps quorum across one disruption
secret.example.yaml Template for the tidaldb-credentials / TIDAL_API_KEY secret (create the real one out-of-band)
kustomization.yaml Ties it together for kubectl apply -k k8s/cluster/ (secret excluded)

The two Services

  • tidaldb-peers (headless, clusterIP: None, publishNotReadyAddresses: true): gives each pod the stable DNS name tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local, which the topology advertises. publishNotReadyAddresses: true is load-bearing — a joiner is not ready until it first converges, but peers must still resolve it to feed it a snapshot + catch-up stream; without this the joiner could never reach a seed to become ready (a deadlock).
  • tidaldb (client-facing, VIP, readiness-gated): the address in-cluster clients hit. Default readiness gating drops not-ready/joining/quarantined/ draining pods, so a client is never routed to a node that is still catching up.

The readiness predicate (cluster-aware /health)

The readinessProbe stays GET /health, but in cluster mode it is now predicate-driven (m11p5 §4). /health returns 503 when the node is:

  • shutting down (SIGTERM drain — leaves both Services before it stops accepting);
  • quarantined (an m11p4 divergent leader-acked suffix — serves status and votes, refuses the data plane until reseeded);
  • removed / decommissioned (a Removed membership record reached it, or a voter's typed removed signal told it so);
  • a joiner (seed-join learner) or install boot (snapshot-reseeded) that has not yet first-converged — convergence means the boot catch-up pull completed at least once AND lag fell to ≤ learner_promote_lag (hysteresis; never lag == 0, which an open-loop write load keeps perpetually false). Sticky-ready after the first convergence.

A restarted, PVC-retained voter is Ready on today's terms — no regression for ordinary pod restarts. GET /cluster/status/local surfaces the inputs (lag_events, quarantined, role, term, reseed_required) for diagnosis, plus the tidaldb_cluster_reseed_required and tidaldb_cluster_divergence gauges.

Deploy the cluster

# 1. Build/publish the release image (runtime user uid 10001 to match the
#    securityContext so the PVC is writable; one image serves every subcommand).
#    Use the release script — NOT a bare `docker build`:
./scripts/build-release.sh <tag> server
#    It HOST cross-compiles macOS-arm64 -> x86_64-unknown-linux-gnu
#    (GCC 15.2 / glibc 2.41), then packages a `debian:trixie-slim` runtime
#    (the binary needs `libmvec.so.1` — ABSENT on bookworm — plus libstdc++6 /
#    libgcc-s1) via the `amd64builder` buildx builder (QEMU), and pushes to
#    registry.threesix.ai.
#
#    CRITICAL — pin the linux/amd64 PLATFORM manifest digest, NOT the OCI index
#    digest and NOT the `unknown/unknown` attestation manifest. Verify before pin:
docker buildx imagetools inspect registry.threesix.ai/tidal/server:<tag> --raw
#    Expect mediaType application/vnd.oci.image.manifest.v1+json, ~6 layers,
#    architecture amd64. (Pinning the index or attestation digest yields an
#    ImagePullBackOff or a no-arch pull.)

# 2. Namespace + the credentials secret (stress/Ref-A shape: name
#    tidaldb-credentials, key TIDAL_API_KEY). Same key on EVERY pod and client.
kubectl create namespace tidaldb-cluster
kubectl -n tidaldb-cluster create secret generic tidaldb-credentials \
  --from-literal=TIDAL_API_KEY="$(openssl rand -hex 32)"

# 3. Apply and wait for the 3-pod voter set.
kubectl apply -k k8s/cluster/
kubectl -n tidaldb-cluster rollout status statefulset/tidaldb --timeout=300s

# 4. Confirm three reachable regions with low lag, and check each node's role.
kubectl -n tidaldb-cluster exec tidaldb-0 -- \
  curl -s localhost:9500/cluster/status | jq '.regions[] | {name, lag_events, reachable}'
for p in tidaldb-0 tidaldb-1 tidaldb-2; do
  kubectl -n tidaldb-cluster exec "$p" -- \
    curl -s localhost:9500/cluster/status/local | jq '{role, term, membership_role, lag_events}'
done

Upgrade by digest (RollingUpdate)

The LIVE image is registry.threesix.ai/tidal/server:m12-writeburst-rc7 (@sha256:171505745b801dcf231b531de6167dbc309a7182957811cbc2228f0a302572b1). It carries the reseed-loop, seed-join-promotion, election-divergence, read-SLA, and write-burst fixes (all shipped). To roll a new build:

# Set the image by its amd64 platform-manifest DIGEST (verified above), then watch
# the RollingUpdate. The StatefulSet rolls highest-ordinal-first, one pod at a time.
kubectl set image statefulset/tidaldb \
  tidaldb=registry.threesix.ai/tidal/server@sha256:<amd64-digest> \
  -n tidaldb-cluster
kubectl -n tidaldb-cluster rollout status statefulset/tidaldb --timeout=600s

Then verify 3/3 Ready and check each pod's boot-reseed outcome:

kubectl -n tidaldb-cluster get pods -l app.kubernetes.io/name=tidaldb
for p in tidaldb-0 tidaldb-1 tidaldb-2; do
  kubectl -n tidaldb-cluster exec "$p" -- \
    curl -s localhost:9500/cluster/status/local | jq '{role, term, lag_events, reseed_required}'
done

The ideal boot-reseed outcome is [(0,NotNeeded),(1,NotNeeded),(2,NotNeeded)] — every pod's PVC-retained WAL caught up via stream, no reseed. A snapshot-install fallback (a pod rejoined behind WAL retention and reseeded via snapshot) is SAFE, not a failure: it converges to lag=0 on its own; readiness simply stays 503 until the install first-converges. Only a pod that loops (quarantine → self-restart → re-detect) is a real problem — see disaster-recovery.md.

Scale up (3 → N): seed-join as a learner, auto-promote

kubectl scale is the whole story — no topology edits. Pods with ordinal ≥ 3 boot with --seed (the StatefulSet's ordinal-branching wrapper adds it automatically) and learn their roster/id/term from a seed, joining as a learner. The leader auto-promotes a learner to a voter once its durable mark is within learner_promote_lag of the leader's frontier (or stops falling behind for K rounds under sustained load).

# Grow to 5 voters. Pods tidaldb-3 and tidaldb-4 seed-join + auto-promote.
kubectl -n tidaldb-cluster scale statefulset/tidaldb --replicas=5
kubectl -n tidaldb-cluster rollout status statefulset/tidaldb --timeout=600s
# Watch the new members converge and promote (promotion_pending shows the lag):
kubectl -n tidaldb-cluster exec tidaldb-0 -- \
  curl -s localhost:9500/cluster/members | jq '.members[] | {id, name, role}'

A scaled pod still mounts the shared topology ConfigMap — a --seed boot requires the local config for the behavioral knob blocks (replication, wal, election, timeouts, grpc_tls); its regions: list is ignored for the seed joiner's roster (the join response is authoritative).

Scale down (N → fewer): remove verb FIRST, then scale

Decommission a member before removing its pod, so the cluster stops counting it toward quorum cleanly. Remove the highest-ordinal members (StatefulSet deletes lowest-ordinal-last on scale-down).

# Going 5 -> 3: decommission tidaldb-4, then tidaldb-3, then scale.
# Remove verb: POST /cluster/members/remove {"region": "<name>"} (any node
# forwards to the leader; one-at-a-time, quorum-commit-gated).
kubectl -n tidaldb-cluster exec tidaldb-0 -- curl -s -X POST \
  -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -d '{"region":"tidaldb-4"}' localhost:9500/cluster/members/remove
kubectl -n tidaldb-cluster exec tidaldb-0 -- curl -s -X POST \
  -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -d '{"region":"tidaldb-3"}' localhost:9500/cluster/members/remove
# Each Remove is delivered to the removed peer (which flips to 503 and stops
# campaigning) before its ship cell retires. Only then is it safe to scale.
kubectl -n tidaldb-cluster scale statefulset/tidaldb --replicas=3

Removing pods without the remove verb leaves zombie members the cluster still expects — never scale down before decommissioning.

Node / pod replacement

  • kubectl delete pod tidaldb-N (PVC retained): the StatefulSet recreates the pod onto its existing PVC. It boots a follower, replays its WAL, and catches up via the StreamSegments stream (boot-time + the 30 s catch-up timer). No operator verb. This is the routine replace — already proven by the p4 leader-kill regression (lag → 0 within ~10 s).
  • PVC + pod delete (fresh reseed): delete the PVC and the pod. The new pod comes up with an empty data dir, latches an install boot, and reseeds via snapshot (FetchSnapshot + stream) from the leader, then rejoins clean. Use this when the data dir is corrupt or the node fell behind a compacted leader. Readiness stays 503 until the snapshot install first converges.

Self-healing reseed (quarantine / behind-compaction)

replication.reseed_self_restart: true (set in the topology ConfigMap) makes a node that durably latches reseed_required (an m11p4 divergence quarantine, or a typed snapshot-required refusal) drain and exit(0); the StatefulSet restarts it and the boot-time install reseeds it via snapshot, clearing the quarantine and the divergence gauge with no operator verb and no PVC wipe. The self-restart is refused (loudly, in /cluster/status/local + the gauge) when the remaining voters cannot sustain quorum without this node — exiting during a 2-voter window would be a total write outage. POST /cluster/reseed latches the marker on demand.

PodDisruptionBudget and graceful shutdown

maxUnavailable: 1 lets a voluntary disruption (node drain, autoscaler, rolling upgrade) take at most one pod at a time, preserving quorum (2 of 3). On SIGTERM a pod flips readiness to 503 (leaving both Services), drains, lets the leader lease/heartbeat windows lapse so a successor is elected cleanly, then checkpoints

  • fsyncs the WAL and saves every hosted shard's HNSW graph before exit.

terminationGracePeriodSeconds: 600 (raised from 60 in m12p6) covers that sequence. At 1536-D the long pole is the graph save: each pod hosts all three shard groups, and a single group's USearch serialize+fsync (~32k vectors/slot) is slow, so the three concurrent saves need ample budget or k8s SIGKILLs mid-save and the next boot rebuilds from raw vectors (a ~5 min single-core stall at 100k). TIDAL_SHUTDOWN_DRAIN_MS=3000 shortens the post-SIGTERM in-flight drain (from the 15 s default) so the save starts promptly inside the grace window; a clean save typically finishes in well under a minute, so 600 s is a ceiling, not the norm — but it makes rolling restarts on a loaded cluster slow, since each pod saves before it exits.

The exit-gate harness

The in-cluster load Jobs in tidal-stress/k8s/ target the cluster by per-pod DNS (tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500), in namespace tidaldb-cluster, reading the tidaldb-credentials secret. They run the capacity ramp, the quorum-throughput gate (stress-job-t2a.yaml, --ack quorum), and the leader-kill chaos drill (stress-job-t2b.yaml).

See also