Splits monolithic cluster.rs into tidal-server/src/cluster/ modules. Adds redeliver-missed relay, bounded HLC drift, lag tracking, and reconcile idempotence. Five new tier-3 test suites (chaos, lifecycle, multiproc, region, routes, runbook) all green. Docs, CHANGELOG, and ROADMAP updated with G4/G5/G6 known gaps.
220 lines
10 KiB
Markdown
220 lines
10 KiB
Markdown
# tidalDB on Kubernetes
|
|
|
|
How to run tidalDB on Kubernetes: a hardened, single-node **standalone**
|
|
deployment as a `StatefulSet` with a durable volume, the three health probes,
|
|
metrics, secret-backed auth, and graceful rolling updates. The manifests live in
|
|
[`k8s/`](../../k8s/) and apply with `kubectl apply -k k8s/`.
|
|
|
|
> ## Deploy the STANDALONE server, one replica
|
|
>
|
|
> tidalDB is single-node-first: the server wraps one embedded engine whose state
|
|
> (WAL + checkpoints + indexes) lives on a data dir. It scales **vertically**
|
|
> (a bigger pod), not by adding replicas — there is no shared-storage multi-writer
|
|
> mode, so `replicas: 1` in the StatefulSet is load-bearing. The experimental
|
|
> multi-region `cluster` subcommand now has a **multi-process** mode (`--region`,
|
|
> one process per region, real gRPC peering, real process isolation — shipped in
|
|
> m8p10), but it is still **experimental**: writes are leader-durable, **not**
|
|
> quorum-acked, and failover is operator-driven (`/cluster/promote`), **not**
|
|
> automatic election. So it is not yet a production-HA story
|
|
> ([cluster runbook](cluster.md)). **The recommended production deployment remains
|
|
> a single STANDALONE pod** until quorum-ack writes and automatic failover land
|
|
> (tracked as post-M8 follow-ups in [ROADMAP.md](../planning/ROADMAP.md)). Run one
|
|
> standalone pod, back it with a durable `PersistentVolume`, and recover from the
|
|
> WAL on restart (see [recovery](../ops/recovery.md)).
|
|
>
|
|
> **Experimental: cluster mode on Kubernetes.** Multi-process cluster mode maps to
|
|
> a `StatefulSet` (or one StatefulSet *per region*) where each pod runs
|
|
> `tidal-server cluster --region <name>` with its own `PersistentVolume` for
|
|
> `--data-dir`, a stable `grpc_addr`/`http_addr` per region (a headless `Service`
|
|
> gives each pod stable DNS), and `TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1`. This is a
|
|
> sketch, **not** a promised manifest set — `k8s/` ships only the hardened
|
|
> standalone deployment. Treat a cluster-mode k8s deployment as experimental and
|
|
> operator-supervised (manual promote + heal), per the cluster runbook.
|
|
|
|
## What's in `k8s/`
|
|
|
|
| File | Purpose |
|
|
|------|---------|
|
|
| `namespace.yaml` | The `tidaldb` namespace |
|
|
| `schema-configmap.yaml` | The schema YAML the server loads (`--schema`); edit for your signals |
|
|
| `statefulset.yaml` | The server: durable PVC, probes, security context, resources |
|
|
| `service.yaml` | Headless `Service` for stable DNS + in-cluster clients |
|
|
| `poddisruptionbudget.yaml` | `maxUnavailable: 0` — a drain can't silently kill the single node |
|
|
| `secret.example.yaml` | Template for the API-key secret (create the real one out-of-band) |
|
|
| `servicemonitor.yaml` | Optional Prometheus-Operator scrape config (apply separately) |
|
|
| `kustomization.yaml` | Ties the core resources together for `kubectl apply -k` |
|
|
|
|
## Prerequisites
|
|
|
|
- A Kubernetes cluster (1.25+) and `kubectl` pointed at it. For local testing,
|
|
[`kind`](https://kind.sigs.k8s.io/) is used in the walkthrough below.
|
|
- A container registry the cluster can pull from (for real clusters), or a local
|
|
image loaded into the node (for `kind`). The image is built from
|
|
[`docker/deploy/Dockerfile`](../../docker/deploy/Dockerfile).
|
|
- A default `StorageClass` (for dynamic `PersistentVolumeClaim` provisioning).
|
|
`kind`, GKE, EKS, and AKS all ship one.
|
|
|
|
## Deploy
|
|
|
|
### 1. Build and publish the image
|
|
|
|
```bash
|
|
# From the repo root — the build context must be the workspace root.
|
|
docker build -f docker/deploy/Dockerfile -t <registry>/tidaldb:<tag> .
|
|
docker push <registry>/tidaldb:<tag>
|
|
```
|
|
|
|
Set that reference in `k8s/statefulset.yaml` (`image:`), pinned by digest in
|
|
production (`@sha256:...`).
|
|
|
|
### 2. Create the namespace and the API-key secret
|
|
|
|
The secret is deliberately **not** in the kustomization so no key lands in git.
|
|
Create it directly:
|
|
|
|
```bash
|
|
kubectl create namespace tidaldb
|
|
kubectl -n tidaldb create secret generic tidaldb-api-key \
|
|
--from-literal=api-key="$(openssl rand -hex 32)"
|
|
```
|
|
|
|
In production, manage it with External Secrets Operator, Sealed Secrets, or
|
|
Vault Agent instead. The StatefulSet injects it as `TIDAL_API_KEY` — clients
|
|
then send `Authorization: Bearer <key>` on every data route. **If the secret is
|
|
empty the server runs unauthenticated and logs a WARN — never do that on a
|
|
shared network.**
|
|
|
|
### 3. Edit the schema (optional)
|
|
|
|
`k8s/schema-configmap.yaml` carries the schema the server loads. Edit it to model
|
|
your signals, text fields, embedding slots, and (optionally) ranking profiles —
|
|
the format is documented in [server-deployment.md](../guides/server-deployment.md).
|
|
The schema is read once at boot; roll the StatefulSet to apply changes.
|
|
|
|
### 4. Apply
|
|
|
|
```bash
|
|
kubectl apply -k k8s/
|
|
kubectl -n tidaldb rollout status statefulset/tidaldb --timeout=180s
|
|
```
|
|
|
|
### 5. Verify
|
|
|
|
```bash
|
|
kubectl -n tidaldb port-forward statefulset/tidaldb 9400:9400 &
|
|
KEY=$(kubectl -n tidaldb get secret tidaldb-api-key -o jsonpath='{.data.api-key}' | base64 -d)
|
|
|
|
curl -s localhost:9400/health # {"ok":true,...}
|
|
curl -s localhost:9400/openapi.json | jq .info # served API contract
|
|
curl -s -X POST localhost:9400/items \
|
|
-H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
|
|
-d '{"entity_id":1,"metadata":{"title":"hello","category":"demo","created_at":"1700000000"}}'
|
|
curl -s -H "Authorization: Bearer $KEY" "localhost:9400/feed?profile=trending&limit=5"
|
|
```
|
|
|
|
## How the health probes map
|
|
|
|
The server exposes three unauthenticated endpoints, wired to the three probe
|
|
types in `statefulset.yaml`:
|
|
|
|
| Probe | Endpoint | Behavior |
|
|
|-------|----------|----------|
|
|
| `startupProbe` | `GET /health/startup` | 200 once the HTTP listener is up; high `failureThreshold` covers slow WAL replay / index load on large data dirs (see [capacity-planning](../ops/capacity-planning.md)) |
|
|
| `livenessProbe` | `GET /health/live` | 200 while the process is alive; restart if it stops answering |
|
|
| `readinessProbe` | `GET /health` | 200 ready / **503 while draining** — on SIGTERM the pod leaves the Service endpoints before it stops accepting |
|
|
|
|
## Rolling updates and graceful shutdown
|
|
|
|
On `kubectl rollout restart` (or any pod delete), the kubelet sends SIGTERM. The
|
|
server flips readiness to 503 (so it leaves the Service), drains in-flight
|
|
requests, then checkpoints and fsyncs the WAL before exit.
|
|
`terminationGracePeriodSeconds: 60` gives that room — raise it if your data dir
|
|
is large. Because there is one replica, a restart is a brief planned outage
|
|
while the new pod replays the WAL; the `PodDisruptionBudget` (`maxUnavailable:
|
|
0`) prevents an *involuntary* drain from taking the node down without operator
|
|
intent.
|
|
|
|
## Persistence and backup
|
|
|
|
The `volumeClaimTemplate` provisions a `PersistentVolumeClaim` (`/data`, 10Gi by
|
|
default — size it from [capacity-planning](../ops/capacity-planning.md)). The WAL
|
|
+ checkpoints there are the source of truth and survive pod restarts. For backup
|
|
and disaster recovery (snapshotting the PVC, restoring a corrupt data dir), see
|
|
[recovery](../ops/recovery.md).
|
|
|
|
## Metrics
|
|
|
|
The pod exposes Prometheus metrics on `:9091/metrics` (**unauthenticated** — it is
|
|
not exposed by the headless Service externally; keep it cluster-internal). Scrape
|
|
it one of two ways:
|
|
|
|
- **Prometheus Operator:** `kubectl apply -f k8s/servicemonitor.yaml` (requires the
|
|
`monitoring.coreos.com` CRDs).
|
|
- **Plain Prometheus:** the pod carries `prometheus.io/scrape`, `prometheus.io/port`,
|
|
and `prometheus.io/path` annotations.
|
|
|
|
Alert rules and a dashboard ship in [`docs/ops/prometheus-alerts.yaml`](../ops/prometheus-alerts.yaml)
|
|
and [`docs/ops/grafana-dashboard.json`](../ops/grafana-dashboard.json); see
|
|
[monitoring](../ops/monitoring.md).
|
|
|
|
## Local walkthrough with `kind`
|
|
|
|
This is the exact flow used to verify the manifests end-to-end:
|
|
|
|
```bash
|
|
# 1. Create a local cluster.
|
|
kind create cluster --name tidaldb
|
|
|
|
# 2. Build the image and load it into the kind node (no registry needed).
|
|
docker build -f docker/deploy/Dockerfile -t tidaldb:deploy .
|
|
kind load docker-image tidaldb:deploy --name tidaldb
|
|
|
|
# 3. Namespace + API-key secret.
|
|
kubectl create namespace tidaldb
|
|
kubectl -n tidaldb create secret generic tidaldb-api-key \
|
|
--from-literal=api-key="$(openssl rand -hex 32)"
|
|
|
|
# 4. Apply and wait for ready.
|
|
kubectl apply -k k8s/
|
|
kubectl -n tidaldb rollout status statefulset/tidaldb --timeout=240s
|
|
|
|
# 5. Verify, then tear down.
|
|
kubectl -n tidaldb port-forward statefulset/tidaldb 9400:9400 &
|
|
curl -s localhost:9400/health
|
|
kind delete cluster --name tidaldb
|
|
```
|
|
|
|
The manifests set `image: tidaldb:deploy` with `imagePullPolicy: IfNotPresent`,
|
|
which is exactly what `kind load` + a local tag need. For a real cluster, swap in
|
|
your registry image.
|
|
|
|
> **`kind create cluster` fails with "could not find a log line that matches …
|
|
> Multi-User System"?** On Docker Desktop the node's `systemd` can die at boot
|
|
> with `Failed to create control group inotify object: Too many open files`
|
|
> (`docker logs <cluster>-control-plane` shows it). The Docker VM's inotify
|
|
> limits are too low; raise them in the VM kernel, then recreate:
|
|
>
|
|
> ```bash
|
|
> docker run --rm --privileged alpine sysctl -w fs.inotify.max_user_instances=8192
|
|
> kind delete cluster --name tidaldb && kind create cluster --name tidaldb
|
|
> ```
|
|
>
|
|
> This is a kind-on-Docker-Desktop prerequisite, unrelated to tidalDB.
|
|
|
|
## Troubleshooting
|
|
|
|
| Symptom | Likely cause |
|
|
|---------|--------------|
|
|
| Pod `Pending` | No default `StorageClass`, or the PVC can't bind — `kubectl -n tidaldb describe pvc data-tidaldb-0` |
|
|
| Pod `CrashLoopBackOff` at boot | Bad schema YAML in the ConfigMap, or a data dir from an incompatible schema — check logs; see [recovery § schema mismatch](../ops/recovery.md) |
|
|
| Pod never `Ready`, but `Running` | Readiness probe failing — `kubectl -n tidaldb logs statefulset/tidaldb`; a large data dir may need a longer `startupProbe` |
|
|
| `401 Unauthorized` on data routes | Wrong/empty `tidaldb-api-key` secret; clients must send `Authorization: Bearer <key>` |
|
|
| Writes lost after restart | Data dir not on the PVC — confirm `--data-dir /data` and the `data` volume mount |
|
|
|
|
## See also
|
|
|
|
- [Server deployment guide](../guides/server-deployment.md) — config, auth, the served OpenAPI spec
|
|
- [Build a feed app](../guides/build-a-feed-app.md) — what to run against this server
|
|
- [Cluster runbook](cluster.md) — the experimental multi-region mode (not production HA)
|
|
- [Monitoring](../ops/monitoring.md) · [Capacity planning](../ops/capacity-planning.md) · [Recovery](../ops/recovery.md)
|