# tidal-stress Open-loop capacity ramp + chaos harness for the tidalDB cluster. Drives the `thepeach` feed workload (signals + vector embeddings) against a live cluster and reports a per-stage capacity verdict. - **Worklog** (what's been run, what we learned): [WORKLOG.md](WORKLOG.md) - **Process** (how to run the next checkpoint): [PROCESS.md](PROCESS.md) --- ## Current target cluster (as of 2026-06-13) The cluster moved to the **m11p5 single-StatefulSet architecture**. The old 3-StatefulSet / static-ClusterIP model (namespace `tidaldb`, IPs `10.43.99.11-13`) is **retired** — any manifest or doc still naming those IPs is stale. | Fact | Value | |------|-------| | Namespace | `tidaldb-cluster` | | Pods | `tidaldb-{0,1,2}` (one StatefulSet, 3 replicas) | | Peer DNS | `tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500` (HTTP), `:9601` (gRPC) | | Client VIP | `tidaldb.tidaldb-cluster.svc.cluster.local:9500` (readiness-gated) | | Server image | `registry.threesix.ai/tidal/server@sha256:173e803…` (`:m11p5`) | | Stress image | `registry.threesix.ai/tidal/stress@sha256:3a75c311…` (`:m11p3`) | | Storage | local-path 5Gi/pod (on-node NVMe) — **NOT** Longhorn (see WORKLOG) | | CPU/pod | limit `2` (the write pool is ~2 workers on the leader) | Targets for any new Job manifest — use pod DNS for `--target` (so status-polling reaches survivors during a kill window) and the VIP for `--leader-url`: ```yaml args: - --target - http://tidaldb-0.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500 - --target - http://tidaldb-1.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500 - --target - http://tidaldb-2.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500 - --leader-url - http://tidaldb.tidaldb-cluster.svc.cluster.local:9500 ``` --- ## Run pattern Every run deploys the generator as an in-cluster Job (port-forward adds API-server serialization latency — never use it for capacity numbers; only the kill loop port-forwards, and only to read `/cluster/status`). ```bash export KUBECONFIG=~/.kube/orchard9-k3sf.yaml kubectl apply -f k8s/.yaml kubectl logs -f job/ -n tidaldb-cluster kubectl delete job -n tidaldb-cluster # re-arm before re-running ``` ## CLI flags (authoritative — from `src/main.rs`) | Flag | Default | Notes | |------|---------|-------| | `--target ` | (required, repeatable) | Region gateway; reads round-robin across all | | `--leader-url ` | none | Pin leader-path writes here to skip the forward hop | | `--api-key` | `$TIDAL_API_KEY` | Bearer; the cluster requires it | | `--ack ` | topology default | Sent as `x-tidal-ack` per write | | `--ramp ` | `peach-100k` | Presets: `smoke`, `quick`, `peach-100k`, `max` | | `--stage-secs ` | 45 | Hold per preset stage; 300–600 for soak | | `--mix ` | `peach` | Presets: `peach`, `reads`, `writes`. Ops: `feed,search,view,like,skip,item,embed` | | `--write-path ` | `leader` | `sharded` removes the single-leader funnel (not replicated) | | `--corpus ` | 10000 | Items+embeddings to seed; 20k for gate runs | | `--users ` | 50000 | Virtual user id space | | `--skip-seed` | false | Set after the first run of a session (corpus persists on PVC) | | `--embedding-dim ` | 128 | Deployed schema = 128; thepeach real = 1536 | | `--hot-skew ` | 1.3 | Power-law concentration onto hot items | | `--poll-status` | false | Poll `/cluster/status` between stages for lag — **always set when measuring lag** | | `--stop-on-knee` | false | Stop at first SLO-breaching stage | | `--dau ` | 100000 | DAU the verdict translates the ceiling against | SLO: feed p99 ≤ 150ms (network-hop allowance over the in-process 50ms SLA); error rate ≥ 1% (429/408/503/5xx/transport) = the knee. ## Layout ``` src/ generator (scheduler, workload model, client, metrics) k8s/ Job manifests — one per checkpoint stress-job.yaml generic ramp stress-job-t2a.yaml T2-A quorum throughput stress-job-t2b.yaml T2-B acked-loss under kills stress-job-t3.yaml T3 automatic-failover gate scripts/ t3-kill-loop-v3.sh HTTP-polling leader-kill loop (no exec into pods) ``` ## Checkpoint status | ID | Gate | Status | |----|------|--------| | T0 | baseline (~90/s replicated, 3669/s sharded) | ✓ done | | T2-A | ≥1000 quorum writes/s | ✓ 2980/s | | T2-B | 0 acked loss across kills | ✓ | | T3 | leader-kill failover <10s p99 ×10 | ✓ max 6157ms (m11p5) | | **T4** | scale 3→5→3 under load, joiner ≤5min | **next** | | T-read | vector-search recall@k + query QPS/p99 | **not built** (see PROCESS) | | T5 | sharded ≥5000 quorum writes/s | blocked on p6 |