Read-SLA fix (rc12→rc13 — cpu-cgroup starvation → multi-second p99 + churning elections): - offload.rs: add SEARCH_GATE semaphore (core_count+1 permits, 50ms shed to 429) so per-shard searches gate on CPU, not reactor threads; concurrent scatter_merge fan-out (join_all) replaces the serial blocking offload_region_read loop - node.rs: scatter_merge → async; per-shard futures run via offload_search (each acquires one SEARCH_GATE permit, moves it into spawn_blocking so the permit is held for the search's full CPU lifetime) - main.rs: explicit tokio runtime with worker_threads floored at 4, independent of the cgroup quota — keeps the control plane (heartbeat/election/apply) on its own workers even when quota < 4 - k8s statefulset: CPU limit 2→3 (was: available_parallelism()=2 → only 2 async workers; search burst starved the reactor) - tidal/wal/compaction.rs: WAL_RETENTION_SEGMENTS 4→16 (64 MiB→256 MiB per-shard catch-up window; a briefly-down follower across a rolling restart streams up instead of forcing snapshot reseed; disk floor 768 MiB/pod, self-trimming) - cluster_reseed.rs: OFFLINE_ITEMS 1800→5600 to exceed the new 16-segment retention window (19 segs > 17); fix sequential quarantine/reseed race via await_status_bool tidalctl S3/R2 backup DR: - tidalctl/Cargo.toml: aws-config, aws-sdk-s3, aws-credential-types, tokio, tempfile - commands/s3.rs: S3Target + export_dir (upload every file, manifest last as atomicity marker) + import_to_dir (download prefix into temp staging dir) - commands/backup.rs: run_backup/run_restore accept Option<&S3Target>; S3 export is additive after local fsync barrier; S3 import stages into TempDir then runs the unchanged verified restore on it - main.rs: --s3-endpoint / --s3-bucket / --s3-prefix flags; all-or-nothing endpoint+bucket validation; usage updated tidal-stress/k8s: recall-rc12-spread-job, soak-nightly-cronjob, soak-monitor, soak-results-pvc, t5-readtput-job manifests |
||
|---|---|---|
| .. | ||
| benches | ||
| k8s | ||
| scripts | ||
| src | ||
| Cargo.toml | ||
| PROCESS.md | ||
| README.md | ||
| WORKLOG.md | ||
tidal-stress
Open-loop capacity ramp + chaos harness for the tidalDB cluster. Drives the
thepeach feed workload (signals + vector embeddings) against a live cluster and
reports a per-stage capacity verdict.
- Worklog (what's been run, what we learned): WORKLOG.md
- Process (how to run the next checkpoint): PROCESS.md
Current target cluster (as of 2026-06-13)
The cluster moved to the m11p5 single-StatefulSet architecture. The old
3-StatefulSet / static-ClusterIP model (namespace tidaldb, IPs 10.43.99.11-13)
is retired — any manifest or doc still naming those IPs is stale.
| Fact | Value |
|---|---|
| Namespace | tidaldb-cluster |
| Pods | tidaldb-{0,1,2} (one StatefulSet, 3 replicas) |
| Peer DNS | tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500 (HTTP), :9601 (gRPC) |
| Client VIP | tidaldb.tidaldb-cluster.svc.cluster.local:9500 (readiness-gated) |
| Server image | registry.threesix.ai/tidal/server@sha256:173e803… (:m11p5) |
| Stress image | registry.threesix.ai/tidal/stress@sha256:3a75c311… (:m11p3) |
| Storage | local-path 5Gi/pod (on-node NVMe) — NOT Longhorn (see WORKLOG) |
| CPU/pod | limit 2 (the write pool is ~2 workers on the leader) |
Targets for any new Job manifest — use pod DNS for --target (so status-polling
reaches survivors during a kill window) and the VIP for --leader-url:
args:
- --target
- http://tidaldb-0.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
- --target
- http://tidaldb-1.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
- --target
- http://tidaldb-2.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
- --leader-url
- http://tidaldb.tidaldb-cluster.svc.cluster.local:9500
Run pattern
Every run deploys the generator as an in-cluster Job (port-forward adds
API-server serialization latency — never use it for capacity numbers; only the
kill loop port-forwards, and only to read /cluster/status).
export KUBECONFIG=~/.kube/orchard9-k3sf.yaml
kubectl apply -f k8s/<job>.yaml
kubectl logs -f job/<job-name> -n tidaldb-cluster
kubectl delete job <job-name> -n tidaldb-cluster # re-arm before re-running
CLI flags (authoritative — from src/main.rs)
| Flag | Default | Notes |
|---|---|---|
--target <url> |
(required, repeatable) | Region gateway; reads round-robin across all |
--leader-url <url> |
none | Pin leader-path writes here to skip the forward hop |
--api-key |
$TIDAL_API_KEY |
Bearer; the cluster requires it |
--ack <leader|quorum> |
topology default | Sent as x-tidal-ack per write |
--ramp <preset|rps:secs,…> |
peach-100k |
Presets: smoke, quick, peach-100k, max |
--stage-secs <n> |
45 | Hold per preset stage; 300–600 for soak |
--mix <preset|op=w,…> |
peach |
Presets: peach, reads, writes. Ops: feed,search,view,like,skip,item,embed |
--write-path <leader|sharded> |
leader |
sharded removes the single-leader funnel (not replicated) |
--corpus <n> |
10000 | Items+embeddings to seed; 20k for gate runs |
--users <n> |
50000 | Virtual user id space |
--skip-seed |
false | Set after the first run of a session (corpus persists on PVC) |
--embedding-dim <n> |
128 | Deployed schema = 128; thepeach real = 1536 |
--hot-skew <f> |
1.3 | Power-law concentration onto hot items |
--poll-status |
false | Poll /cluster/status between stages for lag — always set when measuring lag |
--stop-on-knee |
false | Stop at first SLO-breaching stage |
--dau <n> |
100000 | DAU the verdict translates the ceiling against |
SLO: feed p99 ≤ 150ms (network-hop allowance over the in-process 50ms SLA); error rate ≥ 1% (429/408/503/5xx/transport) = the knee.
Layout
src/ generator (scheduler, workload model, client, metrics)
k8s/ Job manifests — one per checkpoint
stress-job.yaml generic ramp
stress-job-t2a.yaml T2-A quorum throughput
stress-job-t2b.yaml T2-B acked-loss under kills
stress-job-t3.yaml T3 automatic-failover gate
scripts/
t3-kill-loop-v3.sh HTTP-polling leader-kill loop (no exec into pods)
Checkpoint status
| ID | Gate | Status |
|---|---|---|
| T0 | baseline (~90/s replicated, 3669/s sharded) | ✓ done |
| T2-A | ≥1000 quorum writes/s | ✓ 2980/s |
| T2-B | 0 acked loss across kills | ✓ |
| T3 | leader-kill failover <10s p99 ×10 | ✓ max 6157ms (m11p5) |
| T4 | scale 3→5→3 under load, joiner ≤5min | next |
| T-read | vector-search recall@k + query QPS/p99 | not built (see PROCESS) |
| T5 | sharded ≥5000 quorum writes/s | blocked on p6 |