tidaldb/tidal-stress
jx12n 44b768b8c6 feat(m11): sharding × replication + rebalancing (m11p6 L3-L5)
End the "replicated XOR sharded" split: S shard groups, each a
replication group at RF with its own elected leader, leaders balanced
across nodes; any gateway hash-routes.

- One unified write surface: /items,/embeddings,/signals hash-route to
  the owning shard group's leader (ShardRouter FNV-1a) AND replicate at
  RF. x-tidal-ack/x-tidal-seq, quorum await, NotLeader/QuorumTimeout are
  per-group; NotLeader names the group.
- Rebalance verbs (L3): POST /cluster/shards/{id}/transfer (fenced
  leadership move) + /cluster/shards/{id}/replicas (add/remove replica).
  A ?shard= selector threads through every per-shard admin verb and is
  propagated on intra-group forwards (ShardReplica::admin_path). S=1 is
  byte-for-byte (no selector, no shard in NotLeader body).
- Tier-3 exit gate (cluster_sharding.rs): 3 nodes × 3 shards × RF=3 over
  real OS processes — SIGKILL a node under ack=quorum load → only its
  shard-leaderships re-elect, reads never stop, zero acked loss across
  random kill points; plus a rebalance-verb test. Harness:
  MultiProcCluster::start_sharded.
- tidal-stress drives the single path (WritePath::Leader|Sharded gone),
  spreading writes round-robin across gateways or pinning --leader-url.
- Throughput: local 3×3 sustains 3,000 quorum signal-writes/s @ 0% err,
  ~30% CPU, lag ~0 (generator-bound). ≥5,000/s + ≥2.5× scaling is Ref-A.

Known follow-up (tracked): per-group-aware node readiness and cross-node
read fan-out under PARTIAL placement.
2026-06-13 18:23:43 -06:00
..
benches feat(m11): sharding × replication + rebalancing (m11p6 L3-L5) 2026-06-13 18:23:43 -06:00
k8s feat(k8s): m11p5 cluster manifest — local-path PVCs, initContainer, T3 tooling 2026-06-12 22:01:37 -06:00
scripts feat(k8s): m11p5 cluster manifest — local-path PVCs, initContainer, T3 tooling 2026-06-12 22:01:37 -06:00
src feat(m11): sharding × replication + rebalancing (m11p6 L3-L5) 2026-06-13 18:23:43 -06:00
Cargo.toml feat(m11): continuous correctness (m11p9) — fault classes, invariant checkers, soak gates, nightly pipeline 2026-06-13 15:23:59 -06:00
PROCESS.md docs(tidal-stress): README + worklog + continue-the-process flow 2026-06-13 09:23:29 -06:00
README.md docs(tidal-stress): README + worklog + continue-the-process flow 2026-06-13 09:23:29 -06:00
WORKLOG.md docs(tidal-stress): README + worklog + continue-the-process flow 2026-06-13 09:23:29 -06:00

tidal-stress

Open-loop capacity ramp + chaos harness for the tidalDB cluster. Drives the thepeach feed workload (signals + vector embeddings) against a live cluster and reports a per-stage capacity verdict.

  • Worklog (what's been run, what we learned): WORKLOG.md
  • Process (how to run the next checkpoint): PROCESS.md

Current target cluster (as of 2026-06-13)

The cluster moved to the m11p5 single-StatefulSet architecture. The old 3-StatefulSet / static-ClusterIP model (namespace tidaldb, IPs 10.43.99.11-13) is retired — any manifest or doc still naming those IPs is stale.

Fact Value
Namespace tidaldb-cluster
Pods tidaldb-{0,1,2} (one StatefulSet, 3 replicas)
Peer DNS tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500 (HTTP), :9601 (gRPC)
Client VIP tidaldb.tidaldb-cluster.svc.cluster.local:9500 (readiness-gated)
Server image registry.threesix.ai/tidal/server@sha256:173e803… (:m11p5)
Stress image registry.threesix.ai/tidal/stress@sha256:3a75c311… (:m11p3)
Storage local-path 5Gi/pod (on-node NVMe) — NOT Longhorn (see WORKLOG)
CPU/pod limit 2 (the write pool is ~2 workers on the leader)

Targets for any new Job manifest — use pod DNS for --target (so status-polling reaches survivors during a kill window) and the VIP for --leader-url:

args:
  - --target
  - http://tidaldb-0.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
  - --target
  - http://tidaldb-1.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
  - --target
  - http://tidaldb-2.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
  - --leader-url
  - http://tidaldb.tidaldb-cluster.svc.cluster.local:9500

Run pattern

Every run deploys the generator as an in-cluster Job (port-forward adds API-server serialization latency — never use it for capacity numbers; only the kill loop port-forwards, and only to read /cluster/status).

export KUBECONFIG=~/.kube/orchard9-k3sf.yaml
kubectl apply -f k8s/<job>.yaml
kubectl logs -f job/<job-name> -n tidaldb-cluster
kubectl delete job <job-name> -n tidaldb-cluster    # re-arm before re-running

CLI flags (authoritative — from src/main.rs)

Flag Default Notes
--target <url> (required, repeatable) Region gateway; reads round-robin across all
--leader-url <url> none Pin leader-path writes here to skip the forward hop
--api-key $TIDAL_API_KEY Bearer; the cluster requires it
--ack <leader|quorum> topology default Sent as x-tidal-ack per write
--ramp <preset|rps:secs,…> peach-100k Presets: smoke, quick, peach-100k, max
--stage-secs <n> 45 Hold per preset stage; 300600 for soak
--mix <preset|op=w,…> peach Presets: peach, reads, writes. Ops: feed,search,view,like,skip,item,embed
--write-path <leader|sharded> leader sharded removes the single-leader funnel (not replicated)
--corpus <n> 10000 Items+embeddings to seed; 20k for gate runs
--users <n> 50000 Virtual user id space
--skip-seed false Set after the first run of a session (corpus persists on PVC)
--embedding-dim <n> 128 Deployed schema = 128; thepeach real = 1536
--hot-skew <f> 1.3 Power-law concentration onto hot items
--poll-status false Poll /cluster/status between stages for lag — always set when measuring lag
--stop-on-knee false Stop at first SLO-breaching stage
--dau <n> 100000 DAU the verdict translates the ceiling against

SLO: feed p99 ≤ 150ms (network-hop allowance over the in-process 50ms SLA); error rate ≥ 1% (429/408/503/5xx/transport) = the knee.

Layout

src/            generator (scheduler, workload model, client, metrics)
k8s/            Job manifests — one per checkpoint
  stress-job.yaml       generic ramp
  stress-job-t2a.yaml   T2-A quorum throughput
  stress-job-t2b.yaml   T2-B acked-loss under kills
  stress-job-t3.yaml    T3 automatic-failover gate
scripts/
  t3-kill-loop-v3.sh    HTTP-polling leader-kill loop (no exec into pods)

Checkpoint status

ID Gate Status
T0 baseline (~90/s replicated, 3669/s sharded) ✓ done
T2-A ≥1000 quorum writes/s ✓ 2980/s
T2-B 0 acked loss across kills
T3 leader-kill failover <10s p99 ×10 ✓ max 6157ms (m11p5)
T4 scale 3→5→3 under load, joiner ≤5min next
T-read vector-search recall@k + query QPS/p99 not built (see PROCESS)
T5 sharded ≥5000 quorum writes/s blocked on p6