tidaldb/tidal-stress
jx12n 727fbfcb6b fix(m12p6): 6-bug k3s 3-shard cluster repair (rc8+rc9)
Root-caused and fixed five sharding bugs exposed on the real k3s 3-shard
cluster (rc5→rc7), plus a divergent-rejoin reseed loop found in rc9:

1. reseed shard-awareness (Bug 3, keystone): `run_boot_install_for_region`
   visits each hosted group's own shard subdir; per-group leader discovery
   appends `?shard=N` so a divergent shard heals from its own leader (not
   shard-0's WAL/term — cross-shard contamination).
2. leader self-join term (Bug 4): `become_leader_for_term` now calls
   `note_self_won_term` so the elected shard's `joined_term` is set and
   `cluster_promote` routes rebalances correctly (was: topology-era mis-read
   → legacy fenced promote → 500).
3. boot self-heal self-pull guard (Bug 2): `leader_shard != my_shard` gate
   prevents a node pulling its own stream (its stream isn't a registered peer)
   → eliminates the `PeerUnreachable(self)` loop.
4. scatter-merge degraded partial (Bug 1): failed shard logs + continues
   instead of `?`-failing the whole read; bounded read-admission semaphore
   (`offload.rs`) sheds as 429 instead of piling into a 36s p99.
5. WAL retention (Bug 5): `compact_wal_retained` keeps `WAL_RETENTION_SEGMENTS=4`
   most-recent sealed segments; online path gets the same retention clamp.
   Prevents brief-restart forced-reseed.
6. divergent-rejoin reseed loop (Bug 6, rc9): `note_quarantined` latches
   `from_seqno = stream_baseline` (not `frontier + 1`) so `wal_covers`
   returns `needed=true` and the snapshot installs instead of looping.

Also: `TidalDb::close_shared` for deterministic HNSW save on cluster SIGTERM
(HNSW graph was not saved when request-scoped Arc clones were alive at shutdown);
updated profiling doc with full rc8/rc9 fix narrative; k8s recall job YAMLs.
2026-06-16 22:34:21 -06:00
..
benches feat(m12): vector retrieval G1/G2 — recall harness, ANN in RETRIEVE, index tuning 2026-06-14 11:07:09 -06:00
k8s fix(m12p6): 6-bug k3s 3-shard cluster repair (rc8+rc9) 2026-06-16 22:34:21 -06:00
scripts feat(k8s): m11p5 cluster manifest — local-path PVCs, initContainer, T3 tooling 2026-06-12 22:01:37 -06:00
src fix(m12p6): 6-bug k3s 3-shard cluster repair (rc8+rc9) 2026-06-16 22:34:21 -06:00
Cargo.toml feat(m11): continuous correctness (m11p9) — fault classes, invariant checkers, soak gates, nightly pipeline 2026-06-13 15:23:59 -06:00
PROCESS.md docs(tidal-stress): README + worklog + continue-the-process flow 2026-06-13 09:23:29 -06:00
README.md docs(tidal-stress): README + worklog + continue-the-process flow 2026-06-13 09:23:29 -06:00
WORKLOG.md docs(tidal-stress): README + worklog + continue-the-process flow 2026-06-13 09:23:29 -06:00

tidal-stress

Open-loop capacity ramp + chaos harness for the tidalDB cluster. Drives the thepeach feed workload (signals + vector embeddings) against a live cluster and reports a per-stage capacity verdict.

  • Worklog (what's been run, what we learned): WORKLOG.md
  • Process (how to run the next checkpoint): PROCESS.md

Current target cluster (as of 2026-06-13)

The cluster moved to the m11p5 single-StatefulSet architecture. The old 3-StatefulSet / static-ClusterIP model (namespace tidaldb, IPs 10.43.99.11-13) is retired — any manifest or doc still naming those IPs is stale.

Fact Value
Namespace tidaldb-cluster
Pods tidaldb-{0,1,2} (one StatefulSet, 3 replicas)
Peer DNS tidaldb-N.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500 (HTTP), :9601 (gRPC)
Client VIP tidaldb.tidaldb-cluster.svc.cluster.local:9500 (readiness-gated)
Server image registry.threesix.ai/tidal/server@sha256:173e803… (:m11p5)
Stress image registry.threesix.ai/tidal/stress@sha256:3a75c311… (:m11p3)
Storage local-path 5Gi/pod (on-node NVMe) — NOT Longhorn (see WORKLOG)
CPU/pod limit 2 (the write pool is ~2 workers on the leader)

Targets for any new Job manifest — use pod DNS for --target (so status-polling reaches survivors during a kill window) and the VIP for --leader-url:

args:
  - --target
  - http://tidaldb-0.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
  - --target
  - http://tidaldb-1.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
  - --target
  - http://tidaldb-2.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
  - --leader-url
  - http://tidaldb.tidaldb-cluster.svc.cluster.local:9500

Run pattern

Every run deploys the generator as an in-cluster Job (port-forward adds API-server serialization latency — never use it for capacity numbers; only the kill loop port-forwards, and only to read /cluster/status).

export KUBECONFIG=~/.kube/orchard9-k3sf.yaml
kubectl apply -f k8s/<job>.yaml
kubectl logs -f job/<job-name> -n tidaldb-cluster
kubectl delete job <job-name> -n tidaldb-cluster    # re-arm before re-running

CLI flags (authoritative — from src/main.rs)

Flag Default Notes
--target <url> (required, repeatable) Region gateway; reads round-robin across all
--leader-url <url> none Pin leader-path writes here to skip the forward hop
--api-key $TIDAL_API_KEY Bearer; the cluster requires it
--ack <leader|quorum> topology default Sent as x-tidal-ack per write
--ramp <preset|rps:secs,…> peach-100k Presets: smoke, quick, peach-100k, max
--stage-secs <n> 45 Hold per preset stage; 300600 for soak
--mix <preset|op=w,…> peach Presets: peach, reads, writes. Ops: feed,search,view,like,skip,item,embed
--write-path <leader|sharded> leader sharded removes the single-leader funnel (not replicated)
--corpus <n> 10000 Items+embeddings to seed; 20k for gate runs
--users <n> 50000 Virtual user id space
--skip-seed false Set after the first run of a session (corpus persists on PVC)
--embedding-dim <n> 128 Deployed schema = 128; thepeach real = 1536
--hot-skew <f> 1.3 Power-law concentration onto hot items
--poll-status false Poll /cluster/status between stages for lag — always set when measuring lag
--stop-on-knee false Stop at first SLO-breaching stage
--dau <n> 100000 DAU the verdict translates the ceiling against

SLO: feed p99 ≤ 150ms (network-hop allowance over the in-process 50ms SLA); error rate ≥ 1% (429/408/503/5xx/transport) = the knee.

Layout

src/            generator (scheduler, workload model, client, metrics)
k8s/            Job manifests — one per checkpoint
  stress-job.yaml       generic ramp
  stress-job-t2a.yaml   T2-A quorum throughput
  stress-job-t2b.yaml   T2-B acked-loss under kills
  stress-job-t3.yaml    T3 automatic-failover gate
scripts/
  t3-kill-loop-v3.sh    HTTP-polling leader-kill loop (no exec into pods)

Checkpoint status

ID Gate Status
T0 baseline (~90/s replicated, 3669/s sharded) ✓ done
T2-A ≥1000 quorum writes/s ✓ 2980/s
T2-B 0 acked loss across kills
T3 leader-kill failover <10s p99 ×10 ✓ max 6157ms (m11p5)
T4 scale 3→5→3 under load, joiner ≤5min next
T-read vector-search recall@k + query QPS/p99 not built (see PROCESS)
T5 sharded ≥5000 quorum writes/s blocked on p6