tidaldb/tidal-stress/WORKLOG.md
jx12n 46741a3a8c docs(tidal-stress): README + worklog + continue-the-process flow
README.md — tool reference pinned to the m11p5 cluster (tidaldb-cluster ns,
pod-DNS targets, local-path storage); retires the stale 10.43.99.11-13 ClusterIPs.

WORKLOG.md — m11p5 migration + T3 gate (10/10 kills, max 6157ms), the three
run-costing traps (Longhorn storage, exec leader-detection, kubectl wait hang),
reseed behavior, and the vector-DB positioning note. Records a m11p5 finding:
a reseeded node stays NotReady forever on an idle cluster because the
first-converged readiness check needs observed replication traffic — flagged as
an upstream fix, not a manifest bandaid.

PROCESS.md — the repeatable five-beat checkpoint loop, the operational-traps
table, the T4 elasticity flow (now unblocked on m11p5, with the idle-readiness
watch-item), and the T-read vector-search gate that still needs a recall oracle
built before any read-path/vector-DB claim.
2026-06-13 09:23:29 -06:00

7.1 KiB
Raw Permalink Blame History

tidal-stress worklog

Newest first. Chronicle of stress-test runs against the live k3s cluster, the gates they cleared, and the operational lessons each one cost.


2026-06-12/13 — m11p5 migration + T3 automatic-failover gate

What shipped

Rolled the cluster from m11p3/p4 (3 StatefulSets, static ClusterIPs) to m11p5 (1 StatefulSet × 3 replicas, headless-Service DNS, elastic membership). New namespace tidaldb-cluster. Then gated T3 (automatic leader-failover).

Result: T3 PASS — 10/10 leader kills, max election 6157ms (gate <10 000ms), zero acked loss, post-ramp lag=0.

Kill Victim → leader ms Kill Victim → leader ms
1 tidaldb-0 → tidaldb-2 2752 6 tidaldb-2 → tidaldb-0 2656
2 tidaldb-2 → tidaldb-0 4312 7 tidaldb-0 → tidaldb-1 3517
3 tidaldb-0 → tidaldb-2 2180 8 tidaldb-1 → tidaldb-0 2758¹
4 tidaldb-2 → tidaldb-0 3675 9 tidaldb-0 → tidaldb-1 2235
5 tidaldb-0 → tidaldb-2 2168 10 tidaldb-1 → tidaldb-2 6157

¹ Kill 8 measured from Raft logs (election won 03:55:42.758, kill 03:55:40) — the port-forward died when the VIP routed to the killed pod, so the loop's own timer recorded a false TIMEOUT. The election itself was clean.

Stress envelope during the gate: 1500 rps quorum writes, 900s, 1.35M ops, 1.51% error (all 503s = quorum rejections inside the ~24s kill windows, not acked losses), post-ramp lag=0 on all three nodes.

Three things that cost us a run each

  1. Longhorn storage → 7485% error rate. The new cluster first provisioned Longhorn 10Gi PVCs (networked storage). Every quorum write needs follower fsync; over Longhorn that's a network hop on the commit path, which collapsed throughput. Two full T3 runs were invalid before we spotted it. Fix: storageClassName: local-path, 5Gi — matches the on-node NVMe that T2-A was measured on. The first valid run came straight after the switch.

  2. kubectl exec leader-detection died on the killed pod. The v1/v2 kill loops exec'd into a pod to read cluster status; when that pod was the one being killed (or mid-initContainer on reboot), exec returned container not found and the loop logged bogus 1719s "election" times. The status API also returns "none (election in progress)" as the leader field mid-election, which the loop tried to kubectl delete. Fix: v3 uses a port-forward to the VIP + curl /cluster/status with a strict ^tidaldb-[0-9]+$ leader regex — never execs into a tidaldb pod.

  3. kubectl wait --for=condition=Ready hangs across a force-delete+recreate. When a pod is force-deleted and the StatefulSet recreates it with the same name, kubectl wait latches the old object and blocks past its own --timeout. Had to kill the hung wait twice (kills 7 and 9) to let the loop advance. Kills 910 were driven by hand with a fresh port-forward.

The reseed observation (expected, not a bug)

At 1500 rps a ~50s pod outage opens a WAL gap large enough that the returning pod can't catch up by streaming segments — it latches reseed_required (m11p5 §2.4), self-exits, and installs a full snapshot on its next boot. Snapshot install takes ~120150s, during which the pod is 0/1 Running. Quorum holds throughout (2/3 voters always up), and a reseeding pod still participates in elections — kill 4 elected tidaldb-0 in 3675ms while it was mid-reseed. So the slow K8s readiness is orthogonal to election health; the gate measures election time from the Raft log, not pod readiness.

FINDING — reseeded node stays NotReady forever on an idle cluster (m11p5 bug)

Surfaced 11h after the gate, verifying cluster state. tidaldb-1 (kill 10's victim) installed its snapshot cleanly at 03:59:42 (snapshot_seq=1441537) and was a healthy follower from that moment — /cluster/status showed it reachable, lag_events=0; /health returned ok:true. But K8s kept it 0/1 Running for 11.5 hours. The readiness flag did not flip to ready until 15:21:25, when it finally logged joiner first-converged … readiness is now sticky-ready.

Root cause: the install-boot readiness predicate (§4 hysteresis) flips ready only when it observes lag <= learner_promote_lag through the catch-up receiver/ship path. The stress job ended right after kill 10, so the cluster went idle — no new events were shipped, so the convergence-observation path was never exercised, and readiness never flipped. A fully-caught-up node sat out of the client VIP for 11h on a technicality.

Impact: in production this bites a scale-up during a quiet period — the new node installs, catches up, but never joins the client Service until write traffic resumes. It also means T4's joiner-Ready timing is only valid under sustained load; an idle-window join would hang the readiness gate.

This is an upstream tidaldb fix, not a stress-harness workaround: the first-converged check should fire on a caught-up node even with zero inbound replication (e.g. evaluate against a leader heartbeat / committed-index compare, not only against observed ship traffic). Do not bandaid it in the manifest. Filed as a T4 watch-item in PROCESS.md.

Strategic note (vector-DB positioning)

Discussed where tidaldb sits vs other databases. CockroachDB is a different animal (distributed SQL) — not the comp. The real comp is vector DBs (Qdrant/Milvus/Weaviate/pgvector), and the honest read is:

  • tidaldb's differentiator is synchronous quorum durability + <6s auto-failover on the write path — which almost no vector DB has (they treat vectors as a rebuildable derived index and skip write durability). That's a real niche when the signals are source-of-truth (thepeach view/like/skip are).
  • We should not race them on ANN QPS/recall/ingest — a dedicated engine wins that and should.
  • Gap: every number we have is write-path only (--mix writes). We have zero read-path data — no recall@k, no query QPS/p99. That's exactly the axis vector DBs are judged on, and the tool can measure read latency/throughput but cannot measure recall (no ground-truth oracle). Building that is the T-read checkpoint in PROCESS.md.

2026-06-11 — m11p4 regression + m11p3 T2 gate

m11p3 (quorum acks, one replicated log) cleared T2-A (2980 quorum writes/s, p99=49ms, knee ~5k rps) and T2-B (3 leader kills, lag=0 after each, zero acked loss). m11p4 (catch-up timer retry + WAL TSEG segment versioning) held the gate with no regression (2975 writes/s, 0.2%). Full results: docs/ops/stress-test-p3-t2.md, docs/ops/stress-test-p4-regression.md.

This was still on the old 3-StatefulSet / static-ClusterIP architecture — the 10.43.99.11-13 IPs in tmp/stress-test-plan.md date from here and are now stale (see PROCESS.md for the migration).


2026-06-10 — T0 baseline

Replicated /signals ~90/s (WAL fsync + synchronous gRPC ship per write, ~2 workers on a cpu-2 pod); sharded /sharded/signals 3669/s at 0 errors, knee not reached. Full results: docs/ops/stress-test-thepeach.md.