Implements tmp/tidaldb-fleet-hardening (20 planned tasks + 2 found by measurement). Ring 0 — restore verification. .woodpecker.yaml step pods ran at the namespace default of 1500m/2Gi, which OOMKilled a prior pipeline and starved the release gate past its budget. Both push-path steps now declare backend_options.kubernetes.resources as two YAML anchors declared once on their first consuming step. The values are CALIBRATED against measured free node capacity, not against the LimitRange max: `requests: cpu 2` (this roadmap's original figure) fits on NO node and would sit Pending forever, because `ci-build-bounds` grants permission and the nodes supply capacity, and those are not the same thing. The `nightly` cron described in this file for 216 days was never created, so tier-3 chaos, the fault classes, mTLS and the PITR test produced exactly zero signal while reading like standing coverage. nightly-chaos and nightly-security-ops now alias the anchors and have budgets matching the gate (their 120/90 were TIGHTER on the same runner, so they would have failed nightly for a budget reason, not a correctness one). nightly-soak is REMOVED, not scheduled: it drives 1000 rps for 600s gating on p99 <= 250ms, and the best node has 1700m free CPU, so it would fail on starvation rather than regression — manufacturing a nightly false alarm. Its commands move verbatim to docs/runbooks/nightly-soak.md. Ring 1 — four fabrications removed from the wire. - scatter_merge sorted and truncated without re-stamping rank, so /feed and /search returned 1,1,2 under full placement. Reuses merge_cross_shard's existing stamp; asserted on BOTH the multi-group merge path and the single-group [only] fast path that bypasses it. - aggregate_region_row's None arm invented `applied_events: 0` plus a deficit derived from it. applied_events/lag_events are now Option<u64>, null on the wire. leader_last_seq was also unwrap_or(0), so a node that could not reach the LEADER computed 0 - applied = 0 for every region and reported a converged cluster it had never measured — a fabrication pointing the dangerous way. - tidalctl inferred NO REPORT from `applied == 0 && lag > 0`. That heuristic was actively hiding the PVC-wipe shape: a measured zero with a real deficit rendered as "no report" instead of BEHIND. Now read off the wire; converged exits 0, partitioned still exits nonzero. - /sharded/* answered 201/204 for single-copy writes with nothing anywhere saying so. Now requires `x-tidal-ack: local`, rejecting with 400 via the existing invalid_input path. Six call sites migrated, not the two this roadmap predicted — including docs/runbooks/cluster.md §16.3, which told operators to run a quorum-write probe via POST /sharded/items. That probe cannot verify quorum: the surface applies locally with no WAL append. It was used as the safety check between every step of a staged deploy earlier today. Ring 2 — observability. JSON_LOGS was already implemented and the deployment simply never asked for it; the StatefulSet now sets it, plus TIDAL_SERVICE_NAME=tidaldb because enabling it silently renames the VictoriaLogs `service` stream field and would have blinded every query keyed on it. Adds tidaldb_usearch_replicated_vectors_total, incremented on BOTH the origin (wal_blob_first -> Ok(Some)) and the follower apply path — counting only the origin would mean each vector lands on exactly one node, replicas never agree, and the alert built on it pages forever. Found by measurement, not planned: the 401 path discarded every fact about every rejection. Traefik has served 101,858 rejected requests to the public ingress — 87.6% of all its traffic — with no record of who or why anywhere. unauthorized_response now emits reason (missing_token vs invalid_token, the distinction that separates a scanner from a rotation that missed a consumer) and the forwarded client. The token is never logged. Also: scripts/restore-fleet.sh --cluster started the soak monitor while deliberately leaving its gate suspended, orphaning a watcher that has reported "0/30 green nights" for 13 days. The pair now moves together. Doc-guard's three-warning backlog is cleared with real backfill for M4/M6/M12. Verified: fmt clean; clippy 5 crates 0 new warnings (74 vs 74 baseline, counted in a detached worktree at HEAD); lib 2110 passed; cluster_sharding 5; cluster_runbook 10; tidalctl 38; doc-guard 0 warnings. Playwright 32/34 with the two remaining failures asserting the rank fix against the not-yet-rolled image — they are the post-deploy proof.
99 lines
4.1 KiB
Markdown
99 lines
4.1 KiB
Markdown
# Soak: pre-release load and regression gate
|
|
|
|
Run this before tagging a release. It was a `nightly-soak` step in
|
|
`.woodpecker.yaml` until 2026-08-30 and moved here **unchanged** — the commands
|
|
below are the step's commands verbatim, so the coverage survives in a runnable
|
|
form rather than only in git history.
|
|
|
|
## Why this is not a CI step
|
|
|
|
The step never actually ran: the `nightly` cron it was gated on was never
|
|
created, so for 216 days it produced zero signal while reading like standing
|
|
coverage.
|
|
|
|
When the cron was finally configured, this step was **deliberately excluded**,
|
|
and the reason is a measurement rather than a preference. The soak drives
|
|
**1000 rps for 600s** and fails the build if **p99 > 250ms** or **errors > 1%**.
|
|
Measured free CPU on the k3s cluster, 2026-08-30:
|
|
|
|
| node | allocatable CPU | free CPU (requests) | free memory |
|
|
| --- | --- | --- | --- |
|
|
| `k3s-agent-1` | 4000m | 1700m | 4193Mi |
|
|
| `k3s-server-1` | 3000m | 435m | 2309Mi |
|
|
| `k3s-server-2` | 3000m | 550m | 2257Mi |
|
|
|
|
A p99 gate of 250ms cannot be met from 1700m of contended CPU shared with
|
|
production tidalDB. The step would fail nightly on **starvation, not
|
|
regression** — a false alarm every morning, which is the fake-coverage defect
|
|
inverted rather than fixed. A gate that cannot distinguish its own failure mode
|
|
from the thing it is watching for is not a gate.
|
|
|
|
So it runs **here**, by hand, on hardware where its numbers mean something.
|
|
|
|
## Where to run it
|
|
|
|
Anywhere with **≥ 2 dedicated cores** and no co-tenant under load. A developer
|
|
workstation qualifies; the shared k3s nodes do not. If you only have the k3s
|
|
cluster, the honest options are to raise its thresholds to match the hardware
|
|
(and say so in the output) or to skip it and record that you skipped it — not to
|
|
run it and read the result as meaningful.
|
|
|
|
## Run
|
|
|
|
```bash
|
|
export TIDAL_SOAK_RPS=1000
|
|
export TIDAL_SOAK_SECS=600
|
|
export TIDAL_SOAK_MAX_P99_MS=250
|
|
export TIDAL_SOAK_MAX_ERROR_PCT=1
|
|
# Soak target port — an UNCLAIMED slot in the project's reserved dev band
|
|
# (59520-59529): 59520=site, 59521=iknowyou, so the soak uses 59526.
|
|
export TIDAL_SOAK_PORT=59526
|
|
|
|
cargo build -p tidal-server -p tidal-stress
|
|
|
|
PORT="${TIDAL_SOAK_PORT}"
|
|
TARGET="${TIDAL_SOAK_TARGET:-http://127.0.0.1:$PORT}"
|
|
if [ -z "$TIDAL_SOAK_TARGET" ]; then
|
|
mkdir -p /tmp/soak-data
|
|
./target/debug/tidal-server standalone --listen "127.0.0.1:$PORT" \
|
|
--schema tidal-server/config/default-schema.yaml --data-dir /tmp/soak-data &
|
|
SRV=$!
|
|
# Reap the background server + scratch dir on ANY exit (success, gate failure,
|
|
# or the boot-failure exit below) so a re-run is idempotent.
|
|
trap 'kill "$SRV" 2>/dev/null; rm -rf /tmp/soak-data' EXIT
|
|
up=0
|
|
for i in $(seq 1 100); do
|
|
curl -sf "$TARGET/health/startup" >/dev/null 2>&1 && { up=1; break; } || sleep 0.3
|
|
done
|
|
# Fail FAST and unambiguously on a boot failure (bad schema, port already
|
|
# bound) rather than soaking a dead target and mislabeling it as an error-rate
|
|
# regression on the trend line.
|
|
if [ "$up" != "1" ]; then echo "soak target failed to start at $TARGET"; exit 1; fi
|
|
fi
|
|
|
|
./target/debug/tidal-stress --target "$TARGET" \
|
|
--ramp "${TIDAL_SOAK_RPS}:${TIDAL_SOAK_SECS}" --corpus 5000 --mix peach \
|
|
--json-summary soak-summary.json \
|
|
--max-error-pct "${TIDAL_SOAK_MAX_ERROR_PCT}" --max-p99-ms "${TIDAL_SOAK_MAX_P99_MS}" --fail-on-knee
|
|
|
|
cat soak-summary.json
|
|
```
|
|
|
|
> The `$${VAR}` escaping in the original CI step is **not** needed here.
|
|
> Woodpecker preprocesses a bare `${VAR}` before the shell sees it, so the step
|
|
> had to write `$${VAR}` to pass a literal through. In a plain shell the single
|
|
> `${VAR}` form above is correct.
|
|
|
|
## What to expect
|
|
|
|
`soak-summary.json` is the artifact — keep it with the release notes so there is
|
|
a trend line rather than a single opinion. A `--fail-on-knee` failure means
|
|
throughput stopped scaling before the target rps, which is a different finding
|
|
from a p99 breach and worth reporting separately.
|
|
|
|
## Targeting the live cluster
|
|
|
|
Setting `TIDAL_SOAK_TARGET` to the live Ref-A cluster and `TIDAL_SOAK_SECS=3600`
|
|
gives the GA-bar 1-hour 100k-DAU soak. **This soaks production at 1000 rps.**
|
|
Do not set it casually, and never as part of an unattended run.
|