tidaldb/tidal-stress/k8s/t5-readtput-job.yaml
jx12n a946c6128c fix(m12-rc13): read-SLA collapse + WAL_RETENTION_SEGMENTS 16 + tidalctl S3 DR
Read-SLA fix (rc12→rc13 — cpu-cgroup starvation → multi-second p99 + churning
elections):
- offload.rs: add SEARCH_GATE semaphore (core_count+1 permits, 50ms shed to 429)
  so per-shard searches gate on CPU, not reactor threads; concurrent scatter_merge
  fan-out (join_all) replaces the serial blocking offload_region_read loop
- node.rs: scatter_merge → async; per-shard futures run via offload_search
  (each acquires one SEARCH_GATE permit, moves it into spawn_blocking so the
  permit is held for the search's full CPU lifetime)
- main.rs: explicit tokio runtime with worker_threads floored at 4, independent
  of the cgroup quota — keeps the control plane (heartbeat/election/apply) on its
  own workers even when quota < 4
- k8s statefulset: CPU limit 2→3 (was: available_parallelism()=2 → only 2 async
  workers; search burst starved the reactor)
- tidal/wal/compaction.rs: WAL_RETENTION_SEGMENTS 4→16 (64 MiB→256 MiB per-shard
  catch-up window; a briefly-down follower across a rolling restart streams up
  instead of forcing snapshot reseed; disk floor 768 MiB/pod, self-trimming)
- cluster_reseed.rs: OFFLINE_ITEMS 1800→5600 to exceed the new 16-segment
  retention window (19 segs > 17); fix sequential quarantine/reseed race via
  await_status_bool

tidalctl S3/R2 backup DR:
- tidalctl/Cargo.toml: aws-config, aws-sdk-s3, aws-credential-types, tokio, tempfile
- commands/s3.rs: S3Target + export_dir (upload every file, manifest last as
  atomicity marker) + import_to_dir (download prefix into temp staging dir)
- commands/backup.rs: run_backup/run_restore accept Option<&S3Target>; S3 export
  is additive after local fsync barrier; S3 import stages into TempDir then runs
  the unchanged verified restore on it
- main.rs: --s3-endpoint / --s3-bucket / --s3-prefix flags; all-or-nothing
  endpoint+bucket validation; usage updated

tidal-stress/k8s: recall-rc12-spread-job, soak-nightly-cronjob, soak-monitor,
soak-results-pvc, t5-readtput-job manifests
2026-06-17 15:47:37 -06:00

68 lines
2.5 KiB
YAML

# T5 read-throughput (the measurable half): how many /vector_search ops/s the
# 3-node cluster sustains within SLA, spread round-robin across all 3 pods.
# (T5-as-written's 2.5x WRITE-scaling is structurally impossible on 3-node RF3
# full placement — proven in docs/profiling/m12p4-t5-sharded-throughput.md.)
apiVersion: batch/v1
kind: Job
metadata:
name: tidal-t5-readtput
namespace: tidaldb-cluster
labels: { app.kubernetes.io/name: tidal-stress, app.kubernetes.io/part-of: tidaldb }
spec:
backoffLimit: 0
ttlSecondsAfterFinished: 7200
template:
metadata:
labels: { app.kubernetes.io/name: tidal-stress, app.kubernetes.io/part-of: tidaldb }
spec:
restartPolicy: Never
automountServiceAccountToken: false
securityContext: { runAsNonRoot: true, runAsUser: 1000, runAsGroup: 1000, seccompProfile: { type: RuntimeDefault } }
containers:
- name: stress
image: registry.threesix.ai/tidal/stress:m12-rc7-seedretry
imagePullPolicy: IfNotPresent
args:
- --target
- https://tidaldb-0.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
- --target
- https://tidaldb-1.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
- --target
- https://tidaldb-2.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500
- --ca-cert
- /etc/tidaldb/tls/ca.crt
- --verify-recall
- --skip-seed
- --corpus
- "100000"
- --embedding-dim
- "1536"
- --recall-k
- "10"
- --recall-queries
- "1000"
- --read-p99-target-ms
- "10"
- --recall-target
- "0.95"
- --recall-ef-search
- "64"
- --ramp
- "1000:20,2000:20,3000:20,3800:20"
env:
- name: TIDAL_API_KEY
valueFrom: { secretKeyRef: { name: tidaldb-credentials, key: TIDAL_API_KEY } }
- name: TIDAL_STRESS_LOG
value: warn
resources:
requests: { cpu: "1", memory: 512Mi }
limits: { cpu: "3", memory: 2Gi }
securityContext: { allowPrivilegeEscalation: false, readOnlyRootFilesystem: true, capabilities: { drop: ["ALL"] } }
volumeMounts:
- { name: cluster-tls, mountPath: /etc/tidaldb/tls, readOnly: true }
volumes:
- name: cluster-tls
secret:
secretName: tidaldb-cluster-tls
items: [ { key: ca.crt, path: ca.crt } ]