Root-caused and fixed five sharding bugs exposed on the real k3s 3-shard cluster (rc5→rc7), plus a divergent-rejoin reseed loop found in rc9: 1. reseed shard-awareness (Bug 3, keystone): `run_boot_install_for_region` visits each hosted group's own shard subdir; per-group leader discovery appends `?shard=N` so a divergent shard heals from its own leader (not shard-0's WAL/term — cross-shard contamination). 2. leader self-join term (Bug 4): `become_leader_for_term` now calls `note_self_won_term` so the elected shard's `joined_term` is set and `cluster_promote` routes rebalances correctly (was: topology-era mis-read → legacy fenced promote → 500). 3. boot self-heal self-pull guard (Bug 2): `leader_shard != my_shard` gate prevents a node pulling its own stream (its stream isn't a registered peer) → eliminates the `PeerUnreachable(self)` loop. 4. scatter-merge degraded partial (Bug 1): failed shard logs + continues instead of `?`-failing the whole read; bounded read-admission semaphore (`offload.rs`) sheds as 429 instead of piling into a 36s p99. 5. WAL retention (Bug 5): `compact_wal_retained` keeps `WAL_RETENTION_SEGMENTS=4` most-recent sealed segments; online path gets the same retention clamp. Prevents brief-restart forced-reseed. 6. divergent-rejoin reseed loop (Bug 6, rc9): `note_quarantined` latches `from_seqno = stream_baseline` (not `frontier + 1`) so `wal_covers` returns `needed=true` and the snapshot installs instead of looping. Also: `TidalDb::close_shared` for deterministic HNSW save on cluster SIGTERM (HNSW graph was not saved when request-scoped Arc clones were alive at shutdown); updated profiling doc with full rc8/rc9 fix narrative; k8s recall job YAMLs.
30 KiB
M12 cluster deploy + recall findings (k3s, 1536-dim)
⇒ STATUS (2026-06-17): the m12p4 3-shard layer is now FIXED in rc8
Operator deployed M12 to the real 3-shard k3s cluster (rc5→rc7), clean-slate
re-seeded a COMPLETE 100k/1536 corpus, then root-caused and FIXED the five
m12p4 3-shard catch-up/read/reseed bugs (each adversarially verified for
failover-safety before implementation). m12-rc8
(sha256:5e3e5128562d5de98e73809f468ed6654e3445fecaf486f7879d9e9c6c553cb5).
Workspace lib tests green (tidal-server 147/147, tidaldb 1910/1910; compaction
17/17; S=1 reseed 3/3). The 5 fixes:
-
Bug 3 (keystone) — reseed boot-install was never shard-aware in S>1.
run_boot_installran ONCE on the PARENT/data/db, but the per-group marker- divergent WAL live in
/data/db/shard-NNNNN/→ it found no marker, no-op'd, and the divergent shard re-quarantined + self-restarted forever (186 loops). Fix: newrun_boot_install_for_regionruns the install per hosted shard subdir, andrun_boot_installnow takes aShardId—candidates_forresolves THAT group's per-shardgrpc_addr,discover_leaderappends?shard=N— so a divergent shard heals from ITS OWN leader (the original directory-only idea would have installed shard-0's WAL/term into shard-N: cross-shard contamination, caught in verification).node::shard_subdir→pub(crate)(single formatter).main.rscalls the per-region entry. (reseed.rs,main.rs,node.rs)
- divergent WAL live in
-
Bug 4 — an elected LEADER never recorded its own
joined_term(only the followerjoin_term_checkset it), socluster_promoteread the elected shard as topology-era and routed/cluster/shards/{id}/transferdown the legacy fenced leg → 500. Fix:become_leader_for_termcalls a new monotonicElectionRuntime::note_self_won_term(term). (node.rs,election_driver.rs) -
Bug 2 — boot self-heal self-pulled
my_shard(a node can't serve its own stream to itself) →PeerUnreachable(self)forever. Fix: guardleader_shard != my_shard. The "never reconverges" half was Bug 3 (a quarantined node clears only via the now-working reseed). (node.rs) -
Bug 1 — cross-shard read collapse (p99 36s, 500-storm).
scatter_mergefail-fasted (?) on one shard's error andoffload_readhad no admission bound. Fix:scatter_mergeserves the surviving groups as a degraded partial on a per-shard error (logs it; the recall harness still measures true recall); a bounded read-admission semaphore sheds 429 instead of piling onto the blocking pool into a 36s p99. (Recall recovery itself comes from Bugs 2+3 restoring the empty shards.) (node.rs,offload.rs) -
Bug 6 (rc9) — divergent-rejoin reseed loop. Surfaced when a deposed leader (frontier AT/ABOVE the new term's baseline, with old-term data) rejoins: the
decide_join→ReseedRequired/Quarantinelatch usedfrom_seqno = frontier+1, which lands ABOVE the leader's WAL tail, sowal_coversanswersneeded=false→ the reseed install no-ops → the divergent suffix is re-detected → self-restart loop. Fix: latchfrom_seqno = stream_baseline(threaded throughjoin_term_checkfrom the heartbeat), which is<= baselinesowal_coversreturnsneeded=true→ the snapshot installs and re-baselines onto the leader's committed history (discarding the divergent suffix — uncommitted writes the cluster elected past, so no acked-write loss).m12-rc9(sha256:818c923368ecca6e433c42a13672132e9b467c449fa23c035201fe7189f1f52b). (election_driver.rs,node.rs note_quarantined) -
Bug 5 — WAL over-compaction forced reseed on a brief restart. Compaction reclaimed every sealed segment below the checkpoint, so a follower down across one checkpoint hit the compacted-below refusal. Fix: retain the
WAL_RETENTION_SEGMENTS=4most-recent sealed segments past the checkpoint (online + a newcompact_wal_retainedfor shutdown); a too-far-behind follower still correctly reseeds. (tidal/src/wal/compaction.rs,tidal/src/db/lifecycle.rs)
Original handoff detail (pre-fix) retained below for context.
⇒ DEV HANDOFF (2026-06-16, pre-rc8)
Operator deployed M12 to the real 3-shard k3s cluster (rc5→rc7) and clean-slate re-seeded a COMPLETE 100k/1536 corpus. Five fixes landed in-tree (working-tree diff; rc7 image is built from them) and are verified live; the m12p4 sharded catch-up/read layer is NOT production-ready on k3s and is handed back to you.
FIXED + verified (keep these):
| fix | file | what it cures |
|---|---|---|
TidalDb::close_shared(&self) called explicitly from ShardReplica::shutdown |
tidal/src/db/lifecycle.rs, tidal-server/src/cluster/node.rs |
HNSW graph save never ran on cluster SIGTERM (relied on drop(db) hitting refcount 0; lingering request clones prevented it). Now runs. |
concurrent per-shard shutdown (thread::scope) + grace 60→600s + TIDAL_SHUTDOWN_DRAIN_MS=3000 |
node.rs (ClusterNode::shutdown), k8s/cluster/statefulset.yaml |
3 shards saved SERIALLY overran the grace → SIGKILL mid-save. Now all 3 save concurrently in-window (verified: 3× "persisted HNSW graphs to disk"). |
count_alive_other_voters uses the CA-trusting self.blocking_client |
node.rs ~2820 |
reseed self-restart quorum poll built a BARE reqwest client → every TLS peer poll failed → alive_voters=0 → reseed NEVER auto-healed on a TLS cluster. Verified: now alive_voters=2, self-restart fires. |
| seed retry: 8→40 attempts, backoff→500ms (~12s budget) | tidal-stress/src/client.rs retry_write |
a fast bulk seed silently DROPPED ~0.4–4% of the corpus when a leader shed 429. Now seeds 100000/100000. |
| single-shard "blocker #3" was a STALE ConfigMap (shards: commented out live) | k8s/cluster/topology-configmap.yaml (re-applied) |
NOT a code bug. |
OPEN — m12p4 sharded path, needs your repair (gates T5/T4/1M-on-shards BLOCKED):
- Cross-shard
/vector_searchcollapses even on a healthy follower with all 3 shards LOCAL + lag=0: p99 36 s, recall ~0.42 (≈ consulting one shard), 70–100% error at ≥100 rps. The local scatter-merge / read-offload path (vector_search→scatter_mergeoverhosted_dbs, node.rs ~7204) hangs and returns partial. This is the gate-blocker — reads are the SLA. - Catch-up "peer shard sN not in configuration" (
tidal-netPeerUnreachable, error.rs:13): a rejoining node's per-shard gRPC pool never registers the source peer → its catch-up stream can't open → it never reconverges (seen on shards 0/2). - term-0 topology-era divergence trap: a leader that took term-0 bootstrap writes, hit by a term transition under the seed, gets a divergent suffix → self-restarts (correct now), but then STALLS rejoining on (2) + a term-behind refusal. Net: self-restart fires but the node can't actually rejoin the 3-shard cluster (quorum survives 2/3; it does not self-repair to 3/3).
- 3-shard leader-transfer verb 500s for an already-elected shard
(
cluster_promote/promote_localroutes the fenced transfer through the legacy leg, fenced atjoined_term≥1). - WAL compacts past a briefly-restarted follower → forced reseed instead of cheap stream catch-up (periodic-checkpoint compaction is too aggressive for the rollout window).
Recommendation: the single-shard topology has a working read path (prior
session: recall 0.967, p99 7 ms @ 100k) — G1/G2/T4/1M reads are provable there
today. The sharded gates need (1)+(2)+(3) fixed first. Also note the rc6 grace bump
to 600s makes rolling restarts slow on a loaded cluster (the graph save is the long
pole); a periodic graph checkpoint (the checkpoint_graphs helper is reusable from
the periodic thread) would let the grace drop back and survive SIGKILL too.
Date: 2026-06-15. Cluster: orchard9-k3sf, namespace tidaldb-cluster, 3 nodes
(4 vCPU / 16 GiB each), local-path PVCs. Image built from HEAD 8e39ee1 + the
working-tree T4-TLS fixes. Goal: deploy M12 to the real cluster and get the
real recall@10 + read p99 at the 1536-dim production shape (the m11p6 image
404'd /vector_search).
Headline result (real, live cluster)
100k items / 1536-dim, recall@10 vs a brute-force cosine oracle (1000-query pool),
/vector_search pure k-NN against one node:
| metric | value | gate | verdict |
|---|---|---|---|
| recall@10 | 0.9869 | ≥ 0.95 | MET |
| read p99 @ knee | 8.71 ms @ 200 rps | ≤ 10 ms | MET (at 100k) |
| read knee (single node) | ~200 rps | — | saturates fast |
Beyond the knee a single node sheds/errors hard: 62% error at 500 rps (achieved 267), 100% at 1000+. 1536-dim k-NN reads saturate one node at ~250 rps.
Honest scope:
- This is 100k, not 1M. The G1 gate is p99 ≤ 10 ms @ 1M/1536-D — still unproven; at 1M the candidate set grows and p99 rises. 100k clears it at 8.71 ms.
- recall 0.9869 is below the standalone-laptop 0.9997 because the measured node's HNSW carries tombstones from the upsert-reseed (remove+add, below) — a freshly-built index has no tombstones and recalls higher. Re-measure on a fresh index for the canonical number.
- Measured against a single healthy replica with the corpus already resident
(
--skip-seed), because the cluster's write path is unstable at 1536-dim (below) — a full clean seed+query could not complete.
Bug found + fixed: USearch insert was not an upsert (reseed wedge)
VectorIndex::insert's contract is "Replaces any existing vector with the same
ID." The USearch backend (tidal/src/storage/vector/usearch_index.rs) violated
it: it computed is_new but then unconditionally add, which a multi:false
index rejects with Duplicate keys not allowed in high-level wrappers whenever
the key exists.
Impact: a reseeding follower's post-snapshot WAL replay re-applies embeddings the
snapshot already loaded. Every overlapping embedding failed add →
WAL blob replay failed → applied_events never advanced past the snapshot
frontier → the catch-up receiver deadlocked on a backlog it could not apply
(inbound channel full forever). With all three replicas behind the compaction
horizon simultaneously, none would self-restart (quorum protection) → the cluster
was unrecoverable.
Fix: make insert a true upsert — remove the existing key, then add. Added a
regression test (usearch_insert_replaces_existing_key). Verified in production:
operator-forced rolling reseed (followers first from the ahead leader, leader
last) completed with dupkey=0 and converged 3/3 to lag=0, no data loss.
Tradeoff: remove+add tombstones the old slot, so a reseed leaves the index
tombstoned (the recall dip above). A periodic compaction/rebuild reclaims it.
Open upstream issues (NOT bandaided — flagged for the devs)
-
1536-dim write path destabilizes the cluster. The recall seed's quorum write-burst (100k embeddings) trips the leader's partition detector: it marks the followers
partitioned:true, reachable:false(the followers actually applied the data and self-report healthy) → leader believes it lost quorum → writes stall AND reads on the leader time out (HTTP 408, the cross-region gather waits on followers it thinks are gone). The view does not self-heal; only a leader restart + re-election clears it. Root cause is almost certainly the replication ship/ack path not keeping pace with 1536-dim apply (each apply is an HNSW insert at ef_construction=400), tripping heartbeat/ack timeouts. This blocks ingesting the production-shape corpus and is the top blocker to "M12 production-ready at 1536-D." -
HNSW index is rebuilt on every boot, not persisted. Open logs
using USearch (HNSW) backend ... count=100000then spends ~5.5 min single-core rebuilding the graph before binding HTTP. 1M would imply a ~50–70 min boot. The long outage is what let the WAL compact past restarting nodes and triggered the reseed cascade in the first place. Persist/mmap the graph so boot is a load, not a rebuild. (Mitigation applied: startup probe budget raised to 20 min; pod memory limit 2Gi→4Gi — 100k/1536 peaks ~1.9 GiB.) -
The 3-group
shards:topology runs as a single shard at runtime./cluster/statusreports one shard (shard 0, RF=3) on every node despite the topology'sshards:block. Relevant to the T5 sharded-throughput gate — the sharding is not actually engaging.
Deploy mechanics (for repeat)
- Cross-compile on macOS-arm64 → x86_64-linux (GCC 15.2 = glibc 2.41 → trixie
runtime; binary NEEDs
libmvec.so.1, absent on bookworm). Runtime imagedebian:trixie-slim+libstdc++6/libgcc-s1, built via theamd64builderdocker-container buildx builder (QEMU). Pushed toregistry.threesix.ai. - Two-tier PKI (
k8s/cluster/certs.yaml, working-tree) applied cleanly:ca.crtis now a real CA (CA:TRUE); leaf carries the wildcard SAN + ready-only client Service SAN (needed for T4 scale-to-5). - Deployed image:
registry.threesix.ai/tidal/server:m12-rc2(sha256:1cfa30e5938002a881545b60f21cdfa50af6f87be8190d1ed4b1ca5a83d5bc3b) = m12 + the upsert fix.
rc3 update (2026-06-15): three fixes + verification
m12-rc3 (sha256:bceadce5419d5fcd78d3491f9d53cbd4055ee131f5ec7e9efa75b0f36474d1f7)
= rc2 + three changes (built/deployed; combined cargo check green; per-area
unit/chaos tests pass):
-
G1 index tuning — DONE & verified at 100k. Grid search (in-process, clean index) at 1536-D:
ef_search200→64 holds recall@10 0.998 at ~half the p99 (2434µs→1221µs); Int8 disqualified (recall collapses to ~0.71). Default set:ANN_DEFAULT_EF_SEARCH=64(candidate_gen.rs). End-to-end on the cluster (100k/1536, follower target): read-knee moved 200→500 rps, p99 7.06ms, recall@10 0.967 (tombstone-depressed; clean index ≈0.998) — G1+G2 MET at the 100k knee. 1M still unproven (blocked by the ingest issue below). -
False-partition status view — FIXED & verified. Under the 100k/1536 write burst,
partitioned_peersstayed 0 throughout (previously the leader falsely flagged followers partitioned). The detector now uses a gRPC last-contact liveness signal; genuine-partition detection intact (chaos tests pass). -
HNSW persist-on-boot — IMPLEMENTED, NOT yet effective on this cluster. Two real gaps found in verification: (a) the save is wired into
Db::shutdown_innerbut the k8s SIGTERM path doesn't reach it — no{data_dir}/vector/file appeared after akubectl delete pod(fix in progress: wire the save into the cluster serve SIGTERM shutdown); (b) this cluster reseeds on almost every restart (term-mismatch / accumulated PVC churn), and the reseed path rebuilds from the snapshot's raw vectors, bypassing graph-load. Both must be closed for the persist fix to pay off here.
NEW blockers surfaced by rc3 verification (the real 1536-D ingest wall)
The false-partition fix removed the view symptom, but the underlying defect — the leader's HTTP/gather plane degrades under 1536-D load — has two more faces that still block production-shape ingest:
- Write-forward timeout under burst. Seeding 100k to a follower target: the
follower→leader HTTP forward (
/items,/embeddings) succeeds ~1950 writes then every forward times out (leader unreachable ... 503, latency 15000ms) and the seed stalls. On an IDLE cluster the same forward takes ~47ms — it is purely load-induced HTTP-plane saturation on the leader. (Fix in progress.) - Leader read-gather hang. Even idle,
/vector_searchon the leader times out (HTTP 000 @ 8s) while followers serve it in ~50ms — the leader's cross-region/shard read gather hangs. Reads must currently be sent to followers.
Net: at 1536-D the leader cannot ingest a sustained quorum write load nor serve vector reads; followers serve reads fine. This is the top blocker to T4/T5/1M and to a clean (tombstone-free) recall number. Root cause is the apply/HTTP-plane contention; the status-view fix was necessary but not sufficient.
rc4 (2026-06-15): ingest admission-control + SIGTERM bounded-drain
m12-rc4 (sha256:28d0c45e59c487cf0d5e6a574c4c9997b82c6a358d67e74f326ff2c0c0925193)
adds the two gap fixes:
- Ingest admission control + forward retry (fixes the write-forward timeout).
Leader
/items+/embeddingsapply now goes through the boundedClusterWritePool(was unboundedspawn_blocking) → sheds as fast 429+retry_after instead of timing out; the follower→leader forwarder retries 3× honoring 429/backoff within a 28s budget (< the 30s route timeout). Idempotent upserts make retry safe. 147 tidal-server tests pass. - SIGTERM bounded-drain (fixes the persist-save not running). The cluster
with_graceful_shutdownblocked forever on sibling keep-alive connections → 60s grace → SIGKILL → noDrop→ no graph save. rc4 adds a 15s drain deadline- an
&self-reachable deterministic close (ShardReplica.db→ArcSwapOption) socheckpoint_embedding_graphsalways runs. Real multi-process SIGTERM e2e test asserts the shutdown checkpoint is written.
- an
STILL-OPEN catch-up/term bugs (surface under restart churn; NOT yet fixed)
The rc4 rollout exposed two more pre-existing replication bugs that make every rolling restart a slog on these (heavily-churned) PVCs — distinct from anything the four fixes above address:
- Term-gate rejects same-term catch-up chunks. A reseeding/catching-up node
pulls a catch-up stream and the chunks are refused:
catch-up chunk refused by the term gate ... chunk_term=10 rejection=Stale { current_term: 10 }—chunk_term == current_termshould NOT be stale, but it is, so the pull retry-loops forever (the node never converges via the stream). Workaround: restart the node so it takes the snapshot-install path instead of the stream. - Reseed-on-restart churn. A node restart that coincides with a term bump (e.g. the leader restarting during a rolling update) term-mismatches on rejoin and does a full snapshot reseed (rebuild-from-snapshot, ~5-7min) rather than a cheap WAL catch-up — even when only a few events behind. This both slows every rollout AND bypasses the new HNSW graph-load (reseed rebuilds from raw vectors).
These two are the next upstream targets: a correct term-gate (accept == term)
plus a streaming catch-up that actually converges would remove the reseed churn,
let restarts be fast clean reloads, and make the HNSW-persist fix pay off.
rc4 cluster verification (2026-06-15)
- Ingest admission-control + forward-retry — VERIFIED. The 100k/1536 seed
directed at a FOLLOWER (forwards to leader) now commits steadily (leader
applied climbed monotonically ~52k events over 10 min) with zero
forward to leader failed503s, zero false partition (partitionedstayed 0), and only brief graceful 429 backpressure — exactly the design. The pre-fix build hard-stalled at ~1500 writes in 90s. (Throughput is slow only because it upserts the existing corpus; correctness/liveness is the point.) This unblocks the 1536-D ingest path for T4/T5/1M. - HNSW persist-on-boot — NOT yet effective on the cluster. Despite the rc4
SIGTERM bounded-drain fix (which makes the shutdown path RUN — the e2e test
asserts
checkpoint.metais written), a clean follower restart on the real 1536-D cluster still showedusing USearch backend(rebuild) and no{data_dir}/vector/file — the graph save does not materialize via the SIGTERM→bounded-drain→ArcSwapclose→Drop→shutdown_innerpath, even though Agent's direct-shutdown()unit test (dim=4000) writes the file. This is a reconciliation gap between the bounded-drain/ArcSwap-close and the graph-save-in-shutdown_inner(Drop may not run the full shutdown body), on top of the reseed-churn that would bypass graph-LOAD anyway. Implemented + unit-tested, but not production-verified.
Net state after the fan-out (G1 + 2 cluster fixes)
| item | status |
|---|---|
| G1 ef_search=64 tuning | DONE — verified 100k (knee 500 rps, p99 7.06ms, recall 0.967≥0.95); 1M unproven |
| Fix #1a false-partition view | DONE — verified (partitioned_peers=0 under the write burst) |
| Fix #1b ingest saturation (forward-503) | DONE — verified (seed commits steadily, 0 forward-503) |
| Fix #2 HNSW persist | implemented + unit-tested; cluster save doesn't materialize (open) |
term-gate same-term Stale reject |
open (workaround: restart→snapshot-install) |
| reseed-on-restart churn | open (term-mismatch → full snapshot reseed) |
Deployed: m12-rc4
(sha256:28d0c45e59c487cf0d5e6a574c4c9997b82c6a358d67e74f326ff2c0c0925193).
rc5 + CLEAN-SLATE 3-shard rebuild (2026-06-16)
Operator-authorized clean-slate wipe (StatefulSet + 3 PVCs deleted, kustomization
re-applied) → fresh 3-shard cluster on m12-rc5
(sha256:8792b283d6c4d41c0b303afa777fa1a7583a94cca932d73580132c3a0e60e01b).
This removed the accumulated-PVC-churn confound and enabled the 3-shard topology
that the single-shard layout could not host in place.
rc5 = rc4 + the persist-save fix. Root cause of "Fix #2 doesn't materialize":
ShardReplica::shutdown relied on drop(db) reaching TidalDb::Drop (refcount
0) to run shutdown_inner→checkpoint_embedding_graphs. Under load a
request-scoped clone keeps the Arc alive, so Drop never fires and the graph
is never saved. Fix: a new pub fn TidalDb::close_shared(&self) (idempotent via
the closed CAS) that ShardReplica::shutdown calls EXPLICITLY on the shared
handle — the checkpoint now runs regardless of refcount. Verified:
m12p6_graph_persistence (save/load roundtrip) + cluster_graph_persistence
(multiproc SIGTERM reaches the close) both green; native cargo check clean.
Blocker #3 (single-shard at runtime) — RESOLVED, was a stale ConfigMap. The
LIVE tidaldb-cluster-topology ConfigMap had the shards: block COMMENTED OUT
(the on-disk manifest had it enabled but was never re-applied). The fresh deploy
re-applied it: /cluster/status now reports 3 shards, every pod carries
shard-00000/01/02 subdirs, and a 100k seed hash-balances ~even across them
(≈42k events/shard). NOT a code bug.
Ingest is FAST on a fresh 3-shard cluster. The 100k seed (200k writes: items + embeddings) committed at ~680 events/s aggregate, lag 0 throughout — vs the prior cluster's ~80–200/s. The old slowness was upsert-tombstoning over an existing corpus + single-shard, not a fundamental ceiling. ⇒ a 1M seed is ~25 min, not hours.
Still-open (noted, not yet fixed)
- 3-shard leader-transfer verb 500s on a fresh mixed-era cluster. On the
parallel fresh boot tidaldb-0 won leadership of ALL 3 shards.
POST /cluster/shards/{id}/transferto rebalance returns 503 ("promotion target returned 500"): for a shard already in an elected term (≥1) the fenced transfer routes through the internal legacy-promote leg, whichpromote_localfences atjoined_term ≥ 1→ 500. Operator-convenience verb (not a correctness issue); the cluster serves fine with one pod leading all shards. Deferred — the fix is in the delicate election machinery the code repeatedly warns against duplicating. - term-gate same-term
Stale+ reseed-on-restart churn: to be re-evaluated on the clean cluster (hypothesis: largely PVC-churn cruft).
Fix #2 persist — root-caused on the cluster, rc6 (2026-06-16)
Restarting a follower (tidaldb-2, ~32k vectors/slot/shard) on rc5 showed the
persist SAVE now runs (rc4 wrote nothing): /data/db/shard-00000/vector/
held a complete 141 MB item__content_vector.usearch.tmp — close_shared
executed checkpoint_graphs. But the .tmp was never renamed to the final
.usearch, and shards 1/2 had no vector/ dir at all → the boot correctly
REBUILT (load_persisted_slot rejects a .tmp). Root cause: ClusterNode::shutdown
saved the 3 shards serially (node.rs:4161), and a single 1536-dim slot's
USearch serialize+fsync is the long pole, so 3 shards overran the 60s SIGTERM
grace and k8s SIGKILLed mid-save.
rc6 = rc5 + (a) ClusterNode::shutdown now saves the hosted shards
concurrently via thread::scope (each ShardReplica::shutdown is &self
and touches only its own db/shard); (b) terminationGracePeriodSeconds 60→600
and TIDAL_SHUTDOWN_DRAIN_MS=3000 so the save starts promptly and has ample
budget. m12-rc6
(sha256:8fc56cb8e9014812df3de81fbdb4cfed6fdecd3a19ec7c2eb836a3defebdf068).
Open: a SIGKILL/crash still skips the save (boot rebuilds) — the production-robust
follow-up is a periodic graph checkpoint (the checkpoint_graphs helper is
already reusable from the periodic thread; not yet wired because the save is
heavy and needs an interval + change-detection design).
Seed completeness — single-leader 429 saturation
The fresh parallel boot left tidaldb-0 leading ALL 3 shards (the election race;
the per-shard transfer verb 500s, above). A 100k recall seed then dropped
4396/100000 items: every write funnels through tidaldb-0's ONE bounded write
pool, which sheds 429 under sustained load; the stress seed retries 429 only 8×
(~900ms total backoff) and then drops. Real product behavior (backpressure), but
it means a single-leader cluster can't absorb a fast bulk seed. Mitigations:
balanced leaders (3 pools) and/or lower --seed-concurrency. The rc6 rollout
terminates tidaldb-0 LAST, so leadership fails over OFF it — a free rebalance.
rc7 (2026-06-16): self-restart auto-heal + the wedge it exposed → 2nd wipe
The rc6 rollout (the FIRST restart of a loaded 3-shard cluster) exposed three real catch-up/reseed bugs (NOT cruft — clean cluster):
- WAL compacts past a briefly-restarted follower (shard 1: "WAL compacted below seqno 1, earliest 63873") → forced reseed instead of cheap stream catch-up. Aggravated by (3).
- Reseed self-restart ALWAYS refused on a TLS cluster —
count_alive_other_votersbuilt a BAREreqwest::blocking::Client(no cluster CA), so everyhttps://peer:9500/cluster/status/localpoll failed TLS verify →alive_voters_excluding_self=0→ refuse (would "break quorum") → a node that needed a reseed wedged forever. FIXED (rc7): use the CA-trustingself.blocking_client. VERIFIED live: the same path now logsalive_voters_excluding_self=2and the self-restart FIRES. - Shard-2 catch-up "peer shard s2 not in configuration" — the per-shard gRPC pool lacks the source peer while reseed-pending; likely downstream of (2) (incomplete group membership), to re-confirm now that (2) auto-heals.
Persist save VERIFIED on the cluster (rc6): a self-restart logged
persisted HNSW graphs to disk for ALL THREE shards (shard-00000/01/02)
concurrently within the grace window — the rc6 parallel-save + 600s grace works.
But: the rc7 self-restart then LOOPED on the churned cluster — a divergent
term-0 suffix (tail_term=0 vs elected term=1, frontier 65108) that the boot
snapshot-install did not clear → quarantine→self-restart→re-detect→loop. This is
mixed-era churn (the term-0 topology-bootstrap leadership orphaning uncommitted
data across a leadership change). Resolved by a 2nd operator-authorized
clean-slate wipe → pristine rc7 baseline
(sha256:d42fe27c4d6a63e5c5103886d4bfa25998d992c4dfd3b536f01469a1edab94d0).
Lesson: commit fully (lag=0) before any leadership-changing restart; the term-0
topology era is the divergence trap. The boot-install-not-clearing-a-divergent-
suffix loop is a remaining upstream item for the devs.
The 3-shard topology is NOT production-ready on real k3s (dev handoff)
With the seed-retry fix the corpus seeds COMPLETE (100000/100000 in 488.5s),
but exercising the seeded 3-shard cluster surfaced that the 3-shard
catch-up/read paths are broken — distinct from (and on top of) the persist /
self-restart fixes, which DO work. These are the m12p4 sharding code (kind-tested
only); they need upstream work before T5/T4/1M can pass on 3 shards:
- Cross-shard
/vector_searchcollapses under load. A recall ramp on a busy LEADER (tidaldb-0, owns 2 groups) returned p99 33,000 ms, 81–100% error, recall 0.3082 — the ≈⅓ recall is the tell that the read consulted ONE shard. Root: the read target (a leader) self-restarted mid-ramp (see 3) and, while degraded, its sibling-group dbs were unavailable so the gather fell to a remote fan-out that timed out. Reads MUST be sent to a non-leader; even then the gather is fragile. - Catch-up "peer shard sN not in configuration" recurs (now shard 0 AND 2): a rejoining node's per-shard gRPC pool never registers the source peer, so its catch-up stream can't open → it can't converge. This blocks clean rejoin.
- A leader hits the term-0 divergence trap under the seed and self-restarts
(
exitCode 0— the Bug-B fix firing), then STALLS rejoining on (2) plus a term-behind refusal (puller term 0 behind source's 1; rejoin). Net: a node that self-restarts to heal cannot actually rejoin the 3-shard cluster. Quorum survives (2/3), the corpus is intact on the healthy two, but the cluster does not self-repair to 3/3.
Net: the rc5–rc7 fixes (persist, parallel-save, self-restart TLS-trust, seed-retry) are real and verified. But the m12p4 3-shard catch-up + cross-shard read layer needs upstream repair before the sharded-topology gates (T5, and T4/1M on shards) can run. The single-shard topology had a working read path (prior session: recall 0.967, p99 7 ms @ 100k) — the read gates G1/G2 (and T4/1M reads) are provable there today; T5 inherently needs the (currently broken) sharded path.