tidaldb/tidal-server/tests
jordan fab5467b8f test(cluster): reproduce the multi-group reseed silent hole
The served-evidence marker fix (afdda7c) closes the SINGLE-group case, proven by
mp_follower_reseeds_via_snapshot_after_compaction passing with its content probe.
It does not close the multi-group case, and nothing in the suite covered that: the
one reseed gate was single-group, and the harness leaves reseed_self_restart at
false, so a per-group self-restart that never reaches a fixpoint was invisible.

New mp_multi_group_node_converges_after_reseeding_several_groups reproduces the
production shape from k8s/cluster/topology-configmap.yaml: 3 nodes x 3 groups,
full placement, production election timers, reseed_self_restart TRUE. It stops one
node so its group leadership moves and a survivor ends up leading two groups (the
live tidaldb-1 arrangement), writes past WAL_RETENTION_SEGMENTS, gracefully
restarts the survivors to compact, then brings the node back.

The test also stands in for the ORCHESTRATOR. reseed_self_restart drains and
exits(0) expecting a reboot; the harness has no supervisor and `is_alive` only
checks that the handle is retained, so an exited node just stays down. Sustained
HTTP unreachability is the exit signal and `restart` is the reboot, counted
against a finite ceiling. The content probe stays supervised too, because the
first run settled, then re-latched and exited, and an unsupervised probe merely
panicked on a connection error and hid it.

Observed failure, the local twin of the production incident:

  [multi] node 2 settled after 0 orchestrator restart(s)
  [multi] node 2 exited AFTER settling; orchestrator reboot #1
  missing item 500 (reboots=1) ... reseed_required: false, lag_events: 0,
                                   applied_events: 3798, election_tail_term: 2

The node reports no marker and zero lag while an item written before its outage is
absent. That is the same silent hole tidaldb-0 showed at lag_events: 0.

Marked #[ignore] with the reason and the invocation, so the nightly chaos gate
keeps its signal instead of going permanently red on a known-open defect. Removing
the attribute is the gate for the fix.

Also adds write_heavy_item_retrying: which survivor inherits a stopped node's
groups varies per run, so a write may be local for one group and a cross-group
forward for another, and a forward inside an election window legitimately answers
a retryable 503. Retrying keeps the fixture deterministic without masking a hard
failure.
2026-08-21 02:15:35 -06:00
..
support feat(m12p6): persist HNSW graph + bounded SIGTERM drain — boot loads, no rebuild 2026-06-15 13:09:20 -06:00
cluster_chaos.rs feat(m11): continuous correctness (m11p9) — fault classes, invariant checkers, soak gates, nightly pipeline 2026-06-13 15:23:59 -06:00
cluster_cross_shard_reads.rs feat(m12p4): sharded ingestion — scatter-gather pool + cross-shard unified reads (L4) 2026-06-14 15:17:35 -06:00
cluster_e2e.rs feat(m8p10): multi-process cluster mode — scatter-gather, reconcile relay, chaos/UAT suites 2026-06-10 14:07:33 -06:00
cluster_election.rs feat(m11): Raft leader election over WAL stream (m11p4) 2026-06-11 23:30:24 -06:00
cluster_faults.rs feat(m11): continuous correctness (m11p9) — fault classes, invariant checkers, soak gates, nightly pipeline 2026-06-13 15:23:59 -06:00
cluster_graph_persistence.rs fleet remediation: make the workspace gate runnable, then fix what it caught 2026-08-16 12:38:14 -06:00
cluster_grpc.rs feat(m11): cluster security (m11p7) + perf instrumentation floor 2026-06-13 01:25:35 -06:00
cluster_lifecycle.rs feat(m11): membership, snapshot install, and reseed (m11p5) 2026-06-12 19:55:54 -06:00
cluster_membership.rs feat(m12): election-divergence-fix + soak-eval streak + release tooling 2026-06-18 13:08:53 -06:00
cluster_multiproc.rs feat(m8p10): multi-process cluster mode — scatter-gather, reconcile relay, chaos/UAT suites 2026-06-10 14:07:33 -06:00
cluster_quorum.rs feat(m11): continuous correctness (m11p9) — fault classes, invariant checkers, soak gates, nightly pipeline 2026-06-13 15:23:59 -06:00
cluster_region.rs feat(m12p4): sharded ingestion — scatter-gather pool + cross-shard unified reads (L4) 2026-06-14 15:17:35 -06:00
cluster_reseed.rs test(cluster): reproduce the multi-group reseed silent hole 2026-08-21 02:15:35 -06:00
cluster_routes.rs feat(m11): cluster security (m11p7) + perf instrumentation floor 2026-06-13 01:25:35 -06:00
cluster_runbook.rs feat(m8p10): multi-process cluster mode — scatter-gather, reconcile relay, chaos/UAT suites 2026-06-10 14:07:33 -06:00
cluster_security.rs feat(m11): cluster security (m11p7) + perf instrumentation floor 2026-06-13 01:25:35 -06:00
cluster_sharding.rs feat(m11): sharding × replication + rebalancing (m11p6 L3-L5) 2026-06-13 18:23:43 -06:00
middleware.rs feat(m11): cluster security (m11p7) + perf instrumentation floor 2026-06-13 01:25:35 -06:00
reseed_install.rs fix(m12): break the post-reseed false-ReseedRequired loop (durable term marker + readiness gating + restart coordinator) 2026-06-18 21:06:18 -06:00
standalone_offload.rs feat(m11): cluster security (m11p7) + perf instrumentation floor 2026-06-13 01:25:35 -06:00
standalone.rs feat(m11): cluster security (m11p7) + perf instrumentation floor 2026-06-13 01:25:35 -06:00
vector_search.rs feat(m12): vector retrieval G1/G2 — recall harness, ANN in RETRIEVE, index tuning 2026-06-14 11:07:09 -06:00