m11p1 — decoupled ack/ship path: staged writes (seqno+WAL+relay-push, microseconds) separate from group-commit fsync; ShipQueue batches+windows outbound segments; receiver coalesces inbound chunks before applying. Adds first tidaldb_cluster_* metrics. m11p2 — leader WAL is now THE replicated log: fsynced batches feed a bounded WalShipFeed and ship byte-identical to followers; WAL seqnos survive restarts (relay-reset hazard gone). Item metadata and embeddings journal kind-1/2 blob records on the same stream as signals; the m8p10 HTTP broadcast is deleted. StreamSegments catch-up is follower-pulled via server-streaming RPC, triggered on gap detection, follower boot, and leader heal nudge. Promote carries a stream baseline so peers skip pre-stream history.
4.0 KiB
4.0 KiB
m11p1 — Replication Performance Floor (✅ COMPLETE 2026-06-11)
Phase spec and exit gate: docs/roadmap-to-cluster.md §4/m11p1. Status summary: ROADMAP Milestone 11. Changelog: CHANGELOG.md.
What shipped
- Staged two-phase replicated writes (
SignalRelay::stage_write/complete_write): seqno bump + WAL submission + relay-log push stay atomic under the seqno lock (microseconds, rollback on staging failure); the group-commit fsync wait happens outside it, so concurrent writers share fsyncs. The old path held the lock across the fsync, serializing every writer onto a solobatch_timeout+ fsync — the measured ~90 writes/s ceiling and ~178ms latency (2–8 pool workers × ~11ms, 16-deep queue). - Per-peer windowed batch shipping (
ShipQueue): dedicated sender threads per peer coalesce contiguous relay-log runs into multi-event batches (≤replication.batch_max_events),replication.windowin flight per peer, parked-run retry everyreplication.retry_ms. Senders ship only the leader's contiguous durable frontier; an fsync failure after staging poisons the relay (writes rejected, frontier frozen) instead of risking leader/follower divergence. - Follower group-commit coalescing: the segment receiver drains its
inbound backlog (
Transport::try_recv_segment) and applies it through one shared group commit (SignalLedger::apply_replicated_events) in range-disjoint groups (duplicate/subset re-ships cannot double-fold). Without this, each segment paid a solo follower fsync: ~1.8k events/s apply ceiling measured, unbounded lag under the new leader rates. - First cluster observability: per-region topology
metrics_addrwires the engine's/metricslistener (cluster mode previously had none);tidaldb_cluster_*series for ship RTT/batch size (per-peer labeled), WAL fsync latency + group-commit fill, write-pool depth/rejections, and relay committed/durable frontiers. WAL group-commit knobs became config (wal.batch_size,wal.batch_timeout_ms).
Exit gate evidence (local Ref-A-equivalent: 3 real processes, release build)
| Gate | Target | Measured |
|---|---|---|
| Replicated signal throughput | ≥ 2,000/s | 4,534/s within SLO on the ramp (knee ~5.5k/s); 2,739/s sustained 10 min (1.65M writes) |
| Error rate | < 1% | 0.35% (10-min sustain), zero client shed |
| Replication lag | p99 ≤ 2s | ≤ 103 events (~40ms) across the 10-min sustain |
| Replication metrics in Prometheus | visible | tidaldb_cluster_* live on the per-region listener |
| Signal write p99 | ≤ 50ms | p50 21ms / p90 35ms / p99 ~199ms locally — the macOS F_FULLFSYNC floor (7.4ms mean, 10–50ms tail per tidaldb_cluster_wal_fsync_us; 0% under 1ms). Validate on Ref-A (Linux fdatasync), where the fsync floor is far lower. Ref-A re-run pending infra access. |
Tier-3 suites (multiproc 5, chaos 3, lifecycle 2, runbook 9, e2e 2) green over real OS processes with real TCP-relay partition injection.
Incidental fixes
- Ship-sender failure logging is transition-based (1st + every 50th consecutive
failure at WARN, recovery at INFO): the per-retry WARN flood filled the test
harness's undrained stderr pipe and froze every logging thread in the leader
(spurious 408s). The harness now uses
Stdio::null()for node logs (TIDAL_TEST_NODE_LOGS=inheritto stream them while debugging).
Carried hazards (tracked for later phases)
- The relay log remains an unbounded in-memory
Vec(nowRelayEvent, ~48B per entry vs ~150B before) — m11p2 replaces it with durable-WAL read-back (StreamSegments). - Relay seqno resets on leader restart while followers keep their applied high-water marks — pre-existing; m11p4's durable terms/fencing owns it.
- Follower crash between WAL append and frontier advance can double-fold on redelivery after the dedup window rotates — pre-existing crash window; m11p3's no-acked-loss ledger checker will surface and bound it.