# m11p1 — Replication Performance Floor (✅ COMPLETE 2026-06-11) Phase spec and exit gate: [docs/roadmap-to-cluster.md §4/m11p1](../../roadmap-to-cluster.md). Status summary: [ROADMAP Milestone 11](../ROADMAP.md). Changelog: [CHANGELOG.md](../../../CHANGELOG.md). ## What shipped 1. **Staged two-phase replicated writes** (`SignalRelay::stage_write` / `complete_write`): seqno bump + WAL submission + relay-log push stay atomic under the seqno lock (microseconds, rollback on staging failure); the group-commit fsync wait happens outside it, so concurrent writers share fsyncs. The old path held the lock across the fsync, serializing every writer onto a solo `batch_timeout` + fsync — the measured ~90 writes/s ceiling and ~178ms latency (2–8 pool workers × ~11ms, 16-deep queue). 2. **Per-peer windowed batch shipping** (`ShipQueue`): dedicated sender threads per peer coalesce contiguous relay-log runs into multi-event batches (≤ `replication.batch_max_events`), `replication.window` in flight per peer, parked-run retry every `replication.retry_ms`. Senders ship only the leader's contiguous **durable frontier**; an fsync failure after staging **poisons** the relay (writes rejected, frontier frozen) instead of risking leader/follower divergence. 3. **Follower group-commit coalescing**: the segment receiver drains its inbound backlog (`Transport::try_recv_segment`) and applies it through one shared group commit (`SignalLedger::apply_replicated_events`) in range-disjoint groups (duplicate/subset re-ships cannot double-fold). Without this, each segment paid a solo follower fsync: ~1.8k events/s apply ceiling measured, unbounded lag under the new leader rates. 4. **First cluster observability**: per-region topology `metrics_addr` wires the engine's `/metrics` listener (cluster mode previously had none); `tidaldb_cluster_*` series for ship RTT/batch size (per-peer labeled), WAL fsync latency + group-commit fill, write-pool depth/rejections, and relay committed/durable frontiers. WAL group-commit knobs became config (`wal.batch_size`, `wal.batch_timeout_ms`). ## Exit gate evidence (local Ref-A-equivalent: 3 real processes, release build) | Gate | Target | Measured | |------|--------|----------| | Replicated signal throughput | ≥ 2,000/s | 4,534/s within SLO on the ramp (knee ~5.5k/s); **2,739/s sustained 10 min** (1.65M writes) | | Error rate | < 1% | 0.35% (10-min sustain), zero client shed | | Replication lag | p99 ≤ 2s | ≤ 103 events (~40ms) across the 10-min sustain | | Replication metrics in Prometheus | visible | `tidaldb_cluster_*` live on the per-region listener | | Signal write p99 | ≤ 50ms | p50 21ms / p90 35ms / p99 ~199ms locally — the macOS `F_FULLFSYNC` floor (7.4ms mean, 10–50ms tail per `tidaldb_cluster_wal_fsync_us`; 0% under 1ms). Validate on Ref-A (Linux `fdatasync`), where the fsync floor is far lower. Ref-A re-run pending infra access. | Tier-3 suites (multiproc 5, chaos 3, lifecycle 2, runbook 9, e2e 2) green over real OS processes with real TCP-relay partition injection. ## Incidental fixes - Ship-sender failure logging is transition-based (1st + every 50th consecutive failure at WARN, recovery at INFO): the per-retry WARN flood filled the test harness's undrained stderr pipe and froze every logging thread in the leader (spurious 408s). The harness now uses `Stdio::null()` for node logs (`TIDAL_TEST_NODE_LOGS=inherit` to stream them while debugging). ## Carried hazards (tracked for later phases) - The relay log remains an unbounded in-memory `Vec` (now `RelayEvent`, ~48B per entry vs ~150B before) — m11p2 replaces it with durable-WAL read-back (`StreamSegments`). - Relay seqno resets on leader restart while followers keep their applied high-water marks — pre-existing; m11p4's durable terms/fencing owns it. - Follower crash between WAL append and frontier advance can double-fold on redelivery after the dedup window rotates — pre-existing crash window; m11p3's no-acked-loss ledger checker will surface and bound it.