WAL segment format: 8-byte TSEG header (magic + version byte + 3 reserved)
prepended to every new segment. Legacy headerless segments (m0-m11p3) read
as implicit v0 — no migration. Unknown magic/version surfaces as
WalError::SegmentFormatUnknown at open time; foreign files are never
repaired or truncated (fixes the silent data-loss path from the p3 rollout
incident where torn-tail repair zeroed a follower's unreadable segments).
Catch-up transport: FAILED_PRECONDITION ("snapshot required") and stream
errors that skip the shard now arm a timer retry (re-arm-on-skip is the
load-bearing liveness fix — without it a skipped pull never re-fires and
the follower stays permanently behind). Single retry pending per shard;
CatchupRunner owns the Arc'd state shared between the retry tasks and the
transport. Test: tidal-net/tests/catchup_retry.rs covers the retry path.
Stress: k8s stress-job-t2a/t2b yaml + ops/stress-test-p3-t2 runbook.
25 KiB
Changelog
All notable changes to tidalDB will be documented in this file.
[Unreleased]
Added
Catch-up self-healing + WAL segment format versioning (m11p4) — timer-retried pulls, TSEG segment header, structured "snapshot required"
- Failed catch-up pulls retry on a timer. The pull trigger was event-only:
a follower whose
StreamSegmentspull failed (e.g. the leader's gRPC server not yet ready during a rolling restart) waited for the next PUSHED segment to re-expose the gap — in an idle cluster that push never comes, and the follower stayed lagged forever (the 2026-06-11 p3 rollout: both followers stuck at lag=136507). A failed pull now arms a one-shot timer (replication.catchup_retry_ms, default 30000; transportcatchup_retry_interval) that re-pulls from the CURRENT applied frontier. Pulls stay single-flight and rate-limited; the timer's wake-up re-arms when consumed by the rate limit or an in-flight pull, so the gap always keeps a standing wake-up until a pull completes. Verified over real sockets:catchup_retry.rsreproduces the incident (pull fails, leader appears, zero pushes) and proves timer-only self-heal — and that a clean completion arms nothing. - WAL segment files are format-versioned. New segments open with an
8-byte header (
TSEGmagic + version byte + reserved); pre-m11p4 headerless segments stay readable as implicit version 0 — no migration. A segment this binary cannot identify (unknown header version, unrecognized leading bytes, unparseable.segfilename) surfaces as the newWalError::SegmentFormatUnknownat open — previously it scanned as empty (segments=0) and recovery's torn-tail repair could TRUNCATE the foreign file to zero. Foreign-format files are never repaired, truncated, or skipped. Downgrade across m11p4 requires a WAL reseed (runbook §8). - Unservable catch-up is a structured refusal.
SegmentSource::collect_fromreturns typedSegmentReadError::{Unavailable,Failed};TidalDb::read_wal_batchesreturns the typedWalError(was stringifiedTidalError). TheStreamSegmentshandler mapsUnavailabletoFAILED_PRECONDITION—"segments not available from seq N; snapshot required"— and the follower logs it distinctly (catch-up unservable … needs a snapshot (m11p5) or an operator reseed) instead of burying it as a transient. The on-disk segment format is now documented (tidal/src/wal/segment.rsmodule docs + spec 01 §2.2).
Quorum-acked writes (m11p3) — ack=leader|quorum, durable ship acks, commit index, zero-acked-loss ledger gate (closes G4)
ack=quorumis an opt-in durability contract for every replicated write (/signals,/items,/embeddings): success means a majority of the replica set durably holds the write (leader +floor(n/2)followers, each storage-applied and own-WAL-fsynced), so an acked write survives the permanent loss of any single node — including the leader. Deployment default via topologyreplication.ack; per-request override via thex-tidal-ackheader (forwarded verbatim by gateways).ack=leader(the default) is the m0–m11p2 contract unchanged.- Durable frontier reports. Followers PUSH their durably-applied frontier
to the leader once per apply round (new
ReportAppliedRPC, fired by the segment receiver through the newTransport::notify_applied) — batch-level and fully decoupled from ship acks, so the commit index stays fresh even when outbound ships stall (gap-parked follower, pull-based catch-up, quiet leader). Ship acks keep their m11p2 instant floor-hint semantics — both inputs are durable-true because a follower's frontier only ever advances after its storage apply + own-WAL fsync. (The first design held each ship ack until its segment's apply; measured under open-loop load, that couples ship cadence to apply latency and one gap-parked follower spirals into total quorum collapse — the report push is what shipped.) - Follower blob applies are group-committed. m11p2 applied replicated
items/embeddings one record at a time — one solo follower fsync per item,
capping item apply at the fsync floor (~100/s on macOS) and stalling the
quorum frontier behind any item burst. The receiver now hands each apply
round's blob records to the engine as ONE batch (
apply_replicated_blobs: validate all → stage all WAL appends → wait all → upsert storage), and the WAL writer flushes queued blobs under ONE group fsync. Measured: corpus seeding 2,000 items + embeddings 39.3s → 1.8s (22×). - Commit index. The leader folds durable acks (and heal resumes) into
per-peer durable marks; the commit index is the k-th largest (k =
floor(n/2)), leadership-scoped (promote resets it to the stream baseline; demotion fails in-flight waiters — a demoted leader never claims quorum). Handlers await it through a watch-channel bridge — fully async, zero threads parked per waiter (the thread-per-wait design measurably collapsed at 1k rps open-loop by exhausting the blocking pool and starving the very completions that advance the index). Quorum capacity measured on a real 3-process localhost cluster (release, writes mix): 3,600 quorum signal-writes/s within SLO (p50 ~45ms, zero errors, replication lag ≤3 events at ramp end; knee not reached) — 79% of m11p1's 4,534/s leader-ack figure, vs the ≥50% gate. - Honest timeout semantics. Quorum not confirmed within
replication.quorum_timeout_ms(default 2000) → a retryable 503 naming the laggard regions, the commit index, and needed/confirmed counts. The write is in the leader's log and may still commit: retries are at-least-once (items/embeddings retries are idempotent upserts; signal retries can double-count — decided in-phase: no idempotency-key machinery, documented in runbook §8 with the session-write precedent for callers that need exact-once). x-tidal-seqon every cluster write response: the write's seqno in the replicated log (relayed through forwards) — an exact durability cursor againstcommit_indexin/cluster/status/local(which also gainsack).tidaldb_cluster_relay_durable_seqnow reports the commit index (relay_last_seq − relay_durable_seq= quorum lag); new countertidaldb_cluster_quorum_timeouts_total.- The ledger gate (exit gate, run for real): tier-3
mp_quorum_ledger_zero_acked_loss_across_killpointsSIGKILLs the leader under concurrent quorum load and proves zero acknowledged loss on the promoted max-applied survivor — frontier invariant (max acked seq ≤ survivor applied) plus per-item content probes. 167/167 kill points passed on the final design (batches of 64 + 91 + 12; kill timings spread 120–598ms across fresh 3-process clusters; an earlier 100/100 run had validated the superseded ack-holding design before it was replaced — see docs/planning/milestone-11/phase-3.md). Plus tier-3 partition semantics (one follower down: quorum commits; both down: fast 503 naming laggards whileack=leaderflows; heal: recovers) and in-process gRPC coverage (override headers, forwarded quorum writes, blob writes, 400 on bad mode). - Rolling-upgrade order (mixed-version caveat): a pre-m11p3 leader
neither serves the
ReportAppliedRPC nor recognizesx-tidal-ack— it silently applies LEADER-ack semantics to aquorumrequest (a durability downgrade the caller cannot see). Upgrade the leader first:ack=quorumis then honored immediately (commit-index freshness rides the m11p2 ship-ack floor hints until the followers upgrade too). Replication and heal are unaffected by either order. Runbook §8 records the procedure.
One replicated log (m11p2) — items/embeddings ride the WAL, StreamSegments catch-up, HTTP broadcast deleted
- The leader's WAL is now THE replicated log. The group-commit writer hands
every fsynced batch to a bounded in-memory ship feed (
wal::feed::WalShipFeed); the ship queue pushes those already-encoded bytes verbatim — byte-identical on the leader's disk, the wire, and the follower's apply path — and stream seqnos are WAL seqnos, so they survive restarts (the m8p10 relay-reset hazard is gone). The m11p1 in-memory relay log, its durable frontier, and its poisoning machinery left the server write path entirely (/signalsstages straight through the engine's group commit; a WAL fsync failure now surfaces per-write exactly like single-node). - Item metadata and embeddings are replicated mutations.
/itemsand/embeddingsjournal kind-1/2 blob records (WAL headerflagsbyte = batch kind; one record, one seqno) BEFORE storage, on the same stream as signals. Followers apply them kind-aware — WAL-first into their own log, then idempotent storage upserts — and recovery replays them. The m8p10 HTTP item/embedding broadcast (marker-gated fan-out, O(items) heal re-broadcast, authed side-POSTs — the source of both 2026-06-10 live bugs) is deleted; those bug classes are now impossible by construction. Cluster/itemsreturns a plain 201 and/embeddingsa plain 204 (no broadcast-report body). StreamSegmentsimplemented — catch-up is follower-pulled. The ship feed's tail is bounded; a peer that falls behind it is skipped ahead, and the follower pulls the hole itself via the (previously declared-unimplemented) server-streaming RPC over the leader's durable, BLAKE3-verified segments. Pulls trigger on detected gaps, on follower boot (self-driving restart catch-up), and on the leader's heal nudge (POST /cluster/catchup, internal, forwards the operator's own bearer credential). Pulled chunks flow through the same inbound apply path as live ships. Ship acks piggyback the follower's applied seqno (ShipSegmentResponse.applied_seqno), so retries of already-applied data prune and heal isresume_from+ nudge — no redelivery scan, no O(items) traffic.- Promote carries a stream baseline. A promoted leader's stream starts at its
promote-time flushed frontier (persisted in
data_dir/stream_baseline); the fan-out body and every catch-up chunk announce it, so peers jump their frontier past pre-stream history instead of parking on a phantom gap. Ship queues are leadership-gated (activate_from/deactivate): a follower's replicated applies never echo back at its peers. - Multi-process cluster mode now requires
--data-dir(validated at startup): the durable WAL is the replication stream. Standalone single-node deployments are untouched — blob journaling and the ship feed are gated on cluster peers, so single-node item writes keep fjall-only durability with zero extra fsyncs. - New tier-3 suite
mp_items_ride_the_log_and_catchup_stream: items written on the leader AND through a follower gateway converge everywhere via the log (feed parity 1e-6), and a follower stopped through item+embedding+signal writes restarts and converges via its boot-timeStreamSegmentspull with no heal verb and no HTTP item traffic.
Cluster replication performance floor (m11p1) — ack/ship decoupled, batched + windowed shipping, first tidaldb_cluster_* metrics
- The replicated
/signalswrite path no longer serializes every writer onto a solo group-commit fsync nor ships to followers on the request path. Writes are staged (seqno + WAL submission + relay log push, microseconds, atomic with rollback) and completed (shared group-commit fsync + in-memory fold) in two phases, so concurrent writers coalesce into one fsync; follower shipping moved to per-peer sender threads that coalesce contiguous runs into multi-event batches with a windowed in-flight budget (topology knobsreplication.{batch_max_events,window,retry_ms}). The 204 contract is unchanged (leader durability only) — it now returns at leader fsync. Measured on a real 3-process localhost cluster (release build, thepeach mix): 4,534 replicated signal-writes/s sustained within SLO vs ~90/s before (~50×), replication lag bounded at ≤377 events (~80ms) through the whole ramp. - Durable-frontier shipping + relay poisoning. Senders only ship the leader's contiguous fsynced prefix (an event a follower holds but the leader could lose is silent divergence); a staged write whose fsync fails poisons the relay — further cluster writes are rejected and the ship frontier freezes (the CockroachDB/Postgres fsync-failure posture).
- Follower group-commit coalescing. The segment receiver drains its inbound
backlog (
Transport::try_recv_segment) and applies it through ONE shared group commit (SignalLedger::apply_replicated_events), in range-disjoint groups so duplicate/subset re-ships cannot double-fold. Without this the follower apply ceiling was ~events-per-segment / fsync-cost(~1.8k events/s measured) and lag grew without bound under m11p1 leader rates. - First cluster metrics + cluster
/metricslistener. Cluster mode previously had no metrics endpoint at all. New per-region topologymetrics_addrwires the engine's Prometheus listener; newtidaldb_cluster_*series: ship RTT + batch-size histograms, per-peer ship counters/gauges (peer_shardlabels), WAL fsync latency + group-commit fill histograms, write-pool depth/rejections, relay committed + durable frontiers. WAL group-commit knobs are deployment config (wal.{batch_size,batch_timeout_ms}topology block;wal_batch_size/wal_batch_timeoutbuilder methods). - Engine API:
TidalDb::signal_staged/StagedSignal::wait,SignalRelay::{stage_write,complete_write,durable_seq,snapshot_range},ShipQueue(per-peer windowed batch senders with pause/resume),WalWriter::append_signal_staged,WalSender::append_record_staged,range_payload/encode_run. Relay log entries are nowRelayEvent(raw event records, re-encoded deterministically at ship time) instead of pre-encoded single-event bytes. - Ship-sender failure logging is transition-based (first + every 50th
consecutive failure WARN with the running count, recovery INFO) — the
per-retry WARN flood could fill an undrained log pipe and stall the process.
The multiproc test harness now discards child logs via
/dev/null(a piped fd nobody drains deadlocks the node once the kernel buffer fills), withTIDAL_TEST_NODE_LOGS=inheritto stream them while debugging.
M9 — Community Sync & Revocation
- Local embeddable profiles can opt into community personalization and safely leave/purge their contributions. New types:
SignalScope,CommunityId,Membership,MembershipEpoch,PolicyMetadata. Community signal reconciliation viaCrdtSignalStatewith commutative/associative/idempotent merge laws; membership-epoch revocation purges a departed member's contributed signals.
M10 — Governance & Agent Rights
- Community rules and agent-scoped permissions control what signals influence ranking: policy-metadata enforcement and agent-rights scoping wired into the signal-write and ranking paths.
Cluster mode (m8p10): true multi-process region nodes + full tier-3 UAT — M8 COMPLETE
- Multi-process cluster mode:
tidal-server cluster --region <name> [--data-dir <p>](envTIDAL_REGION) runs one process per region (RegionClusterState), each owning oneTidalDband oneGrpcTransportthat binds this region'sgrpc_addrand dials every sibling's realgrpc_addr— real process/host isolation. The topology requires per-regiongrpc_addrANDhttp_addrin this mode; single-process mode (no--region) is unchanged as the dev/demo default. Both modes stay behind the experimental gate (--experimental-cluster/TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1) with mode-specific WARN text. - New routes / behaviors on the multi-process surface:
GET /cluster/status/local(per-node status),GET /cluster/statusaggregates ALL regions with areachablefield (unreachable ⇒reachable:false+ worst-case lag),POST /cluster/reconcile {region}(cross-process CRDT snapshot exchange; idempotent — a repeat is an exact no-op on scores),POST /hardnegs {user_id,item_id}(user-scoped hide; converges via reconcile, filtered from/feed?user_id). Writes on a non-leader forward to the leader;?region=reads forward to the owning region; default reads are LOCAL in multi-process mode. Items/embeddings are leader-applied + HTTP-broadcast to peers with a per-peer report:POST /items→ 201{replicated_to, failed},POST /embeddings→ 200 with the same report on the leader path (NOT 204 — a 204 cannot carry the report; the forwarded/internal path stays 204)./sharded/*fans out across processes (degraded semantics preserved).promotefans out to all peers ({ok, leader, acked, failed}). /cluster/healis the single recovery verb: redelivers missed signal segments (gap-aware) AND re-broadcasts item metadata + embeddings to the healed region (idempotent upserts). After a partition the per-peer gRPC circuit breaker (threshold 5, reset 30s) is open, so re-issue/cluster/healuntil/cluster/statusshows lag 0.TIDAL_HLC_SKEW_MS(multi-process): signed ms offset applied to the process's HLC (reconcile LWW stamping only — not signal-decay timestamps); a test/ops escape hatch.- Tier-3 UAT over real OS processes (feature
cluster-e2e):cluster_multiproc(5),cluster_chaos(3, REAL network-partition injection via a root-free in-harness TCP relay proxy — the toxiproxy-style alternative the ROADMAP sanctions; iptables/pfctl remain an operator option),cluster_lifecycle(2, ±500ms clock skew + rolling upgrade with zero acknowledged-write loss),cluster_runbook(9, everydocs/runbooks/cluster.md§5–§11 operation), pluscluster_e2e(2, single-process smoke). Measured (localhost): replication p99 ~110–133ms (< 2s SLA), failover ~31–34ms (< 10s SLA), reconcile 0–1ms/side (< 100ms SLA).
Fixed
Four production bugs surfaced and fixed by the m8p10 tier-3 UAT
- Silent data loss on out-of-order ships. The replication applied high-water-mark swallowed sequence gaps when eager ships arrived out of order. The applied frontier is now contiguous with a bounded ahead-buffer, so a gap can never be skipped.
- Non-idempotent reconcile (score creep).
take_crdt_snapshotattributed replicated signal streams per-node, so reconciling already-converged nodes crept the decayed scores (0.5 → 0.375 → …). Contributions are now attributed to one canonical replication shard (ShardId::SINGLE), makingmergeidempotent — reconcile of converged nodes is an exact fixpoint. - Lag gauge conflated leader streams across a promotion. A converged node reported a
permanent phantom lag after a
/cluster/promotemoved leadership to a different shard. The lag gauge now tracks the leader high-water-mark per source shard (ReplicationLagGauge::leader_seqno_for), computing lag against the current leader. - Items missed during a broadcast were never backfilled. A node down/partitioned
during an item/embedding broadcast was permanently missing that data even at signal
lag 0.
/cluster/healnow re-broadcasts item metadata + embeddings to the healed region (idempotent upserts), making heal the single recovery verb.
Changed
Cluster mode (m8p8): real gRPC replication
tidal-server'sClusterStatenow wires each follower region to a realtidal-netGrpcTransport(self-loop over loopback gRPC) instead of in-process crossbeam channels — replication between regions traverses real gRPC/TCP (serialization, circuit breaker, HTTP/2). Closes M8 gap G1.- Follower gRPC ports are auto-allocated from the topology (
grpc_addroptional per region) and self-heal a transient bind race by retrying on a fresh port. - Blocking gRPC ships triggered by
POST /signalsandPOST /cluster/healare offloaded to a dedicated thread so they never block the axum reactor. ClusterStateis constructed off the async reactor (GrpcTransport::newblocks on its own runtime).- The experimental opt-in gate and
docker/cluster/Dockerfileare updated: single-process cluster mode is honest that it replicates over real gRPC but runs all regions in one process (no host/process isolation). (True multi-process region nodes shipped subsequently in m8p10 — see the m8p10 entry above.)
Docker build fixes (all three images now that tidal-server pulls tidal-net)
- Install
protobuf-compilerin the builder stage —tidal-net's build script runstonic-build, which needsprotocto compile the WAL-shipping.proto. - Pin the builder base to
bookworm(rust:1.91-bookworm/rust:1.91-slim-bookworm) so its glibc matches thedebian:bookworm-slimruntime; the default trixie base emitted alibmvec.so.1dependency absent on bookworm, aborting the binary at startup. Verified:docker runof the cluster image serves a functional 3-region cluster (write replicates to both followers over gRPC; region-pinned reads serve replicated data; SIGTERM exits 0). - New tests:
tidal-server/tests/cluster_grpc.rs(in-process gRPC replication + HTTP offload path) and a hardened tier-3cluster_e2e.rs(multi-process smoke- promote over real OS processes, feature-gated).
[0.1.0] - 2026-02-23
Added
Core Database Engine
TidalDbembeddable database withephemeral()andwith_data_dir()open modesSchemaBuilderfor defining signal types, decay parameters, and ranking profilesTidalDbBuilderfluent builder with schema, data directory, metrics, and rate limiter configuration
Signal System
- Typed signal recording with exponential decay scoring
- Hot-tier (DashMap) and warm-tier (BucketedCounter) signal storage
- Windowed aggregation:
OneHour,TwentyFourHours,SevenDays,AllTime - Signal velocity tracking
- WAL-backed signal durability with crash recovery
- Periodic signal checkpointing to fjall (every 30s)
- WAL compaction after each checkpoint
Retrieval (RETRIEVE query)
- 5-stage pipeline: universe, filter, score, diversify, return
- Filter expressions:
Eq,In,Gt,Lt,And,Or,Not,InCollection,InProgress,MinSignal,MaxSignal,NearLocation - Built-in ranking profiles:
trending,for_you,new,popular,recent, and 20+ more - Custom ranking profiles via
SchemaBuilder - Diversity enforcement (max N per category/creator)
- Sort modes:
Relevance,Trending,Newest,MostLiked,MostViewed,MostFollowed,AlphabeticalAsc/Desc,Shortest/Longest,LiveViewerCount,DateSaved, and more
Search (SEARCH query)
- BM25 full-text search via Tantivy
- Approximate nearest-neighbor (ANN) semantic search via USearch HNSW
- Reciprocal Rank Fusion (RRF) combining BM25 + ANN scores
- Creator search with
entity_kind(EntityKind::Creator) similar_to(EntityId)for content-based recommendations- Scope pre-filters:
Trending,CohortTrending,Following,Category,Collection - Autocomplete suggestions via
db.suggest()
Entity Model
- Three built-in entity types:
Item,User,Creator - Metadata storage as
HashMap<String, String> - Embedding slots (up to 4 per entity type) via USearch
- Relationships:
Follows,Blocks,Hide,Mute,InteractionWeight
Sessions
- Session lifecycle:
open_session,close_session - Cross-session preference vector updates (EMA blend)
- Session snapshots with signal state and preference vectors
- Session serialization format v0x03 with backward compatibility
Social Graph
- Creator follower/following indexes
- Cohort membership (user segments)
- CoEngagementIndex for co-viewing patterns with LRU eviction
- Social graph filter for "followed creator" content scoping
Collections
- Named collections with
Private,Shared,Publicvisibility create_collection,add_to_collection,remove_from_collection,list_collectionsFilterExpr::InCollectionfor collection-scoped retrieval- Saved searches with
save_search,list_saved_searches,retrieve_saved_search
Observability
enable_metrics(addr)-- Prometheus-format/metricsendpoint +/healthzJSON- 15+ metrics: signal writes, WAL lag, checkpoint age, degradation level, index health
tidaldb_checkpoint_failures_totalcounter for checkpoint monitoringTidalDb::diagnostics()-- structured health snapshot- WAL diagnostics and recovery tools
Safety
- Signal weight NaN/Inf validation (returns
TidalError::InvalidInput) - Metadata size bounds: 64 keys max, 8KB value max, 64KB total max
- Export request limit: 500K signals max per request
FilterExprcomplexity limit: 256 nodes max- Data directory lock (
tidaldb.lock) prevents dual-process corruption - Schema fingerprint persistence detects decay parameter changes on reopen
- Bounded
closed_sessionscache (10K max, LRU eviction) - Metrics server non-loopback bind warning
CLI (tidalctl)
tidalctlbinary for database inspection and diagnostics
RLHF / ML Export
db.export_signals(ExportRequest)-- WAL-based signal export for training datadb.user_session_summary(user_id, since_ns)-- aggregated session statistics
Stability
tidalDB 0.1.0 is pre-1.0. No API or data format stability guarantees are made for 0.x releases. Upgrade guides will be provided for each minor version bump. Do not upgrade 0.x to 0.y on a live data directory without reading the release notes.
Format based on Keep a Changelog