tidaldb/CHANGELOG.md
jx12n 8a0950260f feat(m8p10): multi-process cluster mode — scatter-gather, reconcile relay, chaos/UAT suites
Splits monolithic cluster.rs into tidal-server/src/cluster/ modules. Adds redeliver-missed
relay, bounded HLC drift, lag tracking, and reconcile idempotence. Five new tier-3 test suites
(chaos, lifecycle, multiproc, region, routes, runbook) all green. Docs, CHANGELOG, and ROADMAP
updated with G4/G5/G6 known gaps.
2026-06-10 14:07:33 -06:00

11 KiB
Raw Blame History

Changelog

All notable changes to tidalDB will be documented in this file.

[Unreleased]

Added

M9 — Community Sync & Revocation

  • Local embeddable profiles can opt into community personalization and safely leave/purge their contributions. New types: SignalScope, CommunityId, Membership, MembershipEpoch, PolicyMetadata. Community signal reconciliation via CrdtSignalState with commutative/associative/idempotent merge laws; membership-epoch revocation purges a departed member's contributed signals.

M10 — Governance & Agent Rights

  • Community rules and agent-scoped permissions control what signals influence ranking: policy-metadata enforcement and agent-rights scoping wired into the signal-write and ranking paths.

Cluster mode (m8p10): true multi-process region nodes + full tier-3 UAT — M8 COMPLETE

  • Multi-process cluster mode: tidal-server cluster --region <name> [--data-dir <p>] (env TIDAL_REGION) runs one process per region (RegionClusterState), each owning one TidalDb and one GrpcTransport that binds this region's grpc_addr and dials every sibling's real grpc_addr — real process/host isolation. The topology requires per-region grpc_addr AND http_addr in this mode; single-process mode (no --region) is unchanged as the dev/demo default. Both modes stay behind the experimental gate (--experimental-cluster / TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1) with mode-specific WARN text.
  • New routes / behaviors on the multi-process surface: GET /cluster/status/local (per-node status), GET /cluster/status aggregates ALL regions with a reachable field (unreachable ⇒ reachable:false + worst-case lag), POST /cluster/reconcile {region} (cross-process CRDT snapshot exchange; idempotent — a repeat is an exact no-op on scores), POST /hardnegs {user_id,item_id} (user-scoped hide; converges via reconcile, filtered from /feed?user_id). Writes on a non-leader forward to the leader; ?region= reads forward to the owning region; default reads are LOCAL in multi-process mode. Items/embeddings are leader-applied + HTTP-broadcast to peers with a per-peer report: POST /items → 201 {replicated_to, failed}, POST /embeddings200 with the same report on the leader path (NOT 204 — a 204 cannot carry the report; the forwarded/internal path stays 204). /sharded/* fans out across processes (degraded semantics preserved). promote fans out to all peers ({ok, leader, acked, failed}).
  • /cluster/heal is the single recovery verb: redelivers missed signal segments (gap-aware) AND re-broadcasts item metadata + embeddings to the healed region (idempotent upserts). After a partition the per-peer gRPC circuit breaker (threshold 5, reset 30s) is open, so re-issue /cluster/heal until /cluster/status shows lag 0.
  • TIDAL_HLC_SKEW_MS (multi-process): signed ms offset applied to the process's HLC (reconcile LWW stamping only — not signal-decay timestamps); a test/ops escape hatch.
  • Tier-3 UAT over real OS processes (feature cluster-e2e): cluster_multiproc (5), cluster_chaos (3, REAL network-partition injection via a root-free in-harness TCP relay proxy — the toxiproxy-style alternative the ROADMAP sanctions; iptables/pfctl remain an operator option), cluster_lifecycle (2, ±500ms clock skew + rolling upgrade with zero acknowledged-write loss), cluster_runbook (9, every docs/runbooks/cluster.md §5§11 operation), plus cluster_e2e (2, single-process smoke). Measured (localhost): replication p99 ~110133ms (< 2s SLA), failover ~3134ms (< 10s SLA), reconcile 01ms/side (< 100ms SLA).

Fixed

Four production bugs surfaced and fixed by the m8p10 tier-3 UAT

  • Silent data loss on out-of-order ships. The replication applied high-water-mark swallowed sequence gaps when eager ships arrived out of order. The applied frontier is now contiguous with a bounded ahead-buffer, so a gap can never be skipped.
  • Non-idempotent reconcile (score creep). take_crdt_snapshot attributed replicated signal streams per-node, so reconciling already-converged nodes crept the decayed scores (0.5 → 0.375 → …). Contributions are now attributed to one canonical replication shard (ShardId::SINGLE), making merge idempotent — reconcile of converged nodes is an exact fixpoint.
  • Lag gauge conflated leader streams across a promotion. A converged node reported a permanent phantom lag after a /cluster/promote moved leadership to a different shard. The lag gauge now tracks the leader high-water-mark per source shard (ReplicationLagGauge::leader_seqno_for), computing lag against the current leader.
  • Items missed during a broadcast were never backfilled. A node down/partitioned during an item/embedding broadcast was permanently missing that data even at signal lag 0. /cluster/heal now re-broadcasts item metadata + embeddings to the healed region (idempotent upserts), making heal the single recovery verb.

Changed

Cluster mode (m8p8): real gRPC replication

  • tidal-server's ClusterState now wires each follower region to a real tidal-net GrpcTransport (self-loop over loopback gRPC) instead of in-process crossbeam channels — replication between regions traverses real gRPC/TCP (serialization, circuit breaker, HTTP/2). Closes M8 gap G1.
  • Follower gRPC ports are auto-allocated from the topology (grpc_addr optional per region) and self-heal a transient bind race by retrying on a fresh port.
  • Blocking gRPC ships triggered by POST /signals and POST /cluster/heal are offloaded to a dedicated thread so they never block the axum reactor.
  • ClusterState is constructed off the async reactor (GrpcTransport::new blocks on its own runtime).
  • The experimental opt-in gate and docker/cluster/Dockerfile are updated: single-process cluster mode is honest that it replicates over real gRPC but runs all regions in one process (no host/process isolation). (True multi-process region nodes shipped subsequently in m8p10 — see the m8p10 entry above.)

Docker build fixes (all three images now that tidal-server pulls tidal-net)

  • Install protobuf-compiler in the builder stage — tidal-net's build script runs tonic-build, which needs protoc to compile the WAL-shipping .proto.
  • Pin the builder base to bookworm (rust:1.91-bookworm / rust:1.91-slim-bookworm) so its glibc matches the debian:bookworm-slim runtime; the default trixie base emitted a libmvec.so.1 dependency absent on bookworm, aborting the binary at startup. Verified: docker run of the cluster image serves a functional 3-region cluster (write replicates to both followers over gRPC; region-pinned reads serve replicated data; SIGTERM exits 0).
  • New tests: tidal-server/tests/cluster_grpc.rs (in-process gRPC replication + HTTP offload path) and a hardened tier-3 cluster_e2e.rs (multi-process smoke
    • promote over real OS processes, feature-gated).

[0.1.0] - 2026-02-23

Added

Core Database Engine

  • TidalDb embeddable database with ephemeral() and with_data_dir() open modes
  • SchemaBuilder for defining signal types, decay parameters, and ranking profiles
  • TidalDbBuilder fluent builder with schema, data directory, metrics, and rate limiter configuration

Signal System

  • Typed signal recording with exponential decay scoring
  • Hot-tier (DashMap) and warm-tier (BucketedCounter) signal storage
  • Windowed aggregation: OneHour, TwentyFourHours, SevenDays, AllTime
  • Signal velocity tracking
  • WAL-backed signal durability with crash recovery
  • Periodic signal checkpointing to fjall (every 30s)
  • WAL compaction after each checkpoint

Retrieval (RETRIEVE query)

  • 5-stage pipeline: universe, filter, score, diversify, return
  • Filter expressions: Eq, In, Gt, Lt, And, Or, Not, InCollection, InProgress, MinSignal, MaxSignal, NearLocation
  • Built-in ranking profiles: trending, for_you, new, popular, recent, and 20+ more
  • Custom ranking profiles via SchemaBuilder
  • Diversity enforcement (max N per category/creator)
  • Sort modes: Relevance, Trending, Newest, MostLiked, MostViewed, MostFollowed, AlphabeticalAsc/Desc, Shortest/Longest, LiveViewerCount, DateSaved, and more

Search (SEARCH query)

  • BM25 full-text search via Tantivy
  • Approximate nearest-neighbor (ANN) semantic search via USearch HNSW
  • Reciprocal Rank Fusion (RRF) combining BM25 + ANN scores
  • Creator search with entity_kind(EntityKind::Creator)
  • similar_to(EntityId) for content-based recommendations
  • Scope pre-filters: Trending, CohortTrending, Following, Category, Collection
  • Autocomplete suggestions via db.suggest()

Entity Model

  • Three built-in entity types: Item, User, Creator
  • Metadata storage as HashMap<String, String>
  • Embedding slots (up to 4 per entity type) via USearch
  • Relationships: Follows, Blocks, Hide, Mute, InteractionWeight

Sessions

  • Session lifecycle: open_session, close_session
  • Cross-session preference vector updates (EMA blend)
  • Session snapshots with signal state and preference vectors
  • Session serialization format v0x03 with backward compatibility

Social Graph

  • Creator follower/following indexes
  • Cohort membership (user segments)
  • CoEngagementIndex for co-viewing patterns with LRU eviction
  • Social graph filter for "followed creator" content scoping

Collections

  • Named collections with Private, Shared, Public visibility
  • create_collection, add_to_collection, remove_from_collection, list_collections
  • FilterExpr::InCollection for collection-scoped retrieval
  • Saved searches with save_search, list_saved_searches, retrieve_saved_search

Observability

  • enable_metrics(addr) -- Prometheus-format /metrics endpoint + /healthz JSON
  • 15+ metrics: signal writes, WAL lag, checkpoint age, degradation level, index health
  • tidaldb_checkpoint_failures_total counter for checkpoint monitoring
  • TidalDb::diagnostics() -- structured health snapshot
  • WAL diagnostics and recovery tools

Safety

  • Signal weight NaN/Inf validation (returns TidalError::InvalidInput)
  • Metadata size bounds: 64 keys max, 8KB value max, 64KB total max
  • Export request limit: 500K signals max per request
  • FilterExpr complexity limit: 256 nodes max
  • Data directory lock (tidaldb.lock) prevents dual-process corruption
  • Schema fingerprint persistence detects decay parameter changes on reopen
  • Bounded closed_sessions cache (10K max, LRU eviction)
  • Metrics server non-loopback bind warning

CLI (tidalctl)

  • tidalctl binary for database inspection and diagnostics

RLHF / ML Export

  • db.export_signals(ExportRequest) -- WAL-based signal export for training data
  • db.user_session_summary(user_id, since_ns) -- aggregated session statistics

Stability

tidalDB 0.1.0 is pre-1.0. No API or data format stability guarantees are made for 0.x releases. Upgrade guides will be provided for each minor version bump. Do not upgrade 0.x to 0.y on a live data directory without reading the release notes.


Format based on Keep a Changelog