tidaldb/CHANGELOG.md
jx12n d0a52e4530 feat(m11): catch-up timer retry + TSEG segment version header (m11p4)
WAL segment format: 8-byte TSEG header (magic + version byte + 3 reserved)
prepended to every new segment. Legacy headerless segments (m0-m11p3) read
as implicit v0 — no migration. Unknown magic/version surfaces as
WalError::SegmentFormatUnknown at open time; foreign files are never
repaired or truncated (fixes the silent data-loss path from the p3 rollout
incident where torn-tail repair zeroed a follower's unreadable segments).

Catch-up transport: FAILED_PRECONDITION ("snapshot required") and stream
errors that skip the shard now arm a timer retry (re-arm-on-skip is the
load-bearing liveness fix — without it a skipped pull never re-fires and
the follower stays permanently behind). Single retry pending per shard;
CatchupRunner owns the Arc'd state shared between the retry tasks and the
transport. Test: tidal-net/tests/catchup_retry.rs covers the retry path.

Stress: k8s stress-job-t2a/t2b yaml + ops/stress-test-p3-t2 runbook.
2026-06-11 17:05:20 -06:00

396 lines
25 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Changelog
All notable changes to tidalDB will be documented in this file.
## [Unreleased]
### Added
**Catch-up self-healing + WAL segment format versioning (m11p4) — timer-retried pulls, `TSEG` segment header, structured "snapshot required"**
- **Failed catch-up pulls retry on a timer.** The pull trigger was event-only:
a follower whose `StreamSegments` pull failed (e.g. the leader's gRPC server
not yet ready during a rolling restart) waited for the next PUSHED segment
to re-expose the gap — in an idle cluster that push never comes, and the
follower stayed lagged forever (the 2026-06-11 p3 rollout: both followers
stuck at lag=136507). A failed pull now arms a one-shot timer
(`replication.catchup_retry_ms`, default 30000; transport
`catchup_retry_interval`) that re-pulls from the CURRENT applied frontier.
Pulls stay single-flight and rate-limited; the timer's wake-up re-arms when
consumed by the rate limit or an in-flight pull, so the gap always keeps a
standing wake-up until a pull completes. Verified over real sockets:
`catchup_retry.rs` reproduces the incident (pull fails, leader appears,
zero pushes) and proves timer-only self-heal — and that a clean completion
arms nothing.
- **WAL segment files are format-versioned.** New segments open with an
8-byte header (`TSEG` magic + version byte + reserved); pre-m11p4
headerless segments stay readable as implicit version 0 — no migration. A
segment this binary cannot identify (unknown header version, unrecognized
leading bytes, unparseable `.seg` filename) surfaces as the new
`WalError::SegmentFormatUnknown` at open — previously it scanned as
empty (`segments=0`) and recovery's torn-tail repair could TRUNCATE the
foreign file to zero. Foreign-format files are never repaired, truncated,
or skipped. Downgrade across m11p4 requires a WAL reseed (runbook §8).
- **Unservable catch-up is a structured refusal.** `SegmentSource::collect_from`
returns typed `SegmentReadError::{Unavailable,Failed}`;
`TidalDb::read_wal_batches` returns the typed `WalError` (was stringified
`TidalError`). The `StreamSegments` handler maps `Unavailable` to
`FAILED_PRECONDITION``"segments not available from seq N; snapshot
required"` — and the follower logs it distinctly (*catch-up unservable …
needs a snapshot (m11p5) or an operator reseed*) instead of burying it as a
transient. The on-disk segment format is now documented
(`tidal/src/wal/segment.rs` module docs + spec 01 §2.2).
**Quorum-acked writes (m11p3) — `ack=leader|quorum`, durable ship acks, commit index, zero-acked-loss ledger gate (closes G4)**
- **`ack=quorum`** is an opt-in durability contract for every replicated write
(`/signals`, `/items`, `/embeddings`): success means a **majority of the
replica set durably holds the write** (leader + `floor(n/2)` followers, each
storage-applied and own-WAL-fsynced), so an acked write survives the
permanent loss of any single node — including the leader. Deployment default
via topology `replication.ack`; per-request override via the **`x-tidal-ack`**
header (forwarded verbatim by gateways). `ack=leader` (the default) is the
m0m11p2 contract unchanged.
- **Durable frontier reports.** Followers PUSH their durably-applied frontier
to the leader once per apply round (new `ReportApplied` RPC, fired by the
segment receiver through the new `Transport::notify_applied`) — batch-level
and fully decoupled from ship acks, so the commit index stays fresh even
when outbound ships stall (gap-parked follower, pull-based catch-up, quiet
leader). Ship acks keep their m11p2 instant floor-hint semantics — both
inputs are durable-true because a follower's frontier only ever advances
after its storage apply + own-WAL fsync. (The first design held each ship
ack until its segment's apply; measured under open-loop load, that couples
ship cadence to apply latency and one gap-parked follower spirals into
total quorum collapse — the report push is what shipped.)
- **Follower blob applies are group-committed.** m11p2 applied replicated
items/embeddings one record at a time — one solo follower fsync per item,
capping item apply at the fsync floor (~100/s on macOS) and stalling the
quorum frontier behind any item burst. The receiver now hands each apply
round's blob records to the engine as ONE batch (`apply_replicated_blobs`:
validate all → stage all WAL appends → wait all → upsert storage), and the
WAL writer flushes queued blobs under ONE group fsync. Measured: corpus
seeding 2,000 items + embeddings 39.3s → 1.8s (22×).
- **Commit index.** The leader folds durable acks (and heal resumes) into
per-peer durable marks; the commit index is the k-th largest (k =
`floor(n/2)`), leadership-scoped (promote resets it to the stream baseline;
demotion fails in-flight waiters — a demoted leader never claims quorum).
Handlers await it through a watch-channel bridge — **fully async, zero
threads parked per waiter** (the thread-per-wait design measurably collapsed
at 1k rps open-loop by exhausting the blocking pool and starving the very
completions that advance the index). Quorum capacity measured on a real
3-process localhost cluster (release, writes mix): **3,600 quorum
signal-writes/s within SLO** (p50 ~45ms, zero errors, replication lag ≤3
events at ramp end; knee not reached) — 79% of m11p1's 4,534/s leader-ack
figure, vs the ≥50% gate.
- **Honest timeout semantics.** Quorum not confirmed within
`replication.quorum_timeout_ms` (default 2000) → a **retryable 503** naming
the laggard regions, the commit index, and needed/confirmed counts. The
write is in the leader's log and may still commit: retries are
at-least-once (items/embeddings retries are idempotent upserts; signal
retries can double-count — decided in-phase: no idempotency-key machinery,
documented in runbook §8 with the session-write precedent for callers that
need exact-once).
- **`x-tidal-seq`** on every cluster write response: the write's seqno in the
replicated log (relayed through forwards) — an exact durability cursor
against `commit_index` in `/cluster/status/local` (which also gains `ack`).
`tidaldb_cluster_relay_durable_seq` now reports the commit index
(`relay_last_seq relay_durable_seq` = quorum lag); new counter
`tidaldb_cluster_quorum_timeouts_total`.
- **The ledger gate (exit gate, run for real):** tier-3
`mp_quorum_ledger_zero_acked_loss_across_killpoints` SIGKILLs the leader
under concurrent quorum load and proves zero acknowledged loss on the
promoted max-applied survivor — frontier invariant (max acked seq ≤
survivor applied) plus per-item content probes. **167/167 kill points
passed on the final design** (batches of 64 + 91 + 12; kill timings spread
120598ms across fresh 3-process clusters; an earlier 100/100 run had
validated the superseded ack-holding design before it was replaced — see
docs/planning/milestone-11/phase-3.md). Plus tier-3 partition semantics
(one follower down: quorum commits; both down: fast 503 naming laggards
while `ack=leader` flows; heal: recovers) and in-process gRPC coverage
(override headers, forwarded quorum writes, blob writes, 400 on bad mode).
- **Rolling-upgrade order (mixed-version caveat):** a pre-m11p3 leader
neither serves the `ReportApplied` RPC nor recognizes `x-tidal-ack` — it
silently applies LEADER-ack semantics to a `quorum` request (a durability
downgrade the caller cannot see). Upgrade the leader first: `ack=quorum`
is then honored immediately (commit-index freshness rides the m11p2
ship-ack floor hints until the followers upgrade too). Replication and
heal are unaffected by either order. Runbook §8 records the procedure.
**One replicated log (m11p2) — items/embeddings ride the WAL, `StreamSegments` catch-up, HTTP broadcast deleted**
- **The leader's WAL is now THE replicated log.** The group-commit writer hands
every fsynced batch to a bounded in-memory **ship feed** (`wal::feed::WalShipFeed`);
the ship queue pushes those already-encoded bytes verbatim — byte-identical on
the leader's disk, the wire, and the follower's apply path — and stream seqnos
are WAL seqnos, so they **survive restarts** (the m8p10 relay-reset hazard is
gone). The m11p1 in-memory relay log, its durable frontier, and its poisoning
machinery left the server write path entirely (`/signals` stages straight
through the engine's group commit; a WAL fsync failure now surfaces per-write
exactly like single-node).
- **Item metadata and embeddings are replicated mutations.** `/items` and
`/embeddings` journal kind-1/2 **blob records** (WAL header `flags` byte =
batch kind; one record, one seqno) BEFORE storage, on the same stream as
signals. Followers apply them kind-aware — WAL-first into their own log, then
idempotent storage upserts — and recovery replays them. The m8p10 HTTP
item/embedding broadcast (marker-gated fan-out, O(items) heal re-broadcast,
authed side-POSTs — the source of both 2026-06-10 live bugs) is **deleted**;
those bug classes are now impossible by construction. Cluster `/items` returns
a plain 201 and `/embeddings` a plain 204 (no broadcast-report body).
- **`StreamSegments` implemented — catch-up is follower-pulled.** The ship feed's
tail is bounded; a peer that falls behind it is skipped ahead, and the
follower pulls the hole itself via the (previously declared-unimplemented)
server-streaming RPC over the leader's durable, BLAKE3-verified segments.
Pulls trigger on detected gaps, on follower boot (self-driving restart
catch-up), and on the leader's heal nudge (`POST /cluster/catchup`, internal,
forwards the operator's own bearer credential). Pulled chunks flow through the
same inbound apply path as live ships. Ship acks piggyback the follower's
applied seqno (`ShipSegmentResponse.applied_seqno`), so retries of
already-applied data prune and heal is `resume_from` + nudge — no redelivery
scan, no O(items) traffic.
- **Promote carries a stream baseline.** A promoted leader's stream starts at its
promote-time flushed frontier (persisted in `data_dir/stream_baseline`); the
fan-out body and every catch-up chunk announce it, so peers jump their
frontier past pre-stream history instead of parking on a phantom gap. Ship
queues are leadership-gated (`activate_from`/`deactivate`): a follower's
replicated applies never echo back at its peers.
- **Multi-process cluster mode now requires `--data-dir`** (validated at
startup): the durable WAL is the replication stream. Standalone single-node
deployments are untouched — blob journaling and the ship feed are gated on
cluster peers, so single-node item writes keep fjall-only durability with
zero extra fsyncs.
- New tier-3 suite `mp_items_ride_the_log_and_catchup_stream`: items written on
the leader AND through a follower gateway converge everywhere via the log
(feed parity 1e-6), and a follower stopped through item+embedding+signal
writes restarts and converges via its boot-time `StreamSegments` pull with no
heal verb and no HTTP item traffic.
**Cluster replication performance floor (m11p1) — ack/ship decoupled, batched + windowed shipping, first `tidaldb_cluster_*` metrics**
- The replicated `/signals` write path no longer serializes every writer onto a
solo group-commit fsync nor ships to followers on the request path. Writes are
**staged** (seqno + WAL submission + relay log push, microseconds, atomic with
rollback) and **completed** (shared group-commit fsync + in-memory fold) in two
phases, so concurrent writers coalesce into one fsync; follower shipping moved
to per-peer sender threads that coalesce contiguous runs into multi-event
batches with a windowed in-flight budget (topology knobs
`replication.{batch_max_events,window,retry_ms}`). The 204 contract is
unchanged (leader durability only) — it now returns at leader fsync. Measured
on a real 3-process localhost cluster (release build, thepeach mix):
**4,534 replicated signal-writes/s sustained within SLO** vs ~90/s before
(~50×), replication lag bounded at ≤377 events (~80ms) through the whole ramp.
- **Durable-frontier shipping + relay poisoning.** Senders only ship the
leader's contiguous fsynced prefix (an event a follower holds but the leader
could lose is silent divergence); a staged write whose fsync fails **poisons**
the relay — further cluster writes are rejected and the ship frontier freezes
(the CockroachDB/Postgres fsync-failure posture).
- **Follower group-commit coalescing.** The segment receiver drains its inbound
backlog (`Transport::try_recv_segment`) and applies it through ONE shared
group commit (`SignalLedger::apply_replicated_events`), in range-disjoint
groups so duplicate/subset re-ships cannot double-fold. Without this the
follower apply ceiling was ~`events-per-segment / fsync-cost` (~1.8k events/s
measured) and lag grew without bound under m11p1 leader rates.
- **First cluster metrics + cluster `/metrics` listener.** Cluster mode
previously had no metrics endpoint at all. New per-region topology
`metrics_addr` wires the engine's Prometheus listener; new `tidaldb_cluster_*`
series: ship RTT + batch-size histograms, per-peer ship counters/gauges
(`peer_shard` labels), WAL fsync latency + group-commit fill histograms,
write-pool depth/rejections, relay committed + durable frontiers. WAL
group-commit knobs are deployment config (`wal.{batch_size,batch_timeout_ms}`
topology block; `wal_batch_size`/`wal_batch_timeout` builder methods).
- Engine API: `TidalDb::signal_staged`/`StagedSignal::wait`,
`SignalRelay::{stage_write,complete_write,durable_seq,snapshot_range}`,
`ShipQueue` (per-peer windowed batch senders with pause/resume),
`WalWriter::append_signal_staged`, `WalSender::append_record_staged`,
`range_payload`/`encode_run`. Relay log entries are now `RelayEvent` (raw
event records, re-encoded deterministically at ship time) instead of
pre-encoded single-event bytes.
- Ship-sender failure logging is transition-based (first + every 50th
consecutive failure WARN with the running count, recovery INFO) — the
per-retry WARN flood could fill an undrained log pipe and stall the process.
The multiproc test harness now discards child logs via `/dev/null` (a piped
fd nobody drains deadlocks the node once the kernel buffer fills), with
`TIDAL_TEST_NODE_LOGS=inherit` to stream them while debugging.
**M9 — Community Sync & Revocation**
- Local embeddable profiles can opt into community personalization and safely leave/purge their contributions. New types: `SignalScope`, `CommunityId`, `Membership`, `MembershipEpoch`, `PolicyMetadata`. Community signal reconciliation via `CrdtSignalState` with commutative/associative/idempotent merge laws; membership-epoch revocation purges a departed member's contributed signals.
**M10 — Governance & Agent Rights**
- Community rules and agent-scoped permissions control what signals influence ranking: policy-metadata enforcement and agent-rights scoping wired into the signal-write and ranking paths.
**Cluster mode (m8p10): true multi-process region nodes + full tier-3 UAT — M8 COMPLETE**
- Multi-process cluster mode: `tidal-server cluster --region <name> [--data-dir <p>]`
(env `TIDAL_REGION`) runs **one process per region** (`RegionClusterState`), each
owning one `TidalDb` and one `GrpcTransport` that binds this region's `grpc_addr`
and dials every sibling's real `grpc_addr` — real process/host isolation. The
topology requires per-region `grpc_addr` AND `http_addr` in this mode;
single-process mode (no `--region`) is unchanged as the dev/demo default. Both
modes stay behind the experimental gate (`--experimental-cluster` /
`TIDAL_ALLOW_EXPERIMENTAL_CLUSTER=1`) with mode-specific WARN text.
- New routes / behaviors on the multi-process surface: `GET /cluster/status/local`
(per-node status), `GET /cluster/status` aggregates ALL regions with a `reachable`
field (unreachable ⇒ `reachable:false` + worst-case lag), `POST /cluster/reconcile
{region}` (cross-process CRDT snapshot exchange; idempotent — a repeat is an exact
no-op on scores), `POST /hardnegs {user_id,item_id}` (user-scoped hide; converges
via reconcile, filtered from `/feed?user_id`). Writes on a non-leader **forward to
the leader**; `?region=` reads forward to the owning region; default reads are
**LOCAL** in multi-process mode. Items/embeddings are leader-applied + HTTP-broadcast
to peers with a per-peer report: `POST /items` → 201 `{replicated_to, failed}`,
`POST /embeddings`**200** with the same report on the leader path (NOT 204 — a
204 cannot carry the report; the forwarded/internal path stays 204). `/sharded/*`
fans out across processes (degraded semantics preserved). `promote` fans out to all
peers (`{ok, leader, acked, failed}`).
- `/cluster/heal` is the single recovery verb: redelivers missed signal segments
(gap-aware) AND re-broadcasts item metadata + embeddings to the healed region
(idempotent upserts). After a partition the per-peer gRPC circuit breaker (threshold
5, reset 30s) is open, so re-issue `/cluster/heal` until `/cluster/status` shows
lag 0.
- `TIDAL_HLC_SKEW_MS` (multi-process): signed ms offset applied to the process's HLC
(reconcile LWW stamping only — not signal-decay timestamps); a test/ops escape hatch.
- Tier-3 UAT over real OS processes (feature `cluster-e2e`): `cluster_multiproc` (5),
`cluster_chaos` (3, REAL network-partition injection via a root-free in-harness TCP
relay proxy — the toxiproxy-style alternative the ROADMAP sanctions; iptables/pfctl
remain an operator option), `cluster_lifecycle` (2, ±500ms clock skew + rolling
upgrade with zero acknowledged-write loss), `cluster_runbook` (9, every
`docs/runbooks/cluster.md` §5§11 operation), plus `cluster_e2e` (2, single-process
smoke). Measured (localhost): replication p99 ~110133ms (< 2s SLA), failover
~3134ms (< 10s SLA), reconcile 01ms/side (< 100ms SLA).
### Fixed
**Four production bugs surfaced and fixed by the m8p10 tier-3 UAT**
- **Silent data loss on out-of-order ships.** The replication applied high-water-mark
swallowed sequence gaps when eager ships arrived out of order. The applied frontier
is now contiguous with a bounded ahead-buffer, so a gap can never be skipped.
- **Non-idempotent reconcile (score creep).** `take_crdt_snapshot` attributed replicated
signal streams per-node, so reconciling already-converged nodes crept the decayed
scores (0.5 0.375 …). Contributions are now attributed to one canonical
replication shard (`ShardId::SINGLE`), making `merge` idempotent reconcile of
converged nodes is an exact fixpoint.
- **Lag gauge conflated leader streams across a promotion.** A converged node reported a
permanent phantom lag after a `/cluster/promote` moved leadership to a different shard.
The lag gauge now tracks the leader high-water-mark per source shard
(`ReplicationLagGauge::leader_seqno_for`), computing lag against the current leader.
- **Items missed during a broadcast were never backfilled.** A node down/partitioned
during an item/embedding broadcast was permanently missing that data even at signal
lag 0. `/cluster/heal` now re-broadcasts item metadata + embeddings to the healed
region (idempotent upserts), making heal the single recovery verb.
### Changed
**Cluster mode (m8p8): real gRPC replication**
- `tidal-server`'s `ClusterState` now wires each follower region to a real
`tidal-net` `GrpcTransport` (self-loop over loopback gRPC) instead of in-process
crossbeam channels replication between regions traverses real gRPC/TCP
(serialization, circuit breaker, HTTP/2). Closes M8 gap G1.
- Follower gRPC ports are auto-allocated from the topology (`grpc_addr` optional
per region) and self-heal a transient bind race by retrying on a fresh port.
- Blocking gRPC ships triggered by `POST /signals` and `POST /cluster/heal` are
offloaded to a dedicated thread so they never block the axum reactor.
- `ClusterState` is constructed off the async reactor (`GrpcTransport::new`
blocks on its own runtime).
- The experimental opt-in gate and `docker/cluster/Dockerfile` are updated:
single-process cluster mode is honest that it replicates over real gRPC but
runs all regions in one process (no host/process isolation). (True
multi-process region nodes shipped subsequently in m8p10 see the m8p10 entry
above.)
**Docker build fixes** (all three images now that `tidal-server` pulls `tidal-net`)
- Install `protobuf-compiler` in the builder stage `tidal-net`'s build script
runs `tonic-build`, which needs `protoc` to compile the WAL-shipping `.proto`.
- Pin the builder base to `bookworm` (`rust:1.91-bookworm` /
`rust:1.91-slim-bookworm`) so its glibc matches the `debian:bookworm-slim`
runtime; the default trixie base emitted a `libmvec.so.1` dependency absent on
bookworm, aborting the binary at startup. Verified: `docker run` of the cluster
image serves a functional 3-region cluster (write replicates to both followers
over gRPC; region-pinned reads serve replicated data; SIGTERM exits 0).
- New tests: `tidal-server/tests/cluster_grpc.rs` (in-process gRPC replication +
HTTP offload path) and a hardened tier-3 `cluster_e2e.rs` (multi-process smoke
+ promote over real OS processes, feature-gated).
## [0.1.0] - 2026-02-23
### Added
**Core Database Engine**
- `TidalDb` embeddable database with `ephemeral()` and `with_data_dir()` open modes
- `SchemaBuilder` for defining signal types, decay parameters, and ranking profiles
- `TidalDbBuilder` fluent builder with schema, data directory, metrics, and rate limiter configuration
**Signal System**
- Typed signal recording with exponential decay scoring
- Hot-tier (DashMap) and warm-tier (BucketedCounter) signal storage
- Windowed aggregation: `OneHour`, `TwentyFourHours`, `SevenDays`, `AllTime`
- Signal velocity tracking
- WAL-backed signal durability with crash recovery
- Periodic signal checkpointing to fjall (every 30s)
- WAL compaction after each checkpoint
**Retrieval (RETRIEVE query)**
- 5-stage pipeline: universe, filter, score, diversify, return
- Filter expressions: `Eq`, `In`, `Gt`, `Lt`, `And`, `Or`, `Not`, `InCollection`, `InProgress`, `MinSignal`, `MaxSignal`, `NearLocation`
- Built-in ranking profiles: `trending`, `for_you`, `new`, `popular`, `recent`, and 20+ more
- Custom ranking profiles via `SchemaBuilder`
- Diversity enforcement (max N per category/creator)
- Sort modes: `Relevance`, `Trending`, `Newest`, `MostLiked`, `MostViewed`, `MostFollowed`, `AlphabeticalAsc/Desc`, `Shortest/Longest`, `LiveViewerCount`, `DateSaved`, and more
**Search (SEARCH query)**
- BM25 full-text search via Tantivy
- Approximate nearest-neighbor (ANN) semantic search via USearch HNSW
- Reciprocal Rank Fusion (RRF) combining BM25 + ANN scores
- Creator search with `entity_kind(EntityKind::Creator)`
- `similar_to(EntityId)` for content-based recommendations
- Scope pre-filters: `Trending`, `CohortTrending`, `Following`, `Category`, `Collection`
- Autocomplete suggestions via `db.suggest()`
**Entity Model**
- Three built-in entity types: `Item`, `User`, `Creator`
- Metadata storage as `HashMap<String, String>`
- Embedding slots (up to 4 per entity type) via USearch
- Relationships: `Follows`, `Blocks`, `Hide`, `Mute`, `InteractionWeight`
**Sessions**
- Session lifecycle: `open_session`, `close_session`
- Cross-session preference vector updates (EMA blend)
- Session snapshots with signal state and preference vectors
- Session serialization format v0x03 with backward compatibility
**Social Graph**
- Creator follower/following indexes
- Cohort membership (user segments)
- CoEngagementIndex for co-viewing patterns with LRU eviction
- Social graph filter for "followed creator" content scoping
**Collections**
- Named collections with `Private`, `Shared`, `Public` visibility
- `create_collection`, `add_to_collection`, `remove_from_collection`, `list_collections`
- `FilterExpr::InCollection` for collection-scoped retrieval
- Saved searches with `save_search`, `list_saved_searches`, `retrieve_saved_search`
**Observability**
- `enable_metrics(addr)` -- Prometheus-format `/metrics` endpoint + `/healthz` JSON
- 15+ metrics: signal writes, WAL lag, checkpoint age, degradation level, index health
- `tidaldb_checkpoint_failures_total` counter for checkpoint monitoring
- `TidalDb::diagnostics()` -- structured health snapshot
- WAL diagnostics and recovery tools
**Safety**
- Signal weight NaN/Inf validation (returns `TidalError::InvalidInput`)
- Metadata size bounds: 64 keys max, 8KB value max, 64KB total max
- Export request limit: 500K signals max per request
- `FilterExpr` complexity limit: 256 nodes max
- Data directory lock (`tidaldb.lock`) prevents dual-process corruption
- Schema fingerprint persistence detects decay parameter changes on reopen
- Bounded `closed_sessions` cache (10K max, LRU eviction)
- Metrics server non-loopback bind warning
**CLI (`tidalctl`)**
- `tidalctl` binary for database inspection and diagnostics
**RLHF / ML Export**
- `db.export_signals(ExportRequest)` -- WAL-based signal export for training data
- `db.user_session_summary(user_id, since_ns)` -- aggregated session statistics
### Stability
tidalDB `0.1.0` is pre-1.0. **No API or data format stability guarantees** are made for `0.x` releases. Upgrade guides will be provided for each minor version bump. Do not upgrade `0.x` to `0.y` on a live data directory without reading the release notes.
---
*Format based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/)*