tidaldb/docs/ops/monitoring.md
jx12n d5d1e7d81a feat(m11): observability+ops (m11p8) + perf-sweep wave 2 T2
m11p8 closes G-O + §1.4-3:
- Cluster metrics: breaker state, forwards, self-heal on /metrics; multi-shard sibling render (shard="N")
- Grafana cluster row + 8-rule Prometheus alert group
- Request-id / TraceLayer on both cluster routers; id rides forward hop
- Truthful status: flushed leader applied_events frontier; post-promote ShardId(0) keying fix
- Self-driving heal: tick_self_heal re-arms stuck-peer backlog every ~3s
- WAL PITR: wal.archive_dir, archive-before-delete gap-free
- tidalctl backup/restore with BLAKE3 content-hash verification
- Rolling-upgrade build_version handshake (N/N+1, never rejects) + Woodpecker release gate

perf-sweep wave 2 T2: one-get-per-type pre-pass in ranking executor
- signal_values.rs pre-fetches all signal kinds before scoring loop
- Eliminates per-item repeated DashMap lookups: −18.8% for_you, −31% under writes
- Byte-identical output verified with A/B test harness
2026-06-13 09:17:49 -06:00

272 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Monitoring
This document covers tidalDB's built-in Prometheus metrics endpoint, all exposed metrics, and recommended alerting thresholds.
---
## Setup
Enable the metrics HTTP server via the builder:
```rust
let db = TidalDb::builder()
.with_data_dir("/var/lib/tidaldb")
.with_schema(schema)
.enable_metrics("127.0.0.1:9090")
.open()?;
// Discover the bound address (useful when using port 0):
if let Some(addr) = db.metrics_addr() {
println!("metrics at http://{}/metrics", addr);
println!("health at http://{}/healthz", addr);
}
```
**Security:** The metrics endpoint has no authentication. Bind to `127.0.0.1` (loopback) only. If you need to scrape from a remote Prometheus server, use your infrastructure's network controls (SSH tunnel, reverse proxy with auth, or VPN) rather than binding to `0.0.0.0`. tidalDB logs a WARN-level message if you bind to a non-loopback address.
**Feature flag:** The metrics HTTP server requires the `metrics` feature, which is enabled by default. Build with `--no-default-features` to disable the HTTP server entirely. Base metrics (`uptime_seconds`, `health_ok`, `info`, `checkpoint_failures_total`) are always compiled regardless of the feature flag.
---
## Endpoints
| Path | Content-Type | Description |
|:-----|:-------------|:------------|
| `/metrics` | `text/plain` | Prometheus text exposition format |
| `/healthz` | `application/json` | JSON health check: `{"status":"ok","uptime_seconds":123.456,"version":"0.1.0","build_hash":"..."}` |
---
## Prometheus Scrape Configuration
```yaml
scrape_configs:
- job_name: 'tidaldb'
static_configs:
- targets: ['127.0.0.1:9090']
scrape_interval: 15s
```
---
## Metrics Reference
All metrics use the `tidaldb_` prefix. Metrics marked with "(feature-gated)" are only emitted when the `metrics` Cargo feature is enabled (default: enabled).
### Build and Health
| Metric | Type | Description | Labels |
|:-------|:-----|:------------|:-------|
| `tidaldb_uptime_seconds` | gauge | Seconds since the database was opened. Monotonically increasing. | `partition_id="0"` |
| `tidaldb_health_ok` | gauge | Whether the database is healthy. `1` = ok, `0` = degraded or closed. | `partition_id="0"` |
| `tidaldb_info` | gauge | Build and version information. Always `1`. | `version`, `build_hash`, `partition_id="0"` |
**Normal range for `tidaldb_health_ok`:** Always `1` during normal operation. Drops to `0` during shutdown or if an internal health check fails. Alert immediately if `0` during expected uptime.
### Signal System (feature-gated)
| Metric | Type | Unit | Description |
|:-------|:-----|:-----|:------------|
| `tidaldb_signal_writes_total` | counter | count | Total signal writes since database open. Includes all signal types across all entities. |
| `tidaldb_signal_hot_entries` | gauge | count | Number of entries currently in the signal ledger hot tier (DashMap). Each entry is one `(entity_id, signal_type_id)` pair. |
| `tidaldb_signal_write_latency_us` | histogram | microseconds | Signal write latency distribution. Bucket boundaries: 1, 5, 10, 25, 50, 100, 250, 500, 1000, 5000, 10000 microseconds. |
**Normal range for `signal_hot_entries`:** Proportional to `active_entities * signal_types_per_entity`. The hot tier is trimmed at 5M entries (`DEFAULT_MAX_SIGNAL_ENTRIES`). Alert if approaching 80% of budget (4M entries).
**Normal range for `signal_write_latency_us`:** p50 should be < 50us, p99 should be < 1ms. If p99 exceeds 5ms, investigate WAL write latency or DashMap contention.
### WAL and Checkpoint (feature-gated)
| Metric | Type | Unit | Description |
|:-------|:-----|:-----|:------------|
| `tidaldb_wal_lag_bytes` | gauge | bytes | Total bytes of WAL segment files not yet compacted. Updated after each checkpoint cycle. |
| `tidaldb_wal_compacted_segments_total` | counter | count | Total WAL segments deleted by compaction since database open. |
| `tidaldb_checkpoint_age_seconds` | gauge | seconds | Seconds since the last successful signal checkpoint. Derived from `last_checkpoint_ns` at render time. |
| `tidaldb_checkpoint_failures_total` | counter | count | Total number of failed periodic signal checkpoints. **Not feature-gated** -- always emitted. |
**Normal range for `checkpoint_age_seconds`:** Should stay below 60 seconds (checkpoint runs every 30 seconds, with some jitter from the 500ms poll interval). Alert if > 300 seconds (5 minutes) -- the checkpoint thread may be stuck or the storage engine is under pressure.
**Normal range for `wal_lag_bytes`:** Depends on signal write rate. At 1K signals/sec, expect ~1.2 MB of WAL per 30-second checkpoint cycle. Alert if > 1 GB -- compaction may be failing.
**Normal range for `checkpoint_failures_total`:** Should be 0. Any non-zero value means signal durability is at risk -- the hot tier is not being persisted. Investigate storage errors (disk full, I/O errors).
**Cluster mode:** per-region replication lag is exported as Prometheus series since m11p1 (`tidaldb_cluster_peer_ship_queue_depth` and `tidaldb_cluster_peer_acked_seqno`, labeled by `peer_shard`; the quorum lag is `tidaldb_cluster_relay_last_seq tidaldb_cluster_relay_durable_seq`) — see [Cluster Replication](#cluster-replication-feature-gated-m11p1m11p5) below. It is *also* observable via the cluster HTTP surface (`GET /cluster/status`, each region's `lag_events`); a growing value flags a partitioned or wedged follower. See the [cluster runbook](../runbooks/cluster.md#6-cluster-management-api).
### Index Health (feature-gated)
| Metric | Type | Unit | Description |
|:-------|:-----|:-----|:------------|
| `tidaldb_tantivy_segment_count` | gauge | count | Number of Tantivy index segments for the items text index. |
| `tidaldb_tantivy_indexed_docs` | gauge | count | Number of documents indexed in the items Tantivy text index. |
| `tidaldb_usearch_index_size_bytes` | gauge | bytes | Estimated total byte size of all USearch vector index files (f16). |
| `tidaldb_usearch_vector_count` | gauge | count | Number of vectors stored across all USearch indexes. |
| `tidaldb_bitmap_index_cardinality` | gauge | count | Total entity IDs across all four bitmap indexes (category + format + creator + tag). |
Index health metrics are refreshed every 10 seconds by the checkpoint thread (3x more frequently than checkpoints) so operators get near-real-time visibility.
**Normal range for `tantivy_segment_count`:** Should stay below 20 during normal operation. Tantivy merges segments in the background. If segment count grows unbounded, the text syncer thread may have stalled.
**Normal range for `usearch_vector_count`:** Should match the number of entities with embeddings written via `write_item_embedding()` or `write_creator_embedding()`.
### Session Lifecycle (feature-gated)
| Metric | Type | Unit | Description |
|:-------|:-----|:-----|:------------|
| `tidaldb_active_sessions` | gauge | count | Number of currently active agent sessions. |
| `tidaldb_closed_sessions_total` | counter | count | Total agent sessions closed (explicitly or by sweeper) since database open. |
| `tidaldb_session_auto_closed_total` | counter | count | Total sessions auto-closed by the TTL sweeper due to exceeding `max_session_duration`. |
**Normal range for `active_sessions`:** Depends on your application's agent concurrency. Each open session consumes memory for signal state tracking. Alert if this grows unbounded -- agents may be leaking sessions (opening without closing).
### Rate Limiting and Degradation (feature-gated)
| Metric | Type | Unit | Description |
|:-------|:-----|:-----|:------------|
| `tidaldb_rate_limited_total` | counter | count | Total signal write requests rejected due to per-agent rate limits since database open. |
| `tidaldb_degradation_level` | gauge | level | Current graceful degradation level. `0` = full quality, `1` = reduced candidates, `2` = coarse aggregates, `3` = no diversity enforcement. |
**Normal range for `degradation_level`:** Should be `0` during normal operation. Any value > 0 means the load detector has triggered degradation to protect latency. Investigate system load (CPU, memory pressure, I/O saturation).
### Cluster Replication (feature-gated, m11p1m11p8)
Emitted only on cluster nodes (the series activate when cluster mode takes the
metrics handle; standalone deployments keep their exact metric surface). In
multi-process cluster mode the listener binds the topology's per-region
`metrics_addr` (or, for a seed-joined node, the `--metrics` flag).
| Metric | Type | Unit | Description |
|:-------|:-----|:-----|:------------|
| `tidaldb_cluster_ship_rtt_us` | histogram | microseconds | Replication batch ship round-trip time, all peers. Bucket boundaries: 100µs10s. |
| `tidaldb_cluster_ship_batch_events` | histogram | events | Events per shipped replication batch (batching effectiveness; 1 = no coalescing). |
| `tidaldb_cluster_wal_fsync_us` | histogram | microseconds | WAL group-commit fsync wall time. The load-bearing number behind `wal.batch_timeout_ms` tuning. |
| `tidaldb_cluster_group_commit_events` | histogram | events | Events per WAL group-commit batch (fsync amortization; 1 = every write pays a solo fsync). |
| `tidaldb_cluster_write_pool_depth` | gauge | count | Queued cluster write jobs awaiting a pool worker. |
| `tidaldb_cluster_write_pool_rejections_total` | counter | count | Write submissions shed with backpressure (HTTP 429). |
| `tidaldb_cluster_relay_last_seq` | gauge | seqno | Leader stream high-water mark. Since m11p2 the stream is the WAL itself, so this reports the WAL flushed frontier. |
| `tidaldb_cluster_relay_durable_seq` | gauge | seqno | **The quorum commit index** (m11p3): highest seqno a majority of the replica set durably holds. `relay_last_seq relay_durable_seq` is the cluster's quorum lag. |
| `tidaldb_cluster_quorum_timeouts_total` | counter | writes | `ack=quorum` writes that timed out awaiting the commit index (each returned a retryable 503 naming the laggards). |
| `tidaldb_cluster_peer_acked_seqno` | gauge | seqno | Per peer (`peer_shard` label): contiguous frontier accepted by the peer's transport. |
| `tidaldb_cluster_peer_ship_queue_depth` | gauge | events | Per peer: `relay_last_seq acked` — flushed events not yet accepted (or self-reported applied) by this peer. |
| `tidaldb_cluster_peer_ship_batches_total` | counter | count | Per peer: batches shipped. |
| `tidaldb_cluster_peer_ship_events_total` | counter | count | Per peer: events shipped. |
| `tidaldb_cluster_peer_ship_failures_total` | counter | count | Per peer: failed batch ship attempts (transient + permanent). |
| `tidaldb_cluster_election_term` (m11p4) | gauge | term | This node's current election term (0 = the pre-election "topology era"). |
| `tidaldb_cluster_election_role` (m11p4) | gauge | enum | Election role: `0` follower, `1` pre-candidate, `2` candidate, `3` leader. |
| `tidaldb_cluster_elections_started_total` (m11p4) | counter | count | Elections (pre-vote rounds) this node has started. A nonzero `rate()` under a stable cluster flags election churn (flapping links). |
| `tidaldb_cluster_leader_changes_total` (m11p4) | counter | count | Leadership changes this node has observed. |
| `tidaldb_cluster_divergence_quarantined` (m11p4) | gauge | bool | Divergent-suffix quarantine latch (`1` = fenced from the data plane). Clears only after a genuine reseed (m11p5). |
| `tidaldb_cluster_reseed_required` (m11p5) | gauge | bool | Durable reseed-marker latch (`1` = a snapshot reseed is pending the next boot). |
| `tidaldb_cluster_snapshot_staged` (m11p5) | gauge | seqno | Seq of the currently-staged snapshot artifact (`0` = none staged). |
| `tidaldb_cluster_snapshot_fetches_total` (m11p5) | counter | count | `FetchSnapshot` streams served as the leader-side snapshot source. |
| `tidaldb_cluster_snapshot_pin_force_drops_total` (m11p5) | counter | count | Staged-artifact retention pins force-dropped past the hard cap (a dead joiner that never released — alert: a stuck reseed is freezing compaction). |
| `tidaldb_cluster_remove_delivery_giveups_total` (m11p5) | counter | count | Removal-delivery graces that expired before the removed peer acked the `Removed` record (the node was down/unreachable during decommission). |
| `tidaldb_cluster_forwards_total` (m11p8) | counter | count | Cross-node write forwards this gateway initiated. A high `rate()` at one node = client routing imbalance, not a fault. |
| `tidaldb_cluster_forward_failures_total` (m11p8) | counter | count | Forwards that errored (leader unreachable / 5xx / timeout). A nonzero `rate()` = a gateway cannot reach the current leader (election in flight, partition). |
| `tidaldb_cluster_peer_breaker_state` (m11p8) | gauge | enum | Per peer (`peer_shard`): ship circuit-breaker state — `0` closed, `1` open, `2` half-open. `== 1` for minutes = replication to that peer is stalled. |
| `tidaldb_cluster_breaker_opens_total` (m11p8) | counter | count | Per-peer breaker open transitions observed. A rising counter under steady load = a flapping peer link. |
| `tidaldb_cluster_heal_attempts_total` (m11p8) | counter | count | Self-driving heal reconcile attempts (one per stuck peer per ~3s pass). |
| `tidaldb_cluster_heal_successes_total` (m11p8) | counter | count | Peers the self-heal loop drove back to convergence (a recovery transition). |
| `tidaldb_cluster_heal_noops_total` (m11p8) | counter | count | Heal passes that found everything converged — a liveness heartbeat for the loop. |
| `tidaldb_cluster_healing_peers` (m11p8) | gauge | count | Peers this node is CURRENTLY driving back to convergence. **`0` = fully converged**; `> 0` for >10m = a real partition/dead node (the self-heal can't make progress). |
**The per-node `/metrics` listener (m11p8).** Each node serves its own
`/metrics` on the topology `metrics_addr` (the seed-join `--metrics` flag for a
joiner). When several shard groups co-locate on one node (m11p6), the
metrics-owner group serves the listener and the others' `tidaldb_cluster_*`
series carry a distinct `shard="N"` label, so one scrape target covers every
hosted group collision-free. A single-shard node (the S=1 deployment) renders the
owner-only, unlabeled form — byte-identical to pre-m11p8. The Grafana cluster
dashboard ([grafana-dashboard.json](grafana-dashboard.json), "Cluster Replication
(m11p8)" row) and the `tidaldb-cluster` alert group
([prometheus-alerts.yaml](prometheus-alerts.yaml)) cover every series above.
> **Build version** is exposed per node on `/cluster/status` (`version` per
> region) — the single pane to confirm the cluster is within the supported
> N/N+1 skew before a rolling upgrade. On the wire it rides
> `HeartbeatRequest.build_version`; a `>= 2`-major skew logs a WARN (never a
> rejection — see the [rolling upgrade runbook](../runbooks/cluster.md#14-rolling-upgrade--version-skew-m11p8)).
> **Membership conf-version + learner promotion** are exposed as
> `/cluster/status/local` JSON fields (`membership_version`,
> `promotion_pending (lag=N)`), not registered Prometheus series — poll the
> status surface for scale-up progress and a stuck learner.
### Security signals (m11p7) — log/tracing, not metrics
m11p7's security observability is **structured logs**, not new `tidaldb_cluster_*`
series:
- **Admin audit** — the `tidal_audit` tracing target (and, when set, the
`TIDAL_AUDIT_LOG` JSONL file) carries one record per admin verb:
`{principal, verb, target, term, outcome}`. Scrape this target into your log
pipeline; alert on `outcome` values starting `error:` or `rejected`, and on any
`promote` / `member_remove` whose `principal` is unexpected.
- **Cert rotation** — an INFO log `TLS material rotated…` (gRPC) /
`inter-node HTTP server cert hot-swapped` (HTTP) marks each hot-swap; a WARN
`TLS rotation poll failed` / `peer-channel rebuild failed` means a rotation did
NOT take (a half-written cert mid-roll, or a bad CA) — the node keeps the prior
cert. Alert on the WARN.
- **TLS handshake rejections** — a foreign pod / bad client cert / plaintext probe
logs at DEBUG (`rejected inbound gRPC handshake…` / `HTTP TLS handshake
rejected`). A burst is a probing/misconfiguration signal; raise the log level to
observe it.
- **Rate-limit denials** — reuse the existing **429** signal: a per-principal
deny is a 429 with `Retry-After` on the protected routes (distinct from the
write-pool 429 — the body says `rate limit exceeded`).
---
## Recommended Alerts
| Alert Name | Condition | Severity | Meaning |
|:-----------|:----------|:---------|:--------|
| TidalDB Down | `tidaldb_health_ok == 0` | Critical | Database is unhealthy or shut down. Immediate investigation required. |
| Checkpoint Stale | `tidaldb_checkpoint_age_seconds > 300` | Warning | Checkpoint has not run in 5+ minutes. Signal durability at risk. Check storage I/O and disk space. |
| Checkpoint Failures | `tidaldb_checkpoint_failures_total > 0` | Warning | At least one checkpoint has failed. Signal state may not be durable. Check disk space and storage errors. |
| WAL Disk Pressure | `tidaldb_wal_lag_bytes > 1000000000` | Warning | WAL exceeds 1 GB uncompacted. Compaction may be stuck or checkpoint is failing. |
| Signal Backlog | `tidaldb_signal_hot_entries > 4000000` | Warning | Signal ledger over 80% of the 5M entry budget. Cold entry trimming will begin at 5M. |
| Degraded Ranking | `tidaldb_degradation_level > 0` | Warning | Load-based degradation is active. Ranking quality is reduced to protect latency. Scale up or reduce load. |
| Session Leak | `deriv(tidaldb_active_sessions[5m]) > 0.03 AND tidaldb_active_sessions > 100` | Warning | Active session count trending up. `tidaldb_active_sessions` is a gauge, so `deriv()` (slope) is correct — `rate()`/`increase()` are invalid on gauges and never fire. Agents may not be closing sessions. |
| High Rate Limiting | `rate(tidaldb_rate_limited_total[5m]) > 100` | Info | Sustained rate limiting. Review agent rate limit configuration or reduce write volume. |
| Tantivy Segment Bloat | `tidaldb_tantivy_segment_count > 30` | Warning | Tantivy has many unmerged segments. Text syncer may be stalled. |
| Cluster Peer Ship Stall | `deriv(tidaldb_cluster_peer_acked_seqno[2m]) == 0 AND tidaldb_cluster_peer_ship_queue_depth > 0` | Critical | A peer stopped accepting batches while events queue behind it (partition, dead peer, or paused sender). Check `/cluster/status` and heal. |
| Cluster Quorum Lag | `tidaldb_cluster_relay_last_seq - tidaldb_cluster_relay_durable_seq > 10000` | Critical | A majority of the replica set is not confirming durability (down/partitioned followers, or follower apply throughput exhausted). `ack=quorum` writes will 503; the bodies name the laggards. |
| Quorum Timeouts | `rate(tidaldb_cluster_quorum_timeouts_total[5m]) > 1` | Warning | `ack=quorum` writes are timing out (retryable 503s). Sustained timeouts = a laggard region or an over-budget `replication.quorum_timeout_ms` for the deployment's RTT. Elevated timeouts while `relay_durable_seq` (the commit index) holds steady and followers report healthy = the `ReportApplied` frontier pushes are being lost (packet loss / a pre-m11p3 leader) — followers WARN `applied-frontier report failed` on the first failure of a streak. |
| Cluster Write Shedding | `rate(tidaldb_cluster_write_pool_rejections_total[5m]) > 100` | Warning | Sustained 429 shedding on the cluster write path. Raise `write_workers`/capacity or reduce offered write rate. |
| Election Churn | `rate(tidaldb_cluster_elections_started_total[5m]) > 0.1` | Warning | A node keeps starting elections — flapping links, an over-tight `election.leader_lease_ms`, or a partial partition. Cross-check `tidaldb_cluster_leader_changes_total`. |
| Divergence Quarantine | `tidaldb_cluster_divergence_quarantined == 1` | Critical | A node fenced itself from the data plane after detecting a divergent suffix. Since m11p5 it auto-reseeds on its next boot; if the latch persists, the reseed is not completing — check `tidaldb_cluster_reseed_required` and the snapshot path. |
| Reseed Pending | `tidaldb_cluster_reseed_required == 1` for `> 10m` | Warning | A node has latched the reseed marker but has not completed a snapshot reseed. Expected briefly after a quarantine or a behind-a-compacted-leader restart; a persistent latch means the snapshot fetch is failing (no reachable leader, capability gate, or staging fault). |
| Snapshot Pin Force-Drop | `increase(tidaldb_cluster_snapshot_pin_force_drops_total[1h]) > 0` | Warning | A staged-snapshot retention pin was force-dropped past the hard cap — a joiner started a reseed and never released (died mid-fetch). The dropped pin protects compaction; the stranded joiner must be re-driven or removed. |
### Grafana Dashboard Suggestions
**Row 1: Health overview**
- `tidaldb_health_ok` (stat panel, green/red)
- `tidaldb_uptime_seconds` (stat panel)
- `tidaldb_degradation_level` (stat panel, thresholds at 1/2/3)
- `tidaldb_info` labels (stat panel showing version + build hash)
**Row 2: Signal throughput**
- `rate(tidaldb_signal_writes_total[5m])` (time series, signals/sec)
- `tidaldb_signal_write_latency_us` histogram (heatmap or quantile panel)
- `tidaldb_signal_hot_entries` (gauge, threshold at 4M/5M)
**Row 3: Durability**
- `tidaldb_checkpoint_age_seconds` (time series, threshold line at 300)
- `tidaldb_checkpoint_failures_total` (stat panel, should be 0)
- `tidaldb_wal_lag_bytes` (time series)
- `rate(tidaldb_wal_compacted_segments_total[5m])` (time series)
**Row 4: Index health**
- `tidaldb_tantivy_indexed_docs` (stat panel)
- `tidaldb_tantivy_segment_count` (gauge)
- `tidaldb_usearch_vector_count` (stat panel)
- `tidaldb_usearch_index_size_bytes` (stat panel, bytes format)
- `tidaldb_bitmap_index_cardinality` (stat panel)
**Row 5: Sessions**
- `tidaldb_active_sessions` (time series)
- `rate(tidaldb_closed_sessions_total[5m])` (time series)
- `tidaldb_session_auto_closed_total` (stat panel)
- `rate(tidaldb_rate_limited_total[5m])` (time series)