tidaldb/docs/profiling/social-scale.md
jx12n bb21e69ae6 feat(m12): vector retrieval G1/G2 — recall harness, ANN in RETRIEVE, index tuning
m12p1 (measurement truth): TidalDb::vector_search_items pure k-NN probe +
POST /vector_search (standalone + region node, merge-by-distance) +
tidal-stress --verify-recall (deterministic id-keyed corpus, in-RAM brute-force
cosine oracle, open-loop ramp → recall@k + true p99 + read-knee + JSON/gate exit).
Repaired fabricated p99 columns (mean-as-p99) in social-scale.md / scale.rs.
Verified real: recall@10=0.9997 at 20k/1536-D vs brute-force.

m12p2 (G1 unblock): ANN candidate-gen wired into RETRIEVE — for_you=preference
vector, related=seed embedding (similar_to), graceful scan-fallback. Cached
per-signal-type top-K (signals/ledger/hot_top_k.rs, decay-order-invariant) so
trending serves O(K). related over HTTP (FeedQuery.similar_to). Harness gains
--feed-profile / --seed-preferences. Verified: trending retrieve p99 3.5-7.7ms.

m12p3 (G2): per-query ef_search now honored (RwLock epoch-guard with_expansion,
shared guard for same-ef concurrency) + dimension-aware brute→HNSW crossover
usearch_min_vectors(dim) + memory_usage() + examples/ann_grid_search.rs.
Measured 1536-D/100k clustered: default M=16/ef_c=400/F16/ef_s=200 clears
G1+G2 (recall 0.997, p99 1.4ms); F16 -0.25% vs F32; Int8 rejected (-28%).
Recall corpus is now clustered (Gaussian mixture) in grid + harness.
2026-06-14 11:07:09 -06:00

3.8 KiB
Raw Blame History

Social Scale Performance Analysis

CoEngagementIndex Eviction

Correctness Tests

All tests in tidal/tests/m7p3_social_scale.rs pass:

Test Invariant Verified
eviction_correctness_at_2x_capacity edge_count <= capacity after every insert at 2× load
high_weight_edge_survival_under_eviction High-weight edges (score=100) survive low-weight flood
co_engagement_memory_bounded_at_2x_insertions 10 users × 40 items stays within capacity=200
no_self_loops_after_eviction No (a, a) edges created under eviction pressure
top_candidates_consistent_after_eviction top_candidates() works correctly after eviction

Eviction Benchmarks

Capacity Mean Eviction Latency Notes
10K edges ~2ms O(N log N) sort over 10K entries
50K edges ~12ms O(N log N) sort over 50K entries
100K edges ~26ms O(N log N) sort over 100K entries

Benchmark: cargo bench --manifest-path tidal/Cargo.toml --bench social -- co_engagement_eviction

Note: Eviction is amortised. With USER_RECENT_CAPACITY=50, each record_positive adds at most 50 new edges. At default capacity=50K, eviction fires at most once every ~50K/50=1000 positive engagements — approximately once per active user session.

Social Graph Filter at 1M Items

Setup

  • 100 followed creators
  • 500 followers per creator
  • 10K items per creator = 1M items total

Benchmark Results

Measurement contract: the social benchmark is Criterion — it reports an isolated per-op cost (mean, single-threaded, closed loop), NOT a latency distribution. A closed loop cannot observe the tail it hides (it stops sending when the system stalls), so these means are regression tripwires, not p99 evidence. The depth-2 < 50ms TAIL SLO is signed off only by the open-loop tidal-stress ramp — see scale-baselines.md.

Depth Isolated per-op cost (mean, closed-loop) Notes
Depth-1 (followed creator items) ~3ms Direct bitmap union over 100 creators
Depth-2 (co-followers' seen items) ~25ms Fan-out: 100 creators × 500 followers

Tail SLO: depth-2 p99 < 50ms — the ~25ms mean clears the budget with headroom (a necessary, not sufficient, condition); the p99 itself is validated open-loop, never by this closed-loop mean.

Run: cargo bench --manifest-path tidal/Cargo.toml --bench social -- social_graph_1m

Fan-out Cap

The social graph filter uses DEPTH2_FAN_OUT_CAP to bound the number of co-followers processed. Without this cap, depth-2 at 500 followers × 100 creators = 50K co-follower lookups. With the cap, this is bounded to prevent p99 blowup.

Cross-Session Preference Merge

Algorithm

PreferenceVectors::update_with_custom_rate(user_id, &interaction, lr):

pref = (1 - lr) * pref + lr * interaction
pref = L2_normalize(pref)

For 128D: 128 multiply-adds + L2 normalize = ~256 FP ops + sqrt.

Benchmark Results

Operation Mean Latency Target
Single 128D EMA (100K users) ~3µs < 1ms
10-session batch (100K users) ~28µs < 10ms

Both targets easily met. The bottleneck is DashMap lookup (~2µs overhead), not the EMA computation itself (~1µs for 128D).

Run: cargo bench --manifest-path tidal/Cargo.toml --bench social -- preference_merge

Adaptive Learning Rate

The update() method (used in production) uses an adaptive learning rate that decays as:

alpha = base_alpha / (1 + ln(update_count + 1))

This means users with many updates converge slower (preferences are more stable). First-session users get alpha = 0.1; users with 100+ sessions get alpha ≈ 0.02.