p0: specify Beachhead Validation, advancing all three features to specified
P0 was the only milestone gating the product track and all three of its features sat in `draft` with no spec, while M9/M10/P1/PG1 are released. Engine work was running ahead of the validation that decides whether any of it is wanted. - p0-target-segment-recruitment: screening criteria per beachhead persona, a funnel sized to yield the 20-50 pilot cohort, outreach limits (no accuracy or onboarding promise the prototype cannot meet), consent and data handling, opaque participant ids only, segment balance, and a pre-pilot baseline-feed question so the readout has a control. - p0-concierge-pilot-loop: the 14-day daily loop with the manual source-QA gate the ROADMAP permits, the briefing card contract, the normative session boundary, instrumentation bound to existing pg1 surfaces (signal-type counters, feedback-loop histogram, /diagnostics snapshots) rather than a new counting path, weekly interviews on the fixed beachhead question set, an intervention ledger so concierge help cannot silently inflate quality, abort conditions, and the frozen handoff dataset. - p0-validation-readout: pre-registered GO/NO-GO/EXTEND rule over five gates with an explicit dropout/missed-day/partial-observation policy, double-coded interviews against the beachhead required answers, a sensitivity re-run that downgrades GO to EXTEND if the verdict flips, a falsification section, and a named reviewer who must argue the NO-GO case before publication. Every threshold traces to a ROADMAP P0 acceptance criterion or the beachhead doc; the three that neither document fixes (D2 retention floor, value-confirmed fraction, noise kill-frame ceiling) are marked TBD (owner: product) instead of being invented. Next directive for all three is create_design.
This commit is contained in:
parent
cc0b894480
commit
0634d2f4dc
@ -1,16 +1,15 @@
|
|||||||
id: p0-concierge-pilot-loop
|
|
||||||
slug: p0-concierge-pilot-loop
|
slug: p0-concierge-pilot-loop
|
||||||
title: Concierge Pilot Loop
|
title: Concierge Pilot Loop
|
||||||
description: Daily briefing workflow with manual QA process and interview cadence — run for 2 weeks with pilot cohort
|
description: Daily briefing workflow with manual QA process and interview cadence — run for 2 weeks with pilot cohort
|
||||||
phase: draft
|
phase: specified
|
||||||
created_at: 2026-03-03T06:29:53.125784Z
|
created_at: 2026-03-03T06:29:53.125784Z
|
||||||
updated_at: 2026-03-03T06:29:53.125784Z
|
updated_at: 2026-08-16T17:52:33.904168Z
|
||||||
artifacts:
|
artifacts:
|
||||||
- artifact_type: spec
|
- artifact_type: spec
|
||||||
status: missing
|
status: approved
|
||||||
path: .sdlc/features/p0-concierge-pilot-loop/spec.md
|
path: .sdlc/features/p0-concierge-pilot-loop/spec.md
|
||||||
created_at: null
|
created_at: null
|
||||||
approved_at: null
|
approved_at: 2026-08-16T17:52:33.902562Z
|
||||||
rejected_at: null
|
rejected_at: null
|
||||||
rejection_reason: null
|
rejection_reason: null
|
||||||
approved_by: null
|
approved_by: null
|
||||||
@ -69,7 +68,10 @@ blockers: []
|
|||||||
phase_history:
|
phase_history:
|
||||||
- phase: draft
|
- phase: draft
|
||||||
entered: 2026-03-03T06:29:53.125784Z
|
entered: 2026-03-03T06:29:53.125784Z
|
||||||
|
exited: 2026-08-16T17:52:33.904168Z
|
||||||
|
- phase: specified
|
||||||
|
entered: 2026-08-16T17:52:33.904168Z
|
||||||
exited: null
|
exited: null
|
||||||
dependencies: []
|
dependencies: []
|
||||||
archived: false
|
archived: false
|
||||||
schema_version: 3
|
schema_version: 4
|
||||||
|
|||||||
94
.sdlc/features/p0-concierge-pilot-loop/spec.md
Normal file
94
.sdlc/features/p0-concierge-pilot-loop/spec.md
Normal file
@ -0,0 +1,94 @@
|
|||||||
|
# p0-concierge-pilot-loop: Concierge Pilot Loop
|
||||||
|
|
||||||
|
## Problem Statement
|
||||||
|
|
||||||
|
P0 must answer "do users care enough to return?" (`docs/planning/ROADMAP.md`, P0 milestone thesis). That answer requires a running daily briefing in front of a real pilot cohort for long enough to observe return behavior, plus a record trustworthy enough to defend a GO/NO-GO decision for P1.
|
||||||
|
|
||||||
|
The pilot is a concierge operation, not a product. `docs/personal-briefing-beachhead.md` §10 Phase A specifies "daily brief with strong manual QA on source quality and reasons," and the ROADMAP P0 acceptance criterion states the daily briefing prototype "can include manual source QA." Human intervention is therefore permitted — but unrecorded intervention silently inflates quality results and destroys the readout. Without a defined loop, a fixed briefing contract, verified instrumentation, and an intervention ledger, this pilot produces anecdotes rather than evidence.
|
||||||
|
|
||||||
|
## Goals
|
||||||
|
|
||||||
|
1. **Run a 2-week daily briefing** for the pilot cohort (20-50 users, per ROADMAP P0) at a fixed daily cutoff, every day, with no skipped days.
|
||||||
|
2. **Gate every briefing on manual source QA** performed by a named operator before delivery, per ROADMAP P0 ("can include manual source QA").
|
||||||
|
3. **Ship a briefing artifact** each user can act on: ranked items, reason labels, source links, and the five feedback action controls (`more`/`less`/`hide`/`mute`/`save`).
|
||||||
|
4. **Instrument each session** so median feedback actions per session and D2 retention are computable from recorded events, not reconstructed after the fact.
|
||||||
|
5. **Interview every cohort member weekly** against a fixed question set drawn from beachhead §6.2.
|
||||||
|
6. **Record every operator intervention** so quality results can be re-read with concierge assistance discounted.
|
||||||
|
7. **Hand `p0-validation-readout` a complete dataset** covering all 14 pilot days.
|
||||||
|
|
||||||
|
## Non-Goals
|
||||||
|
|
||||||
|
- Recruiting, screening, or onboarding the pilot cohort — owned by `p0-target-segment-recruitment`.
|
||||||
|
- Computing the GO/NO-GO decision or comparing results to thresholds — owned by `p0-validation-readout`.
|
||||||
|
- Building new engine capability. The pilot consumes existing surfaces (signal writes, reason labels, `/diagnostics`); anything missing is served manually by the operator and logged as intervention.
|
||||||
|
- Self-serve onboarding, time-budget mode, and cohort view — P1/P2 scope.
|
||||||
|
- Automating source ingestion or QA. Manual is explicitly in scope for P0.
|
||||||
|
|
||||||
|
## Functional Requirements
|
||||||
|
|
||||||
|
### FR-1: Daily Loop and Fixed Cutoff
|
||||||
|
|
||||||
|
Each pilot day runs: source ingestion -> candidate ranking -> manual source QA -> delivery -> session capture. Delivery occurs at one fixed local-morning cutoff time, identical every day for all users (cutoff clock time: TBD (owner: product)). Beachhead §5.2 defines the morning brief as the daily loop; midday/evening updates are out of scope. A briefing not delivered by cutoff is a missed day (FR-8), never a late day.
|
||||||
|
|
||||||
|
### FR-2: Manual Source QA Gate
|
||||||
|
|
||||||
|
Before delivery, a named operator reviews the ranked candidate set and (a) removes items failing source quality, (b) verifies each reason label is specific and true, (c) verifies no single source dominates the top 10 (beachhead §6.3 "one source dominates repeatedly"). No briefing is delivered without a recorded QA pass. QA is authorized by ROADMAP P0; each QA edit is an intervention under FR-7.
|
||||||
|
|
||||||
|
### FR-3: Briefing Artifact Contract
|
||||||
|
|
||||||
|
Every daily briefing carries 10-20 ranked items (beachhead §5.1.3). Every card carries: rank position, item title, a reason label drawn from the fixed label set (e.g. "Trending in your cohort", "Matches your <topic> priority", "New source for exploration" — beachhead §5.1.4), a one-tap source link, and controls for all five feedback actions. `more`/`less` apply to topic affinity, `hide` to the item's topic, `mute` to the source, `save` to the user's library. Each control emits exactly one feedback action event bound to `(user, briefing_day, rank_position, item, action)`.
|
||||||
|
|
||||||
|
### FR-4: Feedback Action Capture
|
||||||
|
|
||||||
|
Each feedback action is written as a typed signal with the action name as its signal type, so the pilot reads counts directly off the existing pg1 surface `tidaldb_signal_writes_by_type{signal_type="more"|"less"|"hide"|"mute"|"save"}` (`pg1-instrumented-metrics` FR-2). No parallel counting path is built. Per-user recency comes from the existing user signal timestamp map (pg1 FR-6); loop closure evidence comes from the existing feedback-loop latency histogram (pg1 FR-4).
|
||||||
|
|
||||||
|
### FR-5: Per-Session Instrumentation and D2 Retention
|
||||||
|
|
||||||
|
**Session boundary (normative — `p0-validation-readout` FR-2 defers to this definition).** A session opens when a user first opens that day's briefing, and closes at the earlier of an explicit close/navigation-away or an inactivity timeout of TBD (owner: product), frozen before day 1 and never changed mid-run. A re-open after close within the same `briefing_day` is a new session, so a user may have multiple sessions per briefing day (beachhead §9.2 "average sessions per active day"). Every session carries `actor=participant` or `actor=operator`; operator and QA sessions are tagged at open and excluded from every cohort metric.
|
||||||
|
|
||||||
|
Per session the pilot records: `user_id`, `briefing_day`, `actor`, session open timestamp, session close timestamp, items opened, and the ordered list of feedback actions. Sessions are a derived view, not the primitive: every interaction (open, card open, each feedback action, close) is stored with its own raw timestamp, so the whole run can be re-sessionized under an alternative inactivity timeout without re-collecting data (`p0-validation-readout` FR-4 sensitivity re-run). Two derived measures are produced daily and are the only measures the pilot itself computes:
|
||||||
|
|
||||||
|
- **Feedback actions per session** — median across the pilot cohort, against the ROADMAP P0 bar of >= 1 per session for the median user.
|
||||||
|
- **D2 retention** — fraction of users with >= 1 participant session on briefing day N who also have >= 1 participant session on day N+1, computed per day-pair. Counted on distinct `briefing_day` values, so D2 retention is independent of the inactivity timeout; only the per-session median above is timeout-sensitive.
|
||||||
|
|
||||||
|
A `/diagnostics` JSON snapshot (pg1 FR-5) is captured immediately after each daily cutoff and stored with the day's session records, giving an independent counter read per pilot day.
|
||||||
|
|
||||||
|
### FR-6: Interview Cadence and Fixed Question Set
|
||||||
|
|
||||||
|
Each cohort member is interviewed twice: at the end of week 1 and at the end of week 2 (beachhead §10 Phase A, "after each week"). Interviews use one fixed question set derived from beachhead §6.2, asked verbatim and in order: why not just current feeds; was setup too much; can this be trusted; does it feel repetitive or narrow; when you said `less`, did anything actually change; does it respect your time; would you use it for a work decision; are you comfortable with the data it holds. Responses are recorded verbatim; no interpretation happens during the interview.
|
||||||
|
|
||||||
|
### FR-7: Operator Runbook and Intervention Ledger
|
||||||
|
|
||||||
|
The operator runbook covers, per day: ingestion start, QA pass, delivery confirmation, and end-of-day session-record reconciliation. Every deviation from the automated output is an intervention with a ledger row: timestamp, pilot day, affected users, intervention class (`source_removed`, `reason_label_rewritten`, `item_reordered`, `diversity_forced`, `briefing_hand_assembled`, `delivery_manual`), and free-text cause. When QA finds a bad source, the source is muted for the remainder of the pilot, the affected cards are removed, and the removal is logged — so that day's quality numbers can be recomputed with operator-touched cards excluded.
|
||||||
|
|
||||||
|
### FR-8: Abort Conditions
|
||||||
|
|
||||||
|
The pilot stops early and escalates to `p0-validation-readout` if a beachhead §6.3 Critical failure mode is confirmed cohort-wide: (a) feedback actions demonstrably not reflected in the next refresh, or (b) the feed remains noisy after 2 days with no measurable improvement. It also stops if instrumentation loss makes FR-5 uncomputable for more than 2 pilot days, or if >= 3 of 14 days are missed days. An aborted pilot still hands over its partial dataset, labeled aborted with the triggering condition.
|
||||||
|
|
||||||
|
### FR-9: Readout Dataset Handoff
|
||||||
|
|
||||||
|
At day 14 the dataset is frozen — no rows added, corrected, or recoded after handoff — and `p0-validation-readout` receives: the roster and cohort segment tags (from `p0-target-segment-recruitment`); all 14 daily briefing manifests; the raw timestamped interaction stream; all session records derived from it under the FR-5 session boundary, carrying `actor` tags; all feedback action events with per-session counts by type (`more`/`less`/`hide`/`mute`/`save`); the per-day-pair D2 retention counters; the daily `/diagnostics` snapshots; the intervention ledger; the missed-day log; and all interview transcripts with their coding. The frozen inactivity timeout is handed over with the data so the readout can re-sessionize under an alternative value.
|
||||||
|
|
||||||
|
## Non-Functional Requirements
|
||||||
|
|
||||||
|
- **NFR-1**: Instrumentation is verified end-to-end before day 1; no event schema changes during the 2-week run.
|
||||||
|
- **NFR-2**: Session records and the intervention ledger are append-only; corrections are new rows, never edits.
|
||||||
|
- **NFR-3**: Interview transcripts are stored under participant IDs only — no participant names or contact details in pilot artifacts.
|
||||||
|
- **NFR-4**: Operator time per day is bounded and logged, so the concierge cost of one briefing is a known input to the GO/NO-GO decision.
|
||||||
|
|
||||||
|
## Test Strategy
|
||||||
|
|
||||||
|
- **Pre-pilot dry run**: 2 consecutive days with the operator team as stand-in users. Exercise the full loop including QA gate, delivery, and all five feedback actions. Dry-run data is discarded and never merged into pilot results.
|
||||||
|
- **Instrumentation verification before day 1**: fire one of each feedback action and confirm each increments its `tidaldb_signal_writes_by_type` counter and appears in the `/diagnostics` snapshot; confirm a session open and close emit the FR-5 boundary events with the correct `actor` tag, that an inactivity-timeout close and a re-open produce two distinct sessions on one `briefing_day`, and that a synthetic two-day session pattern yields the expected D2 retention value. Recruiting does not release the pilot cohort until this passes.
|
||||||
|
- **Interview coding**: two coders independently tag each transcript against the §6.2 question set and the §6.3 failure modes; disagreements are resolved by re-reading the verbatim answer, not by discussion. Coding scheme is fixed before the first interview.
|
||||||
|
- **Intervention accounting**: each day, every operator-touched card is reconcilable to a ledger row. Every reported quality figure is produced twice — all cards, and operator-touched cards excluded. A divergence between the two is a finding, not an error to smooth over.
|
||||||
|
- **Missed-day handling**: a day with no delivered briefing is recorded as a missed day with cause, and enters retention math as a no-briefing day for every user. Missed days are never dropped, backfilled, or interpolated.
|
||||||
|
- Pre-registered before day 1: cohort size 20-50 (ROADMAP), duration 14 days, median >= 1 feedback action per session (ROADMAP), the FR-5 inactivity timeout, and D2 retention threshold TBD (owner: product) — the ROADMAP states "agreed threshold" without a value.
|
||||||
|
|
||||||
|
## Dependencies
|
||||||
|
|
||||||
|
- `p0-target-segment-recruitment` (upstream) — supplies the enrolled pilot cohort, segment tags, and interest configuration; the pilot cannot start day 1 without it.
|
||||||
|
- `p0-validation-readout` (downstream) — consumes the FR-9 dataset and issues the GO/NO-GO decision for P1.
|
||||||
|
- `pg1-instrumented-metrics` — `tidaldb_signal_writes_by_type` counters (FR-2), feedback-loop latency histogram (FR-4), `/diagnostics` JSON endpoint (FR-5), user signal timestamp map (FR-6).
|
||||||
|
- `docs/personal-briefing-beachhead.md` §5.2 daily loop, §6.2 core user questions, §6.3 failure modes, §10 Phase A.
|
||||||
|
- `docs/planning/ROADMAP.md` P0 acceptance criteria.
|
||||||
@ -1,16 +1,15 @@
|
|||||||
id: p0-target-segment-recruitment
|
|
||||||
slug: p0-target-segment-recruitment
|
slug: p0-target-segment-recruitment
|
||||||
title: Target Segment & Recruitment
|
title: Target Segment & Recruitment
|
||||||
description: Define persona, write recruitment script, build candidate pool of 20-50 target users for concierge pilot
|
description: Define persona, write recruitment script, build candidate pool of 20-50 target users for concierge pilot
|
||||||
phase: draft
|
phase: specified
|
||||||
created_at: 2026-03-03T06:29:53.119870Z
|
created_at: 2026-03-03T06:29:53.119870Z
|
||||||
updated_at: 2026-03-03T06:29:53.119870Z
|
updated_at: 2026-08-16T17:52:33.879134Z
|
||||||
artifacts:
|
artifacts:
|
||||||
- artifact_type: spec
|
- artifact_type: spec
|
||||||
status: missing
|
status: approved
|
||||||
path: .sdlc/features/p0-target-segment-recruitment/spec.md
|
path: .sdlc/features/p0-target-segment-recruitment/spec.md
|
||||||
created_at: null
|
created_at: null
|
||||||
approved_at: null
|
approved_at: 2026-08-16T17:52:33.875621Z
|
||||||
rejected_at: null
|
rejected_at: null
|
||||||
rejection_reason: null
|
rejection_reason: null
|
||||||
approved_by: null
|
approved_by: null
|
||||||
@ -69,7 +68,10 @@ blockers: []
|
|||||||
phase_history:
|
phase_history:
|
||||||
- phase: draft
|
- phase: draft
|
||||||
entered: 2026-03-03T06:29:53.119870Z
|
entered: 2026-03-03T06:29:53.119870Z
|
||||||
|
exited: 2026-08-16T17:52:33.879134Z
|
||||||
|
- phase: specified
|
||||||
|
entered: 2026-08-16T17:52:33.879134Z
|
||||||
exited: null
|
exited: null
|
||||||
dependencies: []
|
dependencies: []
|
||||||
archived: false
|
archived: false
|
||||||
schema_version: 3
|
schema_version: 4
|
||||||
|
|||||||
110
.sdlc/features/p0-target-segment-recruitment/spec.md
Normal file
110
.sdlc/features/p0-target-segment-recruitment/spec.md
Normal file
@ -0,0 +1,110 @@
|
|||||||
|
# p0-target-segment-recruitment: Target Segment & Recruitment
|
||||||
|
|
||||||
|
## Problem Statement
|
||||||
|
|
||||||
|
P0 ("Beachhead Validation") must answer whether a personal briefing feed drives repeat use. That question is only answerable if the people in the study are the people the product is for. There is currently no screening bar, no recruitment funnel, and no cohort record -- so the pilot could be staffed with contacts, developers, or passive scrollers and produce a retention number that means nothing.
|
||||||
|
|
||||||
|
Open questions this feature closes:
|
||||||
|
|
||||||
|
1. Which of the two beachhead personas (`docs/personal-briefing-beachhead.md` §2) are in scope, and what disqualifies a candidate?
|
||||||
|
2. How many candidates must enter the funnel to land a 20-50 person **pilot cohort** (ROADMAP P0 AC-1)?
|
||||||
|
3. What may the outreach promise, given the prototype is concierge-operated with manual source QA (§10 Phase A)?
|
||||||
|
4. What consent is required before behavioral instrumentation and recorded interviews?
|
||||||
|
5. What is stored per participant, where, and under what identifier?
|
||||||
|
|
||||||
|
## Goals
|
||||||
|
|
||||||
|
1. **Screening bar** -- explicit qualify/disqualify criteria derived from persona jobs-to-be-done (§3) and the adoption-killing failure modes (§6.3).
|
||||||
|
2. **Sized funnel** -- named stages with a stop condition that lands the cohort inside the 20-50 range and survives pilot attrition.
|
||||||
|
3. **Honest outreach artifact** -- states the real commitment, promises nothing the prototype cannot deliver, and does not prime the interview outcome.
|
||||||
|
4. **Consent and data handling** -- one consent record per participant covering instrumentation, recording, retention, and deletion.
|
||||||
|
5. **De-identified cohort record** -- git-tracked roster keyed by an opaque `participant_id`, with no personal identity anywhere in the repo.
|
||||||
|
6. **Segment balance** -- a cohort that is not a single persona or a single role bucket.
|
||||||
|
|
||||||
|
## Non-Goals
|
||||||
|
|
||||||
|
- Running the **daily briefing** or the pilot itself -- `p0-concierge-pilot-loop`.
|
||||||
|
- Interview coding, retention analysis, and the **GO/NO-GO decision** -- `p0-validation-readout`.
|
||||||
|
- Any engine, app, or instrumentation implementation (consumed, not built, here); Phase B scale recruitment (200-500 users, §10 Phase B) and paid acquisition channels.
|
||||||
|
|
||||||
|
## Functional Requirements
|
||||||
|
|
||||||
|
### FR-1: Persona Scope and Screening Criteria
|
||||||
|
|
||||||
|
In scope: the primary persona "information-overloaded decision maker" (§2), narrowed per §10 Phase A to strategy / product / analyst / operations / media / investing / policy roles; and the secondary persona "curious consumer with intent". Both classes are in scope because ROADMAP P0 AC-1 names knowledge workers *and* high-intent consumers.
|
||||||
|
|
||||||
|
Qualify -- all must hold, each recorded as a roster field:
|
||||||
|
- **Q1** Consumes content to make decisions rather than for entertainment (primary), or follows multiple topic domains with stated intent (secondary) -- §2.
|
||||||
|
- **Q2** Uses at least one incumbent baseline daily: feeds (X, LinkedIn, YouTube, Reddit, news apps), newsletters, podcasts, or a general AI assistant. Without an incumbent, the §6.2 question "why not just use my current feeds?" has no comparison.
|
||||||
|
- **Q3** Commits to a daily session across the 2-week pilot and to at least one **feedback action** (`more`/`less`/`hide`/`mute`/`save`) per session -- ROADMAP P0 AC-3.
|
||||||
|
- **Q4** Willing to complete onboarding inputs: 5-10 interests, depth, hard excludes, time budget -- §5.1.
|
||||||
|
- **Q5** Willing to be interviewed at the end of each pilot week, recorded -- §10 Phase A.3.
|
||||||
|
- **Q6** Accepts daily briefing delivery cadence -- guards §6.3 "too many notifications".
|
||||||
|
|
||||||
|
Disqualify -- any one excludes:
|
||||||
|
- **D1** Works on tidalDB, or is a personal/professional contact of the pilot operator. Return driven by social obligation is not a retention signal.
|
||||||
|
- **D2** Evaluating as a developer or buyer rather than as a reader -- §8.3.1 makes a developer platform an explicit beachhead non-goal.
|
||||||
|
- **D3** Unavailable for use on two of the first three pilot days. **D2 retention** (AC-5) and the §6.3 failure mode "feed still noisy after 2 days" both require day-2 exposure.
|
||||||
|
- **D4** Declines instrumentation or interview-recording consent.
|
||||||
|
- **D5** Duplicate of an already-enrolled person reached via a second channel.
|
||||||
|
|
||||||
|
### FR-2: Recruitment Funnel Stages and Sizing
|
||||||
|
|
||||||
|
Stages, in order: `sourced` -> `contacted` -> `responded` -> `screened` -> `qualified` -> `consented` -> `enrolled`. Each transition is recorded with a UTC date on the participant record; each exit records the stage and a reason code (`D1`-`D5`, `no_response`, `declined`, `unreachable`).
|
||||||
|
|
||||||
|
Per-stage conversion targets are `TBD (owner: product)` -- neither the beachhead doc nor the ROADMAP supplies a prior-art rate, and an invented rate would misplan `sourced` volume. The funnel is instead sized backwards from the ROADMAP range: enroll **50** (top of the 20-50 range) so the cohort can lose up to 30 participants over the 2-week pilot and still hold the **>= 20** floor. `sourced` volume is re-planned once the measured `contacted -> enrolled` rate exists.
|
||||||
|
|
||||||
|
Stop condition: `enrolled == 50`, or the recruitment window closes with `enrolled >= 20` and FR-6 balance satisfied. Below 20, recruitment continues and the pilot does not start.
|
||||||
|
|
||||||
|
### FR-3: Outreach Artifact Requirements
|
||||||
|
|
||||||
|
One versioned script per channel. Must state: this is a concierge-operated prototype with manual source QA (§10 Phase A.2); 2-week duration; a daily session scoped to the chosen `5/10/20` minute time budget (§5.1); that behavioral instrumentation is active; that two interviews will be recorded; compensation terms (`TBD (owner: product)`); withdrawal at any time; and the deletion right.
|
||||||
|
|
||||||
|
Must NOT promise: accuracy, coverage or completeness, source breadth, freshness guarantees, or replacement of any named incumbent feed. Must not cite "first useful briefing in under 3 minutes" -- that is a P2 self-serve onboarding target (§6.2, ROADMAP P2), not a P0 concierge target. Must not describe surfaces absent from the prototype (cohort view, time-budget mode) as present.
|
||||||
|
|
||||||
|
Must NOT prime the outcome: the script may not use the phrases "less noise", "more useful", or "saves time". Those are exactly the claims `p0-validation-readout` tests in interviews (ROADMAP P0 AC-4); seeding them in outreach invalidates the finding.
|
||||||
|
|
||||||
|
### FR-4: Consent and Data Handling
|
||||||
|
|
||||||
|
A consent record is obtained before `enrolled`, covering: (a) behavioral instrumentation -- session opens, item impressions, opens, every **feedback action**, and **D2 retention** counters; (b) interview audio recording and transcription; (c) retention period and deletion on request (§6.2 "Is my data private?"). Withdrawal removes the participant from subsequent analysis and deletes recordings; already-aggregated de-identified counters are retained, and the consent text says so.
|
||||||
|
|
||||||
|
Signed forms, recordings, and contact details live in an out-of-repo operator store. Git-tracked artifacts hold only `consent: true`, the UTC date, and the consent script version.
|
||||||
|
|
||||||
|
### FR-5: Cohort Record and Participant Identity
|
||||||
|
|
||||||
|
A git-tracked, de-identified roster at `docs/planning/p0/cohort-roster.md`, one row per candidate: `participant_id` (opaque, stable, assigned at `screened`, e.g. `P0-017`), persona class (`primary`|`secondary`), role/domain bucket, incumbent baselines (Q2), declared interest domains, time budget, recruitment channel, per-stage transition dates, consent flag and date, exit stage and reason if not enrolled.
|
||||||
|
|
||||||
|
No name, email, employer, handle, or any other personal identity appears in any git-tracked artifact. The `participant_id` -> person mapping exists only in the operator store. Every downstream instrumentation event, interview transcript, and readout table keys on `participant_id`.
|
||||||
|
|
||||||
|
### FR-6: Segment Balance Requirements
|
||||||
|
|
||||||
|
Both persona classes are non-empty at enrollment (AC-1 names both). The primary persona is the majority, per §10 Phase A "target one segment"; the secondary-persona floor and the exact split are `TBD (owner: product)`, and must be fixed in the roster header **before** the funnel opens, never adjusted to match results. No single role/domain bucket may account for the entire cohort. Declared interest domains must span at least two of the §2 topic areas (career, health, finance, AI, hobbies) so the §6.3 "one source dominates" failure mode is observable across differing interests. Balance is checked at `enrolled == 20` and again at window close; on violation, recruitment continues in the deficient segment only.
|
||||||
|
|
||||||
|
### FR-7: Exit Criteria and Handoff
|
||||||
|
|
||||||
|
Hand off to `p0-concierge-pilot-loop` only when all hold: (1) `enrolled` is within 20-50 and FR-6 is satisfied; (2) every enrolled participant has a `participant_id` and a consent record; (3) every enrolled participant has submitted onboarding inputs (interests, excludes, time budget) so day-0 briefing generation is unblocked; (4) every enrolled participant has a recorded pre-pilot baseline-feed answer; (5) the instrumentation dry-run has passed. The handoff artifact is the frozen roster plus the version-pinned screening and consent scripts.
|
||||||
|
|
||||||
|
## Non-Functional Requirements
|
||||||
|
|
||||||
|
- **NFR-1**: No personal identity in any git-tracked artifact; the roster is safe for anyone with repo access to read.
|
||||||
|
- **NFR-2**: Screening criteria, balance targets, and the funnel stop condition are frozen before the first `contacted` event. Later changes are versioned with a written rationale, never applied silently mid-funnel.
|
||||||
|
- **NFR-3**: Recruitment channel is recorded per participant so channel bias is measurable in `p0-validation-readout`.
|
||||||
|
- **NFR-4**: The roster is plain markdown -- no tooling, database, or schema required to read or edit it.
|
||||||
|
|
||||||
|
## Test Strategy
|
||||||
|
|
||||||
|
Cohort quality is not automatable; each check below is a human review whose written result is appended to the roster.
|
||||||
|
|
||||||
|
- **Screening audit** -- a reviewer other than the recruiter re-scores every enrolled participant against Q1-Q6 and D1-D5 using roster fields alone. Full census, not a sample: at 20-50 participants sampling has no justification. Any enrollment not defensible from recorded fields is corrected or removed before the pilot starts.
|
||||||
|
- **Duplicate and ineligible detection** -- roster checked for duplicate `participant_id`, same person across channels (operator-store cross-check), and any row that satisfies a disqualifier.
|
||||||
|
- **Instrumentation dry-run** -- with at least one enrolled participant (or the operator using a reserved `participant_id`), complete one full **daily briefing** session and confirm the pipeline records session start, impressions, and each of `more`/`less`/`hide`/`mute`/`save`, and that the `pg1-instrumented-metrics` `/diagnostics` JSON endpoint reflects those writes. Recruitment does not close until this passes -- instrumentation is verified before the cohort is committed, because the §6.3 failure mode "feedback actions appear ignored" is indistinguishable from lost telemetry after the fact.
|
||||||
|
- **Pre-pilot baseline control** -- before day 0, each participant answers in their own words what they use today and how much time it costs. Recorded verbatim, keyed by `participant_id`. This is the control against which `p0-validation-readout` codes the AC-4 "less noise / more useful / saves time" claims; asked pre-pilot so it cannot be contaminated by exposure.
|
||||||
|
- **Pre-registration** -- screening criteria, segment split, funnel stop condition, and the **D2 retention** threshold (`TBD (owner: product)`; ROADMAP AC-5 states only "agreed threshold") are written down and frozen before outreach begins.
|
||||||
|
|
||||||
|
## Dependencies
|
||||||
|
|
||||||
|
- `docs/personal-briefing-beachhead.md` -- §2 personas, §3 jobs-to-be-done, §6.1-6.3 pressure test and failure modes, §10 Phase A.
|
||||||
|
- `docs/planning/ROADMAP.md` P0 acceptance criteria -- cohort size 20-50, feedback-action rate, **D2 retention** threshold.
|
||||||
|
- `pg1-instrumented-metrics` -- `/diagnostics` JSON endpoint and feedback-action counters exercised by the instrumentation dry-run.
|
||||||
|
- **Downstream: `p0-concierge-pilot-loop`** consumes the enrolled **pilot cohort**, the frozen roster, and the pre-pilot baseline answers; it cannot start until FR-7 is met.
|
||||||
|
- **Downstream: `p0-validation-readout`** consumes the baseline answers and the persona/channel fields as controls for the **GO/NO-GO decision**, reached via `p0-concierge-pilot-loop`.
|
||||||
@ -1,16 +1,15 @@
|
|||||||
id: p0-validation-readout
|
|
||||||
slug: p0-validation-readout
|
slug: p0-validation-readout
|
||||||
title: Validation Readout
|
title: Validation Readout
|
||||||
description: Analyze retention metrics and qualitative interviews; produce go/no-go decision for P1 Concierge Alpha build
|
description: Analyze retention metrics and qualitative interviews; produce go/no-go decision for P1 Concierge Alpha build
|
||||||
phase: draft
|
phase: specified
|
||||||
created_at: 2026-03-03T06:29:53.132476Z
|
created_at: 2026-03-03T06:29:53.132476Z
|
||||||
updated_at: 2026-03-03T06:29:53.132476Z
|
updated_at: 2026-08-16T17:52:33.932118Z
|
||||||
artifacts:
|
artifacts:
|
||||||
- artifact_type: spec
|
- artifact_type: spec
|
||||||
status: missing
|
status: approved
|
||||||
path: .sdlc/features/p0-validation-readout/spec.md
|
path: .sdlc/features/p0-validation-readout/spec.md
|
||||||
created_at: null
|
created_at: null
|
||||||
approved_at: null
|
approved_at: 2026-08-16T17:52:33.930395Z
|
||||||
rejected_at: null
|
rejected_at: null
|
||||||
rejection_reason: null
|
rejection_reason: null
|
||||||
approved_by: null
|
approved_by: null
|
||||||
@ -69,7 +68,10 @@ blockers: []
|
|||||||
phase_history:
|
phase_history:
|
||||||
- phase: draft
|
- phase: draft
|
||||||
entered: 2026-03-03T06:29:53.132476Z
|
entered: 2026-03-03T06:29:53.132476Z
|
||||||
|
exited: 2026-08-16T17:52:33.932118Z
|
||||||
|
- phase: specified
|
||||||
|
entered: 2026-08-16T17:52:33.932118Z
|
||||||
exited: null
|
exited: null
|
||||||
dependencies: []
|
dependencies: []
|
||||||
archived: false
|
archived: false
|
||||||
schema_version: 3
|
schema_version: 4
|
||||||
|
|||||||
108
.sdlc/features/p0-validation-readout/spec.md
Normal file
108
.sdlc/features/p0-validation-readout/spec.md
Normal file
@ -0,0 +1,108 @@
|
|||||||
|
# p0-validation-readout: Validation Readout
|
||||||
|
|
||||||
|
## Problem Statement
|
||||||
|
|
||||||
|
The P0 pilot produces behavioral events, feedback actions, retention counters, and interview transcripts, but nothing converts them into a decision. Without a pre-registered analysis, the P1 Concierge Alpha build decision degrades into advocacy: whoever ran the pilot narrates the numbers that flatter it, kill criteria (`docs/personal-briefing-beachhead.md` §12) are re-litigated after the fact, and ambiguous cases ("D2 was close") resolve as GO by default.
|
||||||
|
|
||||||
|
This feature defines the analysis, the thresholds, and the readout artifact **before** the pilot data is read, so that the GO/NO-GO decision is a computation over a results table rather than a judgement call. It must also answer the questions the numbers alone cannot: how many operator interventions were required to keep briefing quality acceptable, and what evidence would falsify the conclusion.
|
||||||
|
|
||||||
|
## Goals
|
||||||
|
|
||||||
|
1. **Deterministic verdict** -- One of GO / NO-GO / EXTEND derived mechanically from the pilot inputs.
|
||||||
|
2. **Pre-registration** -- Every threshold and computation rule fixed and committed before outcome data is read.
|
||||||
|
3. **Traceability** -- Each threshold cites a ROADMAP P0 acceptance criterion or a beachhead §9/§12 metric, or is explicitly `TBD (owner: product)`.
|
||||||
|
4. **Honest cost accounting** -- The manual QA and operator intervention required to sustain quality is reported as a first-class result, not a footnote.
|
||||||
|
5. **Closed loop into the roadmap** -- A NO-GO or EXTEND verdict lands as a roadmap change, not a dropped thread.
|
||||||
|
|
||||||
|
## Non-Goals
|
||||||
|
|
||||||
|
- Recruiting the pilot cohort (`p0-target-segment-recruitment`).
|
||||||
|
- Operating the daily briefing or conducting the interviews (`p0-concierge-pilot-loop`).
|
||||||
|
- Building any P1 surface, ranking change, or engine feature.
|
||||||
|
- Deciding P2/P3 scope, pricing, or go-to-market.
|
||||||
|
|
||||||
|
## Functional Requirements
|
||||||
|
|
||||||
|
### FR-1: Frozen Input Snapshot
|
||||||
|
|
||||||
|
Consume from `p0-concierge-pilot-loop` FR-9, as a single immutable snapshot frozen at day 14 -- no rows added, corrected, or recoded after handoff: (a) the raw timestamped interaction stream (every open, card open, `feedback action`, and close, each with `briefing_day` and `actor`) as the primitive, plus the derived session records and the frozen inactivity timeout value, (b) per-session `feedback action` counts by type (`more`/`less`/`hide`/`mute`/`save`), (c) `D2 retention` counters from the pg1 metrics pipeline (daily `/diagnostics` JSON snapshots plus the interaction stream), (d) the operator-intervention log and the missed-briefing-day log, (e) interview transcripts. Record a content hash per input file in the readout. Analysis reads only the snapshot; later pilot data does not amend a published verdict.
|
||||||
|
|
||||||
|
### FR-2: Session and Cohort Definitions
|
||||||
|
|
||||||
|
A *session* is the session boundary emitted by `p0-concierge-pilot-loop` FR-5 (opens on first open of that day's briefing; closes on explicit close/navigation-away or the inactivity timeout that feature freezes before day 1); this analysis defines no timeout of its own. A re-open within the same `briefing_day` is a distinct session (beachhead §9.2 sessions-per-active-day). Sessions carrying `actor=operator` are excluded from every cohort metric. The *pilot cohort* denominator is participants who completed activation: first `daily briefing` delivered **and** opened. Enrolled-but-never-activated participants are reported separately and excluded from retention and feedback metrics.
|
||||||
|
|
||||||
|
### FR-3: D2 Retention Computation
|
||||||
|
|
||||||
|
Each participant has a personal Day-0 = the `briefing_day` of their first opened briefing. `D2 retention` = fraction of activated participants with >= 1 `actor=participant` session on `briefing_day` Day-0 + 1. Because the metric counts distinct `briefing_day` values rather than session objects, it is invariant to the FR-2 inactivity timeout; only FR-4 is timeout-sensitive. Rules, applied in order:
|
||||||
|
|
||||||
|
1. **Missed briefing.** If no briefing was delivered on a participant's Day-0 + 1 (missed-briefing-day log), that day is void: Day-0 re-anchors to their next delivered-and-opened briefing. Count and report re-anchored participants.
|
||||||
|
2. **Reminders.** Product-cadence push/email (beachhead §5.2) counts as a normal return. A session that follows an operator or support nudge within the same day does **not** count, per beachhead §6.4.5 ("returns on Day 2 without a reminder from support or onboarding prompts").
|
||||||
|
3. **Dropouts.** Intent-to-treat: a participant who stops using the product stays in the denominator as a non-return. Withdrawals are excluded only with a named reason in the operator-intervention log, and the metric is reported both with and without exclusions (see FR-6).
|
||||||
|
4. **Partial observation.** A participant whose Day-0 + 1 falls after pilot end is excluded from D2 and reported as unobserved.
|
||||||
|
|
||||||
|
### FR-4: Feedback-Action Rate Computation
|
||||||
|
|
||||||
|
For each activated participant, compute their median `feedback action` count per session across all their observed sessions (zero-action sessions included). The cohort statistic is the median of those per-participant medians. The ROADMAP criterion "at least one meaningful feedback action per session for the median user" is met iff that cohort statistic is >= 1. Report the distribution, not only the median, and report per-action-type counts so a cohort that only ever taps `save` is visible.
|
||||||
|
|
||||||
|
Because this statistic counts sessions, it is sensitive to the inactivity timeout frozen by `p0-concierge-pilot-loop` FR-5. Record that timeout value in the pre-registration block, then re-derive sessions from the FR-1 raw interaction stream under a stated alternative timeout and recompute this statistic as part of the FR-6 sensitivity check: if G2 flips, the verdict is EXTEND, never GO. Re-sessionization uses the frozen snapshot only -- it never triggers re-collection.
|
||||||
|
|
||||||
|
### FR-5: Interview Coding
|
||||||
|
|
||||||
|
Code every transcript against the beachhead §6.2 required answers. For each of the eight rows, mark `proven` / `not proven` / `not raised`, with a verbatim quote required for `proven`. Separately, code the three ROADMAP value axes against the participant's own named baseline feed (§6.1: existing feeds, newsletters, AI assistants) as a forced choice per axis -- `better` / `same` / `worse`:
|
||||||
|
|
||||||
|
- **less noise**, **more useful**, **saves time**.
|
||||||
|
|
||||||
|
A participant counts as *value-confirmed* only when all three axes are `better`. Also code the §12.4 kill frame: does the participant describe the product as "another noisy feed" or equivalent. Every transcript is double-coded; unresolved disagreement resolves to the conservative code (`same`, and kill-frame present).
|
||||||
|
|
||||||
|
### FR-6: Pre-Registered Decision Rule
|
||||||
|
|
||||||
|
Thresholds are fixed in the readout's pre-registration block before FR-1 data is read. Named parameters:
|
||||||
|
|
||||||
|
| Param | Meaning | Value | Source |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `N_MIN` | activated participants required | 20 | ROADMAP P0 ("Recruit 20-50 target users") |
|
||||||
|
| `T_FEEDBACK` | cohort median actions/session | 1 | ROADMAP P0 |
|
||||||
|
| `T_D2` | `D2 retention` floor | `TBD (owner: product)` | ROADMAP P0 ("agreed threshold"); beachhead §9.2 names no number |
|
||||||
|
| `T_VALUE` | fraction of interviewed participants value-confirmed | `TBD (owner: product)` | ROADMAP P0; beachhead §9.3 names no number |
|
||||||
|
| `T_NOISE` | max fraction coding the §12.4 kill frame | `TBD (owner: product)` | beachhead §12.4 |
|
||||||
|
|
||||||
|
Gates: **G1** activated >= `N_MIN`. **G2** FR-4 statistic >= `T_FEEDBACK`. **G3** value-confirmed fraction >= `T_VALUE`. **G4** `D2 retention` >= `T_D2`. **G5** kill-frame fraction < `T_NOISE`. Verdict, evaluated top-down, first match wins:
|
||||||
|
|
||||||
|
1. **NO-GO** -- G2 or G5 fails (kill-class: beachhead §12.2 feedback ignored, §12.4 another noisy feed).
|
||||||
|
2. **NO-GO** -- this is the second iteration cycle (beachhead §12) and any gate fails.
|
||||||
|
3. **EXTEND** -- first cycle and only G1, G3, or G4 fails. One additional cycle maximum.
|
||||||
|
4. **EXTEND** -- all gates pass but any gate flips under the sensitivity re-run: dropout exclusions applied vs not applied (FR-3 rule 3), and the alternative inactivity timeout (FR-4).
|
||||||
|
5. **GO** -- all gates pass and the sensitivity re-run agrees.
|
||||||
|
|
||||||
|
Verdicts are computed with G1 evaluated first for reporting: when G1 fails, G2-G5 are still computed and published, labelled provisional.
|
||||||
|
|
||||||
|
### FR-7: Readout Artifact
|
||||||
|
|
||||||
|
Publish `docs/planning/p0-validation-readout.md` containing: pre-registration block (parameters, commit hash, timestamp preceding first outcome read); input snapshot hashes; evidence table with one row per gate (`gate`, `definition`, `threshold`, `observed`, `PASS/FAIL`, `source doc`); the verdict and the rule clause that produced it; participant accounting (enrolled, activated, re-anchored, withdrawn, unobserved); the intervention ledger -- every manual source-QA and operator action required to keep briefing quality acceptable, with a per-participant-day rate; the §6.2 coding matrix with quotes; and a **Falsification** section stating explicitly what observation would overturn the conclusion (e.g. for GO: which gate is nearest its threshold and what cohort composition change would flip it).
|
||||||
|
|
||||||
|
### FR-8: Roadmap Feedback
|
||||||
|
|
||||||
|
The verdict must land in the roadmap in the same change as the readout. GO: P0 section marked complete with a link to the readout; P1 unblocked. EXTEND: the P0 ROADMAP section gains the failing gates and the scope of the one permitted additional cycle; P1 stays blocked. NO-GO: the P0 section records the kill-class failure, links the readout, and states the pivot question for the next ponder; P1 Concierge Alpha is not entered. In every case `docs/planning/PRODUCT_ROADMAP.md` and the ROADMAP P0 block cite the readout, and P1 does not enter preparation while the verdict is EXTEND or NO-GO.
|
||||||
|
|
||||||
|
## Non-Functional Requirements
|
||||||
|
|
||||||
|
- **NFR-1**: The analysis is a committed script over the frozen snapshot -- no spreadsheet-only steps, no hand-entered aggregates.
|
||||||
|
- **NFR-2**: Interview transcripts are stored de-identified; the readout carries participant codes, never names, employers, or contact details.
|
||||||
|
- **NFR-3**: The readout states its own limits: cohort of 20-50 self-selected participants over 2 weeks supports a directional GO/NO-GO, not a population estimate.
|
||||||
|
- **NFR-4**: All five gates are reported even when an earlier gate already forces the verdict.
|
||||||
|
|
||||||
|
## Test Strategy
|
||||||
|
|
||||||
|
- **Reproducibility**: run the analysis script twice against the same input snapshot hashes; the gate table and verdict must be byte-identical. A third-party re-run from the snapshot plus the script must reach the same verdict without consulting the analyst.
|
||||||
|
- **Instrumentation trust**: reconcile the `D2 retention` counters against the raw event stream independently; a discrepancy above 0 participants blocks the readout until the source of the divergence is identified (this check is inherited from `p0-concierge-pilot-loop`'s pre-recruitment instrumentation verification).
|
||||||
|
- **Pre-registration check**: verify the commit containing the parameter table predates the first commit or access that reads outcome data. If it does not, the verdict is downgraded to EXTEND and the cycle re-run with parameters fixed.
|
||||||
|
- **Coding soundness**: every transcript double-coded on the §6.2 rows and the three value axes; report raw agreement per axis; conservative tie-break applied and counted.
|
||||||
|
- **Fixture verdicts**: exercise the decision rule against synthetic results tables covering each clause of FR-6 (kill-class fail, second-cycle fail, first-cycle G4-only fail, sensitivity flip, all-pass) and confirm the expected verdict.
|
||||||
|
- **Bias review**: a named reviewer who did not operate the pilot argues the NO-GO case in writing against the assembled evidence before the verdict is published; their objections and the responses are appended to the readout. Recommended split: `@tidal-researcher` argues NO-GO, `@tidal-visionary` owns the verdict, the pilot operator does not vote.
|
||||||
|
|
||||||
|
## Dependencies
|
||||||
|
|
||||||
|
- `p0-target-segment-recruitment` -- defines the `pilot cohort` and its segment criteria; supplies enrollment records and the activation denominator.
|
||||||
|
- `p0-concierge-pilot-loop` -- supplies every FR-1 input: behavioral events, `feedback action` counts, `D2 retention` counters, operator-intervention log, interview transcripts.
|
||||||
|
- `pg1-instrumented-metrics` -- `/diagnostics` JSON endpoint and retention/feedback counters the behavioral analysis reads.
|
||||||
|
- `docs/planning/ROADMAP.md` P0 acceptance criteria and `docs/personal-briefing-beachhead.md` §6.2, §9, §12 -- the only permitted sources of thresholds.
|
||||||
Loading…
Reference in New Issue
Block a user