The decision, made inspectable.
v1 scored the territory with one weighted number and listed five "riskiest assumptions" with no numbers attached. This page replaces both with a model you can argue with: every belief has a prior and a reason, every piece of evidence has a likelihood ratio and a source, every posterior is computed in front of you — and the survey, pilot and legal check are pre-registered as toggles, so you can see what each test would change before it is run.
Source: research/DECISION_MODEL.md (method, calibration rubric, register) · state as of 17 Aug 2026: literature + qualitative evidence + the survey as analysed (default scenario); pilot and legal check pending · every slider is live; nothing is saved unless you copy the JSON at the bottom.
Three layers: robustness, beliefs, gates
Was "anxiety" a robust choice?
Same six criteria and weights as the Define matrix, but each cell is a low / mode / high range and the weights are jittered. 5,000 Monte-Carlo draws report how often each option ranks first — and the matrix is re-scored after the interview evidence.
What do we believe, and why?
Eight assumptions (A1 split by arousal state, A2 split into self-report vs HRV, plus retention). Each has a prior with a rationale, evidence rows with likelihood ratios, and pending tests with pre-registered LRs.
What would we do?
Proceed / pivot / kill are posterior thresholds fixed in advance. Concept viability is the joint probability of a concept's critical assumptions (independence assumed, weakest link named). The verdict updates as you move anything above.
Calibration. Raw strength (decisive 8–10 · strong 3–5 · moderate 1.8–3 · weak 1.2–1.8; reciprocals when disconfirming) is shrunk toward 1 for evidence quality (GRADE-style: meta-analysis full weight; industry-funded or single-source ×0.5 in log space; qualitative n≤5 ×0.35–0.5) and again when items are correlated (same recruit channel, overlapping primary studies). Priors come from base rates where any exist, otherwise 0.50. Method precedent: Fairfield & Charman (2017) and Humphreys & Jacobs (2015) on explicit Bayesian analysis with qualitative evidence.
Anxiety wins under uncertainty — and after the interviews
Move the weights, widen the jitter, switch between the v1 cells (as published in Define) and the v2 cells (anxiety's asset fit 4 → 3 and wearable synergy 5 → 4 after the interview finding and the HRV-API facts). The choice is robust; what changed is which anxiety product, not whether anxiety.
Which weight could flip the ranking?
If no bar reaches 0, no single ±40% weight change makes anxiety lose. The vertical mark is the base gap.
| Criterion (w) | Dementia | ADHD | Anxiety v1 | Anxiety v2 |
|---|---|---|---|---|
| Clinical evidence (.25) | 2 / 3 / 4 | 1 / 2 / 3 | 4 / 5 / 5 | 4 / 5 / 5 |
| Reachable market (.20) | 1 / 1 / 2 | 2 / 3 / 4 | 4 / 5 / 5 | 4 / 5 / 5 |
| Willingness to pay (.20) | 1 / 1 / 2 | 2 / 3 / 4 | 3 / 4 / 5 | 3 / 4 / 5 |
| Asset fit (.15) | 1 / 1 / 2 | 2 / 3 / 4 | 3 / 4 / 5 | 2 / 3 / 4 |
| Regulatory ease (.10) | 2 / 3 / 4 | 1 / 2 / 3 | 3 / 4 / 5 | 3 / 4 / 5 |
| Wearable synergy (.10) | 1 / 2 / 3 | 2 / 3 / 4 | 4 / 5 / 5 | 3 / 4 / 5 |
| Mode total | 1.80 | 2.65 | 4.55 | 4.30 |
Sleep, chronic pain and depression entered the funnel as extras but were not scored in full — stated so the funnel width is honest.
Eight beliefs, each with a prior, evidence and a posterior
The survey rows are pre-set to the analysed outcomes (LRs per the pre-registered plan); pilot, landing and legal rows are pending. Toggle an evidence row off to see how much it carried. Drag a likelihood ratio if you disagree with the calibration. Set a pending test to positive or negative to see the value of running it. Grey pin = prior, dark pin = posterior; red tick = kill/pivot line, teal tick = proceed line.
Which concept the beliefs currently favour
Joint probability that a concept's critical assumptions all hold (independence assumed — an over-simplification stated as such). The weakest link is what the next test should target. Note that the neutral-sound variant of C1 does not need A1a or A5.
Proceed · pivot · kill — as thresholds, not opinions
Fixed before the pilot; the verdict below is computed from the posteriors above. PROCEED requires the efficacy pilot to have run and P(A2a) ≥ 0.90, P(A2b) ≥ 0.70 — set so that the meta-analytic prior alone cannot clear them — plus P(A3) ≥ 0.55, P(A4) ≥ 0.60 and rights (or the neutral variant). Kill floor: P(A2a) < 0.50 after the pilot or P(A4) < 0.35 after survey + landing page.
What the tutor can challenge
Every number is a claim: the priors (why 0.50 for A1a and 0.60 for A1b?), the shrinkage for the interviews (0.35 — too harsh or too kind?), the independence assumption in the concept joint probabilities. That is the point — the argument moves from "I believe" to "here is the LR I would change, and here is what it does."
Update discipline
An LR changes only with a written reason and a date; the previous value stays in the change log; a threshold never moves after the data are seen. Survey thresholds are pre-registered in surveys/questionnaire-v2.md; the pilot protocol in DESIGN_THINKING_FRAMEWORK.md §5.2.