TDMUSIC · progress briefing
← Back to briefing
EMPATHIZE → DEFINE · SURVEY INSTRUMENT v2 · n ≥ 100 (150 preferred) · 7–8 min · EN + 中文

Every question earns its place: who · why · when · where · what · how · how much.

v1 was a 3-minute screener (GAD-2 + behaviour + Van Westendorp). v2 keeps every v1 item verbatim and adds the demand side (why, when and where the need occurs, and what is hired today), the product side (why someone would use it, instead of what, and why not), and a state-dependent content test — so that each answer feeds a named construct, a named hypothesis and a named node in the decision engine.

Source: surveys/questionnaire-v2.md (paste-ready EN for Google Forms · 中文 for 问卷星) · validated instruments reproduced verbatim with sources · adapted scales cited · Chinese wording of non-official items marked [VERIFY WORDING] for back-translation.

/ 01Instrument logic · 5W1H

Seven blocks, one funnel — need before concept, concept before price

The order is a design control: validated need measures come before the concept so the concept cannot prime them; the concept description is content-neutral so the real-music vs neutral-sound question can discriminate; price comes last.

Validated (verbatim)

GAD-2 · PSS-4 · Van Westendorp

Kroenke et al. 2007 · Cohen & Williamson 1988 · van Westendorp 1976. Wording untouched; sources and licences in the markdown.

Adapted (cited)

TAM · Kano · ODI · DOI · CIT

Davis 1989 / Venkatesh & Davis 2000 · Kano et al. 1984 · Ulwick 2002 · Rogers 2003 · Flanagan 1954. Adaptation disclosed; α reported.

Controls

Randomise · check · track

Scenario & Kano order randomised · one attention check + speeder flag · channel-of-origin links · back-translation of 中文 items · no optional stopping.

/ 02The questions

Block by block — question, wording, why it is there, what it feeds

Expand a block. Each item shows the English wording, the Chinese wording, the response format, the construct it measures (the why), and the model node it feeds. Full paste-ready text lives in surveys/questionnaire-v2.md.

A1Age ≥ 18single · screen-out
Are you 18 years of age or older? / 您年满 18 周岁吗?
Why: ethics gate.
A2Age bandsingle
18–24 · 25–34 · 35–44 · 45–54 · 55+
Why: persona age range (primary persona 28–40).
A3Marketsingle · drives currency
Where do you currently live? / 您目前常住在哪里?
US/Canada · UK/EU · Mainland China · HK/TW/SG · Other
Why: overseas-first GTM — US/EU vs China willingness to pay differ ~2.3× (Calm $69.99 vs Tide ¥218).
feeds → segmentation for F (price), A4
A4Work modesingle · v1 Q1 verbatim
Full-time office · Hybrid · Fully remote · Not working · Student
Why: the primary persona is a hybrid-work professional; tests whether the need clusters there.
A5Wearable ownershipmulti · v1 Q2 verbatim
Apple Watch · Oura Ring · Other (Garmin, Fitbit, Whoop, Huawei/Xiaomi) · None
Why: the key segmentation variable for the concept — over-recruit to ≥ 40%.
feeds → A3 segment
A6Wearable data engagementsingle · owners only
How often do you look at your stress, HRV, readiness or "body battery" score? / 您多久看一次自己的压力、HRV、"准备度/身体电量"之类的分数?
Every day · Few times a week · Few times a month · Rarely · Never / don't know what that is
Why: A5 × A6 = "owner who already watches the metric" — the highest-value early-adopter cell.
feeds → A3 prior
B1GAD-2 (validated, verbatim)2-row matrix · 0–3
"Over the last 2 weeks, how often have you been bothered by: (1) Feeling nervous, anxious, or on edge; (2) Not being able to stop or control worrying" — Not at all / Several days / More than half the days / Nearly every day.
"在过去两个星期,有多少时候您受到以下问题所困扰?(1) 感到紧张、不安或烦躁;(2) 无法停止或者控制忧虑" — 完全没有 / 几天 / 一半以上天数 / 几乎每天
Why: the standard anxiety screener (Kroenke et al. 2007; public domain). Sum ≥ 3 = positive screen. Splits the sample into need-intensity segments; not a diagnosis.
feeds → segment "GAD-2 positive"; personas
B2PSS-4 perceived stress (validated, verbatim)4 items · 0–4 · items 2 & 3 reversed
"In the last month, how often have you felt… (1) unable to control the important things in your life? (2) confident about your ability to handle your personal problems? (3) that things were going your way? (4) difficulties were piling up so high that you could not overcome them?" — Never / Almost never / Sometimes / Fairly often / Very often.
"在过去一个月里,您有多经常… (1) 感到无法控制生活中的重要事情?(2) 对自己处理个人问题的能力感到有信心?(3) 觉得事情正朝着您希望的方向发展?(4) 觉得困难堆积如山、无法克服?" [VERIFY WORDING]
Why: the product is positioned as stress-relief wellness, not anxiety treatment — PSS-4 (Cohen & Williamson 1988) measures the actual addressable need in the general population; correlation with GAD-2 gives convergent validity.
feeds → need intensity; segment cross-tabs
B3WHEN — timingmulti · max 3
In the past month, when did you most often feel stressed, anxious or tense enough that you wanted to calm down? / 过去一个月里,您什么时候最常感到压力/焦虑/紧张到想让自己平静下来?
Morning before work · During work hours · Commuting · Evening wind-down · In bed trying to fall asleep · Waking at night · Weekends · No pattern
Why: the arousal-state segmentation (acute vs wind-down) needs a time signature; the interviewees' worst window was late night / in bed.
feeds → A1a/A1b segmentation; JTBD timing
B4WHERE — placemulti · max 2
Office · Home working · Home off-hours · In transit · Public places · Other
Why: context of use decides product form (headphones vs speaker, discreet vs visible, watch on wrist or not).
B5WHY — triggersmulti · max 3
Work pressure / always reachable · Money · Family / relationships · Health worries · Uncertainty / big decisions · Overstimulation · Nothing specific · Other
Why: trigger taxonomy for personas; interviews gave "always-on dread", client rumination, cash-flow — the survey tests how general those are.
B6Somatic / cognitive patternmulti
Racing heart · Chest tightness / shallow breathing · Racing thoughts · Restlessness · Trouble sleeping · Irritability · Can't tolerate silence / noise · None
Why: physiological plausibility — racing heart and racing thoughts are the two things an HR-adaptive audio product can plausibly act on; sleep routes to the insomnia sub-segment.
feeds → pilot measures rationale (A2)
B7Frequency of "need to calm down" momentssingle
Several times a day · Once a day · Few times a week · Once a week · Few times a month · Rarely
Why: need intensity → market sizing input and habit potential (A6).
B8What is hired todaymulti
In those moments, what do you usually actually do? / 在那些时刻,您通常实际会做什么?
Music/audio · Breathing/meditation · Move · Talk to someone · Scroll phone · Eat/drink/smoke/alcohol · Sleep aid/medication · Nothing · Work harder · Other
Why: Jobs-to-be-Done — the competitive set is behaviours, not apps (Christensen et al. 2016). Music's share here is the baseline for the whole thesis.
B9Outcome importance × satisfaction (ODI)7 rows × 2 ratings · 1–5
Outcomes: calm down within minutes · stop racing thoughts · fall asleep faster · almost no effort · discreet, not "therapy" · know it is actually working · not make things worse (no jarring sounds, no lyrics that pull me in).
几分钟内平静 · 停止思绪飞转 · 更快入睡 · 几乎不费力 · 不像"治疗" · 知道它确实起效 · 不会让情况更糟
Why: Ulwick's opportunity score = Importance + max(Importance − Satisfaction, 0) ranks under-served outcomes — the quantitative basis for the POV and for which HMW to prioritise. Pre-registered expectation: "know it's working" and "not make it worse" are under-served; "fall asleep faster" is well served by incumbents.
feeds → POV · HMW-3 · feature priority
C1Frequency of using music/audio to calmsingle · v1 Q4 verbatim
Daily · Few times a week · Few times a month · Rarely · Never
C2What do you actually playmulti
Songs with lyrics I know · Instrumental versions of songs I know · Classical / piano / ambient · Generated soundscapes (Endel, Brain.fm) · Nature / white, brown, pink noise · Guided meditation / calm voice · Podcasts · Lo-fi / chill · Nothing — silence
Why: revealed preference baseline for A1 — behaviour, not opinion.
C3State-dependent preference — two scenarios, order randomisedsingle per scenario · within-subject
Scenario 1 · acute: "Imagine your heart is racing and your mind won't stop — late at night, or right after a stressful call. If you could press one button, which would you most want to hear?"
Scenario 2 · wind-down: "Now imagine you're mildly tense, unwinding after work — nothing urgent."
情景 1(急性):"想象您心跳很快、脑子停不下来……如果只能按一个按钮,您最想听到什么?" · 情景 2(放松):"只是有点紧张,下班后在放松——没有什么紧急的事。"
(a) a calm, slowed-down version of a song I already love · (b) featureless neutral sound, no melody, no lyrics, that quietly adapts to me · (c) a calm human voice · (d) nature / white noise · (e) silence · (f) don't know
Why: the interviews (n=3, acute) rejected real music; the netnography shows both preferences coexist. Hypothesis: preference is state-dependent. Within-subject design → McNemar test on choice (a) across scenarios. Pre-registered: Sc1 (a) ≥ 50% → LR 3 for A1a, ≤ 35% → LR 0.33; Sc2 (a) ≥ 50% → LR 2 for A1b, ≤ 35% → LR 0.5.
feeds → A1a · A1b (decision engine) · the pilot A/B design
C4Reactions to musical features when very anxious9-row matrix · −2 … +2
Lyrics/vocals · A melody I recognise · A strong beat · Sudden or sharp sounds (bird calls, high notes) · A low continuous hum · Very slow tempo · Sweet "healing" music · A song that means something to me · Sound that slowly changes as I calm down
Why: turns the interview complaints (sharp transients, nauseating hum, fake sweetness, beat entrainment, lyric rumination) into a quantified acoustic spec for the neutral-sound variant; the last row tests adaptivity itself independent of content.
feeds → neutral-sound product spec · A1a
C5Value of real artists / known songssingle · 1–5
How much would it matter to you that calming audio is made by real artists or based on songs you know? / 舒缓音频由真人艺术家制作、或基于您熟悉的歌曲,这对您有多重要?
Why: the asset-leverage check from the user's side — how much the catalogue matters to them.
feeds → A1b (LR 1.3 if ≥ 40% top-2)
D1Willingness to connect a wearablesingle · 5-pt + "would consider getting one"
If a calming audio app could read your heart rate from your watch or ring during a session — to adapt the sound and show you the effect afterwards — would you connect it? / 如果一款舒缓音频 App 能在播放过程中读取您手表/指环上的心率……您会连接它吗?
Why: the core A3 item. Pre-registered: top-2 ≥ 60% of owners → LR 2; ≤ 40% → LR 0.5.
feeds → A3
D2Trust in the readingsingle · 1–5 · owners
Why: perceived credibility — Wareable reviews describe stress scores that "mismatch lived experience"; low trust weakens the "show the measured result" promise.
D3What you do after a "high stress" readingmulti · owners
Nothing, just note it · Breathe / meditate · Walk / move · Change plans · It makes me more anxious · I don't look at it · Other
Why: the action gap — the netnography's biometric rows were all about reliability, none about behaviour change; a closed-loop product lives in this gap.
D4Value of a measured resultsingle · 1–5 · all
After a calming session, how valuable would it be to see a measured result — e.g., "your heart rate settled from 92 to 68"? / 一节舒缓结束后,看到"您的心率从 92 降到 68"这样的测量结果,对您有多大价值?
Why: HMW-3 "make relief measurable so users trust it" — is measurability a want or a nice-to-have?
feeds → A3 (LR 1.3 if ≥ 50% top-2)
D5Privacy barriersingle · 1–5
Why: DOI compatibility — heart data to a music app is a new kind of sharing.
feeds → A3 (LR 0.8 if ≥ 4 for > 50%)
ConceptContent-neutral descriptionshown once
"Imagine an audio app that connects to your smartwatch or ring. When you feel tense, you press play. It reads your heart rate while you listen and gently adapts what you hear — slowing, softening, simplifying — to help you settle. When you stop, it shows you what changed. The sound can be calm versions of music you know or neutral, feature-free sound, whichever works for you in the moment. It would be a subscription."
Why neutral: if the concept said "real songs" it would prime C3/E2 and bias the A1 test.
E1TAM — usefulness, ease of use, intention6 items · 7-pt Likert
PU1 helps me calm down faster than what I do now · PU2 improves how I manage stress · PEOU1 easy to use in the moment · PEOU2 connecting and starting takes little effort · BI1 I would try it · BI2 I would use it at least a few times a week
Why: the Technology Acceptance Model (Davis 1989; Venkatesh & Davis 2000) is the standard, citable way to ask "why would you use it" — usefulness and ease predict intention. Two items per construct; α ≥ .70 reported. Pre-registered: BI top-2 ≥ 40% overall / ≥ 55% owners → LR 1.8 for A3.
feeds → A3
E2Kano feature classification6 features × (functional + dysfunctional) · order randomised
K1 adapts in real time to heart rate · K2 shows measured before/after · K3 real songs you know, re-arranged calm · K4 neutral feature-free sound engineered artifact-free · K5 breathing pace cue woven in · K6 works offline / no account — answers: like · expect · neutral · can live with · dislike
Why: Kano (1984) separates what is expected from what delights from what is indifferent or reverse. The C1 differentiation question is exactly whether the biometric loop (K1/K2) is attractive/one-dimensional, and whether real songs (K3) are attractive — or reverse for the acute segment.
feeds → A3 (K1/K2), A1a (K3 reverse), MVP scope
E3Concept form — which would you actually use?single · forced choice
(C1) standalone subscription app · (C2) same feature built into my wearable's own app · (C3) a "calm" album/playlist by artists I like on Spotify/网易云, no app · None
Why: the three concepts from Ideate, tested head-to-head as forms, not features.
feeds → C1 / C2 / C3 weighting
E4Relative advantagesingle · 5-pt
Compared with what you use now to calm down, this would be… much worse → much better.
Why: relative advantage is the strongest predictor of adoption in Diffusion of Innovations (Rogers 2003).
feeds → A3 (LR 1.5 if ≥ 50% top-2)
E5WHY NOT — barriersmulti · max 3
Don't believe it works · Sharing heart data · Another subscription · Already have Calm/Spotify · Wearing watch/ring uncomfortable · Too much effort · Worried it makes anxiety worse (watching my heart rate) · No wearable · Nothing · Other
Why: the risk register and onboarding design come from here; "another subscription" > 40% → LR 0.8 for A4.
F1–F3Apps tried · still paying · why churnedv1 Q5–Q7 verbatim (+ Brain.fm)
Why: revealed WTP and the churn story — the category's 4.7% D30 retention is the structural risk.
F4–F7Van Westendorp Price Sensitivity Meter4 open numeric · monthly · order fixed · currency by A3
Too expensive to consider · Expensive but would consider · A bargain · So cheap you'd question quality.
Why: the canonical PSM (van Westendorp 1976) yields the acceptable range and the optimal price point without asking a leading "would you pay X". Pre-registered: range contains $6.99 → LR 1.5 for A4; OPP < $4 → LR 0.7.
feeds → A4 · paywall price in the prototype
F8Purchase intent at a shown pricesingle · 5-pt · US$6.99 / €6.99 / ¥30 after 14-day trial
Why: stated intent, top-2-box convention, discounted ×0.5 for hypothetical bias (Morwitz et al. 2007). Pre-registered: ≥ 30% (US/EU) → LR 2; ≤ 15% → LR 0.5.
feeds → A4
F9Payment model preferencesingle
Monthly · Annual · One-time · Included in my wearable's subscription · Provided by employer / health plan · Free with ads · Would not pay
Why: C1 (B2C) vs C2 (bundled / B2B2C) vs employer (B2B) — bundled ≥ 35% raises C2.
G1Attention checksingle · required · v1 Q12
Why: exclusion rule (with speeder flag < ⅓ median time).
G2Ideal calming audio toolopen · v1 Q13
G3Critical incidentopen
Think of the last time you needed to calm down quickly. What did you do, and how well did it work? / 请回想最近一次您需要快速平静下来的时候。您做了什么?效果如何?
Why: Critical Incident Technique (Flanagan 1954) — a concrete recent behaviour is more reliable than a general attitude; coded with the netnography theme codes.
G4Optional pilot opt-in emailstored separately
Why: recruits Apple Watch / Oura owners for the n ≥ 20 efficacy pilot without breaking anonymity.
/ 03Analysis plan · pre-registered

Question → construct → threshold → decision node

Everything below is fixed before fielding. Anything not on this table is exploratory and reported as such. Likelihood ratios enter the decision engine exactly as written.

QConstructStatisticPre-registered thresholdFeeds
B1 · B2Anxiety (GAD-2), perceived stress (PSS-4)% GAD-2 ≥ 3; PSS-4 mean by segment; r(GAD-2, PSS-4)Need intensity · personas
B9ODI opportunityImp + max(Imp − Sat, 0), ranked≥ 6 = under-servedPOV · HMW priority
C3State-dependent content preferenceMcNemar on (a) across scenariosSc1 ≥ 50% → LR 3 A1a; ≤ 35% → LR 0.33 · Sc2 ≥ 50% → LR 2 A1b; ≤ 35% → LR 0.5A1a · A1b
C4Acoustic specMean per feature; % negativeNeutral-sound spec
C5Value of real artists% top-2≥ 40% → LR 1.3A1b
D1Willingness to connectTop-2 among owners≥ 60% → LR 2; ≤ 40% → LR 0.5A3
D4 · D5Value of result · privacyTop-2 · % ≥ 4D4 ≥ 50% → LR 1.3 · D5 > 50% → LR 0.8A3
E1TAM PU / PEOU / BIα ≥ .70; BI top-2; PU→BI βBI ≥ 40% (≥ 55% owners) → LR 1.8A3
E2KanoCategory per feature; Better/WorseK1/K2 A or O → LR 1.5 A3 · K3 reverse in GAD-2+ → LR 0.7 A1aA3 · A1a · MVP scope
E3 · F9Concept form · payment modelSharesC2 ≥ C1 or bundled ≥ 35% → raise C2 · C3 > 40% → content-firstC1 / C2 / C3
E4 · E5Relative advantage · barriersTop-2 · frequenciesE4 ≥ 50% → LR 1.5 A3 · "another subscription" > 40% → LR 0.8 A4A3 · A4
F4–F7Van WestendorpOPP · IPP · [PMC, PME]Range ∋ $6.99 → LR 1.5; OPP < $4 → LR 0.7A4
F8Purchase intentTop-2 (×0.5 discount)≥ 30% → LR 2; ≤ 15% → LR 0.5A4
Sample size
n ≥ 100

95% CI ± 9.8 pp on a 50% proportion (± 8 pp at n = 150). McNemar ~80% power to detect a 20-pp shift between scenarios at n = 100.

Segments
≥ 40 · ≥ 30

≥ 40 wearable owners; ≥ 30 GAD-2-positive respondents. Cells < 30 reported as directional only.

Field window
2–3 wks

Stop at n ≥ 150 or the deadline — never because the numbers look good (no optional stopping).

/ 04Validity threats & controls

What could fool us, and what we do about it

Hypothetical bias in WTP

Triangulate PSM + purchase intent (discounted) + landing-page conversion (behavioural).

Social desirability / acquiescence

Reverse items (PSS-4), "prefer silence" and "would not pay" always offered, content-neutral concept.

Order & priming

Need measures before concept; scenario and Kano order randomised; price last.

Self-selection by channel

Separate tracked links (LinkedIn, r/SampleSize, r/ouraring, WeChat, EMBA cohort); report by channel; weight if skew > 20 pp.

Translation

GAD-2 official 中文; PSS-4 / TAM / PSM Chinese items back-translated and piloted (n = 3) before launch.

Insider bias

The founder-researcher does not administer the survey personally; links are anonymous; the analysis notebook is kept for audit.

Ethics · wellness framing

Anonymous · ≥ 18 · no diagnosis · no advice

GAD-2 / PSS-4 shown with the standard "this is not a diagnosis" note and a pointer to general mental-health resources at the end; the optional email is stored separately from answers; data kept for the dissertation only.