← Back to Windtunnel

Why one AI answer is an anecdote.

Model answers are sampled from a distribution: ask the same question twice and the wording, the brands named, and the ranking can all differ. A single answer cannot tell you where you stand. Windtunnel treats the model itself as the measuring instrument and reads rates off repeated samples instead of anecdotes.

01 / Repeated sampling

Every audit prompt runs five times.

Each audit prompt cell is sampled k=5 times per selected model route, so a rate like mention rate is computed over many stored answers, not one. A Message Lift test uses one AI model at k=5, because the comparison, not breadth, is the point there. Every stored answer is timestamped and traceable, so any figure can be re-derived from its source.

Repetition is also how honesty survives a bad day: a model that names you in two of five answers produces a very different rate from one that names you in four, and only repeated samples can tell those models apart. See repeated sampling for the term's definition.

02 / Uncertainty

Which numbers carry an interval.

Rate metrics with enough support carry a Wilson 95% confidence interval: mention rate and the shortlist-family rates. Some metrics are reported as point estimates on purpose: share of voice, average first position, and the stability index are computed in ways that break the assumptions an interval needs, and inventing one would be exactly the false precision this method exists to avoid.

Shortlist lift reports a bootstrap interval that resamples scenario clusters within the run, because recommendation scenarios in one run share a context and must not be treated as independent draws.

03 / The n≥30 gate

Small numbers get labels, not headlines.

Aggregate findings render only when at least 30 eligible samples support them. Below that gate, figures are labeled directional and never promoted into claims. Cell-level observations, such as a single lost-shortlist prompt, are always directional-only, no matter the sample size.

04 / Where a conclusion stops

A result is a snapshot, not a law.

  • Prompts. Conclusions hold for the tested prompt set, printed verbatim in every study.
  • Routes. Every figure names the API/model route and mode that produced it. A consumer chat session can differ.
  • Date. Models change without notice. A figure describes the day it was collected.
  • Evidence class. Measured and simulated results are labeled and never mixed.
Simulation / Two distinct test types

The response is not the shortlist.

Buyer response

A synthetic panel produces free-text reactions, not numeric ratings. Embedding similarity against versioned anchor statements maps reactions onto a five-point Likert probability distribution. Response lift compares point-estimate means on that construct. No invented confidence intervals, purchase probabilities, or sales predictions.

AI recommendation

Each message is supplied as untrusted context to the same brand-neutral shopping situations. Five ranked JSON recommendations are parsed deterministically. Top-five and top-choice rates describe the simulated recommendation outcome; Shortlist lift is the difference in percentage points, with scenario-cluster bootstrap intervals.

Aggregate results require at least 30 eligible samples. Persona slices remain directional. Both test types disclose exact A/B prompts and use shared contexts: only the message changes. Simulation rows never enter measured audit aggregates.

Inspect the published studies ↗