Windtunnel method
Repeated sampling
Repeated sampling means every audit prompt runs k=5 times per model and mode, so rates describe the distribution of answers rather than a single anecdote.
Five answers per prompt, before any counting.
Model answers are sampled from a distribution: wording, named brands, and rankings vary from answer to answer. Windtunnel therefore runs every audit prompt cell k=5 times per selected model route and stores each answer verbatim, timestamped, before any metric is computed.
Metrics are then rates or averages over those samples, with intervals where the math supports them. A Message Lift test uses one AI model at k=5, because there the comparison between two messages, not breadth of coverage, is the question. The fuller rationale lives in the methodology page.
Where you can see it.
Every study printout states its repetition count: the Insta360 study, for example, is five prompts times five repetitions, twenty-five stored answers. The methodology page explains how repetition interacts with intervals and the n≥30 gate.
Three things repeated sampling is not.
- Not a human panel. Every sample comes from a disclosed API model route; nothing in the distribution is a person’s opinion.
- Not more prompts. Repetition multiplies answers per prompt; it does not broaden the question set. Coverage comes from the prompt matrix.
- Not a guarantee of convergence. Some answers genuinely vary between runs; the rates and their intervals carry that uncertainty instead of hiding it.
Labels travel with numbers: figures traced to stored answers from real model routes are MEASURED; Message Lift results are SIMULATED and comparative only; anything below the n≥30 gate or drawn from a single prompt cell is DIRECTIONAL. No figure on this site ships without its label, sample size, and date.