Method and limitations
Evidence you can inspect.
Every result shows when it was produced, which model route was tested, how many answers were reviewed, and where the conclusion stops.
- Repeated sampling. Audit prompts run multiple times because model answers vary. Rates describe that distribution.
- Measured versus simulated. Visibility evidence comes from stored model answers. Both Message Lift test types are comparative simulations. The two evidence classes never mix.
- Current/New prompt parity. Both messages use the same prompts and settings. Only the message changes.
- Provider and model disclosure. Every figure names the API/model route and mode that produced it.
- Sample size and uncertainty. Rates show n and confidence intervals where the metric supports one.
- Results that say no. A study that only ever confirms the brief is not measuring anything. Where the evidence does not support a claim, the report says so, and that is the finding.
- Snapshot limitations. Conclusions apply only to the tested prompts, routes, modes, and date.
FAQ
Is Windtunnel a GEO or AEO tool?
GEO and AEO describe work intended to improve how brands appear in AI answers. Windtunnel is the measurement and testing layer for that work: it measures AI visibility, citations, competitive ranking, and brand perception, then tests messages. It does not publish content, build links, or promise rankings. Agencies use the audit to set a baseline and the test to show what their work changed.
How does Message Lift work?
Message Lift begins with a stored AI answer from the same audit. One Current message and one New message then go through shared contexts, so only the message changes. Buyer-response and AI-recommendation results are both labeled simulated.
Why run the same audit prompt more than once?
LLM answers are probabilistic. One answer is an anecdote. Repeated answers support a rate, with a confidence interval where the math supports one.
Do results match a consumer chat interface exactly?
No. Windtunnel uses disclosed API/model routes and settings. Those can differ from a personalized consumer chat session. The site makes no consumer-interface equivalence claim.
How is this different from a free “what does ChatGPT say about my brand” checker?
A one-off chat check is one anecdote. Windtunnel samples every prompt repeatedly, reports rates with their sample size and a confidence interval where the math supports one, and stores each answer so every figure traces to its source.
Why do AI visibility scores change from day to day?
Answers are sampled from a distribution, so a single run is an anecdote, not a fact about your brand. Windtunnel reports rates over repeated samples, states the sample size, and timestamps every stored answer.
Does this work for Singapore and APAC markets?
Market context is part of every prompt. Results are labeled by market, route, and date, so a Singapore answer is never pooled with one from another market.