AI visibility measurement & message testing

Measure how AI recommends your brand.
Test which message moves you up its shortlist.

We measure how AI describes and recommends your brand, then test your current message against a new one under the same conditions. You get the evidence, the comparison, and the limits.

How an audit works
QuestionsRepeated answersEvidence

A controlled process. An inspectable result.

Inside the report

See what the evidence looks like.

A dashboard of findings, with the conditions behind every number. Explore this real report excerpt, then inspect the full studies.

Windtunnel · Brand audit Real study excerpt

Insta 360

19 Jul 2026 · n=118 · OpenAI + DeepSeek · ungrounded

01 · Presence

Mention rate94.9%Wilson 95% CI 89.3–97.7

Share of voice

  • DJI33.2%
  • Go Pro32.7%
  • Insta 36031.5%
  • Akaso2.5%

03 · Perception organic sentiment · n=112

Organic sentiment: 80% positive · 19% mixed · 1% negative. Rounded displayed percentages; n=112.

Message Lift two framings, one partnershipSimulatedDirectional

  • “…collaborating with Leica to enhance image and video quality”

    9 Jul 2026 · OpenAI + DeepSeek · ungrounded · n=5 per engine per message+0.15 to +0.57

  • “New product with Leica certified lens and Leica designed filters”

    19 Jul 2026 · OpenAI · ungrounded · n=5 per message−0.31

Two separate studies against different baselines, not a head-to-head test. One framing of the partnership moved expressed intent. The other did not.

Self-initiated study, not client work. Every figure traces to a stored answer. Inspect the audit example →

Reconstructed report interface using the previously published study figures. The audit and the two simulated tests are separate evidence sets, not a single combined result. Insta360 was not a client and did not commission or endorse this work.

The engagement

From the question
to the next decision.

Find the gap between the story you tell and the answers buyers receive. Test a candidate message before deciding what to change.

  1. 01

    Discover relevant questions.

    Define the category, competitors, markets and buying situations that matter.

    Scope
  2. 02

    Audit repeated AI answers.

    Measure how the brand appears across disclosed prompts and engine modes.

    Measured
  3. 03

    Inspect the evidence report.

    Trace findings to stored answers, sample sizes and explicit limitations.

    Measured
  4. 04

    Test a candidate message.

    Compare Current and New in shared contexts. These results are simulations.

    Simulated
  5. 05

    Re-audit later.

    A later engagement measures the new state. It is a re-audit, not a subscription.

    Follow-up
02 / What we measure

What we measure, and what each number means.

The metric names matter less than the decisions behind them. Each pillar answers one question a brand team can act on.

From the audit report · Pillar 01

01Presence

Am I in AI’s consideration set?

We ask open, unbranded category questions and measure whether AI volunteers your brand.

Share of voice across 180 unbranded answers: of every brand AI named, yours took 24%.

  • Your brand24%
  • Rival A31%
  • Rival B22%
  • Rival C15%
  • Others8%
Illustrative examplen=180

From the audit report · Pillar 02

02Position

When compared, do I win?

We compare you with named competitors and measure which brand AI recommends.

Asked to choose between you and each rival 60 times, this is how often AI picks you.

  • vs Rival A35%
  • vs Rival B48%
  • vs Rival C27%
Illustrative examplen=60 each
  • Organic Recommendation Rate
  • Comparative Win Rate

From the audit report · Pillar 03

03Perception

How does AI describe my brand?

We examine the qualities, sentiment, and objections AI attaches to your brand.

Nine associations were coded from 25 stored answers. These four carry the decisions.

  • Capture everything20/25
  • Third-person shots18/25
  • Leader, unsupported13/25
  • Newest story, missing0/25
MeasuredInsta360 · n=2511 Jul 2026 · DeepSeek · ungroundedSingle analyst · directional · not a client
  • Sentiment
  • Attribute associations

From the audit report · Pillar 04

04Proof

Is the story true and sourced?

We check factual claims against the brand fact sheet and trace citations to their sources.

Every checkable claim AI makes about you is matched against your fact sheet.

Supported
76%
Unsupported
16%
Contradicted
8%
Illustrative examplen=145 claims
  • Accuracy Rate
  • Citation Share
04 / Why now

The click no longer tells the whole story.

AI can compare, describe, and recommend brands before a buyer reaches a site, sometimes without sending a click at all. Your analytics start after that moment.

Traditional attribution

Where the old scorecard starts too late.

  • 01No click
  • 02No conventional referral
  • 03AI shortlist not visible
  • 04Outdated description before the visit
These metrics remain useful. They begin after AI may have already framed, compared, or omitted the brand.
Message Lift / Controlled comparison

Shared contexts.
Only the message changes.

Current message

What you say today.

Shared contexts.
Only the message changes.

New message

The candidate you are considering.

AI recommendation

Compare top-five shortlist and top-choice rates in brand-neutral shopping situations.

Buyer response

Compare simulated free-text reactions, scored on the same 1–5 response construct.

Both test types are simulated. Neither predicts sales or rankings.

Read how testing works ↗
Hotel group / Buyer response

A rewrite is not always an improvement.

Current3.45
New3.41
Simulatedn=30 per message

DeepSeek · ungrounded · 31 Jul 2026. Anonymized independent study; not a client. Means on a 1–5 construct, not a significance or equivalence claim.

Read why the candidate was stopped ↗
Research & method

Evidence you can
look beneath.

Repeated sampling makes variability visible. Stored answers make results traceable. We report uncertainty where the method supports it, and state what each study cannot establish.

How the scoring works / Buyer response

From free-text answer to labeled number.

Nothing in the scoring is a black box. A stored answer becomes a distribution you can inspect and reproduce, and the anchor method comes from independent peer-reviewed research (arXiv:2510.08338), which Windtunnel productizes.

Simulated · Buyer response

Stored answer

A free-text response, immutable and timestamped.

Embedding

The answer becomes a vector: a position in meaning-space.

Cosine vs anchors

Similarity against anchor statements: 5 intent levels × 6 wordings = 30 anchors.

Distribution

A probability mass across the 1 to 5 scale, not a single guess.

Labeled result

Response lift compares mean scores on the 1–5 construct, with n and its simulated label. No confidence interval is invented.

Simulated

AI recommendation follows a different path.

Each message is supplied as untrusted context to the same brand-neutral shopping situations. Exactly five ranked recommendations are parsed deterministically. Top-five and top-choice rates describe the result; Shortlist lift compares Current and New in percentage points, with scenario-cluster bootstrap intervals.

Inspect both scoring methods ↗

Every Windtunnel figure ships in the Glass Box: labeled measured or simulated, gated by sample size, wrapped in its confidence interval where the math supports one, and traceable to the stored raw response behind it.

Method and limitations

Evidence you can inspect.

Every result shows when it was produced, which model route was tested, how many answers were reviewed, and where the conclusion stops.

  • Repeated sampling. Audit prompts run multiple times because model answers vary. Rates describe that distribution.
  • Measured versus simulated. Visibility evidence comes from stored model answers. Both Message Lift test types are comparative simulations. The two evidence classes never mix.
  • Current/New prompt parity. Both messages use the same prompts and settings. Only the message changes.
  • Provider and model disclosure. Every figure names the API/model route and mode that produced it.
  • Sample size and uncertainty. Rates show n and confidence intervals where the metric supports one.
  • Results that say no. A study that only ever confirms the brief is not measuring anything. Where the evidence does not support a claim, the report says so, and that is the finding.
  • Snapshot limitations. Conclusions apply only to the tested prompts, routes, modes, and date.

FAQ

Is Windtunnel a GEO or AEO tool?

GEO and AEO describe work intended to improve how brands appear in AI answers. Windtunnel is the measurement and testing layer for that work: it measures AI visibility, citations, competitive ranking, and brand perception, then tests messages. It does not publish content, build links, or promise rankings. Agencies use the audit to set a baseline and the test to show what their work changed.

How does Message Lift work?

Message Lift begins with a stored AI answer from the same audit. One Current message and one New message then go through shared contexts, so only the message changes. Buyer-response and AI-recommendation results are both labeled simulated.

Why run the same audit prompt more than once?

LLM answers are probabilistic. One answer is an anecdote. Repeated answers support a rate, with a confidence interval where the math supports one.

Do results match a consumer chat interface exactly?

No. Windtunnel uses disclosed API/model routes and settings. Those can differ from a personalized consumer chat session. The site makes no consumer-interface equivalence claim.

How is this different from a free “what does ChatGPT say about my brand” checker?

A one-off chat check is one anecdote. Windtunnel samples every prompt repeatedly, reports rates with their sample size and a confidence interval where the math supports one, and stores each answer so every figure traces to its source.

Why do AI visibility scores change from day to day?

Answers are sampled from a distribution, so a single run is an anecdote, not a fact about your brand. Windtunnel reports rates over repeated samples, states the sample size, and timestamps every stored answer.

Does this work for Singapore and APAC markets?

Market context is part of every prompt. Results are labeled by market, route, and date, so a Singapore answer is never pooled with one from another market.