← All studies

Two partnership messages, two tests, two different answers.

Two self-initiated Buyer response tests on candidate Insta360 partnership messages. The first message moved simulated buyer intent on both engines. The second moved it backward. These are separate studies against different baselines, not a head-to-head comparison.

Simulated9 & 19 Jul 2026n=5 per engine per messageDirectionalNot a client

Insta360 was not a client, did not commission these tests, did not participate, and has not endorsed the findings. Leica was not involved in any way. Both candidate messages are quoted from public materials; the Current-message baselines are stored AI answers from this project’s own audit.

01 / How the two tests relate

Separate baselines, not a head-to-head.

Each test compared one Current message with one New message through shared contexts; only the message changed. The two tests ran at different dates, with different baselines and different engines, so nothing here pits the two candidate messages against each other. Each test answers one question: does this candidate move simulated buyer intent against its own baseline?

02 / Test one · 9 Jul 2026

The partnership framing moved intent on both engines.

“Insta 360 will launch more action camera products collaborating with Leica to enhance image and video quality.”

Test one results by engine
Engine (ungrounded API routes)CurrentNewShift
OpenAI3.283.43+0.15
DeepSeek3.053.63+0.57

Baseline: the stored AI framing that smaller cameras mean inferior image quality. n=5 stored answers per engine per message; Buyer response scored on the 1 to 5 expressed-intent construct with anchor sets from independent peer-reviewed research (arXiv:2510.08338). Every engine’s result is shown.

03 / Test two · 19 Jul 2026

The certification framing moved intent backward.

“New product with Leica certified lens and Leica designed filters.”

Test two results by engine
Engine (ungrounded API routes)CurrentNewShift
OpenAI3.653.35−0.31

Baseline: the stored AI answer recommending action cameras for a holiday traveller in Japan. n=5 stored answers per message, OpenAI only, ungrounded. The candidate scored worse than the baseline; a weak message dies in the test, not in market.

04 / Limits

What these tests do not show.

  • Directional only. Five stored answers per arm sits far below the n=30 aggregate gate. Treat every shift as a signal to investigate, never as a measured effect size.
  • Simulated, not human. A synthetic panel’s free-text reactions, scored with the independent method’s anchor sets. Not observed buyer behavior.
  • Comparative only. Shifts compare two messages inside one test. Never a sales forecast, never a ranking prediction.
  • Separate baselines. The two tests share no baseline, so the two shifts are not comparable to each other and no combined conclusion is drawn.
  • A snapshot. Valid for the tested prompts, routes, and dates. Models change, and so will this.