← Method

Windtunnel method

Stability index

Stability index is the mean pairwise Jaccard of top-5 tracked-brand sets across repeated samples in the same prompt cell and model route, reported as a point estimate with no interval.

How it is computed

How much does the model repeat itself?

Take every pair of repeated samples for one prompt and one model route. From each answer, take the set of tracked brands in its top five. The Jaccard of two sets is their overlap divided by their union. The stability index is the mean of that overlap across all pairs.

A value near 1 means the model returns nearly the same brand set every time; a value near 0 means the shortlist is a lottery. It is rendered as a point estimate, labeled as such, because the pairs share answers and no honest interval can be put around that average.

Where this appears

Where it shows up.

The stability index rides along with audit metrics in reports and feeds the DIRECTIONAL labeling of low-stability cells. The method notes on the landing page explain why repetition, and therefore stability, exists at all.

What it does not mean

Three things a stability index is not.

  • Not a brand score. It measures the model’s consistency under an identical prompt, not your performance in the answer.
  • Not an accuracy measure. A model can be consistently wrong; stability says nothing about whether a claim is true.
  • Not comparable across metrics. It is a diagnostic for reading other numbers cautiously, most of all when a cell’s stability is low and its figures are labeled DIRECTIONAL.

Labels travel with numbers: figures traced to stored answers from real model routes are MEASURED; Message Lift results are SIMULATED and comparative only; anything below the n≥30 gate or drawn from a single prompt cell is DIRECTIONAL. No figure on this site ships without its label, sample size, and date.