question

How stable are AI evaluator recommendations across model versions, prompt phrasings, and repeated sampling?

Q-3status · open2026-07-13

The instrument-stability question. A finding that is true in March and false in May, with no announcement, is not a finding; it is a dated observation. Before we publish anything about what evaluators recommend, it must characterize how stable those recommendations are. This question gates the credibility of every future measurement, including our own commercial offerings.

What this means commercially

Affected buyers
Anyone relying on an AI evaluator's recommendation as if it were a stable fact
Affected categories
All categories surfaced in AI-assisted comparison and screening
Potential product impact
If recommendations are noisy, single-shot "AI visibility" measurements are largely meaningless
Current confidence
Unmeasured by us. Instrument instability is a known risk.
Evidence tier
Narrated (untiered type)

Authors

Upstream Zero

Machine rendering

Q-3.json