question
How stable are AI evaluator recommendations across model versions, prompt phrasings, and repeated sampling?
Q-3status · open2026-07-13
The instrument-stability question. A finding that is true in March and false in May, with no announcement, is not a finding; it is a dated observation. Before we publish anything about what evaluators recommend, it must characterize how stable those recommendations are. This question gates the credibility of every future measurement, including our own commercial offerings.
What this means commercially
- Affected buyers
- Anyone relying on an AI evaluator's recommendation as if it were a stable fact
- Affected categories
- All categories surfaced in AI-assisted comparison and screening
- Potential product impact
- If recommendations are noisy, single-shot "AI visibility" measurements are largely meaningless
- Current confidence
- Unmeasured by us. Instrument instability is a known risk.
- Evidence tier
- Narrated (untiered type)
Authors
Upstream Zero
Machine rendering