claim

Evaluator-stated rationales are observations about the evaluator's narration, not evidence of mechanism

C-0002status · active2026-07-13

Evidence tier

Narrated · 1 of 6

A standing rule we hold ourselves to, stated as a claim so it can be held to account. When an AI evaluator explains why it recommended something, that explanation is data about what the evaluator says, not proof of what drove the output. Model-stated rationales are known to diverge from the factors that actually determine behavior.

Consequence for method: rationale-type observations are capped at Narrated for any mechanism claim, and behavioral experiments (vary inputs, measure recommendation changes) carry the burden instead. External literature on rationale unfaithfulness exists and will be cited when this claim's evidence chain is built; per our own rules, an uncited memory of a literature is still Narrated.


Authors

Upstream Zero

Machine rendering

C-0002.json