C-0001active2026-07-13
Portions of commercial evaluation have become observable through AI evaluators' behavior
Narrated · 1 of 6
What we claim
Every claim we publicly stand behind, each one labelled with how strong the evidence for it actually is. That includes our own founding claims, which currently rest on the weakest label we have, and say so.
The claims
Why is everything at the weakest level? Because no observation has been published yet, and by our own rules a claim with nothing supporting it stays at the bottom of the scale. The list is designed to be doubted: follow any claim to what supports it, check its history, or fetch the machine version. Verification should never depend on trusting us.