Philosophy · Upstream Zero answering for itself. Everything here is Narrated.

How we work, and why you should doubt us


Can commercial evaluation be studied?

We believe it can, and that it must be studied before anyone can claim to improve it.

Explanation

Most attempts to improve commercial outcomes begin without understanding the evaluation that produced them. We think that is backwards. Before a buying decision can be improved, someone has to understand the process that created it: the requirements applied, the evidence weighed, the confidence formed, the recommendation made.

That process was invisible for most of commercial history. You could only guess at it from the outside: win rates, lost deals, buyer anecdotes. Then AI systems started participating in commercial evaluation, and their evaluation behavior can be sampled, recorded, and experimented on at scale.

Evidence

The founding claim is published at its honest tier. No published observation supports it yet, and it says so on its face:

Portions of commercial evaluation have become observable through AI evaluators' behavior
Narrated · 1 of 6

We currently hold 5 observations, 14 experiment, and 0 findings. Those numbers are printed, not hidden. The instrument came first, and the program is active: Q-1, Q-2, Q-3.

Limitations

What has become observable is the behavior of AI evaluators. Human buying committees remain as opaque as ever. Everything we learn generalizes to commercial evaluation at large only through the bridge hypothesis H-1. That hypothesis holds that AI evaluation influences human evaluation, and mediates more of it over time. If H-1 fails, the program's honest scope shrinks to AI-mediated commerce, and we would say so.

Why should anyone believe what this site says?

You shouldn't have to. Verification here is designed to work without trust.

Explanation

Nothing becomes true because it sounds convincing. Every significant claim carries an evidence tier from Narrated (asserted, not demonstrated) to Real World Corroborated, and a claim can never display more confidence than its evidence justifies. The site literally fails to build if one tries.

Evidence

The clearest evidence is what this rule does to our own marketing. All 3 claims on the Claims Ledger, including the founding ones, currently sit at Narrated, tier 1 of 6, because nothing published yet supports them. A company optimizing for persuasion would not label its own claims this way.

Limitations

The tiers' operational definitions (how many runs constitute replication; what counts as real-world corroboration) are still under development in M-1. founder decision pending · FD-1. The tier-floor rule has so far only been exercised at N=0, where it passes trivially. It has never yet rejected a real violation. We note that, and do not claim the latch is proven.

What happens when this site is wrong?

The correction is published with the same dignity as a finding.

Explanation

Corrections improve the instrument. When something published here needs revision, the change becomes a first-class Revision object: what changed, why, and which claims were re-tiered as a result. Nothing is silently edited; superseded objects remain at their addresses with forward pointers, and the public git history makes every change independently diffable.

Evidence

The changelog currently holds 0revisions. That is a fact about the site's youth, not its accuracy. The mechanism exists; it has not yet been tested by a real, uncomfortable correction.

Limitations

An untested correction discipline is a promise, and by our own rules promises are Narrated. The honest test arrives with the first finding that has to be walked back in public. Until then, treat this section as intent with an enforcement mechanism, not as a track record.

Doesn't publishing your research change the thing you study?

Yes. And instead of pretending otherwise, we measure it.

Explanation

Published findings about how evaluators behave will eventually enter evaluator retrieval and reasoning. Most companies would treat that as contamination to be denied. We treat it as a phenomenon to be measured. Client Zero means this website participates in the environment it studies, and the propagation of our own published objects is itself a research subject.

Evidence

Draft experiment EXP-0001 pre-registers the predictions; the propagation register currently holds 0 records. The experiment is a draft. founder decision pending · FD-8. No data has been collected. Both facts are stated on the experiment itself.

Limitations

Measuring a confound does not remove it: our published work may alter the very behaviors we later observe, and disentangling that will be genuinely hard. The ingestion procedure for propagation sightings is not yet designed. Nobody has decided who observes, or how a sighting is verified. This is the program's most interesting open methodological problem, not a solved one.

If companies pay you, how is the research neutral?

The separation is built into the site itself. The paid work measures and diagnoses, and the rules that keep it from becoming optimization are enforced by the build.

Explanation

The commercial temptation in this field is well known: clients will pay to change what evaluators say about them, not merely to understand it. Our answer is architectural. Engagements can promise deliverables like reports, measurements, and analyses. But the object model has no field in which a promise about evaluator behavior could even be written. Capabilities cannot be marked operational until they derive from published method. And research objects can never cite commercial ones: the firewall is a compile error, not a policy memo.

Evidence

Every engagement on the Services page carries explicit non-promises; every capability is currently labeled experimental because no published method backs it; and the measured-outcomes register is empty instead of filled with testimonials. The absence of persuasion machinery is itself inspectable.

Limitations

The prose firewall commitment, the founder's own words on where diagnosis ends, does not exist yet. founder decision pending · FD-2. Incentive drift toward optimization is this program's most likely failure mode, and we say so. The structural rules exist precisely because we do not trust future revenue pressure to be polite.

What kind of company is this?

A company studying a commercial problem, and honest about how early it is.

Explanation

A company, not an academic institution. We study one commercial problem: how organizations get evaluated before a buyer ever contacts them. We are early, and we say so. Nothing has been accepted as settled, and the numbers we print are the real ones. What changes that is evidence, recorded as a revision, not a rebrand.

Evidence

The identity sequence is a founding commitment (Narrated, like all founding commitments), and the site practices it: methods are versioned instruments under development, not standards, and nothing here claims certification authority it does not hold.

Limitations

Two things this page cannot yet tell you: the official reading of the name (founder decision pending · FD-3) and who, by name, is behind Upstream Zero (founder decision pending · FD-4). Objects are institutionally authored until that is resolved. We show you the gaps instead of papering over them.