Why did we disappear after follow-up questions?
You were on the list, then a requirement removed you.
Studied under
The public evidence layer
This is Upstream Zero's public evidence layer, not a blog. Every experiment, observation, and hypothesis is preserved and dated, so the chronology of what we have learned about how AI decides who to recommend stays transparent.
We study how AI systems evaluate, compare, recommend, and eliminate vendors during buying decisions. The records below are sorted by evidence level so you can tell, at a glance, what has been tested, what has only been observed, and what remains under investigation.
How to read the evidence
We separate what we observed from what we infer, hypothesize, or have replicated. Nothing here implies cause unless a controlled test supports it, and we print the zeros.
Start with your question
Start from the decision you care about. Each question maps to the research components that study it, and from there to the evidence.
You were on the list, then a requirement removed you.
Studied under
The lead changed, or a competitor took your place.
Studied under
What proof an AI system relies on that you have not shown.
Studied under
The requirements you have to meet, and the proof behind them.
Whether your position actually moves after a change.
Studied under
The full process, from opening list to final recommendation.
Research components
The main way we organize the evidence. Every experiment reinforces one of these enduring components, so the work builds toward defined questions rather than a stream of one-off tests.
How an evaluator reads a buyer's stated need and turns it into the criteria it screens vendors against.
Core question
When a buyer states a requirement, how does the evaluator interpret it, and does that interpretation stay stable across runs?
How the initial set of recommended vendors is assembled before any requirement is added.
Core question
Which vendors does an evaluator surface for an unqualified category question, and what governs who is in the opening set?
How and where a vendor is removed from the recommendation set as requirements are added.
Core question
At which requirement is a vendor eliminated, and what makes an evaluator drop it?
How the leading vendor changes as requirements accumulate.
Core question
Does the vendor that leads the opening set keep its lead, and if not, at which requirement does it lose it?
Whether a company remains in the recommendation set as a buyer adds requirements, requests validation, and moves toward selection.
Core question
As requirements are introduced in sequence, does a company remain in the recommendation set, and what removes it when it does not?
Whether the same inputs produce the same recommendation, and whether unstable outcomes trace to unstable requirement association.
Core question
Are recommendation outcomes stable across repeated runs, and when they are not, is the instability in the recommendation or upstream in how requirements are associated?
How one vendor replaces another in the recommendation set.
Core question
When a vendor is added or removed, which competitor takes its place, and what drives the substitution?
Which sources, proof types, and trust signals an evaluator appears to rely on when it validates a vendor against a requirement.
Core question
What evidence does an evaluator appear to use when deciding whether a vendor credibly meets a requirement?
Featured experiments
Clinical trial management software
Project management software
Project management software
Browse all experiments
Every experiment in the program, including the runs still being reviewed into the format above.
What Closed means. A completed experiment is evidence of what occurred under its recorded conditions. Closed means the run is finished, not that the result is a universal truth. Replication, causal support, cross-evaluator agreement, and real-world corroboration are what let a conclusion travel further.
Observations
Candidate observations drawn from more than one experiment. Each is held at its honest level, names the runs it derives from, and is not a finding.
Derived from E-034, E-040, E-041, E-043, E-045.
Narrated. Five categories, one evaluator for most, no causal test of what makes a category behave one way.
Derived from E-044.
Narrated and held. A single category comparison, and a statement about the evaluator's narration, not about the world.
Narrated. Two experiments, and the difference between them is confounded, so cause is not attributed.
Derived from E-034, E-040, E-041, E-043, E-045.
Narrated and held. Consistent across five categories on mostly one evaluator, not yet confirmed across evaluators.
Narrated. One category on two evaluators. Not a demonstration that the effect is surface-independent in general.
Hypotheses under test
Open questions
Findings
Zero, and we print it. A finding is accepted only after replication and scrutiny we have not yet accumulated. 3 founding claims are published at their tier in the claims ledger.