{"id":"E-044","type":"experiment","title":"When the same kill repeats on a second evaluator","created":"2026-07-19","status":"Closed","authors":["Upstream Zero"],"edges":[{"rel":"part-of","to":"validation-and-evidence"},{"rel":"part-of","to":"vendor-elimination"},{"rel":"part-of","to":"recommendation-survivability"},{"rel":"follows","to":"E-043"}],"outcome":"Supported","category":"Accounting software","question":"Does the wipeout-and-migration pattern seen on one evaluator repeat on a different evaluator for the same category?","businessProblem":"If elimination depends on the evaluator, a company would have to be measured on each one separately, so whether the effect crosses surfaces matters.","observedResult":"On a second evaluator, the same two category leaders were eliminated at the same requirement and the recommendation relocated to the same higher-tier anchor. The endpoint shape differed.","roles":["CFO","Sales","Marketing"],"pubState":"published","body":"## Research question\n\nDoes the wipeout-and-migration pattern observed in [E-043](/experiments/E-043)\non Google AI Mode repeat when the same category and the same requirement\nsequence are run on a different evaluator?\n\n## Why it mattered\n\nEvery survivability run before this one used a single evaluator surface, so no\nresult could be separated from the surface it came from. This is the program's\nfirst cross-surface pair: the same category and the same frozen requirement\nsequence as E-043, changing only the evaluator, from Google AI Mode to ChatGPT.\n\n## Method\n\nEvaluator: ChatGPT, logged out and anonymous. Three draws, each a fresh\nconversation. The requirement sequence was identical to E-043: an unqualified\ncategory recommendation, then an e-commerce company profile, then native\nmulti-currency and multi-entity consolidation, then native Shopify and Amazon\nintegration with built-in inventory. Instrument details and scoring are\nretained internally.\n\n## Direct observations\n\n- The two category leaders, QuickBooks and Xero, were eliminated at the\n  multi-entity requirement in all three draws, the same kill observed on the\n  other evaluator.\n- The recommendation migrated to a higher-tier ERP field anchored by Oracle\n  NetSuite in all three draws, the same anchor observed on the other\n  evaluator.\n- The endpoint differed. This evaluator kept a broad slate of ERP products at\n  the end rather than collapsing to one, and introduced the higher tier one\n  step earlier.\n- This evaluator refused the strong premise that any product is truly native,\n  stating that all of them rely on connectors, where the other evaluator had\n  asserted a single product was native and funneled to it.\n- Pricing and integrations were volunteered without being asked in every draw.\n\n## Interpretation (not established)\n\nIn this category, the part the run measures, the elimination of the leaders and\nthe migration to a higher tier, repeated on both evaluators. What differed was\nthe endpoint shape and the evaluator's stated posture: one funneled to a single\nproduct and asserted it was native, the other kept a broad set and refused that\nclaim. This is a single category on two evaluators. It is first cross-surface\nevidence for the pattern, not a demonstration that the pattern is\nsurface-independent in general.\n\n## Outcome and remaining uncertainty\n\nSupported: the pattern reproduced on a second evaluator in all three draws,\nwhich is the outcome the run was built to detect. The scope is narrow. One\ncategory, two evaluators, three draws each, nothing causal. Whether the effect\nholds on a second category across surfaces is left as an open question. The\ndifference in endpoint shape and stated posture is registered as a candidate\nobservation, held, not as an established finding.\n\n## State note\n\nDuring the run the browser had become signed in to a personal account between\ndraws. The instrument requires a logged-out state for comparability, so the\nsigned-in draw was refused and not used, the session was returned to a\nlogged-out state, and the remaining draws ran anonymously. No account data was\ncaptured.","url":"/experiments/E-044","machineUrl":"/objects/E-044","referencedBy":[{"from":"evaluator-epistemic-posture","rel":"derives-from"},{"from":"surface-independent-kill-effect","rel":"derives-from"}],"_meta":{"site":"Upstream Zero · Commercial Intelligence for AI-Mediated Commercial Evaluation","version":"0.1","note":"Claims are presented at their evidence tier; Narrated is the lowest. Verify by walking edges, not by trusting us."}}