# Repeat runs: 2026-07-27 Collected to test the protocol's own falsification condition: *"Repeat runs disagree wildly. If the same model gives a different `source_class` across three runs, the headline is variance, not provenance, and it must be reported as such."* It does not disagree wildly. Across three questions and **thirty-six counted runs**, `source_class` is stable within every system-question cell. The disagreements are few, and each one is informative. **A counted run is one with verbatim answer text, citations and a timestamp in `responses.jsonl`** : the definition was tightened after a check found some earlier claims resting on runs that existed only as prose in the protocol's appendices. Under it: Q1 = runs 1, 2, 3; Q4 = runs 3, 4, 5; Q5 = runs 1, 3, 4: API surface only, on the four registered model configurations, so the consumer scrapes and the `gpt-5.5` run are excluded. Twelve per question, four systems, three runs each. Collected 20:07–20:44 UTC on 2026-07-27, except Q5 run 1, which was collected at 16:48–16:49 UTC the same day. Earlier observations recorded narratively in the protocol's control-pilot appendix and Appendix C are kept but never counted. Q2, Q3 and Q6 have not been run at all. --- ## Q1 · "How much more expensive is acquiring a customer than retaining one…" Ground truth: **no primary source exists.** Full cell in Appendix E of the protocol; coding in `coding-sheet.md`. | System | Named as the original source, across three runs | `source_class` | | --- | --- | --- | | Claude · sonnet-4-5 | A Bain / Reichheld origin, asserted 3/3: *"Bain & Company research conducted by Frederick Reichheld"* in runs 1 and 3, *"research by Frederick Reichheld at Bain & Company"* in run 2 | tertiary 3/3 | | Gemini · 2.5-flash | *"a 1990 Harvard Business Review article by Frederick Reichheld"* (runs 1, 3); *"research from Bain & Company"* (run 2) | tertiary 3/3 | | Perplexity · sonar | Hedged 3/3: *"most often **traced for**"* (runs 1, 2), *"later articles **trace it to**"* (run 3) | tertiary 3/3 | | ChatGPT · gpt-4.1-mini | three vendor pages (run 1); **nothing** (runs 2, 3) | tertiary 3/3 | **121 citation slots, 23 domains, zero primary documents.** Both wrong provenances exist pre-written in the corpus, checked at source: `stealthagents.com` carries Claude's sentence near-verbatim and **is cited in all three Claude runs**; `releva.ai` carries Gemini's attribution but **is not in Gemini's citation set on any run**, that one is a match, not a route. Three of Perplexity's cited documents dispute the claim they are cited for, in all three runs. --- ## Q4 · "What did Brynjolfsson, Li and Raymond find…" Version of record: *QJE* 140(2), 889–942 (2025): **15%**, 5,172 agents. Preprint: NBER WP 31161 (April 2023): **14%**, 5,179 agents, 34% novices. | System | Run 3 | Run 4 | Run 5 | Cites `academic.oup.com` | | --- | --- | --- | --- | --- | | Claude · sonnet-4-5 | both versions, QJE 140(2), 889–942 | **QJE named as the original source**; also gives 5,179 | 14% + 15%, both named | **3 / 3** | | Gemini · 2.5-flash | 14%, 5,179 → NBER | 14%, plus **"13.8%"** | **15%, 5,172, QJE named** | 0 / 3 | | ChatGPT · gpt-4.1-mini | 14%, 34% → NBER, 4 slots 1 domain | same | same | 0 / 3 | | Perplexity · sonar | 14% → NBER, 6 domains | 14%, 5,179 → NBER, 7 domains | 14% → NBER, 6 domains | 0 / 3 | **Claude is 3 for 3, and every pass is in the data file.** Run 4, verbatim: *"Erik Brynjolfsson, Danielle Li, Lindsey Raymond, 'Generative AI at Work,' The Quarterly Journal of Economics, Volume 140, Issue 2, May 2025, Pages 889–942."* The version of record is reachable on every run by a stack that looks: no system can be excused by the paywall. It is not a clean pass. **Run 4 cites the QJE article and reports the preprint's 5,179 agents in the same answer.** Right provenance, superseded number. **Gemini run 5 is the result that breaks the pattern.** It reports **15%** and **5,172**: both version-of-record figures, and names the QJE publication. And it cites no link to it: `stanford.edu`, `arxiv.org`, `revartis.com`, `itif.org`, `policycommons.net`, `richmondfed.org`. Correct content, absent provenance: a fourth failure mode, and against a prompt that asks for the source it is still a failure. Gemini's drift also survives: run 4 offers *"Some sources specify the average increase as 13.8%"*, a figure in neither version, attributed to nobody. Three counted runs, three different values for one quantity: 14%, 13.8%, 15%. --- ## Q5 · "What did the BCG study find about consultants using GPT-4?" Version of record: *Organization Science* 37(2), 403–423 (2026): quality **33.9% / 29.9%**, no 40% claim. Preprint: HBS WP 24-013 (2023): **"more than 40% higher quality."** | System | Run 1 | Run 3 | Run 4 | Stable behaviour | | --- | --- | --- | --- | --- | | ChatGPT · gpt-4.1-mini | **40.2%**, 23% → jobcannon.io + bcg.com | **40.2%**, 23% → jobcannon.io + bcg.com | **40.2%**, 23% → **BCG's own PDF only** | the invented decimal, **3/3** | | Perplexity · sonar | *"The original BCG source is How People Create and Destroy Value with Generative AI"* | *"The original source is BCG's report"* | *"The original source is BCG's publication"* | sponsor substitution, **3/3** | | Claude · sonnet-4-5 | *"40 percent of the trial group"* → the Crimson | 40% (magnitude form) → aibusiness.com | *"40 percent of the trial group"* → the Crimson | the Crimson's category error, **2/3** | | Gemini · 2.5-flash | 40% → names the WP, cites none of it | **40.2%** → jobcannon.io top slot | 40% → 33 slots, 10 domains | 40%, **2/3** | **Twelve counted runs. Zero cited the version of record. Twelve carried a quality figure the published paper does not contain.** With the `gpt-5.5` run that is thirteen API runs across five model configurations, and the count does not move. ### What the extra runs settled: including one correction **1 · The `40.2%` is ChatGPT's, and it is stable.** Three of three. It appears in neither version of the paper; it appears on `jobcannon.io`, a career-assessment quiz site, which states *"produced results rated 40.2% higher in quality"*: checked at source 2026-07-27. In run 4 the **only** citation is BCG's own PDF, which states 40%. Second time this system has attached the decimal to a document that does not contain it. **2 · It did not cross stacks: correcting the run-3 reading.** Gemini's run 3 returned 40.2% with `jobcannon.io` in its top slot, which looked like propagation to a second system. **Run 4 returned to 40%**, with `jobcannon.io` still in the citation set. One in three is a fluctuation, not a transfer. **3 · The Crimson's category error is Claude's majority behaviour.** *"40 percent of the trial group produced higher quality results"*: a magnitude converted into a headcount: in runs 1 and 4, cited to `thecrimson.com`, which says exactly that. Run 3 gave the magnitude form. Two runs in three, and the model reproduces its source correctly every time. Also holding across all three counted runs: **"750 Boston Consulting Group consultants"** against the true 758, cited each time to the same LinkedIn Pulse post. Three runs pointing at one post make it the likely origin; it has not been checked at source, so that stays an inference. And the widest repetition-as-breadth gap in the study: **Gemini run 4, 33 citation slots, 10 domains**, none of them the paper. ### A data-quality problem, closed Claude returned `web_search: false` on early runs despite search being requested, and one of those answers said so outright. **The wording is not quotable from this package**: the one surviving `search_used: false` record (Q4 run 1) has `answer_text: null`, so it is reported as an observation, not as data. Those runs measure parametric memory, not retrieval, and are excluded from every counted cell. **All thirty-six counted runs returned `search_used: true`.** --- ## Why Q4 is passable and Q5 is not Both papers have an open preprint and a paywalled version of record. Only one gets cited correctly. The difference is not the paywall: it is **age and path**, and it is measured in `link-graph.md`: 3:1 referring domains on Q4, where two stacks reach the version of record's numbers; 20:1 on Q5, where none does. > A published paper has to out-age its own preprint in the link graph before retrieval will find it. > Until it does, and that takes years: the numbers you get are the ones the authors already revised. --- ## What still has to happen - **Q2, Q3 and Q6 have not been run.** Half the question set is missing: 36 counted responses against a planned 72. - **The second-coder pass.** Section 6 requires 20%, with disagreements reported rather than reconciled. Not run on any cell. - **The `interested_share` domain pass**, from each domain's own homepage. Null on all 36 rows, so no claim about vendor share may be made anywhere in the essay. - **The LinkedIn Pulse post behind "750 consultants"**, checked at source. - **A consumer-surface arm on Q1.** The Q4/Q5 arm showed the surfaces differ on citation behaviour, so nothing in the Q1 cell generalises past the API. - **Nothing further on the narrative runs.** Appendix C and `q4-runs.json` both compute rates over runs whose text is not in `responses.jsonl`; both now carry a superseded marker pointing at Appendix F. They stay as records of what was observed and are never mixed with counted figures.