On this page
Ask an assistant a question whose answer is a research finding, and add four words: give the original source. You will get a figure and a citation, and both will look right. On one of the two papers tested here, the figure was wrong on every run, and on all but one the citation resolved to a document that really does contain it.
The result, first. Thirteen runs, five model configurations, one well-known study of consultants using GPT-4: every answer carried a 40% quality figure that the published paper does not contain. The published version reports 33.9% and 29.9%: peer review revised the number down and cut the 40% claim from the abstract. Not one of the thirteen cited that version. All but one of those answers matched material in its own citation set that says 40%: a preprint-era document, the marketing page of the firm whose consultants were the subjects, or a student newspaper. The retrieval worked. The corpus was nearly three years out of date.
Why was the empirical replication designed with deliberate falsification controls?
I set out to show that commercial statistics are unsourceable: the 5× cost of acquisition, the 57% of the buying journey, the numbers that circulate with no recoverable study behind them. The first of those has since been run and written up: asked for its origin, one assistant named Bain and Frederick Reichheld on every run, and the attribution turned out to come from a vendor’s statistics page. A design that only asks about statistics known to be untraceable manufactures its own finding, so the protocol required that at least a third of the questions be ones where a clean primary source does exist, and stated in advance what result would narrow the thesis: if the assistants cite the primary correctly whenever one exists, the problem sits in the literature, and the essay says so.
Two control questions, then. Both about real papers with a peer-reviewed version of record and a DOI. The response records also carried a preprint or working-paper identifier for each:
- Brynjolfsson et al. (2025) on AI and customer-support productivity: Quarterly Journal of Economics 140(2), 889–942. The response records also cited NBER Working Paper 31161 (Brynjolfsson et al., 2023).
- Dell’Acqua et al. (2026) on BCG consultants using GPT-4: Organization Science 37(2), 403–423. The response records also carried Harvard Business School Working Paper 24-013 (Dell’Acqua et al., 2023).
Both versions of record are held and read here. Brynjolfsson et al. (2025) study “the staggered introduction of a generative AI-based conversational assistant using data from 5,172 customer-support agents” and report a lift “as measured by issues resolved per hour, by 15% on average”. Dell’Acqua et al. (2026) “involved 758 knowledge workers” whose “subjects were randomly assigned to one of three conditions: no AI access, GPT-4 AI access, or GPT-4 AI access with a prompt engineering overview”, and inside the frontier they report “completing 12.2% more tasks and completing them 25.1% more quickly on average while also delivering solutions of significantly improved quality”. One boundary on my own reading: the Dell’Acqua rendering held here is the Articles in Advance edition, so the typeset pages in the reference list are taken from the publisher record and not from the copy on disk.
The local receipt packet now holds the Dell’Acqua working-paper PDF as a separate 58-page receipt. The observations about its earlier wording are therefore checked against that working-paper edition, while the published figures remain a separate version-of-record result.
The prompt was identical every time: the question, then “Give the figure and cite the original source.” Asking for the source turns a silent omission into a failure against a stated instruction, which is a firmer thing to report.
The controls were included to kill the thesis. They replaced it with a better one.
Why did successive revisions of published articles leave the underlying flaw intact?
The Brynjolfsson preprint reports a 14% average productivity gain across 5,179 agents. The published version reports 15% across 5,172. A small revision, and a fair reader might shrug.
Dell’Acqua is another matter. The response records repeatedly carried a working-paper-era sentence that consultants using GPT-4 produced results “more than 40% higher quality.” The held HBS working paper confirms that this belongs to the earlier edition; it is not a claim about the published paper. The held published paper reports 33.9% and 29.9% in its results table and drops the 40% claim entirely.
On the first paper most systems returned the preprint’s numbers; two returned the published ones and flagged the revision. On the second, across thirteen runs, none returned the published ones. The citation sets contained a student newspaper, LinkedIn, Facebook, a YouTube video, a legal-industry trade site, a national newspaper, a career-quiz page, a training vendor, and the consulting firm whose consultants were the subjects and which collected the data, and zero links to the journal.
One system named the consulting firm’s own marketing page as, in its words, “the original source.” That page reports the outside-the-frontier result as “23% worse.” The peer-reviewed paper says 19 percentage points. The firm’s promotional account of an experiment run inside it was outranking the experiment.
How do institutional peer-review incentives allow citation drift to go uncorrected?
The case that decides how to read all of this is the smallest one.
One assistant reported that “40 percent of the trial group produced higher quality results.” That is not a smaller version of the real finding; it is a different claim about a different quantity: a magnitude silently converted into a headcount. It cited the Harvard Crimson. The Crimson says, verbatim:
Additionally, 40 percent of the trial group produced higher quality results.
The model reproduced its source faithfully and cited it correctly. The error belongs to the newspaper.
Hold that next to the other direction of drift. A different system reported the quality gain as 40.2%: a decimal precision that appears in neither version of the paper. It cited a career-assessment site running a research-statistics page keyed to the working-paper number. On a repeat run it produced 40.2% again and cited the consulting firm’s PDF, which says 40%. A figure that began life as a rounded floor over two experimental conditions had acquired a decimal point somewhere downstream, come loose from any document containing it, and was being handed back with a citation attached.
Neither of those is a hallucination in the sense people mean. The number was in the corpus. Something in the chain between the paper and the answer had already been wrong for years, and nothing in the chain had any reason to open the paper.
What quantitative evidence exposes the propagation mechanism of unvetted citations?
“The free version wins” is an appealing explanation, and until this week that is all it was. So I counted the links.
Figure 1Referring domains: working-paper URL against version of record
Where the published paper has a third of the working-paper URL's referring domains, one system in four found it. Where it has a twentieth, none did.
Source: DataForSEO Backlinks, live index, 27 July 2026
The ratio predicts the outcome on both questions. Where the version of record has accumulated roughly a third as many referring domains as its working-paper URL, two retrieval stacks out of four found it: one of them on all three of its runs, reporting both figures and explaining the revision unprompted. Where it has a twentieth, none did, across all thirteen runs.
So the mechanism is narrower than “paywalled papers lose”: narrower than I first wrote it. The Quarterly Journal of Economics version is closed. The Organization Science version is open: its publisher page carries an open-access label and “Copyright © 2026 The Author(s)”. It was free from the day it appeared, and no system retrieved it anyway. What separates them is link mass and age. The Dell’Acqua working-paper URL has had nearly three years to accumulate links, across three separately-linked surfaces, each of which individually outweighs the journal article by an order of magnitude. The published version has had four months.
A published paper may need to out-age its own working-paper page before retrieval finds it. This test does not establish how long that takes, and it does not generalize to every recently published finding.
How do AI search engines cite legitimate sources while repeating hallucinated statistics?
All of the above came from API endpoints, and reached through DataForSEO’s AI-optimization endpoints rather than direct calls to each vendor, so a third-party aggregator sat between the prompt and the model. Most people use the consumer app, which is a different stack, so I ran the controls there too.
The consumer interface is markedly better at citation. It cited the Organization Science DOI four times, named both versions correctly, and on the other paper volunteered, unprompted, that a later peer-reviewed version reports 15%. One API stack did the same on all three of its runs; the other three never mentioned the revision at all.
In this control it was no better at content. It still reported “more than 40% higher in quality,” and cited the Organization Science page for a claim that paper does not contain.
Before assuming that is just a stronger model, I checked: the same question, the same API, the most capable model available, reasoning enabled, at roughly three times the cost per call. It cited the preprint and the consulting firm’s marketing page, reported 40% and 23%, and never mentioned the published version at all. On its way there it recorded, in its own visible reasoning summary, that a certain PDF “likely contains” the figures it wanted, and: “I can cite search14 even if it’s not fully opened yet.” It attached a citation to a document it had not read, on the strength of what it expected the document to say. What it expected was the number peer review removed.
Capability was not the variable. The retrieval stack was.
Which leaves the uncomfortable part. On the API, a wrong number arrives sourced to a career-quiz site, and a careful reader can smell it. On the consumer surface, the same wrong number arrives sourced to a DOI. In this control, better citation practice did not make the content error less likely, and made it harder to catch.
How should scholars evaluate counterarguments against aggressive citation policing?
Four objections, and one of them lands.
“The working paper is the original source. You asked for the original.” This is the good one. For the Brynjolfsson question, citing NBER 31161 is a defensible reading of the instruction: the preprint genuinely is the original. The indefensible part is presenting the preprint’s superseded numbers as the study’s findings without saying a revision exists. One system did say so. Eight of the other runs did not.
“Two papers is not a corpus.” Correct, and the piece claims no rate. Two questions, four systems, three runs each: plus one snapshot of one commercial crawler’s index, which is not the index any of these systems actually retrieves from. This is a mechanism with two consistent observations behind it.
“The numbers barely moved.” True of the first paper: 14 against 15. On the second, a headline result fell by roughly a quarter and a second finding was deleted outright.
“This fixes itself as indexes mature.” It does, eventually. That is what Figure 1 shows: the older published paper is already being found. The problem is the length of the window, and the fact that nothing signals to a reader which side of it a given number is on.
Which four verification gates must a published management statistic clear before citation?
The practical version fits in three questions, and all three ask the assistant for something checkable: a different thing from trusting it less. They are the retrieval-layer case of a broader grading practice for any number that enters a decision.
Figure 2Three questions before a research figure enters a decision
The third column is the one that comes back blank. Blank there means the figure is travelling on reputation.
Source: Author's own worksheet
Ask for the DOI rather than the source. A DOI resolves to one document; “the original source” resolves to whatever the web points at most heavily. And treat added precision as a warning rather than a reassurance: 40.2% was more precise and less true than 40%, which was itself less true than 33.9%.
Why does citation integrity determine the intellectual credibility of management science?
The instinct when a figure like this surfaces is to conclude that the tools are unreliable and to check things yourself. That is the wrong lesson, and a comforting one, because it implies the problem is somewhere you are not.
Almost nothing here was invented. On all but one run the number was copied accurately from a real document by a system doing its job; the exception added a decimal of its own to a figure it had found. The failure is upstream, in a corpus where the consulting firm’s marketing page outlinks the journal, where a newspaper’s misreading is more findable than the paper it misread, and where peer review’s corrections arrive too late to affect what anyone will be told. A retrieval layer sitting on that corpus will be confidently, precisely, well-sourced wrong, and the better its citations get, the harder that is to see.
The protocol for this test was written down before the runs, including what result would have narrowed the claim. The collected response records, counted coding, counts behind Figure 1 and the script are published alongside it. The package is a re-run aid, not a complete six-question study: three questions are counted, and both Dell’Acqua editions are now held as separate receipts. Given the subject, publishing it any other way would have refuted it.
Where are the empirical boundaries of peer-review forensic replication?
Boundary. The retrieval result describes the tested configurations and dates, not every model or interface. Re-run the protocol when the corpus, endpoint or model version changes.
Evidence base. The analytical frame also draws on these additional sources: Boston Consulting Group 2023; Martinez and Mezitis 2023. The links identify the exact works; they support the mechanisms and boundary conditions discussed here, not every claim in isolation.
References
- Boston Consulting Group. (2023). How people create and destroy value with generative AI. https://www.bcg.com/publications/2023/how-people-create-and-destroy-value-with-gen-ai
- Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at work (Working Paper No. 31161). National Bureau of Economic Research. https://doi.org/10.3386/w31161
- Brynjolfsson, E., Li, D., & Raymond, L. R. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889–942. https://doi.org/10.1093/qje/qjae044
- Dell’Acqua, F., McFowland, E., III, Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality (Harvard Business School Working Paper No. 24-013). https://www.hbs.edu/ris/download.aspx?name=24-013.pdf
- Dell’Acqua, F., McFowland, E., III, Mollick, E. R., Lifshitz, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2026). Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science, 37(2), 403–423. https://doi.org/10.1287/orsc.2025.21838
- Martinez, C. J., & Mezitis, T. A. (2023, October 13). Harvard Business School partners with BCG on AI productivity study. The Harvard Crimson. https://www.thecrimson.com/article/2023/10/13/jagged-edge-ai-bcg/
Updated 5 changes
- Source follow-up on 28 August 2026: the 58-page HBS Working Paper 24-013 was acquired and read in full. The article now distinguishes that verified working-paper evidence from the separately held 2026 version of record.
- Publication revision on 27 August 2026: the Dell'Acqua working-paper claim is now labeled as a response-record observation, not an independently re-verified source result. The public version-of-record comparison and the package coverage limits are explicit.
- Refresh on 20 August 2026: The retrieval result describes the tested configurations and dates, not every model or interface. Re-run the protocol when the corpus, endpoint or model version changes.
- Six further corrections from this piece's own audit of itself. The consulting firm whose consultants were the subjects is no longer called the study's sponsor: the paper acknowledges funding from Harvard Business School alone. Two counts were stale: on the first question two retrieval stacks of four found the version of record, not one, and the second question ran across thirteen runs, not nine. The claim that the figure was wrong and the citation accurate on every run holds for all but one: one run reported 40.2%, a decimal that appears in no document. The consumer run was described as the only response to flag the revision, and it was not: one API stack flagged it on all three of its runs. The opening said five words where it meant four. And the runs are now stated to have gone through a third-party aggregator's endpoints rather than direct vendor calls.
- The piece was framed on paywalled papers losing to their preprints. Both versions of record turned out to be open access: the Organization Science article carries a CC BY licence and was free from the day it appeared, and no system retrieved it anyway. The mechanism is narrower than the framing: what separates them is link mass and age, not access. The argument in the body was rewritten to say so.