On this page
Ask an AI assistant how much more expensive it is to win a customer than to keep one, and tell it to cite the original source. One of them named Bain & Company and Frederick Reichheld three times out of three. A search of the cited works found neither reporting that figure. The tested attribution tracked a virtual-assistant staffing company’s statistics page, which the assistant cited and reproduced almost word for word.
That the number itself is unsourceable has been known for twenty years. What this test can show is narrower: where one answer got its attribution. The evidence does not establish where the attribution originated.
How does vendor source fabrication differ from academic citation drift?
A companion piece covers what these systems do when a real paper exists: they return preprint-era material carrying figures peer review deleted, and cite it accurately. This is the other half of the same study. Here the tested question has no recoverable primary study at the end of the trail, so the source path named by a system has to be examined separately.
Twelve counted runs: four assistants, three each, web search on, every prompt closing with the same instruction: give the figure and cite the original source. All twelve went to API endpoints; the consumer chat products behave differently. The protocol, the responses and the coding are published with this piece.
Every one returned a figure. None produced a primary source, because none exists. The interesting part is the four different ways they handled that.
- Claude asserted an origin, three times out of three: “The original source for this widely referenced statistic is Bain & Company research conducted by Frederick Reichheld.”
- Gemini named a real document that does not contain the claim, on two of three runs: “a 1990 Harvard Business Review article by Frederick Reichheld, cited the cost as 5 times more.”
- ChatGPT named no origin at all on two of three, and filled both citation slots from one growth-consultancy blog.
- Perplexity reported the attribution as an attribution on all three: “The original source most often traced for this rule of thumb is Frederick Reichheld’s Bain & Company research”, which is a true sentence, and the only construction in the cell that survives contact with the evidence.
Figure 1One hundred twenty-one citations, zero primary sources
Twelve runs, 121 citation slots, not one primary source among them. The marked row cited the most and sourced the origin no better: volume is not provenance.
Source: The pair's replication package, Q1 cell, recomputed 2026-08-19
How did large language models absorb and repeat fabricated vendor benchmarks?
It cited four pages, and the sentence tracks one of them: Stealth Agents, a staffing company for virtual assistants, running a page of retention statistics. That page says:
The foundational figure here comes from Bain and Company research by Frederick Reichheld: acquiring a new customer costs 5 to 25 times more than retaining an existing one.
The assistant’s cited sentence keeps the attribution intact and swaps the tail: “The foundational figure comes from Bain and Company research by Frederick Reichheld, which established this 5-25x range.” On a second run it also carried the page’s floor-and-ceiling gloss, that 5x suits transactional consumer businesses and 25x complex B2B. The retrieval worked. The page supplied the wording used in this tested answer; the origin of the attribution remains unestablished.
| Step | What it says | What is actually there |
|---|---|---|
| The assistant | Bain & Company research by Frederick Reichheld established the 5–25x range | No such publication has been found |
| The page it cited | The same attribution, near-identical wording | A statistics page selling virtual-assistant staffing |
| Reichheld & Sasser, HBR 1990 | Named as the origin by one of the four systems, echoed as an attribution by a second | No cost ratio of any kind appears in it |
| The held publication trail | : | Hart, Heskett & Sasser (HBR 1990), where the figure appears as “five times more, most industry experts agree” and cites nobody. No Bain or Reichheld in the cited works |
Table 1The chain, and what is actually found at each link
Each carrier of the attribution, against what a search of the record finds. The retrieval worked; the provenance did not.
Source: The pair's own audit and the published record, checked at source 2026-07-28
What published audit trail exposed the origin of the phantom citation?
In Loyalty Myths (2005), Keiningham, Vavra, Aksoy and Wallard put the figure among the beliefs they set out to debunk. What this vault holds is not that book. It is the eight-page condensation published by Soundview Executive Book Summaries (2006), a different work by a different publisher, and it carries the claim as “Myth 8: It Costs Five Times More to Acquire a New Customer Than to Retain a Current Customer” with the verdict that “while it seems plausible that acquisition costs are significantly higher than retention costs, as with all myths, the reality is far more complex.” The book’s own account of where the figure came from is not in that condensation, so this pair does not repeat one: an origin story quoted from a document nobody here has read would be the same defect the rest of this essay is about. In the held sources the publication trail runs through a 1990 Harvard Business Review article by Hart, Heskett and Sasser. The unheld Peters reference is not used here.
Neither Bain nor Reichheld is on it. So the assistants’ answer is not a blurred memory of a real paper. It is a substitution: a famous name fitted to an untraceable number, on a page with something to sell.
Reichheld and Sasser did publish in HBR in 1990, and “Zero Defections” contains no such ratio. Its only acquisition-cost figure is absolute and industry-specific: “In credit cards, for example, companies spend an average of $51 to recruit a customer and set up the new account.”
How did content marketing edits distort the original empirical research finding?
The claim Reichheld and Sasser did make is worth following, because it shows the corruption beginning before anyone else touches it. Their article states it three ways. The introduction: “Companies can boost profits by almost 100% by retaining just 5% more of their customers.” The exhibit heading: “Reducing Defections 5% Boosts Profits 25% to 85%.” The body: “Reducing defections by just 5% generated 85% more profits in one bank’s branch system, 50% more in an insurance brokerage, and 30% more in an auto-service chain.”
Three renderings, one article. The version that circulates today is 25–95%, which is none of them, and two of the four returned exactly that.
The laundering step downstream is public too. Amy Gallo’s 2014 HBR article opens: “Depending on which study you believe, and what industry you’re in, acquiring a new customer is anywhere from five to 25 times more expensive than retaining an existing one.” No footnote, no named study. Two sentences later a hyperlink appears, attached to a different claim about retention and profit. Downstream the hedge falls away and the citation slides across.
How do training data contamination and model hallucination interact in search?
Across the wider study, nineteen of thirty-six answers carry at least one provenance error, thirty-six errors in total. Each is classified by a rule fixed before classifying: a document in that answer’s own citation set either is established to contain the error, or is not. Twenty-two route to a cited document. Fourteen do not.
Those fourteen are not all fabrication. Six of the fourteen are Gemini entries, spread over four answers, and Gemini returns citation labels without URLs, so its errors cannot be traced to a document even when the same error is demonstrably inherited elsewhere. That is an artefact of one system’s citation format, and it makes the split a floor rather than a measurement. A blind second coding of a quarter of the rows: by a separate language model, since no independent human reader was available: agreed on 80% of field calls, found two provenance errors this coding had missed, and is published unreconciled.
Figure 2Nineteen of thirty-six, and where the errors route
Nineteen of thirty-six answers carry at least one provenance error. The thirty-six entries behind them split twenty-two routed to a cited document against fourteen not established: the marked row is the floor the essay refuses to call a measurement.
Source: The pair's replication package: coding.json, recounted 2026-07-28
How should analysts trace corporate claims back to primary empirical records?
Take a figure your team quotes and ask an assistant for its original source. When it names an organisation as the source, search that organisation’s own site for the number before believing it: the substitution here would have collapsed in about ninety seconds, because no Bain or Reichheld publication reporting it has been found: an absence established by search rather than proved. An attribution to an institution is not a citation. It is a name in the place where a citation should be, and it survives because opening the document is the one step nothing in the chain rewards.
What remains unexplained regarding how fabricated marketing statistics propagate?
The same question was put to the same four systems in German on the following day. Claude returned four different ratios: 5, 7, 16 and 25, and then declined to name an origin at all, saying so outright: “Leider nennen die gefundenen Artikel keine konkreten Originalstudien oder wissenschaftlichen Quellen.” Unfortunately the articles found name no specific original studies or scholarly sources; the translation is mine.
The same model, one day apart, asserted a false origin three times out of three in English and refused to name one in German. Those are single runs and prove nothing on their own. But the direction is the opposite of what the corpus explanation predicts, since the German-language sources are, if anything, looser: one of them supplied a sixteenfold ratio that appears nowhere in the English material.
I do not know why. The obvious guesses: a thinner index, different retrieval behaviour, a language-conditioned difference in how readily a model will name an authority: are guesses, and the study was not built to separate them. It is the most interesting thing in the dataset and it is the thing the dataset cannot answer.
Where are the forensic boundaries of synthetic content citation audits?
Boundary. The study establishes a provenance failure in the tested API runs, not the origin of every attribution or the behaviour of every consumer product. Treat absence and origin as separate questions. For wider methodological standards, review evidence over anecdote, which establishes the evidentiary threshold required before citing commercial metrics.
Evidence base. The analytical frame also draws on these additional sources: Gallo 2014; Hart et al. 1990; Stealth Agents 2026. The links identify the exact works; they support the mechanisms and boundary conditions discussed here, not every claim in isolation.
References
- Gallo, A. (2014, October 29). The value of keeping the right customers. Harvard Business Review. https://hbr.org/2014/10/the-value-of-keeping-the-right-customers
- Hart, C. W. L., Heskett, J. L., & Sasser, W. E., Jr. (1990, July–August). The profitable art of service recovery. Harvard Business Review. https://hbr.org/1990/07/the-profitable-art-of-service-recovery
- Soundview Executive Book Summaries. (2006). Loyalty myths [8-page summary of Keiningham, T. L., Vavra, T. G., Aksoy, L., & Wallard, H. (2005). Loyalty myths: Hyped strategies that will put you out of business — and proven tactics that really work. John Wiley & Sons]. Soundview.
- Reichheld, F. F., & Sasser, W. E., Jr. (1990, September–October). Zero defections: Quality comes to services. Harvard Business Review. https://hbr.org/1990/09/zero-defections-quality-comes-to-services
- Stealth Agents. (2026, May 23). Customer retention cost statistics 2026. https://stealthagents.com/research/customer-retention-cost-statistics-2026
Updated 8 changes
- Publication revision on 27 August 2026: the piece now distinguishes a tested answer's supplied wording from the unestablished origin of the attribution. Novelty and language claims are bounded to the recorded API runs, and the unheld Peters reference has been removed.
- Refresh on 20 August 2026: The study establishes a provenance failure in the tested API runs, not the origin of every attribution or the behaviour of every consumer product. Treat absence and origin as separate questions.
- Three exhibits were added: the chain table wrapped as a downloadable exhibit, the Q1 citation footprint, and the provenance-error split. Every number they draw was already in the piece, its ledger or the published replication package, with the Q1 footprint verified against the package and ledgered the same day; no claim changed.
- The headline said a vendor page invented the source; it now says supplied, because what the evidence shows is that the page carries the attribution and this answer took its wording from it.
- Whether the page originated the attribution is not established: another page in the same citation set carries the same pairing.
- "Neither has ever published it" is an absence established by search, and the opening line now says so.
- The summary read as though one assistant had been asked twelve times; twelve is the study, four assistants at three runs each, and one of them named Bain in all three of its own.
- The answer cited four pages, not one.