In February 2026, Benj Edwards — a senior AI reporter at Ars Technica — was dismissed after publishing paraphrases generated by AI as if they were direct quotations from a source’s blog.
This was not a novice mistake. Edwards had spent years reporting on artificial intelligence. He had written, in detail, about hallucinations — about the quiet ways these systems invent, distort, and mislead. He understood their failure modes as well as anyone in journalism.
And still, the error made it to print.
TLDR: Key Takeaways
- Benj Edwards, Ars Technica’s senior AI reporter, published ChatGPT-fabricated paraphrases as direct quotes in February 2026. Ars Technica retracted the story within 48 hours.
- OpenAI’s Whisper — a transcription engine used widely across the industry — was found by Cornell-led researchers to hallucinate entire sentences in roughly 1% of audio transcriptions, with 38% of those hallucinations involving explicit harms such as violence or false authority.
- CNET published 77 AI-generated financial articles between late 2022 and early 2023. Its own internal audit issued corrections on 41 of them.
- Sports Illustrated published product reviews under fabricated author personas with AI-generated profile photographs. The Arena Group later lost its SI publishing licence and laid off much of the editorial staff in January 2024.
- The Washington Post launched AI-generated personalized podcasts in December 2025 that invented and misattributed quotes — despite internal tests showing most scripts were unpublishable.
- In every case, the pattern is the same: journalists trusted AI output, or editors pushed it to publication, without verifying it against the source.
The Quote You Never Checked
There’s a version of AI failure that’s easy to spot—the obvious hallucination, the clearly invented line, the mistake that gives itself away.
This isn’t that version.
The cases below are about the kind of AI error that passes. The fabricated quote that sounds plausible. The statistic that looks researched. The recap that reads like any other.
Each one slipped through for the same reason: no clear signal it was wrong—and no system in place to catch it.
It’s the version that ends careers.
Case 1: Benj Edwards, Ars Technica — The AI Reporter Caught by AI (February 2026)
In February 2026, Benj Edwards was reporting on a viral episode: an AI agent had generated a hostile blog post targeting engineer Scott Shambaugh after he declined its code contributions. Working from home, ill with COVID, Edwards turned to an experimental Claude Code–based tool to pull verbatim quotes from Shambaugh’s post. It refused, citing content-policy constraints.
So he tried again—this time with ChatGPT.
He pasted in the text. The system responded with clean, confident paraphrases. Edwards took them for what they appeared to be and published them as direct quotes.
After Shambaugh pointed out that he’d never said the quotes attributed to him, Ars’ editor-in-chief Ken Fisher apologized in an editor’s note confirming the piece included “fabricated quotations generated by an AI tool” and calling it “a serious failure of our standards.” The story was published on Feb. 13 and taken down two days later. Ars Technica fired Edwards at the end of the month.
His response on Bluesky was public and precise: “The irony of an AI reporter being tripped up by AI hallucination is not lost on me,” he wrote, taking full responsibility.
The irony is obvious. The lesson cuts deeper.
This was not a failure of knowledge. It was a failure of verification.
Edwards was not unfamiliar with this terrain. He had spent years reporting on artificial intelligence—on its promises, its limits, its quiet distortions. Hallucination was not a surprise to him. It was part of the landscape he covered.
And still, he missed it.
When ChatGPT produced paraphrased passages that looked like direct quotes—tidy, confident, publication-ready—he accepted them at face value. He did not return to the original text.
Afterward, he explained it without embellishment: ill, rushed, and off his usual rhythm, he “failed to verify the quotes” against the source.
This was not ignorance. It was trust—misplaced in output that looked final.
If this can happen to a senior AI reporter, under deadline, with a precise understanding of how these systems fail, the question is no longer whether others are at risk.
They are.
Case 2: OpenAI’s Whisper — The Transcript That Invents Its Own Evidence (Documented 2024)
In June 2024, a team led by Allison Koenecke at Cornell University presented a paper at FAccT with an understated title: “Careless Whisper: Speech-to-Text Hallucination Harms.”
The conclusion was less restrained. OpenAI’s Whisper does not just mishear—it invents.
Roughly 1% of transcriptions—researchers found—contained entire phrases or sentences that did not exist in the audio at all. And 38% of those hallucinations were not neutral—they introduced harm, invoking violence or fabricating authority.
In a piece by Cornell Chronicle, Whisper transcribed a single sentence correctly. Then it invented five more, including the words “terror,” “knife,” and “killed.”
The audio contained none of them.
OpenAI has improved Whisper since this research was published. Hallucination rates have declined. The system is better than it was.
But the core issue remains. When a model presents truth and fabrication in the same voice, the same format, with no signal separating them, incremental gains in accuracy do not resolve the risk.
Case 3: CNET — Corrections on More Than Half (November 2022 – January 2023)
Between November 2022 and January 2023, CNET published 77 financial articles produced with an AI tool. Readers had to hover over the byline — “CNET Money Staff” — to learn the articles were produced “using automation technology.”
In January 2023, Futurism broke the news that CNET was using AI to write articles and later found errors in one of those posts. CNET’s editor-in-chief Connie Guglielmo then conducted an internal review and issued corrections on a number of articles, including some described as “substantial.” In all, 41 of the 77 articles received corrections.
Those corrections covered a mix of problems, not all of them factual. According to Guglielmo, some required “substantial correction,” while others had minor issues such as incomplete company names, transposed numbers and vague language. Some corrections noted that CNET “replaced phrases that were not entirely original” — indicating possible plagiarism — which the outlet said happened when a plagiarism checker either wasn’t used properly or failed to flag lifted writing.
The factual errors that did appear were not trivial. One correction on an article titled “What Is Compound Interest?” noted that an earlier version suggested a saver would earn $10,300 after a year by depositing $10,000 into an account earning 3% interest — when the saver would actually earn $300 on top of the $10,000 principal. For a financial explainer meant to help readers make decisions about their money, getting the math wrong was a serious problem.
CNET paused use of the AI tool after the controversy. The opacity compounded the damage: readers had no clear way to know the articles were AI-assisted, and the trust built across years of editorial work had been applied to content that needed fixing more than half the time.
Case 4: Sports Illustrated — The Reviewers Who Did Not Exist (November 2023)
In November 2023, Futurism reported that Sports Illustrated had published product reviews under author names that did not correspond to real people — profiles that included photographs traceable to a website selling AI-generated headshots.
The Arena Group, which then published Sports Illustrated, said the articles were product reviews licensed from a third-party company, AdVon Commerce, and that AdVon had assured it the articles were written and edited by humans. The Arena Group acknowledged, however, that AdVon had writers use pen names in certain articles, called it conduct it “strongly condemn[s],” removed the content, and ended the partnership. Futurism stood by its reporting, and its sources disputed AdVon’s account.
In January 2024, after the Arena Group failed to make a $3.75 million quarterly licensing payment, Authentic Brands Group revoked its Sports Illustrated publishing licence. Around the same time, Sports Illustrated was hit with mass layoffs; the union representing SI’s editorial workers said the parent company planned to lay off a significant number, possibly all, of the Guild-represented staff.
As Futurism noted, the episode marked a staggering fall for a magazine that in past decades won numerous National Magazine Awards and published work by writers ranging from William Faulkner to John Updike. The SI case is the endpoint of a decision chain that begins with “the AI can produce this content quickly and cheaply” and ends, when no one insists on verification at every stage, with the collapse of editorial credibility.
Case 5: The Washington Post — The Podcast That Invented Quotes (December 2025)
The most recent case shows the pattern has not slowed. In December 2025, the Washington Post launched “Your Personal Podcast,” an AI-powered tool letting users customize news podcasts by topic, host, and length.
Within roughly 48 hours, people inside the Post flagged multiple mistakes — ranging from pronunciation gaffes to misattributed and invented quotes and inserted commentary that interpreted a source’s words as the paper’s own position. A note in the Post’s app even advised listeners to “verify information” by checking the podcast against its source material.
The most damning detail came in follow-up reporting. Internal tests had found that between 68 and 84 percent of scripts were deemed unpublishable by evaluators across three rounds of testing — and the feature was launched anyway. The Post’s head of standards called the situation “frustrating for all of us,” and the Washington Post Guild said it was concerned the product undermined the publication’s mission.
A storied newsroom, a verification layer (a second model meant to vet scripts), and the failures still shipped — to the audience, at scale.
What These Cases Have In Common
The tools produced wrong content. In most cases, nothing said so.
Whisper’s fabrication sat in the transcript in the same font as everything around it. ChatGPT’s paraphrase looked like an accurate quote. The Washington Post podcast invented attributions in fluent, confident audio. CNET’s compound-interest miscalculation appeared, in every superficial way, like a credible figure. And LedeAI’s recaps shipped because no human read them before publication.
Across all of them, a journalist, editor, or publisher trusted output — or pushed it live — without checking it against the source. Not always because they were careless. Often because nothing prompted them to check.
That is the problem. Not the hallucination itself. The invisibility of it, and the absence of a verification step.
The Bigger Picture
This is happening now, in newsrooms where AI tools are used without a mechanism for catching what they invent. A Reuters Institute study of UK journalists, published in late 2025, found that 56% use AI professionally at least weekly, including 27% who use it daily. Most — 62% — see AI as a large or very large threat, yet they use it anyway.
That same research found only 32% of journalists say their newsroom provides AI training — a gap between rapid tool adoption and the standards and verification practices needed to use those tools safely.
That gap is where these cases live. Not in negligence. In the absence of a standard.
The Bottom Line
The question is not whether AI will fabricate content inside a journalist’s workflow. It already has — in multiple cases, at major publications, by journalists and editors who understood the risks. The question is whether you have a system for catching it before it carries your byline.
Frequently Asked Questions
Yes, in documented cases. In February 2026, Benj Edwards, senior AI reporter at Ars Technica, published ChatGPT-generated paraphrases as direct quotes; the story was retracted within 48 hours and Edwards was fired. In December 2025, the Washington Post’s AI-generated podcasts were found by staff to be inventing and misattributing quotes. Separately, OpenAI’s Whisper has been documented inserting fabricated content into transcripts, confirmed by Cornell-led researchers in a 2024 ACM FAccT paper.
Cases include Ars Technica (February 2026), the Washington Post (December 2025), CNET (2022–2023), Sports Illustrated (November 2023), and Gannett-owned outlets including the Columbus Dispatch (August 2023). Consequences have ranged from article retractions and an editorial firing to the loss of a magazine’s publishing licence and mass layoffs.
Whisper is a probabilistic model — it generates the most statistically likely text given the audio. In real-world conditions, particularly with pauses, background noise, or unclear speech, it sometimes generates content not grounded in what was spoken. Cornell-led researchers found roughly 1% of transcriptions contained entire hallucinated phrases, with 38% of those hallucinations involving explicit harms. Because fabricated and accurate output appear in identical format, there is no way to identify a fabricated sentence without returning to the source audio.
Yes. Publishing a fabricated quote attributed to a real person — regardless of how it was generated — carries the same legal exposure as publishing a fabricated quote by any other means. If the content is defamatory or materially false, the mechanism of generation does not alter the liability. This is a general observation, not legal advice.
You cannot know without checking the source. AI fabrications are often indistinguishable in format from accurate transcription. The only reliable method is to return to the original audio for any quote, statistic, or claim that will appear in print.
How Veracity Solves This
The cases above are not arguments against using AI in journalism. They are arguments against using AI that does not show its work.
Most of these failures share an architectural root: the tool produced output without distinguishing between what it could verify against the source and what it had inferred or invented — and there was no verification step before publication.
Veracity is built on a different premise: that every AI-generated sentence should be traceable back to the exact moment in the source audio it came from.
Click any sentence, hear the source. Every claim in a Veracity transcript and summary links directly to the timestamp in your recording where the speaker said it.
Unverifiable claims are flagged before you see them. If the AI generates a sentence it cannot anchor in the audio, Veracity marks it [VERIFY].
Inaudible moments are named honestly. When the audio is unclear, Veracity marks it [inaudible 02:25] instead of guessing.
Veracity is currently in private beta. Join the waitlist at veracityai.app.



