Author: Nabiha Veracity

  • 5 Times AI Fabricated Journalists’ Content — And What It Cost Them

    5 Times AI Fabricated Journalists’ Content — And What It Cost Them


    In February 2026, Benj Edwards — a senior AI reporter at Ars Technica — was dismissed after publishing paraphrases generated by AI as if they were direct quotations from a source’s blog.

    This was not a novice mistake. Edwards had spent years reporting on artificial intelligence. He had written, in detail, about hallucinations — about the quiet ways these systems invent, distort, and mislead. He understood their failure modes as well as anyone in journalism.

    And still, the error made it to print.

    TLDR: Key Takeaways

    The Quote You Never Checked

    There’s a version of AI failure that’s easy to spot—the obvious hallucination, the clearly invented line, the mistake that gives itself away.

    This isn’t that version.

    The cases below are about the kind of AI error that passes. The fabricated quote that sounds plausible. The statistic that looks researched. The recap that reads like any other.

    Each one slipped through for the same reason: no clear signal it was wrong—and no system in place to catch it.

    It’s the version that ends careers.

    Case 1: Benj Edwards, Ars Technica — The AI Reporter Caught by AI (February 2026)

    In February 2026, Benj Edwards was reporting on a viral episode: an AI agent had generated a hostile blog post targeting engineer Scott Shambaugh after he declined its code contributions. Working from home, ill with COVID, Edwards turned to an experimental Claude Code–based tool to pull verbatim quotes from Shambaugh’s post. It refused, citing content-policy constraints.

    So he tried again—this time with ChatGPT.

    He pasted in the text. The system responded with clean, confident paraphrases. Edwards took them for what they appeared to be and published them as direct quotes.

    After Shambaugh pointed out that he’d never said the quotes attributed to him, Ars’ editor-in-chief Ken Fisher apologized in an editor’s note confirming the piece included “fabricated quotations generated by an AI tool” and calling it “a serious failure of our standards.” The story was published on Feb. 13 and taken down two days later. Ars Technica fired Edwards at the end of the month.

    His response on Bluesky was public and precise: “The irony of an AI reporter being tripped up by AI hallucination is not lost on me,” he wrote, taking full responsibility.

    The irony is obvious. The lesson cuts deeper.

    This was not a failure of knowledge. It was a failure of verification.

    Edwards was not unfamiliar with this terrain. He had spent years reporting on artificial intelligence—on its promises, its limits, its quiet distortions. Hallucination was not a surprise to him. It was part of the landscape he covered.

    And still, he missed it.

    When ChatGPT produced paraphrased passages that looked like direct quotes—tidy, confident, publication-ready—he accepted them at face value. He did not return to the original text.

    Afterward, he explained it without embellishment: ill, rushed, and off his usual rhythm, he “failed to verify the quotes” against the source.

    This was not ignorance. It was trust—misplaced in output that looked final.

    If this can happen to a senior AI reporter, under deadline, with a precise understanding of how these systems fail, the question is no longer whether others are at risk.

    They are.

    Case 2: OpenAI’s Whisper — The Transcript That Invents Its Own Evidence (Documented 2024)

    In June 2024, a team led by Allison Koenecke at Cornell University presented a paper at FAccT with an understated title: Careless Whisper: Speech-to-Text Hallucination Harms.”

    The conclusion was less restrained. OpenAI’s Whisper does not just mishear—it invents.

    Roughly 1% of transcriptions—researchers found—contained entire phrases or sentences that did not exist in the audio at all. And 38% of those hallucinations were not neutral—they introduced harm, invoking violence or fabricating authority.

    In a piece by Cornell Chronicle, Whisper transcribed a single sentence correctly. Then it invented five more, including the words “terror,” “knife,” and “killed.”

    The audio contained none of them.

    OpenAI has improved Whisper since this research was published. Hallucination rates have declined. The system is better than it was.

    But the core issue remains. When a model presents truth and fabrication in the same voice, the same format, with no signal separating them, incremental gains in accuracy do not resolve the risk.

    Case 3: CNET — Corrections on More Than Half (November 2022 – January 2023)

    Between November 2022 and January 2023, CNET published 77 financial articles produced with an AI tool. Readers had to hover over the byline — “CNET Money Staff” — to learn the articles were produced “using automation technology.”

    In January 2023, Futurism broke the news that CNET was using AI to write articles and later found errors in one of those posts. CNET’s editor-in-chief Connie Guglielmo then conducted an internal review and issued corrections on a number of articles, including some described as “substantial.” In all, 41 of the 77 articles received corrections.

    Those corrections covered a mix of problems, not all of them factual. According to Guglielmo, some required “substantial correction,” while others had minor issues such as incomplete company names, transposed numbers and vague language. Some corrections noted that CNET “replaced phrases that were not entirely original” — indicating possible plagiarism — which the outlet said happened when a plagiarism checker either wasn’t used properly or failed to flag lifted writing.

    The factual errors that did appear were not trivial. One correction on an article titled “What Is Compound Interest?” noted that an earlier version suggested a saver would earn $10,300 after a year by depositing $10,000 into an account earning 3% interest — when the saver would actually earn $300 on top of the $10,000 principal. For a financial explainer meant to help readers make decisions about their money, getting the math wrong was a serious problem.

    CNET paused use of the AI tool after the controversy. The opacity compounded the damage: readers had no clear way to know the articles were AI-assisted, and the trust built across years of editorial work had been applied to content that needed fixing more than half the time.

    Case 4: Sports Illustrated — The Reviewers Who Did Not Exist (November 2023)

    In November 2023, Futurism reported that Sports Illustrated had published product reviews under author names that did not correspond to real people — profiles that included photographs traceable to a website selling AI-generated headshots.

    The Arena Group, which then published Sports Illustrated, said the articles were product reviews licensed from a third-party company, AdVon Commerce, and that AdVon had assured it the articles were written and edited by humans. The Arena Group acknowledged, however, that AdVon had writers use pen names in certain articles, called it conduct it “strongly condemn[s],” removed the content, and ended the partnership. Futurism stood by its reporting, and its sources disputed AdVon’s account.

    In January 2024, after the Arena Group failed to make a $3.75 million quarterly licensing payment, Authentic Brands Group revoked its Sports Illustrated publishing licence. Around the same time, Sports Illustrated was hit with mass layoffs; the union representing SI’s editorial workers said the parent company planned to lay off a significant number, possibly all, of the Guild-represented staff.

    As Futurism noted, the episode marked a staggering fall for a magazine that in past decades won numerous National Magazine Awards and published work by writers ranging from William Faulkner to John Updike. The SI case is the endpoint of a decision chain that begins with “the AI can produce this content quickly and cheaply” and ends, when no one insists on verification at every stage, with the collapse of editorial credibility.

    Case 5: The Washington Post — The Podcast That Invented Quotes (December 2025)

    The most recent case shows the pattern has not slowed. In December 2025, the Washington Post launched “Your Personal Podcast,” an AI-powered tool letting users customize news podcasts by topic, host, and length.

    Within roughly 48 hours, people inside the Post flagged multiple mistakes — ranging from pronunciation gaffes to misattributed and invented quotes and inserted commentary that interpreted a source’s words as the paper’s own position. A note in the Post’s app even advised listeners to “verify information” by checking the podcast against its source material.

    The most damning detail came in follow-up reporting. Internal tests had found that between 68 and 84 percent of scripts were deemed unpublishable by evaluators across three rounds of testing — and the feature was launched anyway. The Post’s head of standards called the situation “frustrating for all of us,” and the Washington Post Guild said it was concerned the product undermined the publication’s mission.

    A storied newsroom, a verification layer (a second model meant to vet scripts), and the failures still shipped — to the audience, at scale.

    What These Cases Have In Common

    The tools produced wrong content. In most cases, nothing said so.

    Whisper’s fabrication sat in the transcript in the same font as everything around it. ChatGPT’s paraphrase looked like an accurate quote. The Washington Post podcast invented attributions in fluent, confident audio. CNET’s compound-interest miscalculation appeared, in every superficial way, like a credible figure. And LedeAI’s recaps shipped because no human read them before publication.

    Across all of them, a journalist, editor, or publisher trusted output — or pushed it live — without checking it against the source. Not always because they were careless. Often because nothing prompted them to check.

    That is the problem. Not the hallucination itself. The invisibility of it, and the absence of a verification step.

    The Bigger Picture

    This is happening now, in newsrooms where AI tools are used without a mechanism for catching what they invent. A Reuters Institute study of UK journalists, published in late 2025, found that 56% use AI professionally at least weekly, including 27% who use it daily. Most — 62% — see AI as a large or very large threat, yet they use it anyway.

    That same research found only 32% of journalists say their newsroom provides AI training — a gap between rapid tool adoption and the standards and verification practices needed to use those tools safely.

    That gap is where these cases live. Not in negligence. In the absence of a standard.

    The Bottom Line

    The question is not whether AI will fabricate content inside a journalist’s workflow. It already has — in multiple cases, at major publications, by journalists and editors who understood the risks. The question is whether you have a system for catching it before it carries your byline.

    Frequently Asked Questions

    Has AI fabricated direct quotes attributed to real people in journalism?

    Yes, in documented cases. In February 2026, Benj Edwards, senior AI reporter at Ars Technica, published ChatGPT-generated paraphrases as direct quotes; the story was retracted within 48 hours and Edwards was fired. In December 2025, the Washington Post’s AI-generated podcasts were found by staff to be inventing and misattributing quotes. Separately, OpenAI’s Whisper has been documented inserting fabricated content into transcripts, confirmed by Cornell-led researchers in a 2024 ACM FAccT paper.

    Which publications have been affected by AI content failures?

    Cases include Ars Technica (February 2026), the Washington Post (December 2025), CNET (2022–2023), Sports Illustrated (November 2023), and Gannett-owned outlets including the Columbus Dispatch (August 2023). Consequences have ranged from article retractions and an editorial firing to the loss of a magazine’s publishing licence and mass layoffs.

    How does Whisper fabricate content in transcripts?

    Whisper is a probabilistic model — it generates the most statistically likely text given the audio. In real-world conditions, particularly with pauses, background noise, or unclear speech, it sometimes generates content not grounded in what was spoken. Cornell-led researchers found roughly 1% of transcriptions contained entire hallucinated phrases, with 38% of those hallucinations involving explicit harms. Because fabricated and accurate output appear in identical format, there is no way to identify a fabricated sentence without returning to the source audio.

    Can AI journalism errors create legal liability?

    Yes. Publishing a fabricated quote attributed to a real person — regardless of how it was generated — carries the same legal exposure as publishing a fabricated quote by any other means. If the content is defamatory or materially false, the mechanism of generation does not alter the liability. This is a general observation, not legal advice.

    How do I know if an AI has fabricated content in my transcript?

    You cannot know without checking the source. AI fabrications are often indistinguishable in format from accurate transcription. The only reliable method is to return to the original audio for any quote, statistic, or claim that will appear in print.

    How Veracity Solves This

    The cases above are not arguments against using AI in journalism. They are arguments against using AI that does not show its work.

    Most of these failures share an architectural root: the tool produced output without distinguishing between what it could verify against the source and what it had inferred or invented — and there was no verification step before publication.

    Veracity is built on a different premise: that every AI-generated sentence should be traceable back to the exact moment in the source audio it came from.

    Click any sentence, hear the source. Every claim in a Veracity transcript and summary links directly to the timestamp in your recording where the speaker said it.

    Unverifiable claims are flagged before you see them. If the AI generates a sentence it cannot anchor in the audio, Veracity marks it [VERIFY].

    Inaudible moments are named honestly. When the audio is unclear, Veracity marks it [inaudible 02:25] instead of guessing.

    Veracity is currently in private beta. Join the waitlist at veracityai.app.

  • Why AI Transcription Tools Are Lying to You — And How to Protect Yourself

    Why AI Transcription Tools Are Lying to You — And How to Protect Yourself

    OpenAI’s Whisper fabricates content in approximately 1% of transcriptions—and that is only the transcript layer. This article examines what AI transcription tools are doing to journalists’ source material, why the verification savings are smaller than advertised, and what you can do to protect your work right now.

    TLDR: Key Takeaways

    • OpenAI’s Whisper fabricated content in approximately 1% of transcriptions in the version tested by Cornell researchers in 2023. OpenAI has shipped updates since, but independent testing has shown hallucination problems persist across newer Whisper versions. Whisper powers Otter.ai, Descript, and dozens of other tools journalists use daily
    • GPT-4o hallucinates in around 1.5% of summaries on Vectara’s leaderboard; Claude-3-Opus at 10.1%; reasoning models do worse, not better — DeepSeek R1 hallucinates nearly four times as often as its non-reasoning counterpart
    • 53% of CNET’s AI-generated financial articles contained factual errors — approximately 25 times the typical newspaper correction rate
    • Sports Illustrated published product reviews under fabricated author personas with AI-generated headshots. The Arena Group lost its SI publishing licence and laid off the magazine’s entire staff in January 2024
    • The Reuters Institute found journalists who use AI most frequently are more likely to believe they spend too much time on low-level tasks

    Why are AI transcription tools “lying” to you? Because they don’t just mishear—they sometimes invent words that were never said.

    Modern speech-to-text systems can be highly accurate, but they still break down in real-world conditions—accents, background noise, or unclear audio can trigger errors or even full “hallucinations,” where the AI generates entirely fabricated phrases . And because these systems are built on probabilistic models, they aim to produce plausible text—not guaranteed truth .

    What nobody tells you when you start

    Here is what the tools do not warn you about: the transcript you are reading may contain sentences that were never spoken.

    Not mishearings. Not garbled syllables or words dropped in noise. Entire fabricated passages, invented by the AI and inserted without seam into your transcript, carrying no marker and no hesitation. They sit alongside the accurate lines with the same formatting, the same confidence, the same air of having been said.

    If you record interviews, press conferences, or source conversations and rely on AI tools to capture them, this is not a theoretical risk to keep at the back of your mind. It is a documented, measured, and increasingly well-understood failure that the tools themselves do nothing to signal.

    The transcript is clean. The story moves forward. Somewhere in the text is a sentence nobody said.


    The Whisper Problem: When the AI Invents What Was Said

    In June 2024, researchers led by Allison Koenecke at Cornell University published findings at the ACM FAccT conference — later expanded in an October 2024 Associated Press investigation by Garance Burke and Hilke Schellmann, a co-author of the original study — confirming what some journalists had started to suspect: OpenAI’s Whisper speech-to-text model fabricates content in roughly 1% of transcriptions on the version tested earlier — and while OpenAI has since released updates that reduce the rate, independent researchers have repeatedly demonstrated the underlying problem persists across newer versions.

    1%. A figure that sounds small until you understand the infrastructure it runs on.

    Whisper is not a niche tool at the margins of the industry. It is the engine underneath many of the transcription services journalists use daily. At that scale, a 1% fabrication rate translates into millions of invented passages, distributed across thousands of users, sitting inside transcripts that look indistinguishable from the real thing.

    And these are not subtle errors. The researchers found that Whisper does not mishear words. It invents sentences. In documented cases, it inserted violent language into transcripts of calm conversations. In others, it fabricated racial commentary or medical claims from audio that contained nothing of the sort.

    The AI was not guessing incorrectly. It was creating fiction — and presenting it as fact, in the same format, with the same confidence, as everything around it.

    For a journalist, the consequence is precise: a source could appear to have said something they never said. And your transcript would give you no reason to question it.

    The industry has begun to respond. In December 2025, OpenAI released new transcription models that, by the company’s own measurement, produce around 90% fewer hallucinations than Whisper v2 in noisy conditions. Improvement is real — but even the slightest hallucination can cause a lot of damage.


    It Gets Worse at the Summary Layer

    Most journalists do not stop at the raw transcript. The real promise of these tools is the summary: upload the recording, receive a clean account of what was discussed, pull the best quotes, and begin writing. It is where the time savings are most visible. It is also where the second, larger source of error lives.

    Vectara’s Hallucination Evaluation Model — the most widely cited benchmark for AI summarisation accuracy — shows that even the best-performing large language models hallucinate at rates that cannot be dismissed. On the current leaderboard, GPT-4o hallucinates in around 1.5% of summaries; Claude-3.7-Sonnet at 4.4%; Claude-3-Opus at 10.1%. Gemini-3-Pro hits 13.6% on the harder, refreshed benchmark.

    The most counterintuitive finding: turning on reasoning makes models worse at sticking to the source. DeepSeek’s reasoning model (R1) hallucinates at 14.3% on Vectara — nearly four times the rate of its non-reasoning counterpart (V3) at 3.9%. The same pattern shows up across providers. GPT-5, Claude Sonnet 4.5, Grok-4, and Gemini-3-Pro — all marketed as more intelligent than the models that came before — every one of them exceeds 10% hallucination on the harder benchmark.

    The models are not getting more reliable at grounded summarisation. The ones marketed as the smartest are, in some cases, getting worse.

    AI summarisation tools regularly generate plausible-sounding text that does not match what was actually said. They do it confidently. Without flagging. In a format designed to look authoritative.

    Chris Middleton, technology journalist at Diginomica, put it plainly after discovering that Otter’s AI summary had inserted a specific statistic — “1.2 million” — that was never uttered by any speaker in the meeting:

    “Put simply, I can no longer trust Otter to record basic facts, and then present them to me. Instead, I am getting an AI-powered confection that claims to be a precis of a conversation I have taken part in myself.”

    For a journalist, this is not an inconvenience. It is reality being rewritten: quietly, seamlessly, without a single visible break in the text.


    The Real-World Damage Is Already Here

    This is not a future risk. It has already ended careers, bankrupted publishers, and destroyed editorial credibility that took years to build.

    CNET: A 53% error rate

    In late 2022, CNET quietly published 77 AI-generated financial articles. When Futurism investigated in January 2023, roughly 41 of those articles — 53% — contained factual errors requiring corrections, including basic mathematical mistakes in financial calculations. The error rate was approximately 25 times higher than the typical newspaper correction rate of 2–3%.

    Sports Illustrated: Fabricated personas, fabricated reviews

    In November 2023, Futurism exposed product reviews at Sports Illustrated published under fabricated author personas with AI-generated headshots, licensed from a third-party content company called AdVon Commerce. Futurism’s sources said the articles themselves were AI-generated; SI’s parent company denied that specific claim but acknowledged the fake bylines.The Arena Group’s CEO was fired. In January 2024, Arena Group missed a $3.75 million quarterly licence payment, lost its Sports Illustrated publishing licence, and laid off the magazine’s entire unionised staff.

    A 70-year-old institution, gutted.

    The Washington Post: Shipped anyway

    When the Washington Post launched ‘Your Personal Podcast’ — AI-generated personalised podcasts based on its journalism — in December 2025, internal testing showed 68–84% of scripts were deemed unpublishable. The product shipped anyway. An anonymous editor told Semafor it was “truly astonishing that this was allowed to go forward at all.”

    Ars Technica: The AI reporter caught by AI

    In February 2026, Benj Edwards — then senior AI reporter at Ars Technica, a journalist whose entire beat was understanding these systems — used an AI tool to extract quotes from a blog post by engineer Scott Shambaugh. The AI returned paraphrased versions of Shambaugh’s words. Edwards published them as direct quotes. The story was retracted within 48 hours. Edwards was fired by the end of the month. His response on Bluesky: “The irony of an AI reporter being tripped up by AI hallucination is not lost on me.”

    If a senior AI beat reporter cannot catch the errors, what chance does a general assignment journalist on deadline have?


    The Verification Tax: Why the Time Savings Are Smaller Than Advertised

    Before AI transcription, journalists typically spent three to six hours of manual work for every hour of recorded audio. AI compressed that to near-real-time.

    Journalists using AI transcripts often spend roughly half an hour to over an hour per hour of recorded interview reviewing, correcting, and cross-referencing the output against the original audio. Verifying exact quote accuracy. Confirming speaker identification in multi-person recordings. Correcting mangled proper nouns and technical terms. Checking context to ensure nothing is used misleadingly.

    A Reuters Institute survey of UK journalists found that more frequent AI users are actually more likely to believe they spend too much time on low-level tasks. The researchers’ explanation is pointed: AI use generates new, AI-specific low-level tasks — cleaning data, checking outputs — that did not exist before.

    The finding that follows from the same study is one the industry should sit with longer than it has: the journalists most satisfied with their time spent on creative work are those who do not use AI at all.

    Brian Merchant, author of Blood in the Machine, named the dynamic plainly:

    “Many journalists are finding that the supposed efficiencies generated by such uses are often offset by new tasks, like rechecking transcriptions for accuracy, ensuring AI texts are free of hallucinations, and editing output for clarity.”

    AI transcription did not eliminate the work. It redistributed it. Instead of typing during the interview, you are now scrubbing audio after the interview to check whether the AI captured it correctly. For many journalists — especially freelancers without institutional fact-checking support — the net time savings are far smaller than advertised.

    This is the verification tax. And it is one every journalist currently using these tools is quietly paying, whether they have named it or not.


    How to Protect Yourself Right Now

    Until tools are built that address this problem at the level of architecture, here is what you can do today.

    1. Never trust the summary at face value

    Treat every AI-generated summary as an unvetted lead, not a source of truth. The AP’s official guidelines state this explicitly: generative AI output “should be treated as unvetted source material.” If the AP does not trust it, neither should you.

    2. Spot-check against the audio, not just the transcript

    Checking the summary against the transcript is not enough. The transcript itself may contain fabrications. Return to the audio for any quote you plan to publish, any statistic you plan to cite, any claim that seems surprising or consequential. If it matters enough to put in your story, it matters enough to verify against the recording.

    3. Use the three-point check

    When reviewing any AI-generated transcript, listen to at least three sections against the original audio:

    • The first two minutes — to catch initial hallucinations or speaker misidentification
    • A section from the middle — where AI confidence tends to drift and errors accumulate
    • Any passage where a specific number, name, date, or direct quote appears — these are the highest-risk error points

    It will not catch everything. It will catch the errors most likely to end up in print.

    4. Keep your audio. Always.

    Garance Burke, global investigative reporter at the Associated Press, found during her Whisper investigation that at least one organisation using AI transcription had simply discarded the original audio after generating the transcript. If the AI had hallucinated, there was no way to know. The fabricated text was the only record.

    Never delete your recordings until well after publication. The audio is your source of truth. Without it, you have no way to verify what was actually said — and no defence if something goes wrong.

    5. Be especially careful with proper nouns

    AI transcription models are worst at exactly the things journalists need most: names of people, organisations, locations, legislation, and technical terms. These are high-stakes words where a single letter changes meaning. Cross-reference every proper noun against your notes and the original recording.

    6. Watch for sentences that are too clean

    Real speech is messy. People start sentences and abandon them. They pause. Repeat themselves. Speak in fragments. If a sentence in your AI transcript reads like polished prose — grammatically perfect, neatly structured, quotation-ready — that is precisely when you should be most suspicious. The AI may have cleaned up what was said to the point of changing its meaning. Or it may have invented the sentence entirely. In transcription, fluency is not a sign of accuracy. It is sometimes a sign of fabrication.


    The Bigger Picture: Why This Does Not Get Solved Quickly

    AI transcription and summarisation tools are not going to become perfectly reliable anytime soon. The underlying technology — large language models generating text based on probability rather than fact — is architecturally prone to hallucination. Improvements are incremental. The 1% fabrication rate in Whisper was first measured by Cornell researchers on a 2023 version of the model. OpenAI has shipped updates since — but independent testing on newer versions has continued to find hallucinations, and nobody knows how many fabricated passages went undetected before academic researchers started looking.

    Meanwhile, adoption accelerates. The Reuters Institute’s 2024 survey of UK journalists found 56% use AI at work weekly, including 27% who use it daily. Only 16% have never used it. The tools are faster, cheaper, and more convenient than anything that came before.

    Journalists are not going to stop using them. Nor should they have to.

    But the current generation was built for speed, not for trust. Built to generate text as fast as possible — not to demonstrate that the text is accurate. That is a design choice, not a technical limitation. It is entirely possible to build AI tools that flag uncertainty, link every output back to the source audio, and tell you honestly when they do not know what was said.

    Those tools are coming. In the meantime, the only verification system that works is you.


    The Bottom Line

    Your AI transcription tool is not a reliable witness. It is a fast, confident, and occasionally dishonest assistant — one that saves you time only if you verify its work. And right now, the tools give you no help doing that.

    Every published error starts the same way: someone trusted the output without checking the source.

    In a profession where your credibility is your career — where a byline is not just a name but a record of what you were willing to stand behind — that is a risk that compounds quietly, until it does not.

    Check the audio. Keep your recordings. Question the clean sentences. And never let an AI tool be the only record of what someone said.


    Frequently Asked Questions

    Can AI transcription tools hallucinate entire sentences?

    Yes. Cornell University researchers found that OpenAI’s Whisper fabricates content in approximately 1% of transcriptions. These are not mishearings. In documented cases, the model inserted violent language, racial commentary, and fabricated medical claims into transcripts of calm conversations that contained none of those things.

    How common is hallucination in AI-generated summaries?

    More common than most journalists realise. Vectara’s Hallucination Evaluation Model — the industry’s most widely cited benchmark — shows GPT-4o hallucinates in around 1.5% of summaries, Claude-3.7-Sonnet at 4.4%, Claude-3-Opus at 10.1%, and the newest reasoning models all exceed 10% on the harder refreshed benchmark. DeepSeek’s reasoning model (R1) hallucinates at 14.3%, nearly four times the rate of its non-reasoning counterpart (V3) at 3.9% — meaning the models marketed as the smartest are, in some cases, the least faithful to source material.

    Which AI transcription tools are affected by the Whisper hallucination problem?

    Any tool built on OpenAI’s Whisper model is subject to the 1% fabrication rate documented by Cornell researchers. This includes integrations within Otter.ai, Descript, and dozens of smaller platforms. The summary layer, present in most AI transcription tools regardless of the underlying speech-to-text model, carries its own separate hallucination risk across GPT-4, Claude, and other large language models.

    Do AI transcription tools flag when they have hallucinated?

    No. This is the central problem. Current AI transcription and summarisation tools generate output without distinguishing between content they can trace to the source and content they have inferred or fabricated. There is no warning, no uncertainty marker, no visible break in the text. The hallucinated sentence looks identical to the accurate one. Tools built specifically for journalism — such as Veracity — address this by linking every AI-generated sentence back to a timestamped moment in the source audio and flagging claims that cannot be verified.

    Is AI transcription still worth using in journalism?

    Yes — as a starting point, not as a final record. Raw transcription still saves significant time compared to manual typing. The risk is concentrated in the summarisation layer and in treating AI output as verified fact. The AP’s official guidance is clear: generative AI output should be treated as unvetted source material. Used with that standard in place, AI transcription tools remain valuable. The problem is that most journalists have not been told what the risks actually are, or given tools that help them manage those risks efficiently.

    How do I know if my AI transcript contains hallucinated content?

    You cannot know without checking. Fabrications are designed — by the nature of how these models work — to look indistinguishable from accurate transcription. The only reliable method is to return to the original audio for any quote, statistic, or claim that will appear in print. Treating every AI output as unvetted source material — as the AP officially recommends — is the only defensible standard in the absence of a tool that does this automatically.


    How Veracity Solves This

    Veracity is built around a single principle: every AI-generated sentence must be traceable back to the exact moment in the source audio where it came from. No exceptions, no silent hallucinations.

    Here is what that means in practice.

    Click any sentence, hear the source

    When Veracity transcribes your interview and generates a summary, every claim in that summary is clickable. Click it, and you hear the exact seconds of audio where the speaker said it. No scrubbing through a 40-minute recording. No hoping your memory of the interview matches what’s on the page.

    Unverifiable claims are flagged before you see them

    If the AI generates a sentence it cannot trace back to your audio, Veracity marks it [VERIFY] automatically. You know which sentences need your judgement before you’ve published anything. The hallucination that would have slipped past you in Otter or Descript is the one Veracity refuses to let through unmarked.

    Inaudible moments are flagged honestly

    This is the part no other tool does. Once you’ve written your story, you paste it back into Veracity. The AI reads every sentence you wrote against every source in your story workspace — audio, transcripts, notes — and tells you what’s confirmed, what cannot be verified, and what directly contradicts the record.

    When the audio is unclear — background noise, mumbling, crosstalk — Veracity marks it [inaudible 02:25] instead of guessing. A tool that admits when it doesn’t know is more trustworthy than one that confidently invents.

    Your finished article, defended

    Veracity shows its work, flags its uncertainty, and gives you a way to check anything it produces against the source of truth — your recording.

    Veracity is currently in private beta. If you want early access, join the waitlist at veracityai.app.

  • A Senior Editor Was Fired Over AI-Fabricated Quotes—The Lesson Every Journalist Can’t Ignore

    A Senior Editor Was Fired Over AI-Fabricated Quotes—The Lesson Every Journalist Can’t Ignore

    TLDR: Key Takeaways

    •   15 out of 53 Substack posts by former editor-in-chief Peter Vandermeersch were found to contain AI-generated or fabricated quotes, per Columbia Journalism Review (March 2026)

    •   8 different commentators, academics, and journalists were misquoted in a single CJR writeup, with quotes the AI tools had invented

    •   0.7% to over 10% of sentences hallucinated on basic summarisation tasks by leading AI models, per the Vectara Hughes Hallucination Evaluation Model Leaderboard

    •   94% of news consumers want journalists to disclose AI use, but more than a third lose trust in the story when they see that disclosure, per Trusting News research (2026)

    •   49% of UK journalists already use AI for transcription at least once a month, per Reuters Institute research

    •   The tools Vandermeersch used, ChatGPT, Perplexity, and Google NotebookLM, are the same tools most working journalists now rely on for summarising reports and transcripts

    What’s the real lesson from a senior editor being fired over AI-fabricated quotes? Verification still matters more than speed.

    The incident shows that while AI can assist reporting, it can also generate convincing falsehoods. Publishing without checking—even once—breaks the core rule of journalism: if you didn’t verify it, you shouldn’t print it.

    The case every journalist should be reading

    In late March 2026, Columbia Journalism Review published an investigation into a scandal the European press had been quietly circling for weeks. A Dutch freelance reporter, Menno van den Bos, had contacted CJR and the Tow Center for Digital Journalism with a suspicion: a Dutch-language writeup of CJR’s Journalism 2050 issue contained quotes that did not exist anywhere in the original reporting.

    He was right.

    The writeup was by Peter Vandermeersch, former editor-in-chief of NRC Handelsblad and former CEO of Mediahuis Ireland. At the time of the scandal, Vandermeersch held a thought-leadership role inside Mediahuis as a Journalism and Society fellow. His official remit, in his own words, was exploring the responsible use of AI in newsrooms.

    Van den Bos’s investigation found that fifteen of Vandermeersch’s fifty-three Substack posts contained AI-generated or fabricated quotes. In the CJR writeup alone, he had invented quotes from eight separate commentators, academics, and journalists. None of them had said the things attributed to them. None of the quotes existed anywhere else that van den Bos could find.

    Vandermeersch was suspended from his fellowship. On his blog, he admitted that he had used ChatGPT, Perplexity, and Google NotebookLM to summarise lengthy reports, and that he had trusted the outputs to be accurate. Instead, the systems had, in his words, put words into people’s mouths.

    The sentence he wrote next is the one every working journalist should read twice:

    “It is particularly painful that I made precisely the mistake I have repeatedly warned colleagues about: these language models are so good that they produce irresistible quotes you are tempted to use as an author.”

    He had discovered the problem in his own writing a year earlier. Two of his articles had been flagged for containing AI-generated quotes. He did not correct them at the time. The record hardened quietly around the error, and nothing moved until van den Bos went looking.

    Why this case matters more than the others

    There have been plenty of AI-in-journalism disasters in the last two years. A reporter at the Cody Enterprise in Wyoming resigned after fabricating quotes with AI. CNET corrected dozens of AI-written articles for errors including plagiarism and mathematically wrong financial advice. The New York Times issued a public correction in March 2026 and cut ties with a freelance book reviewer who used AI that incorporated unattributed passages from a Guardian piece. The Times’ own union responded with a letter calling management’s AI standards woefully inadequate.

    The Vandermeersch case is different for one reason.

    He was the expert.

    He was not a rookie writer on his first beat. He was not moonlighting. He was not hiding the work from an editor. He was the person whose full-time title placed him on exactly this watch. The AI still got past him.

    That is the line this case draws across the profession: if the person paid to catch the error cannot catch it by reading, then reading is not the defence. It never was. It has only looked like one.

    What AI hallucination actually looks like inside a newsroom

    AI does not fail the way most journalists expect it to. It does not produce broken sentences or obvious nonsense. It produces fluent, confident, grammatically perfect prose that happens to contain things no one ever said.

    A summary might report that a CEO confirmed a merger timeline when what the CEO actually said was that discussions were ongoing. A quote might be placed in the mouth of the wrong speaker, in language they would plausibly use, on a topic they did in fact discuss. A statistic might arrive in a tidy sentence, referenced with authority, and exist nowhere in the transcript.

    This is how error enters the workflow now: not as noise, but as narrative.

    The Vandermeersch quotes were not caught because they sounded wrong. They were caught because a different journalist, working in a different country, went back to the primary source and checked. Until that moment, the fabrications had done what hallucinations always do in good prose. They sat in the text. They looked like reporting. They hardened into record.

    “AI can help increase productivity, but it can’t take responsibility for what gets published. In a newsroom, every claim has to lead back to a real person who knows what they’re talking about and is willing to stand behind it.” — Nick Toso, CEO of journalist discovery platform Rolli, quoted in TVNewsCheck, February 2026.

    The trap underneath the trap

    There is a reason careful journalists still fall into this, and it is not carelessness. It is architectural.

    AI language models do not flag their own uncertainty. Hallucinations do not arrive with hesitation or qualifiers. They arrive in the exact register a journalist is trained to trust: clean, confident, specific, publishable. This is the technical reality behind Vandermeersch’s own phrase, irresistible quotes you are tempted to use as an author. The hallucinations do not flag themselves.

    And the hallucination rates are not marginal. The Vectara Hughes Hallucination Evaluation Model Leaderboard, the industry’s most widely referenced benchmark for grounded summarisation, shows leading models hallucinating between 0.7% and over 10% on simple summarisation tasks where they are explicitly given source material and asked to stick to it. On harder, domain-specific content, rates climb to nearly 19%. A small percentage sounds like an acceptable margin of error. It is not. One misquote is enough to end a career.

    This is the gap Vandermeersch fell through. Not a cognitive failure. A structural one.

    The reader trust paradox

    Journalists using AI today are caught between two statistics that should not coexist but do.

    Research from the nonprofit Trusting News in late 2025 and early 2026 found that 94% of news consumers want journalists to disclose when they have used AI in their work. In the same research, more than a third of readers said they lost trust in a story when they saw that disclosure. Sixty-two percent said newsrooms should only use AI if they have clear ethical guidelines around its use.

    Disclose and lose trust. Withhold and hide something. The profession is being asked to solve a problem it was never given the tools to solve.

    And the exit from this trap is not using less AI. Reuters Institute research found that 49% of UK journalists already use AI for transcription at least once a month. Adoption is a settled question. The open question is whether the AI output can be defended when someone asks how it was produced.

    The only durable answer is verification. Not verification as a promise, but verification as a visible, structural layer inside the tool itself.

    What real verification has to look like

    When journalists say they “verify AI outputs,” they usually mean they read the summary and compare it against what they remember from the interview. That is not verification. That is memory, competing against a system engineered to sound more confident than memory.

    Real verification means traceability. Every AI-generated sentence, every claim in a summary, every quote in a quote bank, every line in a set of action items, should link back to a specific, timestamped moment in the source audio. If the AI cannot trace a statement to the recording, it should say so, in the output, before the journalist reads past it.

    The Vandermeersch workflow had none of this. Documents went into a chatbot. A summary came out. Quotes were lifted from it. The chain from what the source actually said to what appeared under his byline had no checkpoints at all.

    A verification-first workflow has three.

    The first is source-linked outputs. Every sentence the AI produces must be anchored to the exact moment in the recording it came from. Clickable. Hearable. Not summarised in a sidebar, but linked at the level of the individual claim.

    The second is honest failure. When the AI cannot ground a sentence in the source, the sentence must be visibly flagged before the journalist ever reads it. Silence is not accuracy. Silence is the absence of honesty.

    The third is the piece almost no existing tool offers, and the one that would have caught Vandermeersch entirely: the journalist’s own finished writing has to be checked against the sources. Not the AI’s draft. The journalist’s own words, in their own voice, tested sentence by sentence against every recording and document in their research file. Because the most dangerous moment is not when the AI summarises the interview. It is when the journalist, working from a mix of notes and memory and a slightly-wrong AI summary, writes a sentence that is almost right, and nobody catches it before it goes out.

    Veracity is built on exactly this principle. Every AI-generated sentence traces back to the exact audio moment it came from. Claims the AI cannot verify are flagged before the journalist ever sees them. And once an article is written, journalists can paste it back into Veracity and check every sentence against their recordings before they publish.

    The tape is the source of truth. It always has been.

    How to protect your newsroom right now

    If your newsroom is using any AI transcription or summarisation tool today, here is a practical checklist. None of it is theoretical. All of it would have caught Vandermeersch.

    1. Treat every AI-generated quote as unverified until you hear it. Not probably fine. Unverified. Go back to the original recording, find the moment, confirm the words are there, in that order, spoken by the person attributed. If you cannot confirm all three, the quote does not go in the piece.

    2. Never use a chatbot output as your source of truth for quotes. ChatGPT, Perplexity, and NotebookLM are synthesis tools. They are not evidence. If a quote exists only inside a chatbot response and nowhere in your primary material, the quote does not exist.

    3. Verify every proper noun and every number against the recording. Names, figures, dates, titles, and specific claims are the most common sites of AI error. They are also the details that make a misquote legally and professionally dangerous. Check each one.

    4. Ask your current tool what it cannot verify. If it does not have a mechanism to flag unverifiable claims, that silence is not accuracy. It is the absence of a safety layer. Consider tools that are explicit about what they cannot confirm.

    5. Check your own finished writing against your recordings before you file. This is the step almost no workflow currently includes, and the step that matters most. The AI summary is not the highest-risk moment in the process. The moment you file is. Build verification into that moment, not after it.

    6. When you disclose AI use, disclose it specifically. Trusting News research found that reader trust recovers when disclosures explain precisely what the AI was used for and how the output was verified. Generic disclosure loses trust. Specific disclosure keeps it.

    Frequently Asked Questions

    What happened to Peter Vandermeersch?

    Peter Vandermeersch, a former editor-in-chief of NRC Handelsblad and a Journalism and Society fellow at Mediahuis, was suspended in March 2026 after a Dutch freelance journalist found that fifteen of his fifty-three Substack posts contained AI-generated or fabricated quotes. Columbia Journalism Review reported that he had used ChatGPT, Perplexity, and Google NotebookLM to summarise reports and had trusted the outputs without verifying them against the original sources.

    Do AI tools like ChatGPT and NotebookLM actually invent quotes?

    Yes. Large language models generate text by predicting statistically likely sequences of words, not by retrieving verified facts. When asked to summarise a document or extract quotes, they can produce sentences that never appeared in the source. This is called hallucination. The Vectara Hughes Hallucination Evaluation Model leaderboard shows leading models still hallucinating on grounded summarisation tasks, even when given the source material directly.

    Why is it so hard for journalists to catch AI-fabricated quotes by reading?

    Because hallucinations are not written in the register of error. They arrive in the same fluent, confident, publishable prose the AI uses when it is accurate. Vandermeersch’s own phrase, irresistible quotes you are tempted to use as an author, captures why. Reading an AI output and checking if it sounds right is not a defence, because sounding right is what the hallucinations do best.

    What is AI hallucination in journalism?

    AI hallucination in journalism is when a language model generates a claim, quote, or fact that was not present in the source material. Unlike obvious errors, hallucinations appear as well-structured, confident text, making them difficult to catch without going back to the original recording or document. In journalism, this is particularly dangerous because a hallucinated quote can constitute a factual error, a misquote, a damaged source relationship, or a legal liability.

    How can journalists prevent AI-fabricated quotes in their work?

    The only durable defence is a verification-first workflow. Every AI-generated quote must be traceable back to a specific audio timestamp or document passage. Unverifiable claims must be visibly flagged. And the journalist’s own finished writing should be checked against the original sources before publication. This is precisely the architecture Veracity was built around: every AI-generated sentence is clickable and plays the exact audio it came from, and finished articles can be checked against every recording in the story file before they go out.

    Why don’t existing AI transcription tools catch this?

    Tools like Otter, Fireflies, and Descript transcribe audio and generate summaries, but their AI outputs are disconnected from the source at the sentence level. The raw transcript may be clickable, but the AI-generated summary sitting next to it is a standalone block of text with no per-sentence link back to the recording. There is no structural mechanism to verify each sentence against the audio. The gap between transcription and verification is exactly where cases like Vandermeersch’s happen.

    What is provenance-linked AI and why does it matter for journalism?

    Provenance-linked AI is a design pattern in which every sentence a model generates is paired with a direct reference back to the specific source material it came from. For a journalist, that means every claim in a summary, every quote in a quote bank, and every sentence in a draft links back to an exact audio timestamp or document passage. If the AI cannot link a sentence to a source, the sentence is flagged as unverified. This is the standard every newsroom adopting AI should be holding their tools to, and it is the standard Veracity is built on.

    Is AI transcription safe to use in a newsroom?

    AI transcription is safe as a workflow tool when treated as a starting point for verification, not a finished record. Raw speech-to-text is generally more reliable than AI-generated summaries of that text. The risk rises sharply when journalists pull quotes directly from AI summaries, or build their own writing on top of them, without checking each claim against the original recording. Any tool used in a newsroom should make that check efficient, visible, and automatic.

    A different approach

    This is why Veracity exists: an AI workspace designed specifically for journalists, where every AI-generated sentence traces back to the exact moment in your source audio, unverifiable claims are flagged before you ever see them, and once you have written your article in your own words, you can check every sentence against your recordings before you publish.

    One AI-fabricated quote can end a career. March 2026 made that a proven statement, not a hypothetical one. And the person it ended was the person who knew the risk better than anyone.

    The tape is the source of truth. It always has been.

    Veracity is currently in private beta. If you are a working journalist and want early access, join the waitlist at veracityai.app.

  • 30% of AI Summaries Get the Facts Wrong — and Your Newsroom Is Probably Using One

    30% of AI Summaries Get the Facts Wrong — and Your Newsroom Is Probably Using One

    TLDR: Key Takeaways

    • ~Nearly 30% of summaries generated by early neural abstractive systems contained unsupported or fabricated facts — a problem the field has been grappling with ever since
    • 51% of AI answers to news questions have significant accuracy issues, per BBC research (February 2025)
    • 13% of quotes AI systems sourced from BBC articles were either altered or did not exist in the original
    • More than half (60%) of UK journalists are ‘extremely concerned’ about the possible negative impact of AI on public trust in journalism.
    • 49% of UK journalists already use AI for transcription monthly — but no mainstream tool verifies AI-generated sentences against source audio
    • Zero of the major AI transcription tools designed for general use were built specifically for journalism

    Why do AI summaries get facts wrong so often? Because they’re designed to sound coherent—not to guarantee accuracy.

    Research shows that around 30% of AI-generated summaries drift from the source material, often in subtle ways that are hard to spot . And broader studies suggest the problem can be even worse, with AI misrepresenting news content nearly half the time in some cases . The issue isn’t just blatant errors—it’s missing context, overstated claims, or small distortions that change the meaning.

    The number every journalist needs to know

    Around one in three AI-generated summaries quietly drifts from the truth of the source material. Not in obvious ways, not in lines that announce themselves as wrong—but in subtle shifts of meaning, misplaced emphasis, and statements that sound plausible enough to pass without question. This isn’t a hypothetical flaw. It is a measured failure rate, observed in the very systems journalists are rapidly folding into their daily workflows.

    And yet, across newsrooms, these tools are being absorbed into daily workflow as though accuracy were already a settled question.

    A recording goes in. A summary comes out. Quotes are lifted from neatly packaged “banks.” The story moves forward.
    But somewhere in that chain, a question goes unasked: Did this actually come from the tape?

    Because in most cases, there is no reliable way to know. No clear path back to the moment something was said. No signal when the AI has filled a gap instead of reflecting the record.

    Just polished output—easy to use, harder to trust.


    What AI hallucination actually looks like in a newsroom

    AI does not misspell words or produce garbled sentences. It produces fluent, confident, well-structured text that happens to contain claims nobody actually made.

    A summary might confidently report that a CEO confirmed a merger timeline, when in reality the CEO said only that discussions were ongoing. It might place a quote in the mouth of the wrong speaker. It might introduce a statistic that sounds credible, precise—even quotable—but was never uttered in the recording.

    This is how error enters the workflow now: not as noise, but as narrative.

    Unless the journalist returns to the tape and verifies each claim line by line—a process that erases the very efficiency AI promises—those fabrications can pass quietly into publication. There is no alert, no marker of uncertainty, no visible break in the workflow. The error doesn’t announce itself. It moves through the process unnoticed, until it becomes part of the record.


    Why journalists are using tools built for other professions

    There is no shortage of AI transcription and meeting tools. But almost none were built with journalism in mind. They were designed for sales calls, internal syncs, product interviews—environments where speed matters, but precision is negotiable.
    Journalists have adopted them anyway. Not because they fit the work, but because there has been nothing better to reach for.

    The difference is not marginal—it’s structural. A sales team can absorb a flawed summary because the record is not the product; the outcome is. Deals move forward through follow-ups, clarifications, and ongoing communication. Small inaccuracies are corrected in the next call, the next email, the next interaction.

    Journalism does not have that margin for error. Once a line is published, it stands as a record. A single misquoted sentence can trigger a correction, strain a source relationship, or escalate into a legal risk.

    The stakes are categorically different, but the tools are identical.

    None of these general-purpose tools offer what journalists actually require:

    • The ability to click on any AI-generated sentence and hear the exact moment in the source audio that supports it
    • Explicit flags when the AI cannot trace a claim back to the recording
    • A distinction between a sentence the AI is confident about and one it generated to fill a gap

    They produce polished output—and push the burden of verification entirely onto you. Journalists are left retrofitting tools that were never built for their standards, filling the gaps with time and caution.

    That is not a workflow. It is risk, absorbed quietly.


    The data is worse than most editors realise

    Across 2025 and 2026, independent studies converge on the same conclusion: AI tools are failing on accuracy at a rate that is incompatible with the standards journalism demands.

    BBC Research (February 2025)

    The BBC tested four prominent, publicly available AI assistants against their own published journalism. The findings:

    • 51% of all AI answers to questions about the news had significant issues of some form
    • 19% of AI answers that cited BBC content introduced factual errors — incorrect facts, numbers, and dates
    • 13% of quotes attributed to BBC articles were either altered or did not exist in the original article

    Columbia Journalism Review Testing

    When the Columbia Journalism Review tested how well AI tools perform on real journalism tasks, they found that every tool they evaluated underperformed against the human benchmark in generating accurate long summaries. These were not obscure products—they were the tools journalists use every day, on deadline.

    AssemblyAI Research on Abstractive Summaries

    Research into abstractive AI summarization—the type most transcription tools use to generate interview summaries — shows a ~30% rate of factual misalignment with source material. Abstractive summaries do not simply extract and repeat what was said; they paraphrase and synthesize, which is where the errors compound.

    Center for News, Technology and Innovation

    A 2025 CNTI report on AI transcription and translation in journalism states what most journalists already know: AI outputs cannot be trusted without human review.

    But in the current workflow, that responsibility sits unsupported. The tools generate. The journalist verifies. And the space between the two remains entirely manual.


    Other industries already solved this

    What makes this gap harder to ignore is that another profession, facing the same stakes, has already moved past it.

    In legal technology, where the margin for error is just as narrow, tools are built around traceability. Platforms like Filevine Depo CoPilot link AI-generated insights directly to transcript lines and synced audio, allowing lawyers to click a claim, hear the testimony, and assess context instantly.

    Journalism, by contrast, is still working with tools that generate first and leave verification behind. The technology is already here. Research highlights that trust remains the central challenge, and that AI systems must be designed to support more transparent and verifiable news production. What is shifting is not whether verification matters, but how visibly it is demonstrated to audiences.

    Journalists, for their part, do not need persuading. Verification is not an added step—it is the work. What is changing is the expectation that the tools they rely on should meet the same standard.


    What real verification looks like

    When people talk about “verifying AI outputs,” they often mean scanning the summary and weighing it against what they think they remember from the interview. That is not verification. It is recall under pressure—an unreliable check against a system designed to sound certain even when it is wrong.

    Real verification is not interpretation—it is traceability. Every AI-generated sentence—every claim, every quote, every extracted point—should lead you back to a precise, timestamped moment in the source audio. And where that link does not exist, the system should say so plainly, rather than smoothing over the gap with language that only appears certain.

    This is not a radical idea. It is the same standard journalists apply to their own reporting: show your sources. The only difference is applying that standard to the tools journalists use, not just the articles they write.

    And it cannot end with AI outputs. The highest-risk moment in a journalist’s workflow is not when the tool produces a summary—it is the moment just before publication, when the reporting has been shaped, the narrative set, and the byline attached. That is where errors harden into record: a quote recalled from memory rather than the tape, a number shifted in phrasing, an attribution that reads cleanly but does not hold up.

    Veracity is built on exactly this principle. Every AI-generated sentence traces back to the exact audio moment it came from. Claims the AI cannot verify are flagged before the journalist ever sees them. And once an article is written, journalists can check every sentence against their recordings before they publish.


    The trust crisis is already here

    The Reuters Institute’s research paints a stark picture: 62% of UK journalists are extremely concerned about AI’s impact on public trust in media. And from the audience side, only 12% of people are comfortable with news produced entirely by AI.

    These numbers reflect a profession caught in a contradiction. Journalists are adopting AI faster than ever—49% of UK journalists already use AI for transcription on a monthly basis—but neither they nor their audiences trust the output.

    Trusting News found that even transparent disclosure of AI use can erode audience trust. The signal alone is not enough. What audiences are responding to is not the presence of AI, but the absence of proof.

    The issue is not disclosure itself. It is that AI-assisted journalism still lacks a credible way to demonstrate that what has been published is actually grounded in the source material.

    The question is not whether AI-generated errors will make it into published journalism. The question is how many already have.


    How to protect your newsroom right now

    If your newsroom uses any AI transcription or summarization tool today, here is a practical checklist:

    1. Treat all AI summaries as first drafts, not sources. Do not pull quotes from them. Go back to the timestamped audio and confirm every line. A summary is a navigation aid, not a record.

    2. Verify proper nouns and numbers every time Names, figures, dates, and titles are the most common sites of AI error. Check every one against the original recording or document.

    3. Ask your tool what it cannot verify. If it cannot flag unverifiable claims, that silence is not accuracy—it is a failure of the tool. Use systems that are explicit about uncertainty.

    4. Check AI-generated content before it reaches your editor, not after Errors caught before filing are corrections avoided. Build a verification step into your workflow before submissions, not as a post-publication process.

    5. Demand provenance from your tools Any tool worth using in a newsroom should be able to show you, for any AI-generated sentence, exactly where in your source material that sentence came from. If it cannot, it should not be in your workflow.


    Frequently Asked Questions

    How accurate are AI transcription tools for journalism?

    AI transcription accuracy varies by tool and audio quality, but the accuracy problem is most severe in AI summaries, not raw transcripts. AssemblyAI research shows approximately 30% of abstractive AI summaries contain statements that do not align with the source material. The BBC found that 51% of AI answers to news questions had significant accuracy issues, and 13% of AI-attributed quotes were either altered or fabricated.

    What is an AI hallucination in journalism?

    An AI hallucination in journalism is when an AI tool generates a plausible-sounding claim, quote, or fact that was not present in the source material. Unlike obvious errors, hallucinations typically appear in fluent, well-structured text — making them difficult to catch without manually verifying against the original audio or document. In journalism, this is particularly dangerous because a hallucinated quote or attribution can constitute a factual error, a misquote, or in some cases a legal liability.

    Do AI transcription tools flag inaccurate summaries?

    Most mainstream AI transcription tools do not flag potentially inaccurate summaries. They generate text without distinguishing between claims they can trace to the source and claims they have inferred or fabricated. Tools built specifically for journalism—such as Veracity—take a different approach: every AI-generated sentence is linked to a timestamped moment in the source audio, and claims that cannot be verified are explicitly flagged before the journalist sees them.

    What percentage of AI summaries contain errors?

    Research on abstractive AI summarization—the method used by most AI transcription tools—indicates that approximately 30% of summaries contain factual inaccuracies relative to the source material. Separately, BBC research found that 51% of AI answers to news-related questions had significant issues, and 19% of AI answers that cited BBC content introduced new factual errors.

    How can journalists verify AI-generated quotes?

    The most reliable method is to treat every AI-generated quote as unverified until you have heard the original audio at the correct timestamp. This means going back to your recording, finding the cited moment, and confirming that the AI’s version accurately represents what was said—including context, speaker attribution, and wording. This process is what Veracity automates: every AI-generated sentence is clickable, and clicking it plays the exact source audio it was drawn from.

    Why do AI tools produce inaccurate summaries?

    AI summarization tools use abstractive methods—they paraphrase and synthesize rather than simply extracting text. This produces more readable summaries but introduces the risk of semantic drift, where the AI’s restatement of a claim shifts its meaning. Combined with the fact that most tools were trained on general-purpose data rather than journalistic standards, and that they have no mechanism to self-check against source audio, inaccuracies enter the output unannounced.

    Is AI transcription safe to use in a newsroom?

    AI transcription is safe as a workflow tool when treated as a starting point for verification rather than a final record. Raw transcription (speech-to-text) is generally more reliable than AI summarization. The risk rises significantly when journalists pull quotes directly from AI summaries without checking them against the original recording. The CNTI recommends that human review remain a critical step in any AI-assisted journalism workflow.


    A different approach

    This is why Veracity exists—an AI workspace designed specifically for journalists, where every AI-generated sentence traces back to the exact moment in your source audio, unverifiable claims are flagged before you ever see them, and once you have written your article in your own words, you can check every sentence against your recordings before you publish.

    The tape is the source of truth. It always has been.

    Veracity is currently in private beta. If you are a working journalist and want early access, join the waitlist at veracityai.app.

    Approximately 30% of abstractive AI summaries contain factual inaccuracies—and most newsrooms have no reliable way to catch them before publication. This article examines where these failures occur, what the research shows, and what real verification looks like in a working newsroom.