Tag: Transcriptiontoolsforjournalists

  • Why AI Transcription Tools Are Lying to You — And How to Protect Yourself

    Why AI Transcription Tools Are Lying to You — And How to Protect Yourself

    OpenAI’s Whisper fabricates content in approximately 1% of transcriptions—and that is only the transcript layer. This article examines what AI transcription tools are doing to journalists’ source material, why the verification savings are smaller than advertised, and what you can do to protect your work right now.

    TLDR: Key Takeaways

    • OpenAI’s Whisper fabricated content in approximately 1% of transcriptions in the version tested by Cornell researchers in 2023. OpenAI has shipped updates since, but independent testing has shown hallucination problems persist across newer Whisper versions. Whisper powers Otter.ai, Descript, and dozens of other tools journalists use daily
    • GPT-4o hallucinates in around 1.5% of summaries on Vectara’s leaderboard; Claude-3-Opus at 10.1%; reasoning models do worse, not better — DeepSeek R1 hallucinates nearly four times as often as its non-reasoning counterpart
    • 53% of CNET’s AI-generated financial articles contained factual errors — approximately 25 times the typical newspaper correction rate
    • Sports Illustrated published product reviews under fabricated author personas with AI-generated headshots. The Arena Group lost its SI publishing licence and laid off the magazine’s entire staff in January 2024
    • The Reuters Institute found journalists who use AI most frequently are more likely to believe they spend too much time on low-level tasks

    Why are AI transcription tools “lying” to you? Because they don’t just mishear—they sometimes invent words that were never said.

    Modern speech-to-text systems can be highly accurate, but they still break down in real-world conditions—accents, background noise, or unclear audio can trigger errors or even full “hallucinations,” where the AI generates entirely fabricated phrases . And because these systems are built on probabilistic models, they aim to produce plausible text—not guaranteed truth .

    What nobody tells you when you start

    Here is what the tools do not warn you about: the transcript you are reading may contain sentences that were never spoken.

    Not mishearings. Not garbled syllables or words dropped in noise. Entire fabricated passages, invented by the AI and inserted without seam into your transcript, carrying no marker and no hesitation. They sit alongside the accurate lines with the same formatting, the same confidence, the same air of having been said.

    If you record interviews, press conferences, or source conversations and rely on AI tools to capture them, this is not a theoretical risk to keep at the back of your mind. It is a documented, measured, and increasingly well-understood failure that the tools themselves do nothing to signal.

    The transcript is clean. The story moves forward. Somewhere in the text is a sentence nobody said.


    The Whisper Problem: When the AI Invents What Was Said

    In June 2024, researchers led by Allison Koenecke at Cornell University published findings at the ACM FAccT conference — later expanded in an October 2024 Associated Press investigation by Garance Burke and Hilke Schellmann, a co-author of the original study — confirming what some journalists had started to suspect: OpenAI’s Whisper speech-to-text model fabricates content in roughly 1% of transcriptions on the version tested earlier — and while OpenAI has since released updates that reduce the rate, independent researchers have repeatedly demonstrated the underlying problem persists across newer versions.

    1%. A figure that sounds small until you understand the infrastructure it runs on.

    Whisper is not a niche tool at the margins of the industry. It is the engine underneath many of the transcription services journalists use daily. At that scale, a 1% fabrication rate translates into millions of invented passages, distributed across thousands of users, sitting inside transcripts that look indistinguishable from the real thing.

    And these are not subtle errors. The researchers found that Whisper does not mishear words. It invents sentences. In documented cases, it inserted violent language into transcripts of calm conversations. In others, it fabricated racial commentary or medical claims from audio that contained nothing of the sort.

    The AI was not guessing incorrectly. It was creating fiction — and presenting it as fact, in the same format, with the same confidence, as everything around it.

    For a journalist, the consequence is precise: a source could appear to have said something they never said. And your transcript would give you no reason to question it.

    The industry has begun to respond. In December 2025, OpenAI released new transcription models that, by the company’s own measurement, produce around 90% fewer hallucinations than Whisper v2 in noisy conditions. Improvement is real — but even the slightest hallucination can cause a lot of damage.


    It Gets Worse at the Summary Layer

    Most journalists do not stop at the raw transcript. The real promise of these tools is the summary: upload the recording, receive a clean account of what was discussed, pull the best quotes, and begin writing. It is where the time savings are most visible. It is also where the second, larger source of error lives.

    Vectara’s Hallucination Evaluation Model — the most widely cited benchmark for AI summarisation accuracy — shows that even the best-performing large language models hallucinate at rates that cannot be dismissed. On the current leaderboard, GPT-4o hallucinates in around 1.5% of summaries; Claude-3.7-Sonnet at 4.4%; Claude-3-Opus at 10.1%. Gemini-3-Pro hits 13.6% on the harder, refreshed benchmark.

    The most counterintuitive finding: turning on reasoning makes models worse at sticking to the source. DeepSeek’s reasoning model (R1) hallucinates at 14.3% on Vectara — nearly four times the rate of its non-reasoning counterpart (V3) at 3.9%. The same pattern shows up across providers. GPT-5, Claude Sonnet 4.5, Grok-4, and Gemini-3-Pro — all marketed as more intelligent than the models that came before — every one of them exceeds 10% hallucination on the harder benchmark.

    The models are not getting more reliable at grounded summarisation. The ones marketed as the smartest are, in some cases, getting worse.

    AI summarisation tools regularly generate plausible-sounding text that does not match what was actually said. They do it confidently. Without flagging. In a format designed to look authoritative.

    Chris Middleton, technology journalist at Diginomica, put it plainly after discovering that Otter’s AI summary had inserted a specific statistic — “1.2 million” — that was never uttered by any speaker in the meeting:

    “Put simply, I can no longer trust Otter to record basic facts, and then present them to me. Instead, I am getting an AI-powered confection that claims to be a precis of a conversation I have taken part in myself.”

    For a journalist, this is not an inconvenience. It is reality being rewritten: quietly, seamlessly, without a single visible break in the text.


    The Real-World Damage Is Already Here

    This is not a future risk. It has already ended careers, bankrupted publishers, and destroyed editorial credibility that took years to build.

    CNET: A 53% error rate

    In late 2022, CNET quietly published 77 AI-generated financial articles. When Futurism investigated in January 2023, roughly 41 of those articles — 53% — contained factual errors requiring corrections, including basic mathematical mistakes in financial calculations. The error rate was approximately 25 times higher than the typical newspaper correction rate of 2–3%.

    Sports Illustrated: Fabricated personas, fabricated reviews

    In November 2023, Futurism exposed product reviews at Sports Illustrated published under fabricated author personas with AI-generated headshots, licensed from a third-party content company called AdVon Commerce. Futurism’s sources said the articles themselves were AI-generated; SI’s parent company denied that specific claim but acknowledged the fake bylines.The Arena Group’s CEO was fired. In January 2024, Arena Group missed a $3.75 million quarterly licence payment, lost its Sports Illustrated publishing licence, and laid off the magazine’s entire unionised staff.

    A 70-year-old institution, gutted.

    The Washington Post: Shipped anyway

    When the Washington Post launched ‘Your Personal Podcast’ — AI-generated personalised podcasts based on its journalism — in December 2025, internal testing showed 68–84% of scripts were deemed unpublishable. The product shipped anyway. An anonymous editor told Semafor it was “truly astonishing that this was allowed to go forward at all.”

    Ars Technica: The AI reporter caught by AI

    In February 2026, Benj Edwards — then senior AI reporter at Ars Technica, a journalist whose entire beat was understanding these systems — used an AI tool to extract quotes from a blog post by engineer Scott Shambaugh. The AI returned paraphrased versions of Shambaugh’s words. Edwards published them as direct quotes. The story was retracted within 48 hours. Edwards was fired by the end of the month. His response on Bluesky: “The irony of an AI reporter being tripped up by AI hallucination is not lost on me.”

    If a senior AI beat reporter cannot catch the errors, what chance does a general assignment journalist on deadline have?


    The Verification Tax: Why the Time Savings Are Smaller Than Advertised

    Before AI transcription, journalists typically spent three to six hours of manual work for every hour of recorded audio. AI compressed that to near-real-time.

    Journalists using AI transcripts often spend roughly half an hour to over an hour per hour of recorded interview reviewing, correcting, and cross-referencing the output against the original audio. Verifying exact quote accuracy. Confirming speaker identification in multi-person recordings. Correcting mangled proper nouns and technical terms. Checking context to ensure nothing is used misleadingly.

    A Reuters Institute survey of UK journalists found that more frequent AI users are actually more likely to believe they spend too much time on low-level tasks. The researchers’ explanation is pointed: AI use generates new, AI-specific low-level tasks — cleaning data, checking outputs — that did not exist before.

    The finding that follows from the same study is one the industry should sit with longer than it has: the journalists most satisfied with their time spent on creative work are those who do not use AI at all.

    Brian Merchant, author of Blood in the Machine, named the dynamic plainly:

    “Many journalists are finding that the supposed efficiencies generated by such uses are often offset by new tasks, like rechecking transcriptions for accuracy, ensuring AI texts are free of hallucinations, and editing output for clarity.”

    AI transcription did not eliminate the work. It redistributed it. Instead of typing during the interview, you are now scrubbing audio after the interview to check whether the AI captured it correctly. For many journalists — especially freelancers without institutional fact-checking support — the net time savings are far smaller than advertised.

    This is the verification tax. And it is one every journalist currently using these tools is quietly paying, whether they have named it or not.


    How to Protect Yourself Right Now

    Until tools are built that address this problem at the level of architecture, here is what you can do today.

    1. Never trust the summary at face value

    Treat every AI-generated summary as an unvetted lead, not a source of truth. The AP’s official guidelines state this explicitly: generative AI output “should be treated as unvetted source material.” If the AP does not trust it, neither should you.

    2. Spot-check against the audio, not just the transcript

    Checking the summary against the transcript is not enough. The transcript itself may contain fabrications. Return to the audio for any quote you plan to publish, any statistic you plan to cite, any claim that seems surprising or consequential. If it matters enough to put in your story, it matters enough to verify against the recording.

    3. Use the three-point check

    When reviewing any AI-generated transcript, listen to at least three sections against the original audio:

    • The first two minutes — to catch initial hallucinations or speaker misidentification
    • A section from the middle — where AI confidence tends to drift and errors accumulate
    • Any passage where a specific number, name, date, or direct quote appears — these are the highest-risk error points

    It will not catch everything. It will catch the errors most likely to end up in print.

    4. Keep your audio. Always.

    Garance Burke, global investigative reporter at the Associated Press, found during her Whisper investigation that at least one organisation using AI transcription had simply discarded the original audio after generating the transcript. If the AI had hallucinated, there was no way to know. The fabricated text was the only record.

    Never delete your recordings until well after publication. The audio is your source of truth. Without it, you have no way to verify what was actually said — and no defence if something goes wrong.

    5. Be especially careful with proper nouns

    AI transcription models are worst at exactly the things journalists need most: names of people, organisations, locations, legislation, and technical terms. These are high-stakes words where a single letter changes meaning. Cross-reference every proper noun against your notes and the original recording.

    6. Watch for sentences that are too clean

    Real speech is messy. People start sentences and abandon them. They pause. Repeat themselves. Speak in fragments. If a sentence in your AI transcript reads like polished prose — grammatically perfect, neatly structured, quotation-ready — that is precisely when you should be most suspicious. The AI may have cleaned up what was said to the point of changing its meaning. Or it may have invented the sentence entirely. In transcription, fluency is not a sign of accuracy. It is sometimes a sign of fabrication.


    The Bigger Picture: Why This Does Not Get Solved Quickly

    AI transcription and summarisation tools are not going to become perfectly reliable anytime soon. The underlying technology — large language models generating text based on probability rather than fact — is architecturally prone to hallucination. Improvements are incremental. The 1% fabrication rate in Whisper was first measured by Cornell researchers on a 2023 version of the model. OpenAI has shipped updates since — but independent testing on newer versions has continued to find hallucinations, and nobody knows how many fabricated passages went undetected before academic researchers started looking.

    Meanwhile, adoption accelerates. The Reuters Institute’s 2024 survey of UK journalists found 56% use AI at work weekly, including 27% who use it daily. Only 16% have never used it. The tools are faster, cheaper, and more convenient than anything that came before.

    Journalists are not going to stop using them. Nor should they have to.

    But the current generation was built for speed, not for trust. Built to generate text as fast as possible — not to demonstrate that the text is accurate. That is a design choice, not a technical limitation. It is entirely possible to build AI tools that flag uncertainty, link every output back to the source audio, and tell you honestly when they do not know what was said.

    Those tools are coming. In the meantime, the only verification system that works is you.


    The Bottom Line

    Your AI transcription tool is not a reliable witness. It is a fast, confident, and occasionally dishonest assistant — one that saves you time only if you verify its work. And right now, the tools give you no help doing that.

    Every published error starts the same way: someone trusted the output without checking the source.

    In a profession where your credibility is your career — where a byline is not just a name but a record of what you were willing to stand behind — that is a risk that compounds quietly, until it does not.

    Check the audio. Keep your recordings. Question the clean sentences. And never let an AI tool be the only record of what someone said.


    Frequently Asked Questions

    Can AI transcription tools hallucinate entire sentences?

    Yes. Cornell University researchers found that OpenAI’s Whisper fabricates content in approximately 1% of transcriptions. These are not mishearings. In documented cases, the model inserted violent language, racial commentary, and fabricated medical claims into transcripts of calm conversations that contained none of those things.

    How common is hallucination in AI-generated summaries?

    More common than most journalists realise. Vectara’s Hallucination Evaluation Model — the industry’s most widely cited benchmark — shows GPT-4o hallucinates in around 1.5% of summaries, Claude-3.7-Sonnet at 4.4%, Claude-3-Opus at 10.1%, and the newest reasoning models all exceed 10% on the harder refreshed benchmark. DeepSeek’s reasoning model (R1) hallucinates at 14.3%, nearly four times the rate of its non-reasoning counterpart (V3) at 3.9% — meaning the models marketed as the smartest are, in some cases, the least faithful to source material.

    Which AI transcription tools are affected by the Whisper hallucination problem?

    Any tool built on OpenAI’s Whisper model is subject to the 1% fabrication rate documented by Cornell researchers. This includes integrations within Otter.ai, Descript, and dozens of smaller platforms. The summary layer, present in most AI transcription tools regardless of the underlying speech-to-text model, carries its own separate hallucination risk across GPT-4, Claude, and other large language models.

    Do AI transcription tools flag when they have hallucinated?

    No. This is the central problem. Current AI transcription and summarisation tools generate output without distinguishing between content they can trace to the source and content they have inferred or fabricated. There is no warning, no uncertainty marker, no visible break in the text. The hallucinated sentence looks identical to the accurate one. Tools built specifically for journalism — such as Veracity — address this by linking every AI-generated sentence back to a timestamped moment in the source audio and flagging claims that cannot be verified.

    Is AI transcription still worth using in journalism?

    Yes — as a starting point, not as a final record. Raw transcription still saves significant time compared to manual typing. The risk is concentrated in the summarisation layer and in treating AI output as verified fact. The AP’s official guidance is clear: generative AI output should be treated as unvetted source material. Used with that standard in place, AI transcription tools remain valuable. The problem is that most journalists have not been told what the risks actually are, or given tools that help them manage those risks efficiently.

    How do I know if my AI transcript contains hallucinated content?

    You cannot know without checking. Fabrications are designed — by the nature of how these models work — to look indistinguishable from accurate transcription. The only reliable method is to return to the original audio for any quote, statistic, or claim that will appear in print. Treating every AI output as unvetted source material — as the AP officially recommends — is the only defensible standard in the absence of a tool that does this automatically.


    How Veracity Solves This

    Veracity is built around a single principle: every AI-generated sentence must be traceable back to the exact moment in the source audio where it came from. No exceptions, no silent hallucinations.

    Here is what that means in practice.

    Click any sentence, hear the source

    When Veracity transcribes your interview and generates a summary, every claim in that summary is clickable. Click it, and you hear the exact seconds of audio where the speaker said it. No scrubbing through a 40-minute recording. No hoping your memory of the interview matches what’s on the page.

    Unverifiable claims are flagged before you see them

    If the AI generates a sentence it cannot trace back to your audio, Veracity marks it [VERIFY] automatically. You know which sentences need your judgement before you’ve published anything. The hallucination that would have slipped past you in Otter or Descript is the one Veracity refuses to let through unmarked.

    Inaudible moments are flagged honestly

    This is the part no other tool does. Once you’ve written your story, you paste it back into Veracity. The AI reads every sentence you wrote against every source in your story workspace — audio, transcripts, notes — and tells you what’s confirmed, what cannot be verified, and what directly contradicts the record.

    When the audio is unclear — background noise, mumbling, crosstalk — Veracity marks it [inaudible 02:25] instead of guessing. A tool that admits when it doesn’t know is more trustworthy than one that confidently invents.

    Your finished article, defended

    Veracity shows its work, flags its uncertainty, and gives you a way to check anything it produces against the source of truth — your recording.

    Veracity is currently in private beta. If you want early access, join the waitlist at veracityai.app.