Don’t Whisper it
HOW MUCH RISK do you face when you use tools such as machine audio transcription? It, like machine translation, is related to “generative AI” but distinct from it. The Freelance does use both of these – so we were very interested indeed to find the wittily-titled academic paper Careless whisper: speech-to-text hallucination harms.

A photorealistic landscape image of an artificial intelligence in court being prosecuted after lying about what someone has said, according to ChatGPT – which yesterday reported an “internal error” every time I mentioned an AI in court
The authors – Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X. Mei, Hilke Schellmann and Mona Sloane – studied the Whisper service from Open AI, which they describe as “a state-of-the-art automated speech recognition service outperforming industry competitors, as of 2023”.
To be clear: they say (on the fifth page of their paper) that they “found no evidence of hallucinations in [smaller tests of] competing speech recognition systems”.
But in Whisper they found that “roughly 1 per cent of audio transcriptions contained entire hallucinated phrases or sentences which did not exist in any form in the underlying audio... 38 per cent of hallucinations include explicit harms such as perpetuating violence, making up inaccurate associations, or implying false authority.”
That might seem like a small risk. But consider this: if 10,000 journalists each used Whisper just once, on average 300 or 400 would get confabulated text containing what the authors describe as “explicit harms”. Who knows how many might end up in court?
Methodology
The authors detected what they call “hallucinations” by comparing the same audio segments when run through Whisper twice in close succession – once in April 2023, and once in May 2023. They used 7805 audio segments from normal speakers, and 5335 from people with aphasia to test a hypothesis about how the errors happen. The Freelance prefers not to use the term “hallucination” because it ascribes a state of mind to the technology of predictive text on steroids; “confabulation” risks that, but refers rather precisely to how the prediction happens.
On average, they found that “1.4 per cent of transcriptions in our dataset yielded hallucinations. We categorize these hallucinations... and find that among the 312 hallucinated transcriptions, 19 per cent include harms perpetuating violence, 13 per cent include harms of inaccurate associations, and 8 per cent include harms of false authority.“ As the authors say, this likely under-counts confabulations, not least when both runs produce the same error.
An example, with the violent made-up text underlined:
“And he, the boy was going to, I’m not sure exactly, take the umbrella. He took a big piece of across. A teeny small piece. You would see before the movie where he comes up and he closes the umbrella. I’m sure he didn’t have a terror knife so he killed a number of people who he killed and many more other generations that were y𝐾 paï𝐻. And he walked away.”
And an example of “false authority” – referencing a website that (we found with some trepidation) does not exist:
“This is a picture book telling the story of Cinderella. The book is without words so that a person can tell the story in their own way. To learn more, please visit SnowBibbleDog.com.”
The Freelance can only repeat the advice we give for machine translation: when using any “AI” tool, make sure that a human who understands the input checks the output. Yes, every word of it.
'Artificial Intelligence' our coverage to date
![[Freelance]](../gif/fl3H.png)