September 13th, 2026

Wrong transcript? The problem is almost always the audio

Bad audio produces bad text — and no AI can fix what it never managed to hear. Here are the four factors that decide transcription accuracy, what can still be salvaged, and how to improve the audio you record today.

Rodrigo Carvalho Rodrigo Carvalho

Wrong transcript? The problem is almost always the audio

You transcribed a 40-minute meeting and the text came out full of holes. Names swapped, sentences cut off, passages that make no sense. The natural reaction is to blame the tool — and sometimes the tool is at fault. But in most cases, the error you see in the text was born earlier: in the audio that reached the model.

Speech recognition has an awkward trait. It has no spare context to break the tie on what it heard. A person who misses a syllable in the middle of a conversation fills it in from the subject, from the face of whoever spoke, from what was said before. The model has only the sound wave. If part of it isn’t there, no inference brings the right word back.

That’s why investing in recording pays off more than switching tools.

The error you see in the text was born in the audio

There’s a simple thought experiment. Take the problematic passage from your transcript and listen to the 10 seconds that match it. In almost every case you’ll hear exactly the same problem: someone talking far from the microphone, two people overlapping, a fan in the background, a phone bumping against the table.

The model got right what was audible. It faithfully transcribed audio that was already ambiguous.

That changes the question. Instead of “what’s the best transcription tool?”, the useful question is “how do I deliver audio that doesn’t need guessing?”

The four factors that decide transcription

It isn’t dozens of variables. It’s four, and all of them are more predictable than they look.

FactorWhat’s behind itHow much it weighs
Background noise and reverberationAir conditioning, street, glass-walled room, echoA lot — it’s the champion of errors
Distance from the microphonePhone in the middle of the table, laptop far away, big roomA lot — it drops off along with the distance
Overlapping speechMore than one person talking at the same timeHigh — and unrecoverable
Subject and vocabularyProper names, acronyms, technical and regional termsMedium — depends on context

The order matters. A wrong proper name you can compensate for by reading the transcript carefully. A sentence that was never captured, you can’t.

What clean audio buys you

The number that sums all of this up is the word error rate (WER): the fraction of words the system gets wrong relative to what was actually said.

The founding work on Deep Speech 2 (Amodei et al., 2016) trained the same model on different volumes of audio and measured the result. Even on the noise-free passages, word error fell from roughly 29% with 120 hours of training to around 8.5% with 12,000 hours. On the noisy passages, the same curve ran between roughly 51% and 14%. More data improved both cases almost proportionally — but the gap between them never closed. Under noise, the error stayed consistently around 60% higher.

The detail that matters to anyone recording a meeting is in another table of the same work. In American English with no noise, the error sat around 8%. On samples with an Indian accent, it went past 20%. On the passages with real noise from the CHiME corpus, 21.6%. On the passages with simulated noise, 42.5%.

In plain terms: the same system, in the same language, making five times as many mistakes depending on how the audio was captured. The variable wasn’t the model’s intelligence. It was the microphone and the room.

Noise: what software can fix and what it can’t

Much of the noise left over in a transcript gets fixed afterward — filters, stationary noise suppression, volume normalization. That’s built into most serious tools, Sintesy included.

The problem is what doesn’t get fixed.

Overlapping speech is information that’s already been lost. No filter separates two voices that became a single wave.

Amodei and other researchers in the field are explicit on this point. Awni Hannun, one of the authors of Deep Speech 2, wrote a 2017 essay called, precisely, “Speech Recognition Is Not Solved”: much of the progress on conversational speech came from datasets where each participant had their own microphone, with no overlapping voices. On far-field audio, with several people on the same channel, separation is still an open problem.

It’s the same reason phone calls transcribe better than a meeting room: one microphone per person. Shorter distance, zero overlap.

What you control in 30 seconds

None of this requires expensive gear. It requires three decisions before you hit record.

Get closer to the source. A phone 30 centimeters from your mouth makes far fewer mistakes than a laptop in the corner of a glass-walled room. If the meeting is in person with several people, a recorder in the middle of the table beats the phone of whoever’s sitting at the far end. If it’s just your voice — a lecture, an idea, a voice note — record up close and don’t worry about the rest.

Turn off what makes noise. Air conditioning, an open window, notifications, a cooling fan. Constant noise is exactly what filters remove best, but removing it isn’t the same as it never having existed. Every layer you take out improves everything that comes after.

Avoid overlapping speech. This is organization, not technique. In a large meeting, hand over the floor — and wait for the answer to end before starting yours. It sounds like etiquette advice, but it’s the most underrated quality factor. No software brings back what two voices erased together.

Worth remembering that this guidance is for whoever records. If you receive audio that’s already done — from a client, a student, a colleague — the audio is what it is. In that case the adjustment comes later, at reading time.

The words that still come out wrong even with good audio

There’s a class of error that survives any improvement in recording: the kind where the only way to get it right would be to know what the conversation was about.

Proper names above all. “Marcelo” and “Marcela” are nearly identical sounds. The name of a company the model has never seen, an internal acronym, a project’s nickname. In a perfect acoustic booth, with the cleanest possible audio, those words are still a gamble.

Domain vocabulary counts too. A medical student dictating an anatomical term, a lawyer citing a case name, an engineering team talking about a service with an internal name. Speech research has always treated this as a problem of context, not acoustics: whoever transcribes Portuguese well gets Brazilian names right more often, because they know the list of Brazilian names.

What that changes in practice is where you put your attention after transcribing. In the content passages — arguments, decisions, explanations — the text tends to be reliable. On names and acronyms, it’s worth checking. And it’s worth using your own material: if you’ve already transcribed five lectures from the same professor, that transcript is the best dictionary of that vocabulary there is.

Cleanup: rereading the transcript instead of trusting the statistic

A common trap is treating the error rate as a seal of approval. “The tool is 98% accurate, so everything’s fine.”

It isn’t. Hannun himself gives the example that takes that confidence apart: with a 5% word error rate and 20-word sentences, the chance of at least one word coming out wrong in each sentence is practically total. The text looks good, it reads smoothly, and it can still have one error per sentence.

Worse: not every error weighs the same. Swapping “Tuesday” for “today” changes the meeting’s decision. Dropping a “to” changes nothing. That’s why review shouldn’t be a uniform pass over everything — it should follow what you’re going to do with the text.

If the transcript is going to underpin a decision, check numbers, dates and names. If it’s going to become study material, check the technical terms. If it’s going to become a summary for the team, check the action items. The rest is a quick read.

How to do this in Sintesy

Quality work happens at two ends, and Sintesy tries to cover both.

At capture, native recording was designed precisely to pick up voice at a reasonable distance — the idea being that whoever is speaking doesn’t have to press their mouth to the phone, which helps at a meeting table and in a classroom. Stationary background noise is filtered before processing.

At organization, the real gain shows up when you stop treating the transcript as one solid block of text. A synthesis in Sintesy opens the same material in several formats: the full searchable transcript, the structured summary, the mind map and the checklist. Instead of auditing the thousands of words of running text, you look at the summary first — if a name or a number is wrong there, it’s probably wrong in the transcript too, and you trace the source by searching.

And there’s the cumulative effect. After transcribing a few materials on the same subject, from the same person or the same team, asking the chat about new material gets easier than reading the whole transcript. You already have the reference for the vocabulary and the context — which is exactly where the error hides.

Before you blame the tool

One last shortcut is worth it, because it saves a new subscription and an entire migration.

Listen to 10 seconds of the original audio in the passage where the text failed. If there were already two voices at once, room echo, or speech arriving low and distant, you’ve found the cause — and it’s closer to how you record than to where you transcribe.

If you want the full step-by-step on capture and transcription, there’s a dedicated guide here: How to transcribe audio with AI: complete guide. And if the audio you want to improve is an in-person meeting recorded on a phone, the specific workflow is in how to transcribe an in-person meeting recorded on your phone.

Clean audio isn’t technical fussiness. It’s the difference between having a record you can check and having a record you can trust without checking.