14 September 2026
Why transcripts get speakers wrong, and how we fixed it
We found a real meeting transcript that merged two speakers into one label, and rebuilt speaker detection to catch it.
A four person meeting came back from transcription with three speakers. One person's turns were folded into someone else's label for the full 31 minutes. The summary described that person twice, under two names, in adjacent paragraphs. Nobody caught it until Demian, who was in the room, read it back.
That's the failure that started this. Diarization, the part of transcription that figures out who's talking, is hard. AssemblyAI's own docs call it a best effort, not a guarantee. We had been treating it like one.
Speakers with names, not letters
Transcripts with more than one voice now show Speaker A, Speaker B, and so on in the transcript view. Click a label and give it a real name. The rename carries through the transcript, the summary, and the action points, including summaries generated before you renamed anyone.
That took real work. A summary is written once, in prose, by a model that only knows "Speaker A". Renaming can't rewrite text that's already been written, and writing it again would charge you a second time for the same recording. So we substitute names at read time, everywhere a summary gets served. Rename someone once and every version of that summary reflects it.
Telling us the truth ahead of time helps
The meeting that started all this had four people in the room. Left alone, AssemblyAI found three. Told up front that four people were on the recording, it found all four instead. We tested this on the actual failing recording before we shipped it: no hint gave three speakers, a wrong hint gave six and a garbled transcript, and the right count gave four and a clean one.
So the upload form now asks, once, if you know it: how many people are in this recording? Skip it if you're not sure. A wrong guess is worse than no guess, so we never pre-fill an answer.
And when we can't tell, we say so
Some recordings don't split cleanly, no matter what we're told. When that happens, the summary drops speaker attribution instead of presenting a guess as fact. A wrong attribution in a summary is invisible until it matters. The transcript stays labeled, with a plain note that some voices couldn't be told apart, and you can switch labels off with one click.
None of this changes the words in your transcript. It changes how straight we are with you about who said them.
Upload a recording and see it for yourself: audio in, transcript, summary and action points in your inbox, priced per minute.