← Writing

Why your transcript repeats the same phrase over and over

·6 min read

You open a transcript and halfway down the page it turns into this:

thank you so much thank you so much thank you so much thank you so much thank you so much …

for a screen, or ten. Nobody in the recording said "thank you so much" forty times. Somewhere around that point, the recogniser stopped listening and started repeating itself.

This is one of the stranger failures in automatic transcription, and it isn't specific to one product. It has a name — repetition looping — and it's worth understanding, because the obvious version is easy to see and the quieter versions are not.

What a loop actually is

Many modern speech recognisers, Whisper among them, write text the way language models do: one piece at a time, each piece chosen partly from the audio and partly from what has already been written. Most of the time the audio does the steering. But when the audio stops giving it much to go on — music, a long silence, noise, a stretch of speech it can't make out — what's already been written takes over. And the most likely thing to come after "thank you so much" in a transcript that already contains "thank you so much" is, as far as a language model is concerned, "thank you so much".

Once it's in that groove, it can stay there for a long time. It isn't transcribing the stretch of audio anymore; it's continuing a pattern.

It isn't one vendor's bug

OpenAI's open-source Whisper is the best-documented case, because its code is public and the defences are visible in it. Its transcribe() function treats a window of output as failed if the text compresses too well — a gzip compression ratio above 2.4 by default — because a page of one phrase repeated compresses extremely well. A window that fails that check is decoded again with more randomness, stepping the temperature up from 0.0 through 1.0.

The same file says what makes them more likely. Feeding the previous window's text in as a prompt for the next one is on by default, and the documentation for that option says turning it off makes the model "less prone to getting stuck in a failure loop, such as repetition looping or timestamps going out of sync."

We've seen it in the cloud recogniser we use too. On 4 October 2026 two slices of real uploads came back with exactly this shape. In one, a four-word phrase appeared 1,908 times, and the last word was timestamped at 10,326 seconds on a slice that was 960 seconds long. The other slice had nine minutes with no words in them at all — the same failure from the other side.

The quieter versions

The forty-line wall is easy. These are the ones that get past people:

  • Short loops. A phrase repeated six or eight times reads like someone stuttering, or like a speaker who really did say "yeah, yeah, yeah". It isn't always obvious which it is.
  • Loops that replace a gap. The text doesn't repeat for long, but everything said during the next few minutes is simply missing. The transcript reads cleanly across the join.
  • Timestamps that run away. The words look fine, but their times stop matching the audio, or run past the end of the recording. Click a line to play it and you land somewhere else entirely.
  • A stretch with no words. Not a loop, but the same breakdown: minutes of speech that produced nothing, which looks exactly like minutes of silence.

How to catch it

You don't need to read the whole transcript to find these. Three checks do most of the work:

  1. Look at the density. People speak at a fairly steady rate. A minute of transcript with three times the words of the minutes around it is a loop; a minute with none, while the recording has sound in it, is a gap. If your tool shows timestamps, this is a quick skim.
  2. Check the end. The last word's time should be close to the recording's length. A last word timed well past the end means the timestamps went out of sync somewhere.
  3. Play the suspicious spots. Wherever a phrase repeats more than a few times, play that moment. It takes seconds, and it's the only way to tell a real "yeah, yeah, yeah" from a stuck model.

What to do with the part it swallowed

The important thing to know is that a loop isn't an error you can fix by editing the text. The words the recogniser repeated are not the words that were said, and the words that were said were never written down. Deleting the repeats leaves you with a clean-looking transcript with a hole in it.

So, in order of effort:

  • Run it again. Sometimes a second attempt goes through. With the same tool and settings it can also fail the same way; if it does, move on to the next step rather than retrying again.
  • Cut out what triggers it. Long stretches of music, hold tones or dead air before and after the part you care about give the model nothing to anchor to. Trimming them helps.
  • Transcribe the gap separately. If one section keeps failing, export just those minutes and transcribe them on their own.
  • Listen and type. For a quote that matters, the recording is the record. A few minutes of manual transcription beats any amount of trusting a transcript that already failed there once.

What we do, and where it stops

When the recogniser stalls while separating speakers on a slice of audio and we send that slice again without speaker labels — which is how the 4 October failures first showed up — we check what comes back for loops and for sound with no words. A loop (a phrase repeated eight or more times, faster than anyone speaks) is cut back to one copy, and a notice on the transcript gives the time of each place and says that "the recogniser looped on one phrase instead of transcribing what was said". Those words are missing from the transcript, not from the recording, so you know which moments to play. In the two slices above, that cut 13,473 words to 1,989 in one and 2,153 words to 11 in the other.

We don't run that check on every recording, and the reason is worth saying. Over 468 ordinary slices of real uploads, the same checks would have flagged a loop in 8 and a 30-second-plus stretch with no words in 40. Without listening, we can't tell those apart from a song or a quiet room. A warning that fires on nearly one slice of audio in ten teaches people to ignore it on the one where it's true, so for now it runs only where the failure has been measured. On everything else, the three checks above are yours to do — and a vendor's accuracy figure, measured on someone else's audio, says nothing about where your file failed (what "99% accurate" actually means).

Neither the check nor the notice brings back the missing words. They tell you where to listen.

Judge it on your own recording.

Ours is the only transcript we can vouch for, and it isn't yours. Run something real through it and see where it holds.

Start free