txscribe
← Writing

Preparing interview transcripts for thematic analysis

·6 min read

Thematic analysis starts long before the first code. The transcript you code from quietly decides what you can find: whether hesitations are there to notice, whether you can tell the participant from the interviewer at a glance, whether you can go back and hear how something was said.

Automatic transcription has made the typing part fast. It hasn't made these decisions for you, and some of its defaults are wrong for research. This is a checklist for getting interview transcripts into shape before analysis, written for people doing qualitative work — theses, evaluations, user research — with recordings rather than a transcription budget.

Decide how verbatim you need to be

Before transcribing anything, decide what the transcript has to preserve, and write it down so every interview is done the same way.

  • Verbatim keeps the fillers, false starts and repetitions, and — depending on the conventions you adopt — marks pauses, laughter and overlapping speech. You need it when how something is said is part of your data: discourse or conversation analysis, or work on hesitation and emotion.
  • Intelligent (clean) verbatim keeps every word that carries meaning and drops the "um"s and restarts. Thematic analysis that works at the level of meaning often uses it, but whether it's enough depends on your question — a hesitation or a restart can itself be meaningful.

Machine transcripts land somewhere in between, and not by design: recognisers often tidy away fillers and false starts, and they can't be relied on to mark pauses, laughter or crosstalk consistently. If your method needs those, you'll be adding them by hand while you listen. If it doesn't, the machine's version is a reasonable starting point — but check that it hasn't tidied away something meaningful, like a long hesitation before an answer.

Know who said what

You'll usually want to tell participants' words from the interviewer's — whether you code the questions too is a choice your approach should make, not the transcript — so every turn needs a speaker.

Automatic speaker labels help, but they're a grouping of voices, not an identification. They come back as "Speaker 1" and "Speaker 2", and they're most often wrong at short interjections and where people talk over each other. In a focus group, with several similar voices, expect to spend real time on this.

The fix is mechanical: rename each speaker, then listen through and reassign the turns that landed on the wrong person. In txscribe you rename a speaker once and every turn follows; there's more detail in how to transcribe an interview with speaker names. Check before you start that your language gets speaker labels at all. Not every one does, and without them the transcript comes back unlabelled: there are no speakers to rename, and you'd be marking turns by hand.

Use participant codes, not names

Rename speakers to the codes in your data management plan — P01, P02, INT for the interviewer — rather than to real names. It's the same one step, and it means a transcript you export or share never carries a participant's name in its speaker labels.

It is not anonymisation on its own. Names, places, employers and anything else identifying that people say are still in the text. Pseudonymise those as a separate pass, before the transcript goes into your analysis software or to a co-coder, and keep the key somewhere separate.

Keep a way back to the audio

A code applied to a line of text is an interpretation of something someone said. When you're unsure — was that sarcasm? a question? — the answer is in the recording, not the transcript.

Keep timestamps. Exported with a time at the start of each speaker's turn, a transcript lets you or a co-coder find the moment in seconds, long after the analysis software has separated the text from its audio. In txscribe, Word and text exports carry a timestamp at each turn, and inside the app any line plays back from its own second.

Correcting is familiarisation

In Braun and Clarke's widely used approach to thematic analysis, the first phase is familiarising yourself with the data — reading and re-reading, and listening. Researchers who typed their own transcripts got that for free, and some worry that automatic transcription takes it away.

It doesn't have to. Listening through the recording while you correct the transcript is familiarisation: you hear every interview in full, with the text in front of you, and you're paying the close attention that fixing errors demands. Note first impressions as you go. What you shouldn't do is skip the listening because the transcript looks finished — a clean-looking transcript can still have wrong words and missing ones, and you'll be building themes on top of them.

Two things make that pass faster:

  • A custom word list for the terms your participants use — names of organisations, local places, jargon — so they're more likely to be right first time instead of being wrong in every interview.
  • Consistent formatting across interviews, so that importing into your analysis software is the same every time.

Getting it into your analysis software

Most qualitative analysis tools import Word documents and plain text. A few habits make the import cleaner:

  • One interview per file, named with the participant code and date.
  • Speaker names at the start of each turn, spelled identically in every file, so participant text can be separated from the interviewer's.
  • Timestamps kept, so a coded passage still leads back to the recording.

txscribe exports DOCX and TXT with the speaker and time heading each turn, on every plan. You can also group an interview set into a collection to keep a study's recordings together.

Settle the ethics questions first

Before a single recording goes into any cloud service — ours included — check that your ethics approval, consent forms and data management plan allow it. Participants consented to something specific, and "an online transcription service" may or may not be inside it. Things to find out for any service you're considering:

  • Who processes the data. Our privacy policy lists every processor: Google Cloud services do the transcription and storage, and Google's Gemini on Vertex AI produces summaries, translations and answers from the transcript.
  • Whether it's used for training. We do not use recordings or transcripts to train models.
  • How deletion works, and how long anything lingers. When you delete a transcript, is the recording deleted too — and the translations, summaries and anything else made from it? How long do backup or recoverable copies survive? Get the answer in writing and put it in your data management plan.

If your institution requires data to stay on your own machines, a cloud service isn't the right tool, whatever its policies say.

A checklist

  1. Choose verbatim or clean verbatim, and write down your conventions.
  2. Confirm your ethics approval covers the transcription route.
  3. Transcribe, naming the language rather than leaving it on automatic.
  4. Rename speakers to participant codes and fix misattributed turns.
  5. Listen through while correcting — and treat that as familiarisation, with notes.
  6. Pseudonymise what people said, separately from the speaker labels.
  7. Export one file per interview, with speakers and timestamps, and import.

If you'd like to see what the first draft of a transcript looks like on one of your own recordings, you can try txscribe without an account.

Judge it on your own recording.

Ours is the only transcript we can vouch for, and it isn't yours. Run something real through it and see where it holds.

Start free