October 4th, 2026

How to transcribe research interviews for your thesis: from audio to content analysis

Practical guide to transcribing thesis research interviews without six hours of typing per hour of audio: AI transcription, review, pseudonyms, appendix vs. annex and the categories of content analysis.

Rodrigo Carvalho Rodrigo Carvalho

How to transcribe research interviews for your thesis: from audio to content analysis

You recorded 6, 8, 10 interviews of 40 minutes to an hour each for your thesis — and just found out that manually transcribing a single hour of audio takes a trained researcher 4 to 6 hours of typing. Do the math: that’s weeks of work that produce not a single line of analysis.

Transcription is the part of a thesis that eats the most time and earns the least credit. This guide shows you how to get through it fast and with rigor — and how to come out of it with the categories of your content analysis already in hand.

Before the tool: what your methodology requires of the transcript

Not every study needs the same kind of transcription. This determines how much work you’re in for — and it’s worth deciding before you type a single word.

Clean (or normalized) transcription: fixes verbal tics, repetitions, and slang while keeping the literal meaning of what was said. It’s the accepted standard for thematic analysis and content analysis — most thesis research based on interviews falls here.

Verbatim transcription with notation: records pauses, hesitations, overlapping speech, and intonation. It’s only required when your theoretical framework analyzes speech itself (conversation analysis, discourse analysis focused on enunciation). If that’s your case, the notation — timed pauses, overlap markers — is manual work, done by listening to the audio.

How do you know which one you need? It’s in your theoretical framework. If your technique is Bardin-style content analysis, clean transcription is enough. When in doubt, ask your advisor before transcribing — it’s the difference between weeks and months of work.

Step by step: transcribing the interview with AI

1. Start with quality audio

No AI can rescue a bad recording. Record with your phone close to the interviewee, in a quiet room, and run a 30-second test before you start. The most common causes of a wrong transcript aren’t the model — they’re noise, microphone distance, and overlapping voices. (I covered this in detail in why the problem is almost always the audio.)

2. Import the file

In Sintesy, the path is the main dashboard: the New → Upload file button, or drag the file anywhere on the screen. Supported audio formats include mp3, wav, m4a, aac, ogg, opus, and flac — the same ones phone recorders and recording apps produce.

Two limits matter here: files up to 20 MB on the free plan and up to 2 GB on the Pro plan. A one-hour interview recorded on a phone usually fits comfortably; if your recording is very long, export it in a compressed format like m4a or opus.

On the free plan you get 60 minutes of transcription per day — enough to process one interview per day and review it right after. If you’re up against a deadline, the Pro plan processes everything at once.

3. Separate the researcher from the interviewee

The transcript comes out with speaker separation: every utterance is tagged with who said it. This saves you the most tedious part of manual transcription — and the part that matters most for analysis, because you’ll be citing who said what.

Rename the labels to something that reads like research: P for your questions, E1 for the first interviewee. Do this during the review, together with the next step.

4. Review against the audio — always

An automatic transcript isn’t ready for a thesis as-is. Where it fails is predictable: proper names, place names, technical terms from your field, and regional slang. Efficient reviewing doesn’t mean listening to everything again: read the full transcript marking anything in doubt, and go back to the audio only at those points.

For methodological rigor, use the protocol examining committees love: randomly sample 2 or 3 segments per interview and check them word by word against the audio. If they match, you have quality evidence to describe in your methodology chapter.

5. Export and standardize

Export as DOCX (a Pro plan feature) and keep one file per interview, with a standardized name: interview-01-e1-2026-09-28.docx. You’ll build the analysis on top of this set of transcripts — and clearly named files save hours when your advisor asks for “interview 4”.

Organize your transcripts for analysis day

Three habits separate a usable transcript from a pile of text:

Pseudonyms from day one. Replace names with E1, E2, E3 during the review — don’t leave it for later. Keeping real names “just for now” is the most common way a real name leaks into the thesis appendix. Keep the correspondence table in a separate file, outside your work.

Context in a header. Each transcript opens with date, duration, location, and recording method. When you cite “E1, 34, public school teacher,” that card is already filled in.

Documented ethics. If your research involves human subjects and went through the Brazilian research ethics system (CEP/CONEP committees — CNS Resolution 466/2012, or 510/2016 for the human and social sciences), your informed consent form must authorize the recording. Re-read what your form approves: sending the audio to a transcription service is data processing, and it’s worth confirming that it’s covered by what the participant signed.

From transcripts to categories: Bardin in practice

Bardin’s content analysis runs in three phases — and the good news is that the first one starts while you’re still reviewing.

Phase 1 — Pre-analysis: a floating reading of the material, following the rules of exhaustiveness (nothing in the corpus is left out), homogeneity (all the material answers the same research question), and relevance. In plain terms: read everything, skipping no interview.

Phase 2 — Exploration of the material: the text is cut into recording units — segments that answer your research problem — and grouped by analogy into categories. Categories must follow the principles of mutual exclusion (a segment doesn’t live in two categories), homogeneity (a category gathers segments of the same kind), and relevance (connected to your theoretical framework).

Phase 3 — Treatment and interpretation: the organized categories become the results analysis section, with verbatim quotes from the interviews illustrating each one.

Where AI helps without invading the method: in Sintesy, every transcript has a chat where you ask in natural language — “in which segments does E3 talk about demotivation?” — and the answer comes back with the segments from the transcript. It’s the same principle as the PDF chat for scientific articles, applied to your interviews. Mind maps and the per-source summary show the theme structure of each speaker.

What stays yours: the decision about the categories. AI points out recurrences; the rigor of grouping, naming, and interpreting belongs to the researcher — and that is exactly what your examining committee evaluates.

Does the transcript go in the appendix or the annex?

A classic question from anyone formatting their thesis at 2 a.m. Under the Brazilian ABNT standard (NBR 14724), the rule is a single one — and the same logic applies to APA and ISO styles:

  • Appendix: material produced by you — the transcripts of your interviews, the question guide you built, the category table.
  • Annex: material from third parties — the technical standard you cited, the official document, another author’s article.

So your transcripts go in the appendix — identified by a capital letter (Appendix A, Appendix B, one per interview). The semi-structured question guide too. And remember: quotes in your analysis chapter appear in the body of the text with the pseudonym and speaker reference (E5, 41), never with real names.

One practical decision before you attach everything: in many studies, the advisor asks for only one full transcript as a sample in the appendix, with the rest documented in the methodology. That trims dozens of pages off the thesis. Agree on it with your advisor.

FAQ

How long does it take to transcribe a one-hour interview? By hand, 4 to 6 hours per hour of audio — qualitative research methodology references converge on that range. With AI, the transcript comes out in minutes and a careful review takes 30 to 60 minutes per interview. The bottleneck stops being typing and becomes analyzing — which is the work that earns the grade.

Can I submit the automatic transcript without reviewing it? No. Names, numbers, technical terms, and slang are where AI slips — and they’re exactly what the examining committee notices. Reviewing is part of the protocol, and describing how you reviewed in the methodology chapter is what turns a tool into a method.

Do I need a verbatim transcription, with pauses and hesitations? Only if your framework analyzes speech itself (conversation analysis, for instance). For thematic content analysis, clean transcription is accepted and used in most theses. Confirm with your advisor before you start.

Can AI generate the categories ready-made? It lists recurrences and points out emerging themes with the segments — a map of the terrain. But naming, grouping, and interpreting categories is the core of your role as a researcher. Use it as a magnifying glass, not as an author.

To clear your transcription debt today

The work of a thesis built on interviews isn’t typing — it’s analyzing. Every hour you don’t spend pausing audio in Word is an hour of reading your framework, writing your analysis chapter, or talking with your advisor.

Start with the longest interview, the one you’ve been putting off. Upload the audio, review the names, swap in pseudonyms, and export. Tomorrow, one more. In a week, the heaviest phase of your thesis is behind you — and what’s left is the part that earns the grade.