This post is about how to use AI transcription tools in a way that actually serves the archive — where to let the tool do its job, where to catch what it misses, and how to produce a transcript that is both efficient to create and accurate enough to trust. The tool gives you speed. You supply the meaning.
Free Resource
Free: Start Your Oral History Project
Get our Quick-Start Checklist — 10 steps to your first recorded interview.
No spam. Unsubscribe anytime.
Why AI Transcription Is Worth Using (and Where It Falls Short)
Transcription fatigue is one of the most common reasons oral history projects stall. A family collects recordings — sometimes dozens of hours — and then nothing happens to them, because the prospect of transcribing by hand is overwhelming. AI tools break that logjam. A one-hour interview that would take three to four hours to transcribe manually can produce a working draft in under five minutes. That is a real difference, and it matters for getting preservation work done.
But AI transcription fails in specific, predictable ways — and knowing the failure modes before you start is the most important preparation you can do.
- Dialect. Tools trained primarily on standardized American or British English will misread AAVE, Gullah, Haitian Creole, Appalachian speech, and other regional or community varieties — sometimes badly. Words get substituted for phonetically similar standard-English words. Grammatical structures that don't match the training data get flattened or rewritten.
- Proper nouns and place names. A family farm named Bellwood, a church called Greater New Jerusalem, a neighborhood known locally as "the Bottom" — none of these are in the training data. The tool will guess. It will often guess wrong.
- Overlapping speech. When two people speak at once, most AI tools either pick one voice or produce a garbled blend of both. Group interviews and informal family conversations are particularly vulnerable.
- Background noise and low-volume speech. Elders who speak quietly, recordings made in noisy kitchens or during family gatherings, audio where someone is not facing the microphone — all of these degrade accuracy significantly.
These are not reasons to avoid AI transcription. They are reasons to go into it knowing that a raw AI transcript is a first draft, not a finished record — and to build your review process accordingly.
The Tools That Actually Work for Oral History
There are dozens of transcription services available. Four are worth knowing in concrete terms for oral history work.
Otter.ai. Best for clear English-language recordings. Strong on standard speech, affordable for families (free tier available, paid plans start around $10/month), and the transcripts are editable directly in the browser — you can correct errors without downloading anything. Speaker identification works reasonably well for two-person interviews. Struggles with dialect, noise, and any speech that diverges significantly from mainstream American English.
Whisper (OpenAI, open source). The strongest performer on dialect and non-standard speech of any widely available tool. Whisper was trained on a much broader range of audio than most consumer services, which makes it substantially more accurate on AAVE, code-switching, accented speech, and community varieties. It is free, but the standard version requires some technical comfort to run locally. For non-technical users, Whisper Web (a browser-based version) and whisper.cpp (a simplified local install) make it accessible without requiring programming knowledge. Worth the extra setup step for any recording with dialect content.
Rev.com. Offers both AI and human transcription. The AI service is comparable to other tools; the human transcription service — where a professional transcriptionist listens and types — is significantly more accurate for difficult audio and costs around $1.50 per minute. For a 60-minute interview, that is $90. Expensive for routine use, and worth every dollar for a recording you cannot afford to get wrong: the only recording of a dying elder, a rare dialect, an interview where the audio quality is genuinely poor. The human option is not a luxury — for certain recordings, it is the only responsible choice.
Descript. A strong editing interface — the ability to edit text and have the audio change accordingly is genuinely useful for producing clean access copies. Less accurate on dialect than Whisper, and not recommended as a primary transcription engine for oral history with non-standard speech. Better understood as an editing and export tool than a transcription tool: use something else to generate the first draft, then bring it into Descript to produce a polished version.
Practical recommendation: Start with Otter or Whisper Web for most recordings. Use Rev's human transcription for any recording where accuracy cannot be compromised — a rare account, a once-recorded voice, a language or dialect that AI tools will mangle. The cost of getting that transcript wrong is higher than the cost of paying for it to be right.
How to Prepare a Recording Before You Run It Through AI
This is the most overlooked step in AI transcription workflows. Most transcription errors are not produced by the tool — they are produced by the recording, and many are preventable before the file ever reaches the service.
- Trim dead air at the start and end. Many AI tools calibrate their audio models using the first few seconds of a recording. If those seconds are silence, shuffling, or background noise before anyone speaks, the tool may underperform on the audio that follows. Trim the recording to start within a few seconds of first speech.
- Convert to a standard format if needed. WAV or high-bitrate MP3 (192kbps or above) will produce better results than compressed formats like .m4a from a phone recording app. Most transcription services accept common formats; when in doubt, export as WAV.
- Add a spoken header if re-recording is possible. Before uploading, if you have the ability to add to the recording, record a brief spoken introduction at the start: "This is [your name], interviewing [subject name], on [date], in [city]." This gives the tool calibration audio and produces a header that will appear at the top of the transcript.
- Use noise reduction sparingly, if at all. Noise reduction software can improve intelligibility for speech that hews close to standardized English — but for dialect-heavy recordings, it can actually degrade transcription accuracy by altering the phonetic characteristics the model is trying to parse. If you are going to use noise reduction, test it on a short excerpt first and compare results with and without.
One rule that applies without exception: never process the original file. Always work from a copy. The original recording is the archive. Every edit, trim, or format conversion happens to a duplicate. If anything goes wrong, the original is untouched.
How to Review and Correct an AI Transcript Without Losing Accuracy
A raw AI transcript is a first draft. It is not a finished record, and should not be stored as one. The review process is what separates a useful archival document from a file that will mislead whoever reads it in twenty years.
Listen at 0.75x speed while reading. Do not read a transcript without listening to the recording at the same time. The ear catches errors the eye normalizes. Slow the playback slightly — 0.75x is fast enough to stay efficient and slow enough to catch most substitutions and missed words.
Mark every proper noun for manual verification. Names of people, places, churches, farms, organizations, and neighborhoods should be flagged and verified against other sources — family members, other interviews, local records. The tool guessed. You need to confirm.
Do not skip [inaudible] or [unclear] flags. These markers often appear at exactly the moments that matter most — a name whispered, an answer the subject hesitated over, a phrase in a language the tool didn't recognize. Go back to those moments in the audio. If you still cannot make it out, document it as [inaudible — timestamp] with a note on context. A documented gap is far more useful than a confident wrong word.
Mark dialect corrections in brackets. When the AI has substituted a standard-English word for what was actually spoken, correct the transcript but document the substitution: [word as spoken: "finna" — AI transcribed as "gonna"]. This preserves the evidence of what was actually said and creates a record of where the tool erred.
Do not "correct" dialect into standard English. This is the most consequential error a reviewer can make. Changing "he don't never" to "he never" changes the record — not just the grammar, but the voice, the identity, and in some cases the meaning. If a family member wants a cleaned-up version for sharing, produce it as a separate access copy and label it clearly. The archival transcript preserves the speech as it was spoken.
Add a documentation header to every transcript file. At the top of the document, before the transcript begins: subject name, interviewer, date, location, language(s) spoken, and transcription method. Example: AI-assisted: Otter.ai, reviewed by [your name], July 3, 2026. This header is what makes the transcript a citable archival document rather than an anonymous text file.
The Complete Method for Recording, Documenting, and Archiving What You Capture
The Oral History Documentation Guide gives you transcript templates, metadata frameworks, finding aid worksheets, and a step-by-step system for turning a raw recording into a properly documented archival record.
Specific Challenges — Dialect, Multilingual Recordings, and Low-Quality Audio
Dialect (AAVE, Gullah, Creole, regional speech)
Whisper performs best of the widely available tools on dialect-heavy recordings. When reviewing any transcript produced from dialect speech, check specifically for these distortions, which appear most often:
- -in vs. -ing endings: "workin'" transcribed as "working," or vice versa with the opposite substitution.
- Double negatives flattened: "don't know nothing" rewritten as "don't know anything."
- Copula deletion: "She real quiet" transcribed as "She's real quiet" — the AI inserts the missing "is" because it expects it.
- Idiomatic phrases: Community-specific expressions treated as transcription errors and substituted with phonetically similar standard phrases.
Flag each of these with a bracketed correction note. Never fix them silently — the documented correction is part of the record.
Multilingual and code-switching recordings
AI transcription tools handle code-switching poorly. Most will detect what they determine to be the dominant language of a recording and attempt to apply it uniformly — which means anything said in the secondary language either gets phonetically approximated in the dominant language or drops out entirely. The result can look deceptively complete.
Best practice: run the recording through Whisper with language detection enabled (its multilingual model handles language switches better than most alternatives) and treat the output as a very rough first draft. Review will take longer. Recruit a bilingual listener for the sections the tool could not handle. For a full guide to recording and working with multilingual oral history, see How to Record Oral History in a Second Language.
Low-quality or noisy audio
For genuinely difficult audio — significant background noise, low recording volume, unclear speech — Rev's human transcription service is the only reliable option. AI tools will hallucinate words when audio quality drops below a certain threshold: the output looks clean, with no [unclear] flags, because the model filled in what it expected to hear rather than acknowledging what it could not parse.
Know the difference between a low-confidence transcript and a hallucinated one. A low-confidence transcript is obvious — it is messy, full of [unclear] markers, with obvious gaps and substitutions. A hallucinated transcript looks clean and professional and may be substantially wrong. If you have difficult audio and the AI output looks suspiciously polished, listen carefully. The cleaner it looks, the more carefully you should verify it against the recording. The hallucinated transcript is worse than the messy one, because it is harder to detect.
What AI Transcription Can't Replace
There are things an AI transcription tool will not give you, no matter how good the output is. It will not tell you that there was a long pause before an answer — and that the pause was the most important part of what was said. It will not know what your grandmother meant when she said "back then," or what name she chose not to say when she got to a certain part of the story. It will not recognize that the elder said a place name that is only used by people who grew up there, and that the standard spelling it produced is not the right one.
AI transcription is a time-saving tool for getting words onto a page. It is not an interpretation. It is not a context layer. It is not a memory. The most important documentation work — noting what the elder's tone shifted on, what question seemed to land differently, what seemed left unsaid — happens after the transcript is produced, in the annotations and notes you add alongside it.
The transcript is the access copy. The recording is the archive. Your annotation of what mattered, what was significant, what requires context — that is the part that makes the document oral history rather than just a text file. The AI gets words onto the page. Everything that gives those words meaning comes from you.
The tool gives you speed. You supply the meaning.
Tools for the Work That Comes Before and After
The Family Oral History Workbook ($14.99) gives you question banks, worksheets, and interview guides for families recording elders — so the recordings you run through these tools are worth transcribing.
The Oral History Documentation Guide ($19.99) gives you the complete archival method — transcript templates, metadata frameworks, naming conventions, and a system for turning a raw recording into a properly labeled, findable record.
The Thronateeska Historical Continuity Project is a public education initiative focused on Indigenous continuity, Afro-Indigenous identity, oral history traditions, and ancestral continuity. A transcript is only as good as the care taken to produce it.
More from this project
- How to Transcribe an Oral History Interview Without Losing Your Mind
You have the recording. Now you need a usable transcript. A practical, honest guide to transcription tools, AI auto-transcription, formatting standards, and what to do with the document after — for family historians who have never done this before.
- How to Record Oral History in a Second Language
When an elder speaks Spanish, Haitian Creole, Gullah, Yoruba, Garifuna, or another language — and you don't fully understand it — you record anyway. A practical guide for bilingual and multilingual families navigating oral history across a language gap.
- How to Preserve Oral History Recordings for 100 Years
The format you recorded on today will be obsolete before your grandchildren are born. Long-term preservation is not about buying the right hard drive — it's about building a system that migrates, verifies, and multiplies copies across decades.
- How to Record High-Quality Audio on Your Phone for Family Oral History
The phone in your pocket is good enough — if you know how to use it. A practical guide to environment, positioning, app selection, and settings for recording family oral history on a smartphone.