Audio Transcription Services

Human transcripts of interviews, depositions, dictation and evidentiary audio — in 326 languages.

Last reviewed August 23, 2026 by the Prism Linguistics editorial team

Audio transcription services turn a recording into an accurate written transcript. Prism Linguistics transcribes interviews, depositions, medical dictation, focus groups, podcasts and evidentiary audio in 326 languages, using human transcriptionists rather than software alone. Most files come back within 2 to 5 business days, and quotes arrive within 60 minutes during business hours.

A recording sitting on a drive is hard to work with. A transcript is something you can actually use: an attorney can cite it, a researcher can code it, a producer can pull quotes from it. Getting the style right for that purpose is half the job, and it's the first thing we'll ask you about.

What Our Professional Transcriptionists Handle

Every project goes to a transcriptionist who works with that kind of audio regularly.

Interviews & Focus Groups

Research interviews, oral histories and focus groups, with consistent respondent labels and timestamps ready for qualitative coding software.

Depositions & Legal Proceedings

Depositions, hearings, arbitrations and attorney-client interviews, taken verbatim where the wording has to survive scrutiny. See our legal language services.

Medical Dictation

Physician dictation, case notes, IME reports and recorded consultations, handled with HIPAA-aware care. More on our healthcare language services.

Podcasts & Media

Episode transcripts for show notes, accessibility and SEO, plus raw interview tape for editing. Related: media translation.

Meetings & Corporate Audio

Board meetings, earnings calls, panel sessions and all-hands recordings, turned into a written record that stands up as minutes.

911, Body-Cam & Evidentiary Audio

911 calls, body-worn camera audio, custodial interviews and wiretap material, handled with chain-of-custody care. See law enforcement services.

Multilingual Transcription and Translation Combined

This is where we differ from a typical US transcription company. Foreign-language audio is a two-stage job: someone has to write down exactly what was said in the source language before anything can be translated. That first stage goes to a native speaker of the language on the recording, drawn from linguists working across 326 languages.

From there you have three options: a transcript in the original language only, an English translation only, or both side by side. Attorneys and research teams reporting to an IRB usually want both, because the original is the evidence and the translation is the interpretation. The translation stage runs through the same reviewed workflow as our document translation services, and where a court or agency needs it we attach a signed Certificate of Translation Accuracy, as used in our certified translation work.

A practical example: a Houston law firm holds a 40-minute recorded phone call in Spanish. We deliver the Spanish transcript with timestamps, an aligned English translation and a certification page, so the exhibit is ready for filing without a second vendor.

Verbatim vs. Clean-Read Transcription

Picking the wrong style is the most common reason a transcript gets sent back, so we ask what it's for before we start.

Comparison of verbatim and clean-read transcription styles
Feature Verbatim Clean Read
What it captures Every word as spoken: false starts, repetitions, fillers ("um", "you know"), stammers, with non-verbal sounds noted The speaker's own words and structure, with fillers, stumbles and repetitions removed
Best for Depositions, 911 and body-cam audio, custodial interviews, disciplinary hearings, linguistic research Business meetings, research interviews, podcasts, dictation, oral histories
Reads like A court record: faithful but slow to read Natural written speech: quotable and easy to skim
Relative cost Higher, because it takes longer to produce Standard

A simple rule: if the transcript could end up in front of a judge or a regulator, choose verbatim. If it's going to be read, coded or quoted from, choose clean read. Tell us the purpose and we'll recommend one.

Audio File Formats and Turnaround Times

If it plays, we can usually transcribe it.

MP3 WAV M4A WMA AAC FLAC OGG AMR Voicemail exports Dictation apps Body-cam files

Turnaround depends on runtime, audio quality and style. Most standard files return within 2 to 5 business days. Same-day and next-day delivery are often possible for short, clear recordings, subject to capacity, so tell us your deadline up front and we'll confirm it in writing. For anything large we'll set up a secure upload link rather than email.

Two boundary cases worth knowing. If your source is footage rather than audio, our video transcription service works from the file directly and can timestamp against the picture. If the goal is text on screen for viewers rather than a document, that's subtitling and captions, a different craft with its own reading-speed rules.

Human Transcription Accuracy vs. AI Tools

Honest answer first: for a clear English recording of one person speaking into a decent microphone, automated transcription has become good and cheap, and an app may serve you fine. We'd rather say that than pretend otherwise.

Human transcription earns its cost where automated tools reliably fail: overlapping speech, heavy accents, background noise, code-switching between languages mid-sentence, drug names and case citations, and telling apart six voices in a conference room. It also matters when an error has consequences. An AI tool that confidently mishears a dosage in medical dictation, or a name in a 911 call, produces a transcript that looks clean and is wrong in a way nobody catches.

Our process puts two people on every file: a transcriptionist produces the draft and a reviewer checks it against the audio before delivery. Where a section genuinely can't be made out, we mark it inaudible with a timestamp rather than guess. A visible gap is honest; a confident guess is not.

Confidentiality for Medical and Legal Recordings

Recordings are often more sensitive than documents: a voice identifies a person, and people say things aloud they would never write down. Access is limited to the transcriptionist and project manager on the job, transfer runs over secure links, and files are deleted from active systems after delivery.

For medical audio we work to HIPAA-aware practices and will sign a business associate agreement where your organization requires one. For legal and law-enforcement material, including body-cam and custodial interview audio, we apply chain-of-custody care: no copies beyond the working team, logged handling, and transcripts formatted so they can be exhibited. Non-disclosure agreements are routine on this work, not a special arrangement.

How to Order Audio Transcription

Four steps from recording to transcript.

  1. Send your audio

    Upload the file through the quote form or ask us for a secure link for large recordings. Tell us the runtime, the language, roughly how many speakers there are, and what the transcript is for.

  2. Confirm style, deadline and price

    We recommend verbatim or clean read, confirm timestamps and speaker labels, and reply with a written quote and delivery date, usually within 60 minutes during business hours.

  3. Human transcription and review

    A professional transcriptionist produces the transcript and a second reviewer checks it against the audio. Anything genuinely inaudible is flagged with a timestamp rather than guessed.

  4. Delivery in your format

    You receive the transcript in Word, PDF or your preferred format, with a translated version alongside the original if you ordered multilingual transcription.

Audio Transcription Cost: A Realistic Guide

As a guide, professional human transcription of clear English audio in the US typically runs about $1.50 to $3.50 per audio minute. Strict verbatim and legal formatting sit toward the top of that range or above it. Multilingual transcription is quoted as two stages, transcription plus translation, and usually lands higher depending on the language pair.

Four things move the number for any given file:

  • Runtime. The starting point, and the only figure most people think of.
  • Audio quality. The biggest hidden variable. A noisy recording takes far longer to transcribe than a clean one.
  • Number of speakers. One dictated voice is quick; eight people in a focus group means constant attribution and rewinding.
  • Style and deadline. Full verbatim takes longer than clean read, and rush work carries a premium.

A short sample of your audio tells us more than any rate card. Your quote confirms the exact price before we start.

Audio Transcription FAQs

How much does audio transcription cost per audio minute?
As a guide, professional human transcription in the US market typically runs about $1.50 to $3.50 per audio minute for clear English recordings. Strict verbatim, legal formatting, poor audio, many speakers and rush turnaround push that higher, and multilingual transcription with translation is quoted as two stages. Send us the runtime and a short sample and your quote confirms the exact price.
What is the difference between verbatim and clean-read transcription?
Verbatim transcription captures every word exactly as spoken, including false starts, repetitions and fillers such as 'um' and 'you know', which matters when the exact wording is evidence. Clean-read transcription removes those and gives a readable transcript of what the speaker meant. We recommend one based on what the transcript is for.
Can you transcribe audio that is not in English?
Yes. We transcribe recordings in 326 languages, with the transcription done by a native speaker of the language on the recording. You can order the transcript in the original language, an English translation, or both side by side, which is what most research teams and attorneys ask for.
Is human transcription more accurate than AI transcription tools?
For clear, single-speaker English audio, AI tools have become genuinely good and cheap. Human transcription still wins on overlapping speech, heavy accents, background noise, specialist terminology, speaker attribution and any recording where an error has legal or clinical consequences. We flag inaudible sections honestly rather than guessing, which automated tools do not.
How do you handle confidential recordings like medical dictation or police interviews?
Access is limited to the transcriptionist and project manager assigned to your job. Medical audio is handled with HIPAA-aware practices, and evidentiary audio for attorneys and law enforcement is handled with chain-of-custody care: secure transfer, no copies beyond the working team, and deletion from active systems after delivery. We sign non-disclosure and business associate agreements where your organization requires them.
What audio file formats do you accept, and how fast is turnaround?
Most common formats are fine, including MP3, WAV, M4A, WMA, AAC, FLAC and recordings from phones, dictation apps and body-worn cameras. Typical turnaround is 2 to 5 business days depending on runtime and audio quality. Faster delivery, including same-day for short clear recordings, is often possible subject to capacity, so tell us your deadline when you ask for a quote.

Got a recording to transcribe?

Tell us the runtime, the language and how many speakers. We'll reply with a written quote and a confirmed turnaround.

Get a Free Quote Call +1 (833) 282 8883