← Back to Blog

What Is a Transcript? Definition, Types, and Examples

Rachel Nguyen··10 min read
TranscriptionGuidesVideo ToolsHow-ToDefinitions
Open laptop on a clean desk showing a text document with a transcript, microphone nearby, soft office lighting

What Is a Transcript? Definition, Types, and Examples

The word "transcript" gets attached to a surprising range of things. College grades, Supreme Court hearings, podcast episodes, YouTube videos, job interview recordings: all of them produce something called a transcript. Most searches for a definition land on a dictionary entry that gives one sentence and leaves you with more questions.

This guide covers what a transcript is, the types you're most likely to run into, what one actually looks like in different formats, and how transcripts get made. By the end, you'll know exactly which type you need and how to get it.

A transcript is a written record of spoken content. It captures what was said, either word-for-word or in a cleaned-up form, from a recording, speech, or conversation. Transcripts are used in schools, courts, podcasts, meetings, and anywhere audio or video needs a text version for reference, accessibility, or repurposing.

What "Transcript" Actually Means

The word comes from the Latin "transcriptum," a written copy. Merriam-Webster defines it as "a written, printed, or typed copy; especially a usually typed copy of dictated or recorded material."

In practice, "transcript" covers two broad categories: written records of audio (spoken words turned into text) and written records of academic history (your grades and courses). Both are technically transcripts, though in everyday digital life, most people mean the audio-to-text variety.

Transcription as a profession dates to at least the 1870s, when court reporters used stenotype machines to capture legal proceedings in real time. Manual transcription of a one-hour recording takes 4 to 6 hours, a roughly 4:1 time-to-content ratio that held for over a century. AI transcription tools, which analyze audio using neural networks trained on large speech datasets, process that same hour in 30 to 90 seconds, with accuracy between 85% and 95% for clear audio. Noisier recordings (crowded rooms, phone calls, strong accents) typically yield 70% to 80% accuracy. Academic transcripts in the US became standardized in the 19th century when college enrollment records needed to transfer reliably between institutions. Today, transcripts serve four core domains: legal (court and deposition records), medical (physician notes and dictation), academic (student records), and media (podcasts, video content, meeting notes for accessibility and search).

The Most Common Types of Transcripts

There are five types you'll run into most often.

Academic transcript. A record of every course you've taken, the grade you received, and any honors or degrees conferred. Colleges and employers request these when evaluating your background. Official transcripts are sealed and sent institution-to-institution; unofficial versions are printable copies for your own reference.

Legal transcript. Court reporters capture every word spoken on the record during trials, depositions, and hearings, covering judges, attorneys, and witnesses alike. These documents are legally binding and can be cited in future proceedings. Accuracy requirements are strict: the FCC mandates 99% accuracy for broadcast content, and legal transcripts hold to a similar standard.

Medical transcript. Physicians often dictate patient notes after appointments. A medical transcriptionist (or AI tool) converts those recordings into structured records entered into a patient's file. Misheard medication names or dosages can cause direct harm, so this is one area where human review of AI output is still standard practice.

Audio and video transcript. The type most internet searches are actually looking for. A podcast episode, YouTube video, Zoom call, TikTok, or recorded interview becomes a searchable text document. Podcasters use these for show notes and SEO. YouTubers turn them into subtitles and blog posts. Researchers use them to analyze interview data. Teachers use them to make lecture content accessible.

Meeting transcript. A record of what was discussed, decided, and assigned during a call. Many platforms (Zoom, Teams, Google Meet) generate these automatically on paid tiers. A clean meeting transcript saves the team from rewatching 45 minutes of recording to locate one decision.

Verbatim vs. Clean Transcripts

Every transcript falls somewhere on a spectrum between fully verbatim and fully edited. Knowing which type you need before you start saves time.

Verbatim transcription captures everything: every "um" and "uh," repeated words, false starts, filler sounds, and background noise. Legal proceedings and certain qualitative research methods demand this level of precision because the exact words (pauses included) carry meaning.

Clean (edited) transcription removes filler words, repairs false starts, and smooths informal speech into readable sentences. Podcast show notes, video captions, business meeting summaries, and most content creation use cases call for clean transcripts because readability matters more than preserving every hesitation.

Most AI transcription tools produce a clean version by default, stripping out filler words automatically. If your use case requires verbatim output (qualitative research, legal depositions, linguistic analysis), check whether your tool supports that mode before generating. For a deeper look at transcription types and when each applies, transcription for qualitative research covers the full classification: full verbatim, clean verbatim, edited, phonetic, and Jeffersonian notation.

What a Transcript Looks Like

The format depends on how the transcript will be used.

Plain text transcript. Paragraphs of text, often with speaker labels (e.g., "Host:", "Guest:"). What you'd see in a news article's printed version of a presidential speech, or in a podcast's show notes section.

Timestamped transcript. Each line or paragraph includes a time marker showing when that content was spoken. Useful for navigating long recordings: you can jump directly to the relevant section rather than scrubbing through the audio.

SRT file. A subtitle format where text is broken into short segments, each timed to video playback. Every segment includes a sequence number, start and end timestamp (formatted as HH:MM:SS,mmm), and the spoken line. SRT is the most widely supported caption format across YouTube, Vimeo, TikTok, Instagram, and most video hosting platforms.

VTT file. Developed by the W3C in 2010 for use in web browsers. Structurally similar to SRT, but VTT supports CSS styling, text positioning, and additional metadata that SRT doesn't. HTML5 video players use VTT natively.

JSON or Excel export. For data analysis, some transcription tools export structured transcript data with timestamps, speaker labels, and confidence scores in machine-readable formats. PixScript's Business tier includes JSON export alongside the standard text formats.

For most online video, creators end up generating an SRT or VTT file and uploading it alongside the video to enable captions on their platform of choice.

How Transcripts Are Made

There are two approaches: manual and automated.

Manual transcription means a person listens to the audio and types it out. A skilled transcriptionist handles roughly 60 to 80 words per minute, but accounting for playback pauses, rewinds, and proofreading passes, a one-hour recording takes 4 to 6 hours to transcribe manually. At professional service rates ($1 to $2 per audio minute), transcribing a one-hour interview costs $60 to $120.

AI transcription changed the economics. Modern speech recognition models process audio using transformer-based neural networks, returning a transcript in 30 to 90 seconds for most file lengths. Accuracy on clear audio sits at 85% to 95%. It drops to 70% to 80% on crowded environments, heavy accents, or low-quality recordings. For most content creators, that accuracy level is sufficient without any human review. For legal and medical transcription, AI handles the first pass and a trained specialist reviews and corrects the output.

The practical workflow for most people: upload the file or paste a URL into an AI transcription tool, get a clean timestamped transcript in under 2 minutes, then export in the format the task requires. For a step-by-step breakdown of this process, how to transcribe a video covers the main methods and when each one fits. If your priority is cost, free video transcription covers the tools and free-tier limits across the main options.

How PixScript Can Help

PixScript handles transcription for videos and audio files without any setup. Paste a URL from YouTube (including Shorts), TikTok, or Instagram Reels and get a timestamped transcript in seconds. Or upload an audio or video file directly through the dashboard (MP3, MP4, M4A, WAV, MOV, WEBM, up to 500 MB) for longer recordings and interviews.

Once you have a transcript, PixScript gives you a few paths forward. Export as SRT or VTT for video captions. Export as PDF or plain text for sharing or archiving. Run it through the AI rewrite tool to turn a raw transcript into a polished blog post, social media post, or script. The translation feature converts transcripts into 10 languages on Pro and 50+ on Business.

Pricing: the free tier covers 10 transcripts per month with TXT export and a 5-minute maximum. Pro is $9/month (unlimited transcripts, all export formats, AI features, 30-minute max). Business is $19/month (unlimited length, bulk processing, JSON/Excel export, 50+ translation languages).

If you're working with recorded content regularly, having a transcript workflow reduces a task that used to take hours to something that takes a few minutes.

Frequently Asked Questions

What's the difference between a transcript and a transcription?

Transcription is the process of converting speech to text. A transcript is the finished document that results from that process. The distinction matters in formal and academic contexts: legal and research writing typically use the terms precisely. In everyday conversation, people use the words interchangeably, and that's generally fine.

Is a transcript the same as captions or subtitles?

Captions and subtitles are formatted transcripts that are timed to video playback (SRT or VTT files). A plain transcript is just the text without timing data. You can generate captions from a transcript by adding timestamps and breaking it into timed segments, which is exactly what AI transcription tools do automatically when you request SRT or VTT output.

How accurate are AI transcripts?

For clear audio with a single speaker, AI transcription typically reaches 85% to 95% accuracy. Accuracy drops to 70% to 80% in noisier conditions: multiple speakers talking over each other, strong regional accents, phone call audio, or significant background noise. A quick manual review after generation catches the errors that matter most.

Can you get a transcript of a YouTube video?

Yes. YouTube generates auto-captions for most videos, accessible under the video by clicking the three-dot menu and selecting "Open transcript." For higher accuracy or export options (SRT, VTT, PDF), a tool like PixScript lets you paste the YouTube URL directly and download the transcript in any format you need.

How long does it take to transcribe a 30-minute recording?

An AI transcription tool processes a 30-minute file in 15 to 45 seconds. Manual transcription of a 30-minute recording takes 2 to 3 hours. Professional transcription services charge $1 to $2 per audio minute, so a 30-minute file costs $30 to $60 if outsourced.

One Tool for Any Transcript You Need

Whether you're working with a podcast interview, a meeting recording, a YouTube video, or an uploaded audio file, getting to a usable transcript is now a matter of seconds rather than hours. The type you need (academic, legal, verbatim, clean, SRT, VTT) determines the format, but the underlying process is the same: audio in, structured text out.

PixScript transcribes any video URL or uploaded file and exports in SRT, VTT, PDF, or plain text. Try it free at pixscript.com.