Stroke, Scope, Proof: Why ASR Still Needs a Human Behind It

By Merritt Gilbert, CER, CDR, FPM | Director, BlueLedge

Automatic Speech Recognition (ASR) has found a real place in legal transcript production — but it’s not the star of the show. It’s a tool, and the value it adds depends entirely on the human using it.

ASR, or voice-to-text technology, can convert the audio of a legal proceeding into a draft transcript, often called the “first draft” or “first stroke.” That’s a genuinely useful shortcut. What it can’t do is finish the job. ASR is not perfect, and every ASR-produced transcript still requires a human to review, edit, and format it before it’s finalized and certified. “Certified” isn’t a formality — it means the transcript carries a certificate signed by the reporter or transcriptionist verifying its accuracy. That verification can only come from a person.

Where ASR Fits in the Process

Transcript production breaks down into five steps: capture the record, stroke the first draft, scope, proofread, and finalize. The first two steps look different depending on who’s capturing the record.

Digital reporters attend and record the proceeding, capturing speaker changes, hard-to-understand words, proper noun spellings, and exhibit information — but they don’t leave with a first draft unless ASR is involved. Without ASR, the legal transcriptionist listens to the audio afterward and transcribes the first draft manually. With ASR, that step is automated, and the transcriptionist moves straight to scoping.

That’s one of the key benefits of ASR: it handles the initial transcription, allowing a skilled human to focus on the review, correction, and expertise needed to produce a final transcript.

Garbage In, Garbage Out

ASR is only as good as the audio it’s given. A clear audio recording is the baseline requirement — background noise or distortion will significantly reduce accuracy, and even small changes, like a speaker dropping their voice slightly, can make a portion of the record difficult for the engine to render correctly.

This is one of the most forward-facing reasons why a professional must be involved. A human transcriptionist working from unclear audio can slow down and replay a segment five or ten times to work out what was said. ASR doesn’t have that judgment. It produces its best guess and moves on, which means audio quality problems show up as errors baked directly into the first draft. A transcriptionist scoping an ASR draft still has to do that same repeated listening to verify the rough patches, except now they’re checking the machine’s guess instead of building the transcript from scratch.

As reliance on ASR grows, so does the cost of not catching those errors, whether it is labor costs of a human recognizing and correcting the error, or even worse, the consequences of an inaccurate record within the justice system.

The Three Steps ASR Can’t Touch

Regardless of whether ASR was used, the final three steps look the same in a reliable and verified process:

  • Scope — The transcriptionist listens to the audio recording while editing the first draft, correcting and adding to the existing text.
  • Proofread — The transcriptionist reads through the transcript (without listening to the audio) to catch any remaining grammatical, punctuation, and formatting errors.
  • Final Review — The transcriptionist paginates the index, reviews headings, parentheticals, and by-lines, checks the meta pages, and runs a pre-production checklist for any straggling errors.

Proofreading deserves special attention here, because it’s often misunderstood. Proofing means working with the text only — not listening to the audio again. That means reading every single word, watching for homonyms and mistranscribed words, and paying close attention to context. It also means accepting a hard limit: some errors are only catchable by ear. “I went to her house” instead of “his house” reads perfectly fine on the page. It’s wrong. A proofreader without the audio has no way to know that.

This is exactly why scoping — the step where the transcriptionist listens while editing — can’t be replaced by ASR either. ASR generates text from audio, but it doesn’t verify meaning against context as well as a trained ear does.

As ASR takes over more of the first-stroke workload, the job title “legal transcriptionist” is starting to look a lot more like “scopist.” Scoping an ASR draft requires the same foundation a transcriptionist needs to build a transcript from scratch — formatting rules, proper punctuation, legal terminology, all of it. There’s no shortcut around that knowledge for a scopist; spotting an error in someone else’s draft, or a machine’s, requires knowing what the correct version looks like just as much as writing it does.

The Bottom Line

ASR is a legitimate efficiency gain for digital reporters and transcriptionists. It allows transcriptionists to complete more pages in the same amount of time, helping meet the industryโ€™s growing demand while maintaining a focus on reliability, security, accuracy, and completeness. And a first draft, however it’s generated, is still just a first draft. Scoping, proofreading, and final review are where accuracy gets solidified — and none of those steps can be automated away.

The tool drafts. The human certifies.


See It in Action

Want to see where AI transcription actually falls short? Join BlueLedge on September 24 for Stroke, Scope, Proof: Why AI Still Needs a Human Behind It, a free webinar featuring real examples of AI transcription errors, expert perspectives on where AI fits in the process, and practical guidance for legal transcriptionists navigating the changing profession.

Register for the webinar โ†’

Interested in learning how to scope and proof effectively with ASR drafts? Check out this online course: ASR Mastery: Proofing & Scoping