Craft guide · 12 min read

How to Add and Correct Captions on Video Clips

Automatic captions are right about the words nobody notices and wrong about the ones that carry the point. Here is the short list to check, how to break and time the lines so they actually read, and when to burn them in.

By the Clippingnet teamUpdated 2,645 words

ONE SENTENCE, TWICE WHAT THE MACHINE HEARD Pre Yanka says our see PM doubled on fifty second cuts reads fine, which is the problem 3 words worth checking WHAT WAS ACTUALLY SAID Priyanka says our CPM doubled on fifteen second cuts a name · an acronym · a number the three kinds that break, and the three that carry the point every other word was already right

Why the captions decide it

A clip is watched in places where sound is awkward: an office, a train, a bed at midnight, a queue. That is not a statistic anyone outside the platforms can verify, but it does not need to be one – you already know how you watch. A clip that cannot be followed silently is asking for a favour before it has earned anything.

So captions are not an accessibility checkbox bolted on at the end. On a talking clip they are the primary channel, and the audio is the enhancement. That changes what a mistake costs: a wrong word is not a typo the viewer forgives, it is the sentence failing.

WHERE THE MISSES LAND HEARD A MILLION TIMES · RELIABLE theandwegoingreallythinkaboutbecausepeoplejust RARE · AND THE REASON YOU MADE THE CLIP Priyankaa nameReykjavika placeCPMan acronymfifteena numberretentionthe jargonClippingneta product an accuracy score counts the words on the left · a viewer only notices the words on the right
Illustrative, not a benchmark. The point is the shape rather than the score: a model is least sure about the words it has seen least often, and those are the names, numbers and terms a clip exists to deliver.

And here is the part that catches people. Automatic captions are genuinely good now, which is exactly why they are dangerous. They are fluent, punctuated and confident, so nothing about them looks wrong – and the handful of words they miss are, with grim reliability, the ones you made the clip to say.

What machines actually get wrong

Speech recognition works on probability. A model has seen because and really and think uncountably many times, so it is almost never wrong about them. It has seen your co-founder’s name approximately never.

That is why an overall accuracy figure is close to useless for deciding whether captions are publishable. The misses are not spread evenly across the sentence – they concentrate on low-frequency words, and low-frequency words are exactly the ones a clip exists to deliver. Four categories cover almost all of it:

  1. Names. People, companies, places, handles. A model will produce something that sounds right and is spelt wrong, which is worse than a blank.
  2. Numbers. The classic is fifteen heard as fifty. It is one letter on screen and a completely different claim, and it is the single most damaging error in this list.
  3. Acronyms said as letters. CPM, ROI, API, RPM. Letters spoken in sequence get reassembled into words that look deliberate.
  4. Subject jargon and homophones. Retention becomes attention; bitrate becomes bit rate. Both results are real English, so no spellchecker will ever flag them.

YouTube is unusually candid about this in its own documentation, which says automatic captions “might misrepresent the spoken content due to mispronunciations, accents, dialects, or background noise” and that “the quality of the captions may vary” because they are generated by machine learning. Its own advice is to review them and edit the parts that were not transcribed properly.

The same page lists why captions sometimes never appear at all: audio still processing, an unsupported language, a video that is too long, poor sound quality or unrecognised speech, a long silence at the beginning, or several people talking at once. Two of those are worth noting because they are yours to fix – poor sound and overlapping speakers. If a model cannot hear it, neither could a viewer, and no amount of correcting afterwards repairs that.

Automatic-caption behaviour and the file formats below read from YouTube Help and supported subtitle and caption files on 20 September 2026. Other platforms document this far less clearly than YouTube does, so this page quotes no numbers for them – check what your own account offers.

The ninety-second correction pass

Once you know the errors cluster, correcting captions stops being proofreading and becomes a search. You are not reading the transcript. You are looking for six or seven specific words in it.

the thing that changed everything for us was Pre Yanka's idea. we stopped posting at random and started posting to a schedule. our see PM doubled in a week, and on the fifty second cuts attention went from twelve percent to thirty one percent.

Nothing here looks broken. It is grammatical, it is punctuated, and it is wrong in four places – which is exactly why people post it.

44 words6 worth checking4 were wrongillustrative example

Step through those three stages and the method is the whole argument. The raw version reads perfectly well, which is why it gets posted. The short list is short – a handful of words out of forty-odd. And every single error is on that list, because the list is the category of words a model is least sure about.

In practice: read the transcript once looking only for capital letters, digits and anything that belongs to your subject rather than to English. Fix those. Then read it once more at speed, aloud if you can, to catch anything that changed the meaning while staying grammatical. That is the pass. On a clip of under a minute it is faster than choosing a thumbnail.

Before you caption

Cut the clip first – then fix the words

Correct is not the same as readable

A caption can be perfectly spelt and still fail, because captions are read rather than heard, and reading has its own constraints. Two of them do most of the damage: where the lines break, and when the words arrive.

CORRECT, BUT HARD TO READ THREE LINES, BROKEN ANYWHERE we switched to fifteen secondcuts and the completionrate went up one phrase, split across two lines, because the line ran out TWO LINES, BROKEN WHERE YOU PAUSE we switched tofifteen second cuts and the completion ratewent up two blocks in sequence, each a complete thought AND WHEN IT APPEARS the word is said here on the word, or a frame before it half a second late reads as a mistake
Captions are read, not heard, so a line that is spelt correctly can still fail. Two lines at a time, broken where a speaker would pause, timed to the word rather than after it.

Line breaks go wrong because most tools break where the line runs out rather than where the sense does. Splitting fifteen second / cuts across two lines forces the reader to hold half a phrase in their head, and on a clip that is already moving, they simply do not. Timing goes wrong more subtly: a caption that appears half a second after the word is said feels like a glitch, and viewers who could not tell you why will leave anyway.

  • Two lines at a timeThree or more and the eye stops tracking the picture. If a phrase needs more, make it two blocks.
  • Break where you would pauseKeep a phrase together – fifteen second cuts belongs on one line. Never split a word across lines.
  • On the word, not after itA caption that lands half a second late reads as an error even when every letter is right.
  • Roughly six words a lineWide enough to read at a glance, narrow enough to stay clear of the interface.
  • Sentence case beats capitalsALL CAPS is slower to read and takes more room. Save it for a single word you mean to shout.

There is also a placement question, which matters more on vertical clips than anywhere else – put a caption where the platform draws its own buttons and half of it is simply not there. Long videos to Reels covers the safe area in detail, and the same caution applies on TikTok.

Styling that helps, styling that shows off

Caption styling has drifted a long way past legibility, and most of what is now fashionable actively costs you viewers. The test for any styling decision is narrow: does this make the words faster to read? If not, it is decoration on the one element that cannot afford any.

What earns its place is unglamorous. A heavy weight rather than a thin one. A solid or strongly shadowed backdrop, because footage changes brightness underneath and a white word on a bright sky is gone. A size big enough to read on a phone held at arm’s length, which is larger than it looks on your monitor. And the same position on every clip you make, so a returning viewer’s eye already knows where to go.

What does not: word-by-word karaoke highlighting on every word, which makes the eye chase rather than scan and is exhausting past about fifteen seconds; a different colour for emphasis on half the sentence, which means emphasis on none of it; animated entrances per word; and fonts chosen for personality. One highlighted word in a clip is a tool. Every word highlighted is a tic.

One genuine exception to the restraint: a single word enlarged or coloured at the moment of the punchline works, precisely because nothing else in the clip is doing it. That is also how hooks earn attention – by contrast with everything around them.

Burned in or uploaded?

There are two different things called captions, and they behave nothing alike. Burned-in captions are pixels: rendered into the video during export, impossible to turn off, impossible to change later. A caption file is text sitting alongside the video, which the platform draws on top and the viewer can usually toggle.

Burned inCaption file
Shows up no matter whatyesonly if the viewer has them on
Can be fixed after postingno – re-export and re-uploadyes – edit the file or the text
Readable by search and by screen readersno – it is pixelsyes – it is text
Survives being re-uploaded elsewhereyesno – the file stays behind
Can be translatednoyes

For a clip going into a feed, burned-in is the safer default, for one blunt reason: it cannot fail to appear. A caption file that is off by default on a platform that autoplays muted has protected nobody. Burned-in captions also survive the clip being downloaded, reposted or cross-posted, which is often how a clip actually travels.

Where you have the choice, do both. YouTube accepts caption files in plain formats – it lists SubRip (.srt) and SubViewer (.sbv) among its basic ones and suggests them for beginners because they are editable plain text, with WebVTT (.vtt) and TTML (.ttml) available when you need positioning or styling. Burn the captions in for the people scrolling, and upload the file so the words are also text: searchable, translatable and available to a screen reader.

The one place to think twice about burning in is anything with slides or fine detail on screen, where a caption bar covers the thing the clip is about. Webinar clips run into this constantly, and it is the reason captions are situational rather than mandatory there.

Caption once, not once per clip

The expensive mistake is not making captions badly. It is making them repeatedly.

If you cut eight clips from one interview and correct each clip’s captions separately, you will fix the same misheard name eight times – and get it slightly different at least once. Work the other way round. Correct the transcript of the source video first, before anything is cut, then every clip taken from it inherits words you have already checked. The proper nouns in a fifty-minute recording are a short list too, and it is the same list every episode.

Two habits make that cheaper still. Keep a names file for your show – guests, products, places, recurring jargon – because the same handful of words break every single time and many tools let you feed them in as hints. And fix the audio rather than the captions where you can: a clean track produces captions that need almost nothing, which is the only correction pass that takes no time at all.

This is also the argument for doing captions inside the same pass as the cutting rather than in a separate tool later. Turning long videos into short clips covers that workflow, podcast clips is where it pays off most, and AI clipping versus manual editing is the wider version of the same trade – let the machine do the pass, and keep the judgment.

FAQ

Are automatic captions good enough to post as they are?

Usually not, and the reason is specific rather than snobbish: the words a model is least sure about are the rare ones, which are your names, numbers, acronyms and subject terms. YouTube says as much itself – automatic captions “might misrepresent the spoken content due to mispronunciations, accents, dialects, or background noise” – and recommends you review them and edit anything that was not transcribed properly. (YouTube Help, read 20 September 2026.)

How long should a correction pass take?

On a thirty to sixty second clip, about ninety seconds – because you are not reading every word. You are checking a short list: every proper noun, every number, every acronym and every term that belongs to your subject rather than to general English. On a clip like the one above that is 6 words out of 44.

Do silent or music-only clips need captions?

No. If nobody is talking there is nothing to transcribe, and adding a caption bar to a purely visual clip only covers the picture. Travel clips are the clearest case – a location label does the job instead.

Should I burn captions in or upload a caption file?

Burned in guarantees they appear and travel with the file; a caption file keeps the words editable, translatable and readable by search and screen readers. Burned-in captions are the safer default for a feed clip, and there is nothing stopping you from doing both on a platform that accepts a file.

What caption file formats can I upload to YouTube?

YouTube lists SubRip (.srt), SubViewer (.sbv or .sub), MPsub (.mpsub), LRC (.lrc) and Videotron Lambda (.cap) as basic formats, and SAMI (.smi), RealText (.rt), WebVTT (.vtt) and TTML (.ttml) as advanced ones that carry styling or positioning. It suggests SubRip or SubViewer to start with because they are plain text you can open and fix in any editor. (YouTube Help, read 20 September 2026.)

Why did my video get no automatic captions at all?

YouTube lists several reasons: the audio is still processing, the language is unsupported, the video is too long, the sound quality is poor or the speech is not recognised, there is a long silence at the start, or several people are talking over each other. Poor sound is the common one, and it is worth fixing at the source. (YouTube Help, read 20 September 2026.)

Do captions help a clip get recommended?

Nobody outside the platforms can say what the ranking systems weigh, so treat any confident claim about that with suspicion. What is observable is simpler: a clip that can be followed with the sound off holds people longer than one that cannot, and how long people watch is the thing every platform is unambiguous about caring for.

Related: choosing the best moments, YouTube videos to Shorts and what video clipping is.

Mark the in. Mark the out. Done