What clipping content means
Clipping content is taking a long recording and lifting short, self-contained pieces out of it. The mechanics are elsewhere on this site – what clipping is covers the concept and the general route covers the process end to end.
This page is about the question both of those quietly assume, and which nobody writes about because it is commercially inconvenient: whether your recording contains clips at all. Every tool and nearly every guide starts from the premise that your footage is fine and the only variable is your skill. That is not true. Clippability is a property of the source, it varies enormously, and you can assess it in about two minutes instead of finding out after an hour.
Clippable is not the same as good
This is the part that surprises people, and the diagram at the top of this page is the whole of it. Quality and clippability are independent, and a well-made piece of content can be actively harder to clip because it is well made.
A tightly argued forty-minute talk builds: each point rests on the one before, the examples refer back, the conclusion means nothing without the middle. That is good construction and it is exactly what makes it unclippable – nothing detaches, because everything was deliberately attached.
A loose conversation with six anecdotes in it is worse content by most measures and far better raw material. Each anecdote is already self-contained, because the speaker had to set it up for the person sitting opposite them.
So the first thing to stop doing is treating low yield as a verdict on the content or on yourself. It is a structural fact, it is often predictable before you record, and it is sometimes worth changing the structure – which is different from changing the quality.
Try it: does it survive cold?
One test does nearly all the work. A clip is watched by someone who has heard nothing before it, so the only question that matters about a moment is whether it survives that. Four excerpts, judged cold:
Four excerpts. Would someone who has heard nothing before them understand each one?
-
So that’s why we ended up rewriting the whole thing in March.
Two backward references in one sentence – that’s why and the whole thing. A stranger has neither, and no amount of editing supplies them.
-
Most people think hiring slowly is cautious. It is the opposite: every week a role sits open, the team absorbs the work and quietly stops doing their own.
A complete claim with its reason attached, and nothing in it pointing outside itself. This is what detachable actually looks like.
-
And Maria said the same thing when we talked about it last year.
Who is Maria, what thing, which conversation. Three unanswered references, and the clip cannot answer any of them.
-
We stopped doing quarterly planning entirely. Nothing broke. That surprised me more than anything else.
Self-contained, and the surprise does the work a hook would otherwise have to do. Notice it is also the shortest of the four.
Nothing answered yet.
The two that fail do so for the same underlying reason: a word pointing at something the viewer never saw. That’s why, the whole thing, the same thing, Maria – each is a reference with nothing behind it, and no editing, captioning or reframing supplies what is missing. That is worth knowing because those moments often sound like the best bits when you are listening with the context already in your head.
Run this on a transcript rather than on the video. Reading is several times faster than real time and it strips out the delivery, which is what fools you – a confident voice makes an incomplete sentence feel complete. Choosing the best moments covers picking between the ones that pass.
Where the clips live
Detachability is not spread evenly through a recording, and its distribution is decided by the shape of the content rather than by anything you do afterwards.
Interviews are the richest source, and the reason is structural rather than a matter of who is talking: every answer is addressed to someone who asked a question, so the speaker supplies context automatically. That is why podcast clips and webinar Q and A sections outperform the scripted parts of the same recordings.
Panels cluster. Each new question resets the context, so the minute or two after it contains several liftable answers, and the stretch before the next question contains almost none. Scan the openings.
Cumulative talks are thin, and the one clip they usually do contain is near the end, where the speaker summarises. If you are clipping a talk, look at the last five minutes first – it is the only part written to stand alone.
Recording so there are more of them
Almost all of this is decided before anyone presses record, which means it is fixable at the point where fixing is free. Six habits, none of which costs time on the day:
- Answer in whole sentencesFold the question into the answer. “The reason we stopped is…” travels; “Yeah, exactly” does not, however good what follows it is.
- Name the thing each timeSay the pricing page rather than it. Pronouns pointing at something said ten minutes ago are the single commonest reason a good moment cannot be lifted out.
- Say the number with its unit“Forty per cent of signups” rather than “forty per cent”. A figure with no noun attached is unusable on its own and impossible to caption honestly.
- Land one idea per answerTwo ideas in one answer means the clip carries a passenger. If a second thought arrives, stop and start it again as its own answer.
- Leave a beat either sideA second of silence before and after gives a clean in and out point. Talking over the end of your own sentence is what forces an ugly cut later.
- Say the surprising thing firstPut the conclusion at the front of the answer and the reasoning behind it. That order is what makes an answer liftable, and it is also what makes it a hook.
The first two are worth more than the other four combined. Whole sentences and named things between them account for most of the difference between a recording that yields ten clips and one that yields two – and they cost the speaker nothing except a small amount of self-consciousness in the first session, which passes.
If you are interviewing someone, you can do most of this for them by asking questions that force a complete answer: “what changed when you stopped?” rather than “did that work?”. A closed question produces a clip-proof answer, which is a problem the editing cannot solve afterwards.
Doing it well
Once a recording does contain clips, doing it well is four decisions, each of which has its own page on this site rather than a paragraph here.
Where to start. Almost never at the beginning of the passage – the strong opening line is usually some way in, and the best-written sentence is often the worst place to begin because it is the answer rather than the question. Writing better hooks covers it.
What the words say on screen. Automatic captions are right about the common words and wrong about names, numbers and jargon, which are the ones the clip exists to deliver. Captions for video clips has the correction pass.
What shape it goes out in. A vertical crop deletes about two thirds of the frame, and which two thirds is a decision rather than a default. Horizontal to vertical video covers what it costs.
Whether it holds up. Most softness is decided before the export dialog opens. Why clips look blurry covers the diagnosis.
Notice that none of the four is about the tool. Tools matter for how much footage you can get through, not for whether any of these decisions come out right.
When there is no clip in it
The answer this category never gives is that some recordings contain nothing worth posting, and recognising that quickly is a skill rather than a failure.
The signs are consistent. Nothing survives the cold test in the first twenty minutes. Every candidate needs a sentence of setup that the clip cannot carry. The interesting content is genuinely cumulative, so the good bit is good because of the thirty minutes before it. Or the audio is poor enough that the words cannot be recovered, which caps everything downstream.
When that is the case there are three honest moves, in order. Change the structure next time – add a question-and-answer section, or ask someone to interview you, which converts a talk into the shape that yields. Write the clip instead of finding it: record a thirty-second piece to camera making the same point, built to stand alone. Or accept that this recording is not the raw material and spend the hour somewhere with better odds.
Forcing a clip out of a recording that does not contain one is how channels train an audience to scroll past them, and it costs more than the empty week would have.
FAQ
What does clipping content actually mean?
Taking a long recording and cutting short self-contained pieces out of it for feeds. The mechanics are covered in what video clipping is and the end-to-end route in turning long videos into short clips. This page is about the question those two assume: whether the recording contains clips at all.
Is every video clippable?
No, and that is the most useful thing to know before you spend an hour looking. Clippability depends on whether moments detach from their context, and some content is built so that nothing does – a tightly argued talk where each point rests on the last can be excellent and contain almost nothing liftable.
How do I tell before I start?
Read or skim the transcript rather than watching, and look for sentences that would make sense to someone who has heard nothing before them. If you cannot find three in the first twenty minutes, the recording is probably low-yield and you should decide that now rather than after an hour of scrubbing.
Does that mean my content is bad?
No – the two are independent, and sometimes opposed. Density of detachable moments is a structural property, not a quality judgement. A loose conversation with six anecdotes has high clip yield and may be worse content than the lecture that yields nothing.
Can better tools fix low clippability?
Not really. Every tool in the category reads speech to decide what matters, so a recording with nothing detachable gives all of them the same very little. The tools page covers the category; the constraint here is upstream of all of it.
What is the difference between this and choosing the best moments?
Scope. Choosing the best moments is about picking between candidates inside a recording you have already decided is worth mining. This page is about whether there are candidates, and what makes a recording produce them.
Does this apply to footage with no talking?
The detachability test still applies, but visually rather than verbally – something has to change on screen and be comprehensible without what came before. Travel footage is the clearest case, and it has no transcript to skim, which is why it needs a different scanning method entirely.
Related: choosing the best moments, podcast clips and what video clipping is.