Frequently Asked Questions
Short answers on video previews, analysis, transcription and performance.
Video previews and thumbnails
What is an automated video preview?
A short clip cut automatically from a longer video and shown in place of a static thumbnail. Instead of a single frame, the viewer sees a few seconds of the most compelling motion in the asset, which communicates subject, pace and production quality at once.
Do video previews actually increase views?
Usually, yes — because they replace a frame that was chosen arbitrarily. Most content systems grab a frame from a fixed position in the file, which lands on a transition, a blur or an empty set as often as anything worth clicking. The gain comes from replacing a random choice with a considered one, not from magic.
What makes a good video thumbnail?
A clear subject — usually a face — in focus and reasonably large; enough contrast to survive being displayed small; and a frame that raises a question the video answers. Avoid mid-transition blur, letterboxing and burnt-in caption fragments.
How do I test whether a thumbnail is working?
Measure play rate on the placement — of the people who saw the unit, what share started the video. Change one variable at a time, randomise per visitor rather than per request, run for whole days, and fix your sample size before you start. Watch completion rate as a guardrail: a thumbnail can lift clicks by misrepresenting the video.
Video analysis and AI
What does video content analysis actually do?
It examines the footage itself rather than the metadata around it — scoring segments for motion, faces, speech, audio energy and scene changes. The output is a description of what is inside the asset and where, which is what makes individual moments addressable instead of just whole files.
Why isn't metadata enough?
A title, a category, a date and some upload tags describe the file, not the footage. They say nothing about what happens at minute fourteen or which thirty seconds would hold a browsing visitor. A recommender working from metadata can only match at the level of the whole asset, because as far as it knows the video has no parts.
Where does AI genuinely help with video, and where does it not?
It helps where a judgement is mechanical and repeated many times: frame and segment selection, format conversion, transcription, archive indexing, matching an asset to a slot. It disappoints where the judgement is editorial — deciding what to cover, finding the angle, or anything where being wrong has consequences.
Can highlights be generated from a live broadcast?
Yes. The feed is scored continuously as it arrives, using signals that indicate significance without understanding the sport — sustained motion, crowd audio spikes, scoreboard changes, commentary intensity. The trade-off is that real-time means deciding without hindsight, so every such system sits somewhere on a precision-recall curve.
Transcription and accessibility
How do I get a transcript of a YouTube video?
Paste the video URL into a transcript extractor. You get the full text with clickable timestamps plus any caption languages the video carries. The limitation is that this reads the existing caption track — a video with no uploaded or auto-generated captions has nothing to extract, and needs speech recognition run over the audio instead.
Are captions and transcripts the same thing?
No. Captions are time-synchronised, appear over the video, and should include speaker changes and meaningful non-speech sound. A transcript is the full text as a document, not synced. You want both — they do different jobs, and publishing one does not satisfy the need for the other.
Are automatic captions good enough for accessibility?
Often not without a correction pass. Speech recognition fails predictably on proper nouns, technical vocabulary, speaker attribution and non-speech audio — which are exactly the parts carrying the meaning. Correcting names and technical terms closes most of the gap in a few minutes per video.
Publishing, performance and revenue
Why does most of my video library go unwatched?
Almost always presentation rather than quality. Giving an asset a considered thumbnail, an accurate title and a placement where people are looking takes manual time that scales with the catalogue — so the top few items of the day get it and everything else gets the CMS default.
What should I measure in video analytics?
Play rate per placement, watch time per visitor, completion by quartile, repeat video sessions, and share of library viewed. Treat total views and site-wide video starts with suspicion: both rise with autoplay and traffic while telling you nothing about whether the video worked.
Does an embedded video player hurt page speed?
Nearly always, and often more than the video gains. Defer the player until someone signals intent — a click, or the placement scrolling into view. Until then a poster image and a play affordance cost a fraction of the weight.
Where does publisher video revenue actually come from?
Impressions, fill rate and rate. Most teams push hardest on rate, which has the least headroom because the market sets it. The impression side usually has more room: most sites have far more video than they surface, and most visitors never start one.
Still looking?
Longer treatments of most of these live in the blog, and the step-by-step walkthroughs are collected in the guides. If something here is wrong or missing, tell us.