ClipSonic
Reference

AI video clipping glossary

Plain-English definitions for the terms ClipSonic uses — the three clipping modes, what a hook reason and viral score mean, and the framing and caption vocabulary of short-form video.

AI Viral Clips
A ClipSonic clipping mode that sends the video's transcript to a language model (Groq, OpenAI, or a local model via Ollama). The model returns ranked clip suggestions, each with a title, a hook reason, and a viral score.
Smart Viral Clips
A ClipSonic mode that scores every candidate moment on your own machine against six signals — keyword weight, hook quality, speech pace, dramatic pauses, audio energy, and clean cut points — with no LLM and no API key.
Fixed-Length Clips
A ClipSonic mode that uses no AI at all: you pick a length (30, 45, 60, 90, or 120 seconds) and it splits the whole video into back-to-back clips, snapped to natural silence so cuts don't land mid-word.
Hook reason
A short, plain-English explanation of why the AI picked a given moment as clip-worthy — for example a strong opening line, a punchline, or a complete self-contained thought. Shown next to each AI Viral Clips suggestion.
Viral score
A rating the AI assigns to each suggested clip, estimating how likely the moment is to perform as a short. Used to sort suggestions by potential before you commit to editing.
Highlight detection
The step where ClipSonic finds the best moments in a long video — either by on-device scoring (Smart Viral Clips) or by an AI model reading the transcript (AI Viral Clips).
Face-tracked reframing
Automatically detecting and following faces frame by frame so the speaker stays centered when a widescreen 16:9 video is cropped to a vertical 9:16 short.
9:16 reframing
Converting a widescreen 16:9 video into the 9:16 vertical aspect ratio used by YouTube Shorts, TikTok, and Instagram Reels.
Duo split-screen
A framing mode that tracks two speakers independently and stacks them in the 9:16 frame — used for co-hosted podcasts and interviews.
Gameplay + facecam layout
A framing mode that places a tracked facecam over a full-frame gameplay background, built for streamers and gaming creators.
Animated captions
Word-by-word highlighted subtitles burned into the video. ClipSonic ships six presets — Clean, Karaoke, Bold Pop, Neon Glow, Bounce, and Boxed — each draggable to any position on the frame.
Local rendering
Doing all the video work — cutting, reframing, captioning, and exporting — on your own computer via a bundled ffmpeg engine, with no cloud render queue and no upload of the source video.
Local Whisper transcription
Running the speech-to-text engine (whisper.cpp) directly on your computer, GPU-accelerated on most machines, with word-level timing for captions — no API key and no per-minute cloud fee.
Transcript
The timed text of everything spoken in a video. ClipSonic uses it to find highlights and to time animated captions word by word.
One-pass export
Cutting, framing, and captioning a clip together in a single render, instead of re-encoding the video once per step.
Batch export
Rendering every approved clip from a project in one pass, rather than exporting a clip at a time.
Watermark
A logo or label overlaid on exported clips by many cloud clipping tools' free tiers. ClipSonic never adds one, on any plan.
Clip window
A candidate start/end range in the source video that a highlight-detection mode considers as a possible clip before scoring it.

See these in the app

Download ClipSonic and try all three clipping modes on your own video — free, no account required.