Reference
AI video clipping glossary
Plain-English definitions for the terms ClipSonic uses — the three clipping modes, what a hook reason and viral score mean, and the framing and caption vocabulary of short-form video.
- AI Viral Clips
- A ClipSonic clipping mode that sends the video's transcript to a language model (Groq, OpenAI, or a local model via Ollama). The model returns ranked clip suggestions, each with a title, a hook reason, and a viral score.
- Smart Viral Clips
- A ClipSonic mode that scores every candidate moment on your own machine against six signals — keyword weight, hook quality, speech pace, dramatic pauses, audio energy, and clean cut points — with no LLM and no API key.
- Fixed-Length Clips
- A ClipSonic mode that uses no AI at all: you pick a length (30, 45, 60, 90, or 120 seconds) and it splits the whole video into back-to-back clips, snapped to natural silence so cuts don't land mid-word.
- Hook reason
- A short, plain-English explanation of why the AI picked a given moment as clip-worthy — for example a strong opening line, a punchline, or a complete self-contained thought. Shown next to each AI Viral Clips suggestion.
- Viral score
- A rating the AI assigns to each suggested clip, estimating how likely the moment is to perform as a short. Used to sort suggestions by potential before you commit to editing.
- Highlight detection
- The step where ClipSonic finds the best moments in a long video — either by on-device scoring (Smart Viral Clips) or by an AI model reading the transcript (AI Viral Clips).
- Face-tracked reframing
- Automatically detecting and following faces frame by frame so the speaker stays centered when a widescreen 16:9 video is cropped to a vertical 9:16 short.
- 9:16 reframing
- Converting a widescreen 16:9 video into the 9:16 vertical aspect ratio used by YouTube Shorts, TikTok, and Instagram Reels.
- Duo split-screen
- A framing mode that tracks two speakers independently and stacks them in the 9:16 frame — used for co-hosted podcasts and interviews.
- Gameplay + facecam layout
- A framing mode that places a tracked facecam over a full-frame gameplay background, built for streamers and gaming creators.
- Animated captions
- Word-by-word highlighted subtitles burned into the video. ClipSonic ships six presets — Clean, Karaoke, Bold Pop, Neon Glow, Bounce, and Boxed — each draggable to any position on the frame.
- Local rendering
- Doing all the video work — cutting, reframing, captioning, and exporting — on your own computer via a bundled ffmpeg engine, with no cloud render queue and no upload of the source video.
- Local Whisper transcription
- Running the speech-to-text engine (whisper.cpp) directly on your computer, GPU-accelerated on most machines, with word-level timing for captions — no API key and no per-minute cloud fee.
- Transcript
- The timed text of everything spoken in a video. ClipSonic uses it to find highlights and to time animated captions word by word.
- One-pass export
- Cutting, framing, and captioning a clip together in a single render, instead of re-encoding the video once per step.
- Batch export
- Rendering every approved clip from a project in one pass, rather than exporting a clip at a time.
- Watermark
- A logo or label overlaid on exported clips by many cloud clipping tools' free tiers. ClipSonic never adds one, on any plan.
- Clip window
- A candidate start/end range in the source video that a highlight-detection mode considers as a possible clip before scoring it.
See these in the app
Download ClipSonic and try all three clipping modes on your own video — free, no account required.