A podcast clip with no captions is a clip most people won't watch — the sound is off, and spoken-word content has nothing else carrying it. Here's how to caption clips accurately and without a second app.
Start from a transcript, not manual typing
Typing captions by hand and syncing them is slow and error-prone. Instead, transcribe the clip (or the whole episode) with word-level timestamps, and the caption timing is already done — each word appears exactly when it's spoken. This is the same transcript that powers highlight detection, so if you clipped with AI you already have it.
Style choices for spoken word
- One or two lines, not a paragraph. Big text, few words on screen at once.
- Word-by-word highlight. Keeps a muted viewer's eye moving with the audio — see word-by-word vs line captions.
- High contrast. A solid or outlined style over a busy podcast set beats thin plain text.
Positioning for one vs two speakers
Single speaker: lower third, clear of the face and any lower-third graphics. Two-speaker split-screen: centre the caption block between the stacked panes, or put it along the bottom — anywhere it doesn't cover either person's mouth or eyes.
Pick one caption style per show. Consistency makes your clips recognisable in a feed; switching styles every clip just looks unfinished.
Burn them in
Export with the captions rendered into the video, not as a separate subtitle file. Burned-in captions display identically on every platform and can't be turned off by a viewer who'd otherwise never enable them.
Doing it in ClipSonic
ClipSonic generates captions from the clip's transcript, offers six animated styles with word-by-word highlight, lets you drag the caption anywhere on the frame, and burns them into the export — all in the same pass as the cut and the reframe. More: animated captions and the AI podcast clipper.
Download ClipSonic free and caption a clip from your latest episode.