Subtitles
Transcribe a voiceover and drop timed, styled subtitle text onto your After Effects timeline — a capability AE has no native answer for.
After Effects can place text and read a timeline, but it has no way to listen to a voiceover and turn it into timed captions. Prism adds that. Point Prism at the audio in your composition and it transcribes the speech, then writes a real text layer whose words are timed to the spoken track — native, editable, and yours to restyle. If you only need the timed text, not a caption layer, use Transcribe instead.
What it does
Subtitle generation is one of Prism's AI-native tools. Prism transcribes your spoken track on its server and uses the timestamped result to build a subtitle text layer on the timeline. The result is the same kind of object you would create by hand: a text layer you can re-time, re-word, restyle, or split.
As with every Prism action, the AI reads the live state of your comp before it writes the subtitles as native layers (see how Prism works).
Before you start
- Your composition contains the voiceover or audio track you want captioned.
- Prism is installed and activated, and your AI client points at Prism's MCP server. If not, see setup and connect your AI.
- The comp you want captioned is the active comp in After Effects.
How to use it
Run it from the Tools tab, or ask your AI in plain language — Prism runs the same tool either way.
Audio is sent only to produce the transcript, then discarded — your project stays on your machine (how Prism works).
A worked example
| You say | Prism does |
|---|---|
| "Transcribe the voiceover and add captions, one running line, no stroke." | Reads the audio in your comp, transcribes it, and writes a single subtitle text layer timed to the speech. |
| "Now split those captions per word, karaoke style." | Rebuilds the same layer with one timed word at a time, using the per-word timing from the transcript. |
What you get back: a native subtitle text layer on the timeline, its words aligned to the spoken track and ready to restyle.
If the audio is silent or unreadable: the request surfaces "no speech detected in the audio." Check that the right audio layer is in the active comp and is not muted, then ask again.
Languages and caption format
Prism transcribes a wide range of languages and accents, so most voiceovers — not just English — come back clean. Timing comes back per word, which gives you two ready formats:
| Format | Looks like | Good for |
|---|---|---|
| One running line | A full caption line on screen at a time | Explainers, talking-head, narration |
| Per word | One word revealed at a time | Punchy, karaoke-style reads |
You can ask for either when you generate the subtitles, or switch between them afterward — the layer is ordinary AE text.
Styling and timing
The subtitle layer is ordinary AE text, so anything you would do by hand still applies — change the font, add an animator, reposition, or split lines. One gotcha to know: if glyphs look clipped, ask for subtitle text without a stroke — see troubleshooting for that and other AE text quirks.
Every Prism action is undoable in one step and protected by a restore point, so experimenting with styling is safe (how Prism works).
Notes and limits
Transcription draws on AI credits — your prepaid wallet for heavy, generative AI — at 5 credits per minute of audio (so a 3-minute track costs about 15 credits). For the monthly credit allowance and how metering works, see usage and limits.
Related
- Usage and limits — AI-credit allowances and how metering works.
- Transcribe — the raw transcript and per-word timing, without a caption layer.
- BPM and beat markers — detect tempo and write beat markers to the timeline.
- SVG import — bring in SVG as real text and shapes, not outlines.
- AI-native tools — the family of capabilities AE has no native answer for.