Transcribe
Transcribe turns the audio in a layer into text you can actually work with. Point it at a voiceover, an interview, or a video clip, and it gives you the full transcript with timing for every word. It is the starting point whenever you need to caption, subtitle, or sync motion to what is being said, because once you have timed words, you can place markers, build subtitle layers, or cut to the beat of speech. If what you want is finished caption layers on the timeline, use Subtitles directly: it runs the same transcription at the same rate, so there is no need to pay twice.
How to use it
In your composition, select the audio or video layer you want to transcribe. Any layer with an audio track works. There is no need to export or pre-render the audio first.
Open the panel and switch to the Prism AI tab. Transcribe is one of its six capabilities.
Run the tool. If you know the spoken language you can set it; otherwise leave it on auto-detect. The audio is uploaded and transcribed, then the result comes back into the panel.
You can also just ask your AI in chat to transcribe a layer. It runs the same tool for you, defaults to the selected layer in the active composition, and gets back the timed words to use however you need.
What you get
When the transcription finishes, two things happen:
- On the timeline: Prism creates a new guide layer named
[Transcript] <your layer name>directly above the source layer, and places one timeline marker per word, each at the exact time that word is spoken and timed to the source layer, so everything lines up frame-accurately.
A few specifics worth knowing:
- Your source layer is never changed. All output goes onto a separate guide null layer, so the audio or video stays exactly as it was.
- Re-running refreshes cleanly. If a
[Transcript]guide already exists for that layer, its markers are cleared and rebuilt. You will not get duplicates stacking up. - Word-level timing is the whole point. Every word carries a start and an end, which is what makes it the source material for subtitle layers and for cutting to speech.
- Wide format support. It accepts the common audio and video formats After Effects handles: for example MP3, WAV, M4A, FLAC, OGG, MP4, MOV, and WEBM. Large video files are supported, up to roughly 900 MB per upload.
- Markers outside the layer's visible range are skipped. Anything that would land before the layer starts or after the composition ends is left off; the panel reports how many markers were placed and how many were skipped.
Notes and limits
Transcribe is charged per minute of audio, taken from your AI balance. The current rate is in your dashboard. For how the balance is metered and how plan caps work, see usage and limits.
Related
- Subtitles: turn the same transcription into a timed caption layer on the timeline
- Usage and limits: how your AI balance is metered and where to see the current rates
- Plans and pricing: Free vs Pro, and how the prepaid AI balance works
- The Prism AI: all six capabilities and what each costs
- Troubleshooting: text and format gotchas in After Effects
Compatibility
Which After Effects versions, operating systems, and AI clients Prism supports: AE 2024/2025/2026 on macOS and Windows, with any MCP client.
BPM
Prism detects BPM and beats in your audio and places frame-accurate timeline markers in After Effects, giving every animation a rhythm grid to land on.