Skip to content
Prism
Prism AI

Transcribe

Transcribe turns the audio in a layer into text you can actually work with. Point it at a voiceover, an interview, or a video clip, and it gives you the full transcript with timing for every word. It is the starting point whenever you need to caption, subtitle, or sync motion to what is being said, because once you have timed words, you can place markers, build subtitle layers, or cut to the beat of speech. If what you want is finished caption layers on the timeline, use Subtitles directly: it runs the same transcription at the same rate, so there is no need to pay twice.

The short version
Select an audio or video layer, then run Transcribe to get the full text plus a timestamp for every word.
The result lands as one marker per word on a new guide layer above your source: your original layer is never touched.
Language is auto-detected, or you can set it (for example English, Spanish, French, German) for cleaner results.
Charged per minute of audio, from your AI balance.

How to use it

1
Select the layer

In your composition, select the audio or video layer you want to transcribe. Any layer with an audio track works. There is no need to export or pre-render the audio first.

2
Open the Prism AI tab

Open the panel and switch to the Prism AI tab. Transcribe is one of its six capabilities.

3
Run Transcribe

Run the tool. If you know the spoken language you can set it; otherwise leave it on auto-detect. The audio is uploaded and transcribed, then the result comes back into the panel.

4
Or ask your AI

You can also just ask your AI in chat to transcribe a layer. It runs the same tool for you, defaults to the selected layer in the active composition, and gets back the timed words to use however you need.

What you get

When the transcription finishes, two things happen:

  • On the timeline: Prism creates a new guide layer named [Transcript] <your layer name> directly above the source layer, and places one timeline marker per word, each at the exact time that word is spoken and timed to the source layer, so everything lines up frame-accurately.

A few specifics worth knowing:

  • Your source layer is never changed. All output goes onto a separate guide null layer, so the audio or video stays exactly as it was.
  • Re-running refreshes cleanly. If a [Transcript] guide already exists for that layer, its markers are cleared and rebuilt. You will not get duplicates stacking up.
  • Word-level timing is the whole point. Every word carries a start and an end, which is what makes it the source material for subtitle layers and for cutting to speech.
  • Wide format support. It accepts the common audio and video formats After Effects handles: for example MP3, WAV, M4A, FLAC, OGG, MP4, MOV, and WEBM. Large video files are supported, up to roughly 900 MB per upload.
  • Markers outside the layer's visible range are skipped. Anything that would land before the layer starts or after the composition ends is left off; the panel reports how many markers were placed and how many were skipped.

Notes and limits

Transcribe is charged per minute of audio, taken from your AI balance. The current rate is in your dashboard. For how the balance is metered and how plan caps work, see usage and limits.

  • Subtitles: turn the same transcription into a timed caption layer on the timeline
  • Usage and limits: how your AI balance is metered and where to see the current rates
  • Plans and pricing: Free vs Pro, and how the prepaid AI balance works
  • The Prism AI: all six capabilities and what each costs
  • Troubleshooting: text and format gotchas in After Effects