GlisseBlog › Tools compared
Tools compared

ElevenLabs for documentary narration: getting a performed voice, not a read

How to use ElevenLabs for long-form documentary narration: voice choice, performance direction, consistency across a 20-minute script, and timing the edit to the words.

Updated 2026-09-16 · 2 min read · by the Glisse team

ElevenLabs is the narration tool most documentary creators end up with, because its newer models can be directed: told to pause, slow down, emphasise, and change register. That's the difference between a voice that reads and a voice that performs. Here's how to get the second one.

Choosing a voice

Directing the performance

The performed models accept direction inside the script: a pause here, emphasis on a word, a quieter register for a line. Use it per line rather than per script. The script should already be written for it: short sentences, pauses on the page, dates in their own sentences. See how to write documentary narration.

Consistency over 20 minutes

Generate the whole script in one pass where the tool allows it, with the same settings. Regenerating one paragraph later with different settings produces an audible seam. If you must re-record a paragraph, re-record the whole chapter.

Word timing: the step people skip

The narration file isn't enough for a documentary cut. Editors place the picture change a beat after the key word; the map moves on the place name; the card lands on the pause. That needs word-level timestamps.

Glisse's pipeline treats narration as the first step of every build: the script is performed (ElevenLabs' performed models, with Gemini voices as an alternative), every word is timestamped, and the edit is cut onto the words. Change the script and the film re-cuts to the new take rather than stretching the old pictures. See AI voiceover for documentaries.

Music and sound alongside

Narration sits 6–9 dB above the music bed, with the bed low-passed slightly under speech. Drop the music entirely for the key line. See music for documentaries.

FAQ

Is ElevenLabs good for documentaries?

Yes; its performed models take per-line direction for pauses and emphasis, which documentary narration needs.

Which ElevenLabs voice is best for narration?

A measured mid-register voice that handles lists of dates and one-word lines without rushing or flattening. Test before committing.

How do I sync ElevenLabs narration to my edit?

Get word-level timestamps and anchor cuts to them. Glisse does this automatically: the edit follows the narration's words.

Make documentaries with an editor's taste

Glisse turns a real editor's style into a template. Describe your topic, and AI adapts it into a finished long-form film: licensed archive footage, maps, narration and music. Editors earn every time their template is used.

Join the waitlist →