The AI documentary stack in 2026: the six-tool workflow, and the one-tool version
The full stack creators assemble for AI-assisted documentaries (script, voice, footage, maps, edit, thumbnails), what each piece costs in time, and how the middle of the stack collapses into a template.
Ask any AI assistant "how do I make a documentary with AI" and you get the same stack: an LLM for the script, ElevenLabs for the voice, Veo or Runway for footage, stock libraries for the rest, DaVinci Resolve to cut it, Midjourney for the thumbnail. It's a good answer. It's also 30 hours of work per film. Here's the stack with real time costs, and the shorter version.
The six-tool stack
| Job | Tool | Time per 15-min film |
|---|---|---|
| Research and script | Claude / ChatGPT + your rewrite | 3–5 h |
| Narration | ElevenLabs / Gemini voices | 1 h |
| Footage: archive and stock | Internet Archive, Commons, Storyblocks | 8–12 h |
| Gap shots | Veo (Flow), Runway, Kling | 1–2 h |
| Maps | After Effects / Fusion | 3–6 h |
| Edit, sound, colour | DaVinci Resolve / Premiere | 12–20 h |
| Thumbnail | Midjourney + Canva | 1 h |
| Total | 29–47 h |
Two things stand out. The creative decisions (script, review) are a small share. And the biggest line, the edit, is where taste lives; it can't be handed to a generator without the result turning generic.
Where each piece stops
- Script tools draft essays unless forced into a film shape. See the master prompt.
- Voice tools need direction and word timing to drive a cut.
- Generators produce shots, not films, and only belong in gaps.
- Stock covers connective tissue, not the subject.
- Editors give control and cost the hours.
The collapsed stack
Glisse replaces the middle four rows with a template made by a real documentary editor. Script in; the AI performs the narration with word timing, searches licensed archive collections per line, builds maps for the places and years, cuts in the editor's pacing and runs a sound-design pass. Generated gap shots import where needed.
| Job | Tool | Time |
|---|---|---|
| Research and script | Claude / ChatGPT + rewrite | 3–5 h |
| Build (voice, footage, maps, cut, sound) | Glisse, from an editor's template | under 1 h |
| Review and changes | Glisse, plain-language requests | 1–2 h |
| Gap shots (if any) | Veo / Runway, imported | 0–1 h |
| Thumbnail | Midjourney + Canva | 1 h |
| Total | 6–10 h |
The taste in the film is the editor's; the hours saved are sourcing and assembly. See a documentary editing workflow that ships weekly.
When to keep the full stack
Feature-length festival films, projects with a client who wants frame-level control, or when you're the editor and the edit is the point. Even then, many editors generate the assembly in Glisse and finish in Resolve.
FAQ
What tools do I need to make a documentary with AI?
At minimum: a script tool, a voice, footage sources, maps and an editor. Glisse combines the voice, footage, maps and edit around an editor's template.
Is DaVinci Resolve still needed with AI tools?
For frame-level finishing, colour and bespoke graphics, yes. For the assembly, a template-based build gets there in a fraction of the time.
Which is cheaper, the full stack or Glisse?
The full stack's tools are mostly free or cheap; the cost is 30+ hours. Glisse charges credits per build and saves most of those hours.
Make documentaries with an editor's taste
Glisse turns a real editor's style into a template. Describe your topic, and AI adapts it into a finished long-form film: licensed archive footage, maps, narration and music. Editors earn every time their template is used.
Join the waitlist →