Sora, Veo and the AI slop problem: why generated video can't carry a documentary
Text-to-video models like Sora and Veo make convincing seconds of footage. Here's why films built from them feel hollow, what "AI slop" actually is, and what separates a generated video from an edited one.
Every few months a new text-to-video model makes a clip that fools people for five seconds. Then someone strings forty of those clips into a "documentary", and the result is what the internet now calls slop. The problem isn't the model. It's the missing editor.
What AI slop actually is
Slop is content with no point of view. The tell-tale signs in video:
- Constant pace. Every shot lasts the same 4 seconds because a script was split into sentences and each sentence got a clip.
- Nothing is real. Generated footage of real events. Viewers can't say why, but they stop trusting the narrator.
- A voice that never breathes. Narration read at one speed with no silence for a line to land.
- Music as wallpaper. One bed, one level, start to finish.
- Text that's almost words. Image models still can't spell.
None of those are generation failures. They're editing failures. Slop is what happens when a model's output goes straight to the audience without anyone deciding what the film is.
Where Sora and Veo genuinely help
Used as a small part of a documentary, generated video is valuable:
- A reconstruction of something never filmed.
- A diagram or process shown as motion.
- An establishing shot of a place with no usable footage.
The rule that keeps a film credible: real footage for everything real; generated footage only for what can't be filmed, and ideally labelled.
What separates a film from slop: taste
A documentary is decisions: how long to hold an image, when the picture changes relative to the word, when music stops, where the chapters break, what the cold open is. These decisions are what audiences experience as quality, and no video model makes them. Editors do.
That's the premise behind Glisse. Instead of asking AI to invent a film, it captures a real editor's finished documentary as a template (pacing, structure, sound design, map style, titles) and adapts it to your topic. The AI does the sourcing, maps, narration and cutting; the taste is the editor's. Generated footage appears only where the topic has none of its own.
A quick test for your own videos
Watch two minutes with the sound off. If every shot is the same length and nothing on screen is real, it's slop, whatever model made it. If the pace changes, the footage is real and the picture responds to the voice, it's a film.
FAQ
Can Sora or Veo make a documentary?
They can make clips for one. A documentary needs structure, real footage, narration, maps and sound design, which are editing decisions, not generation.
Why does AI-generated video feel fake in documentaries?
Because audiences expect evidence. Generated footage of real events breaks the implicit promise that what's on screen happened.
How does Glisse avoid AI slop?
By starting from a template made by a human editor and using licensed archive footage first. Generated video fills only the gaps where nothing was filmed.
Make documentaries with an editor's taste
Glisse turns a real editor's style into a template. Describe your topic, and AI adapts it into a finished long-form film: licensed archive footage, maps, narration and music. Editors earn every time their template is used.
Join the waitlist →