The gap between a bad AI visual and a good one is usually the description, not the model. Here is the structure that works and the mistakes that produce mush.
The order that works
Put the information in this order and results improve immediately:
- Shot type. Wide shot, close up, aerial, low angle.
- Subject, concretely. A rusted red bicycle against a wet brick wall.
- Lighting and time. Golden hour, overcast, harsh noon, neon at night.
- Motion. Slow dolly in, subject moves left to right, camera static.
- Look. Grainy 16mm, high contrast, muted palette.
Models weight the front of a prompt heavily, which is why shot type and subject go first and stylistic notes go last.
Concrete nouns beat adjectives
This is the single biggest improvement available.
Weak: a moody atmospheric urban scene with a sad vibe
Strong: a wet alleyway at night, one flickering strip light, steam from a vent, no people
The first describes a feeling and leaves every actual decision to the model. The second describes a picture. Feelings are what you want the viewer to have, not what you should be typing.
One idea per shot
Prompts asking for two events produce neither. "A woman walks through a door into a forest" is two shots, and asking for it in one generation reliably produces something incoherent.
Split it. Shot one is the door, shot two is the forest. You wanted a cut there anyway.
State the motion explicitly
Unstated motion is where models improvise, and they improvise badly. If you do not say what moves, you will get drifting, warping, and objects quietly changing shape.
Say it:
- "Camera static, only the rain moves"
- "Slow push in, subject still"
- "Locked off, smoke drifting left"
"Camera static" is the most useful two words in AI video prompting and almost nobody uses them.
Known failure cases
Avoid or work around these, because they fail on every current model:
- Hands, particularly doing anything specific
- Legible text. Signs, logos, writing of any kind
- Crowds. Faces in the background will be wrong
- Fast complex motion. Dancing, sports, anything with limbs at speed
- Continuity across a cut. The same object will not match between two generations unless you use a reference image
If a shot needs one of these, either design around it or accept that you will re-roll several times.
Use a reference image for consistency
This is how you keep a world coherent across eight scenes. Generate one frame you are happy with, then use it as the reference for everything else. It carries palette, setting and subject across shots in a way that no amount of prompt repetition will.
Prompt templates
Atmospheric loop, works for Canvas: Close up, ink dispersing in water, black background, single light source from above, camera static, slow motion, high contrast
Establishing scene: Wide shot, empty coastal road at dawn, low fog, overcast light, camera static, muted palette, 16mm grain
Subject scene, no face risk: Medium shot from behind, figure in a long coat walking away, rain-soaked street, neon reflections, slow dolly follow
Texture cutaway: Extreme close up, rain hitting a car window at night, out of focus streetlights behind, camera static, shallow depth of field
