Text-to-Video Prompts That Actually Look Good
Most AI video looks generic because the prompt was generic. Here is the prompt anatomy experienced creators rely on (subject, action, setting, camera, lighting, style) to land clean, cinematic results on the first generation.

Text-to-video models are good enough now that the bottleneck has moved from the model to the prompt. Type 'a car driving on a highway' and you get something flat, soft, and forgettable, because the model fills every gap with the statistical average of its training data. That is exactly why so much AI video looks interchangeable. Feed the same model a structured text-to-video prompt and it returns a low-angle tracking shot at golden hour with motion blur and graded color. This guide walks through the anatomy of a strong prompt, hands you a copy-ready formula, contrasts weak versions against strong ones, and covers model-specific tips for Sora 2, Veo 3, and Kling so you burn fewer generations.
Why prompt structure matters
A text-to-video model doesn't picture a scene the way a director does. It predicts frames from your words, weighting whatever is most specific. Hand it something vague like 'a woman walking in a city' and it defaults to the most average interpretation it has seen: even framing, flat light, no point of view. The output is technically correct and visually dead, which is the real reason so much AI video looks identical.
Structure fixes that by telling the model not only what to show but how to show it. Name a camera move, a lighting condition, and a mood, and you collapse millions of plausible 'average' outputs down to the narrow cinematic slice you actually wanted. Two people on the same model can land worlds apart: one is describing a subject, the other is describing a shot.
The anatomy of a strong video prompt
Every prompt that produces clean, cinematic footage is really answering six questions. Flowery language isn't the point; hitting all six clearly is. Picture a shot list compressed into a single sentence.
- 1
Subject
Who or what is on screen, described concretely. 'A weathered fisherman in a yellow raincoat' beats 'a man' because the model now has texture, color, and character to render.
- 2
Action
What the subject is doing, as one clear motion. Movement is the thing video does that a still can't, so name it: 'pulling a net over the side of the boat,' not just 'standing.'
- 3
Setting
Where and when. Time of day, location, and weather anchor the scene. 'On a small fishing boat at dawn, fog rolling over the water' hands the model a world to build.
- 4
Camera move
How the shot is captured: 'slow dolly-in,' 'low-angle tracking shot,' 'handheld follow,' 'static wide.' This one phrase is the biggest lever you have for a cinematic feel.
- 5
Lighting
The light source and its quality: 'golden-hour backlight,' 'soft overcast,' 'neon rim light,' 'harsh midday sun.' Lighting is what separates 'rendered' from 'shot.'
- 6
Style / mood
The overall look and emotional register, like 'cinematic, shot on 35mm, moody and contemplative' or 'bright, hyper-real, energetic.' This ties the whole frame together.

Order helps, coverage matters more
Lead with subject and action, then layer in setting, camera, lighting, and style, since most models weight the front of a prompt. What matters most, though, is that all six show up. Missing layers are exactly where the generic AI look creeps back in.
The copy-ready formula
Once the six layers stick, you can write strong prompts almost on autopilot with a fill-in-the-blanks template. Hold yourself to one shot and one continuous motion.
[Camera move] of [subject, specific] [single clear action], [setting + time of day], [lighting], [style / mood], 9:16
Slot your idea into each blank and the prompt is complete. Filled out, it looks like this:
Low-angle tracking shot of a red Porsche speeding down a coastal highway, late afternoon with the ocean on the right, warm golden-hour backlight and lens flare, cinematic, shot on 35mm film, 9:16
Write the prompt as a shot description, not a wish. Concrete nouns and named techniques (dolly-in, rim light) carry far more weight than 'beautiful.'
— LazyKiwi creator playbook
Weak vs. strong examples
Nothing internalizes the formula faster than seeing one idea written badly and then well. The strong versions aren't padded for length; every added word is pulling its weight.
| Weak prompt | Strong prompt |
|---|---|
| A woman walking in a city | Slow dolly of a young woman in a red coat walking toward camera through a rain-slicked Tokyo alley at dusk, neon backlight, anamorphic, cinematic, 9:16 |
| A zombie smoking | Slow dolly-in on a decaying zombie leaning against a brick wall lighting a cigarette, dim alley at night lit by a flickering neon sign, cold blue rim light, moody horror, shallow depth of field |
| A farmer in a field | Static wide shot of an aging farmer walking through rows of dry crops at sunset, long shadows and warm dusty haze, soft golden backlight, contemplative documentary mood, shot on film |

Add one motion, not three
Describe the subject doing several things at once and the model blends them into mush. Pick the single most important action and let the camera move carry everything else.
Model-specific tips
The formula travels across models, but each flagship rewards slightly different habits. Tuning to whichever one you're on will save you regenerations.
- Sora 2 handles complex, multi-beat descriptions and holds up well on physics and continuity. Lean into rich, narrative prose and named camera language, since it follows full cinematic sentences closely rather than keyword lists.
- Veo 3 shines on photorealism and native audio, so describe the soundscape too ('ambient rain, distant traffic') when sound carries the shot. Be explicit about camera moves and lighting and it returns polished, broadcast-grade frames.
- Kling is strong on expressive human motion and dynamic action, and it's great value for fast iteration. Keep prompts on one clear subject and one strong movement; it can drift in crowded scenes, so simplify the setting and let the motion lead.
per prompt, don't pack a whole scene into one clip
Whatever you're running, generate the same prompt twice before you rewrite it. The variation between seeds tells you whether the problem is your prompt or just an unlucky roll.
Common mistakes to avoid
| Mistake | Fix |
|---|---|
| Vague subject ('a person') | Describe wardrobe, age and defining details |
| No camera direction | Name a shot type every time |
| Adjective soup ('beautiful, epic') | Swap vibe words for techniques and nouns |
| Cramming multiple actions | One clear motion per shot |
| Ignoring lighting | Specify the light source and quality |

Cut these five and your first-generation hit rate climbs fast. You won't nail a perfect first try every time, but a prompt that's specific enough makes 'good' the likely outcome instead of a lucky one.
Key takeaways
- Models render whatever is most specific, so vague prompts produce flat, average footage.
- A complete prompt covers six parts: subject, action, setting, camera move, lighting, and style/mood.
- Use the formula: [camera move] of [subject] [action], [setting], [lighting], [style/mood], 9:16.
- Camera direction and lighting are the two biggest levers for a cinematic look.
- Tune to the model: Sora 2 for continuity, Veo 3 for photoreal plus audio, Kling for fast single-subject motion.
Creator Growth Lead
Marcus Tan
I study what makes short-form content spread and turn it into repeatable playbooks you can run inside the LazyKiwi workbench.
FAQ
Common questions
How long should a text-to-video prompt be?
One to two clear sentences that cover all six components. Longer isn't better; past a couple of sentences, models start weighting the front and ignoring the tail. Pack the detail into concrete nouns and named camera and lighting techniques rather than extra clauses.
Why does my AI video look flat and generic?
Almost always because the prompt is missing camera direction and lighting. Without a named shot type ('low-angle tracking shot') and a light condition ('golden-hour backlight'), the model falls back on even framing and flat light. Add those two and most clips immediately look more intentional.
Should I describe multiple actions in one prompt?
No. Each clip should be a single shot with one continuous motion. Describe several actions at once and the model blends them into unnatural movement. Generate separate clips for separate actions and stitch them together afterward.
Do I need different prompts for Sora 2, Veo 3, and Kling?
The core formula works across all three, but each rewards different habits. Sora 2 handles multi-action sequences, Veo 3 benefits from describing sound and photoreal lighting, and Kling does best with one clear subject and one strong movement. Adjust the emphasis, not the underlying structure.
Make a sentence into a cinematic shot.
Use the prompt formula, pick your model, and generate clean text-to-video that holds up on the first try, all inside LazyKiwi.
Keep reading

Sora 2 vs. Veo 3: Which Video Model Should You Use?
When each flagship model wins.
7 min read
How to Make Scroll-Stopping AI Effect Videos for TikTok, Reels & Shorts
Build a hook people can't scroll past.
8 min read
Turn One Photo Into a Week of Short-Form Content
Multiply one image into many posts.
5 min read