Writing prompts that direct, not describe
The single biggest upgrade in AI video prompting is a change of job title. Most people prompt like a novelist — describing a world and hoping the model films it well. The people getting consistently good clips prompt like a director: they specify the shot, the performance, the light, and the sound, and leave the world-building to the model.
This guide teaches the five-part structure used by every prompt in our library. It isn't official syntax — HappyHorse accepts free text — but it maps onto what the model actually pays attention to.
The five-part structure
Subject → Action → Camera → Light & style → Audio
Here's the night rain tracking shot with its seams showing:
Slow tracking shot following [camera] a woman in a yellow raincoat [subject] walking down a neon-lit Tokyo alley at night [action + world], heavy rain, reflections on wet asphalt, shallow depth of field, anamorphic lens flare, cinematic teal-and-orange grade [light & style]. Audio: rain on pavement, distant traffic, muffled city hum [audio].
Five decisions, one sentence each. Let's take them in order of how often people get them wrong.
1. One subject, made findable
The model allocates detail to what you name first and most specifically. "A woman in a yellow raincoat" beats "a person" because the raincoat is a visual anchor — one saturated object the eye (and the model) can track through every frame. Crowds, pairs, and "various people" divide the budget; unless the format demands otherwise, cast one lead.
2. One action, with a beat
An action the model can complete in 5–8 seconds, ideally with a beginning and an end: walks down the alley, throws a single punch, slowly closes her eyes. The classic mistake is scripting a sequence — "she enters, sits down, orders coffee, and smiles" is four shots, not one. If your action has commas in it, you're writing a scene, and the model will smear all four beats into mush.
3. One camera instruction
Choose from the vocabulary real crews use — the model knows it:
- Static / locked — safest; all motion budget goes to the subject
- Slow push-in / pull-back — adds intent without parallax risk
- Tracking / following — travels with the subject; great for walks and runs
- Drone rise / aerial reveal — earns a wide; name what gets revealed
- Rack focus — depth motion with zero travel; underused and reliable
- Handheld energy — organic micro-shake for UGC-flavored work
One verb only. "Tracking shot that rises into a drone reveal" is two shots fighting inside one clip — the number-one cause of warped geometry mid-clip.
4. Light and style, in that order
Light is structural: name its source and quality ("hard single-source rim light from the left", "soft window light", "golden hour with long shadows"). Style rides on top as grade and medium: "teal-and-orange grade", "film grain", "cel-shaded 90s anime".
What doesn't work is quality-begging: "masterpiece, ultra-detailed, 8k, best quality" is image-model folklore and does nothing here except dilute the words that matter. Say how it looks, never how good it looks.
5. Audio: the line most prompts are missing
End every prompt with an explicit audio sentence. Three moves cover almost everything:
- Layered ambience — "rain on pavement, distant traffic, muffled city hum" (two or three layers reads as a mixed track)
- Synced beats — "one deep bass impact as the punch lands" (tie the sound to the action word)
- Negative direction — "no music" (silence and restraint are instructions too)
For dialogue, put the exact line in quotation marks and describe the delivery — that's a big enough topic that it has its own guide.
Worked example: fixing a flat prompt
Before:
A beautiful cinematic video of a magical forest with amazing lighting, ultra realistic, 8k
Every word is either vague ("beautiful", "magical", "amazing") or begging ("ultra realistic, 8k"). There's no subject, no action, no camera, no sound.
After:
Slow push-in through ancient mossy trees toward a single shaft of dawn light hitting a fern grove, drifting pollen glowing in the beam, cool shadow around the edges of frame, gentle mist at ground level. Audio: forest dawn chorus, one distant woodpecker, soft wind in high branches.
Same forest. But now the model knows where the camera goes (push-in), what it's moving toward (the light shaft — a destination gives a push-in purpose), what floats in the light (pollen — motion for free), and what the forest sounds like. This is the difference between describing and directing.
Common failure modes, quick reference
| Symptom | Likely cause | Fix |
|---|---|---|
| Subject morphs mid-clip | Action sequence too long | One action, one beat |
| Geometry warps at edges | Two camera moves | One camera verb |
| Generic "AI look" | Quality-begging words | Replace with concrete light/grade |
| Random soundtrack | No audio direction | Add an explicit audio line |
| Cluttered, unfocused frame | Multiple named subjects | One lead with a visual anchor |
Practice deliberately
Pick any entry in the prompt library, run it as written, then change one part — swap the camera verb, or rewrite only the audio line — and run it again. Ten single-variable experiments will teach you more than a hundred from-scratch attempts.
When your subject needs to be exactly right — a product, a face, a brand asset — text alone stops being enough. That's when you graduate to the image-to-video workflow.