Image to video: the control workflow
Text-to-video is a slot machine with good odds. Image-to-video is a camera. The moment your work involves something that must look exactly right — a product with a label, a person with a face, a brand with assets — starting from a still image stops being optional and becomes the workflow.
HappyHorse's image mode takes your still as the first frame (or as a reference) and generates motion forward from it. Everything the still contains is inherited: composition, identity, lighting, flaws. That inheritance is the whole trick — and the whole responsibility.
Why professionals default to image-first
- Identity survives. Faces and product geometry are what text-to-video most often ruins. A source image pins them.
- Art direction happens in a stills tool. You can iterate composition in an image generator (or a camera) in seconds per attempt, then spend video credits only on shots already framed right.
- Campaigns need consistency. Five clips grown from five stills of the same photoshoot cut together like a campaign. Five text-to-video rolls of the same prompt cut together like a mood board.
Preparing a first frame that animates well
The model animates what's there, so the still needs room for motion to happen:
- Compose with negative space. A subject filling 100% of frame has nowhere to go. Leave air in the direction of intended motion.
- Light it the way the video should stay lit. Lighting is inherited stubbornly. A flat-lit photo will not become moody footage; bake the mood into the still.
- Straight-on and sharp for products. Glare, reflection, and blur in the source get amplified across frames, not cleaned up.
- Match the still's ratio to the output ratio. Generating 9:16 video from a 16:9 still forces an interpretation you didn't choose.
The golden rule: animate around the anchor
The prompt patterns that work all share one grammar — the subject is explicitly frozen and the world does the moving. Compare the two halves of this instruction from our perfume splash prompt:
"…a thin ring of water splashes up around it … the bottle stays perfectly still and sharp"
That second clause is doing the heavy lifting. Give the model motion to render (splash, light sweep, fog, particles, a rotating pedestal — never the product) and give the anchor an explicit stillness instruction. The sneaker pedestal and skincare morning light prompts are variations of the same move, and it transfers to any object you care about.
For people, the anchor is identity rather than stillness: "the woman from the image" plus a bounded action — a head turn, a gesture, a spoken line — keeps the face consistent. Unbounded actions ("dances energetically") ask the model to re-invent the person, and it will.
The still + motion + audio stack
A complete image-to-video prompt has three layers:
| Layer | Job | Example |
|---|---|---|
| Anchor | What must not change | "The sneaker from the image, camera locked" |
| Motion | What moves around it | "volumetric light beams sweeping, floating dust" |
| Audio | What it sounds like | "airy pulse, soft whoosh as each beam passes" |
If you can't point to all three layers in your prompt, the missing one is where the generation will disappoint you.
When text-to-video is still the right call
Image-first isn't a religion. Reach for pure text when you're exploring (you don't know what you want yet — roll the slot machine), when the subject is generic by design (weather, oceans, crowds, abstract texture), or when the style itself is the subject (watercolor, claymation) and drift is charm rather than defect.
A useful habit: prototype the idea in text mode, then once a composition emerges that works, recreate it as a still and switch to image mode for the keeper takes.
Next
Your subject is now locked and moving. The next lever is the one most creators never touch: directing the soundtrack.