All guidesProduction9 min read

Directing audio and lip sync

HappyHorse generates sound in the same pass as the picture. Not a stock soundtrack laid underneath — the wave's crash is rendered by the same process rendering the wave, which is why sounds land on their actions. This guide covers the three tiers of audio direction, from ambience to scripted multilingual dialogue.

Tier 1: ambience — the one-sentence upgrade

Every prompt should end with an audio sentence, even for "background" content. The pattern that works is two or three named layers, roughly ordered near-to-far:

Audio: rain on pavement, distant traffic, muffled city hum.

Audio: waves breaking, receding foam hiss, deep night silence between sets.

Layered direction reads as a mixed track; a single vague word ("city sounds") reads as a preset. Notice the third layer in each example is spatial — hum, silence, room tone. That far layer is what makes generated audio feel mixed rather than pasted.

Silence is also a direction. ASMR and product work often need "no music" stated explicitly — omit it and the model may helpfully add a soundtrack over the sounds you actually wanted. See the coffee ASMR prompt, where the audio is the content.

Tier 2: synced beats — sound glued to action

Because audio and video share one generation, you can tie a sound to an action with plain wording:

…throws a single punch toward camera … Audio: one deep bass impact, slowed breathing.

…snaps her fingers, and in a whip-pan blur her outfit transforms … Audio: finger snap, whoosh, upbeat pop sting after the reveal.

Two techniques hiding in those examples:

  • Countable language. "One deep bass impact" gives the engine a beat to hit. "Intense dramatic music" gives it a mood with no timing. Count your sounds when timing matters.
  • Sequence words control sync order. "Sting after the reveal" places the sound relative to the visual event. The prompt's word order is your timeline.

This is the cheapest post-production you'll ever do: edits cut on sound beats feel intentional, and your beats now arrive pre-synced.

Tier 3: lip-synced dialogue

The headline feature. Put the exact line in quotation marks and a character will speak it, mouth movement matched, in any of seven supported languages: English, Mandarin, Cantonese, Japanese, Korean, German, and French.

The anatomy of a dialogue prompt, from our founder talking-head:

The woman from the image, framed chest-up … looks into camera and says warmly in English: "We built this for small teams — and it finally feels effortless." Natural hand gesture on "effortless" … Audio: her voice, clear and close-mic'd, faint office ambience.

Six rules that make or break it:

  1. Exact script in quotes. The engine syncs to those words verbatim. Paraphrase in, garbage out.
  2. Name the language. "Says in Japanese:" — then write the line in Japanese. Don't write English and ask for translation.
  3. Keep lines under ~20 words per 8-second clip. Longer scripts get rushed delivery and slippage at the tail.
  4. Direct the delivery. "Warmly", "deadpan", "through a laugh" — one adverb of performance direction goes a long way.
  5. Frame for the mouth. Chest-up or closer, face toward camera. Lip sync on a wide profile shot is wasted budget.
  6. Add a non-verbal beat after the line. A laugh, a nod, a sip — see the street interview. Sync that continues past the words is what makes viewers stop doubting.

Two speakers

Current practical ceiling: one line per speaker per camera angle, with the cut on the dialogue turn — the pattern in the café exchange prompt. Two people trading lines inside one continuous frame still de-syncs often enough that we don't recommend it for deliverable work.

A director's checklist

Before you generate, scan your audio sentence:

  • Two or three layers, near to far?
  • Timed sounds countable and tied to action words?
  • "No music" stated if you don't want music?
  • Dialogue: quoted script, named language, delivery note, close framing?

Half the model is the soundtrack. Direct it like you direct the frame — then go make sure your durations and aspect ratios aren't quietly working against you.