Alibaba

    Happy Horse 1.1

    Multilingual lip-sync specialist for spokesperson content. A text-to-video model built around spoken performance: write the lines into the prompt and Happy Horse 1.1 renders a presenter who says them, with lip-sync across English, French, Spanish, Turkish, Japanese and more.

    The Studio catalog runs Happy Horse 1.0 image-to-video, not the 1.1 text-to-video endpoint on this page — the first button opens our text-to-video workflow, where comparable models are available today.

    Duration
    3-15s
    Max resolution
    1080p
    Audio
    Native
    Aspect ratios
    Nine, from 21:9 to 9:21
    Commercial use
    Allowed

    Features

    What Happy Horse 1.1 does differently.

    01

    Multilingual lip-sync

    Spoken lines are matched to mouth movement in English, French, Spanish, Turkish, Japanese and further languages, so a presenter reads the script rather than mouthing at it.

    02

    Native audio from the prompt

    Dialogue is written into the prompt and produced in the same pass as the picture, which keeps delivery and lip movement on the same clock without a separate dubbing step.

    03

    1080p at three to fifteen seconds

    Renders at 720p or 1080p for any duration between 3 and 15 seconds, long enough for a full spokesperson beat in one take.

    04

    Cinematic camera and grade

    Shallow depth of field, film grain and camera moves are part of the model's output, so a talking head does not have to look like a webcam.

    Examples

    Prompts that show its range.

    Medium shot professional news anchor at sleek desk, cool blue lighting. 0-5s looks into camera 'Good evening. Tonight, a breakthrough...' 5-10s turns to second camera 'full story right after this.' Precise lip-sync, broadcast quality, shallow DoF.

    Try this prompt

    Instructor to camera in a bright studio, speaking in French, warm key light, slow push in, shallow depth of field.

    Try this prompt

    Vertical 9:16 spokesperson in a lobby, speaking in Japanese, handheld camera, film grain.

    Try this prompt

    Happy Horse renders each example on demand rather than serving fixed clips, so prompts are shown instead of video. Run any of them in the Studio to see your own render.

    Workflows

    Every way to run Happy Horse 1.1.

    Billing is per second of output and steps up with resolution, so a short 720p take costs a fraction of a full-length 1080p one. Credits come from one shared balance across image, video and speech, and the exact estimate is shown in the Studio before you generate.

    Built for

    • Spokesperson and talking-head video: presenters, anchors, instructors
    • Multilingual content delivered without a dubbing pass
    • Cinematic shot-by-shot sequences with timed dialogue
    • Social cuts in vertical and square ratios from the same script

    FAQ

    Happy Horse 1.1, answered.

    Limits, audio, licensing and price — the things worth knowing before you spend a credit.

    Write the lines, get the performance.

    Open the Studio, start a text-to-video render, and see the credit estimate before you spend anything.