Native multimodal audio
Sound effects land on the action they belong to, with ambience, music and lip-synced dialogue produced in the same pass as the picture. Audio costs nothing on top of the per-second video rate.
xAI
Native multimodal audio, three modes in one model. One model covering text, image and reference input, from 1 to 15 seconds and up to 1080p, with sound effects, ambience, score and lip-synced dialogue generated alongside the picture rather than dubbed on afterwards.
The Studio runs the base Grok Imagine video endpoint today rather than these v1.5 endpoints, so the buttons open our text-to-video and image-to-video workflows; see /models/grok-imagine for the version you can generate with right now.
Features
Sound effects land on the action they belong to, with ambience, music and lip-synced dialogue produced in the same pass as the picture. Audio costs nothing on top of the per-second video rate.
Fluid dynamics, steam, glass and reflections hold together under motion, and faces keep their micro-expressions and eye tracking instead of settling into a stare.
Text-to-video works from a prompt alone, image-to-video treats your still as the first frame, and reference-to-video takes one to seven images that the prompt addresses as IMAGE_1 through IMAGE_7.
480p, 720p and 1080p are billed at separate rates, so a rough cut can be drafted cheaply at 480p and only the approved shot re-rendered at 1080p.
Examples
(0-5s) Medium shot she speaks warmly gesturing to product. (5-10s) Slow push-in close-up. (10-15s) Over-shoulder vanity. Glossy warm cinematic.
Try this promptBrightly colored athletic running shoe on mossy ground, red-to-yellow gradient, neon green foam sole, extreme low angle, slow spin
Try this promptCamera tracks forward over fjord, mist drifts, small red boat, orchestral strings swell
Try this promptCharacter turns toward camera looks up, rain starts falling, crisp rainfall and fishing town ambience
Try this promptGrok Imagine Video 1.5 renders each example on demand rather than serving fixed sample clips, so the prompts are shown here instead. Run one to see your own render.
Workflows
Audio is included at no extra charge, each input image adds a small surcharge, and reference-to-video is capped at 720p. Credits come from one shared balance across image, video and speech, and the exact estimate is shown in the Studio before you generate.
Built for
FAQ
Limits, audio, licensing and price — the things worth knowing before you spend a credit.
Open the Studio to generate with the Grok Imagine video model available today, and see the credit estimate before you spend anything.