xAI has expanded its Imagine Video 1.5 model with new reference-based generation features, image and voice consistency tools, and native 1080p output.
The update builds on last month's Imagine Video 1.5 launch, adding text-to-video generation that requires no starting image, alongside native 1080p resolution for both text-to-video and image-to-video modes. A new voice consistency feature lets users pair a character image with a voice reference so the same face and voice persist across scenes.
The most significant addition is multi-reference support, allowing up to seven reference images per generation. Each reference can lock a specific element in place, such as a character, a scene, or a product, letting users swap other elements while keeping that one fixed.
Image and voice references are rolling out first to SuperGrok Heavy and SuperGrok Plus subscribers in the US via grok.com and iOS, with wider tier availability expected within days. Text-to-video and native 1080p are already generally available across grok.com, iOS, and Android.
The features are also live in the xAI API under the model grok-imagine-video-1.5, covering image references, text-to-video, and 1080p output.
