metir
metir
Download on App StoreGet it on Google PlayF1 FantasyLoginSign Up
Back to Blog
Grok Imagine
xAI
AI Video Generation
Text-to-Video
AI Provenance

Grok Imagine Video 1.5: 1080p, References, Text-to-Video

xAI upgrades Grok Imagine Video 1.5 with native 1080p, text-to-video, and up to seven locked character references, closing the consistency gap in AI video.

Metir AI TeamAugust 2, 20266 min read
Grok Imagine Video 1.5: 1080p, References, Text-to-Video

xAI has expanded Grok Imagine Video 1.5, the video-generation model it launched on May 31, 2026, with a set of upgrades that push it closer to a real production tool rather than a novelty clip generator. The update adds support for text, image, and voice references, native 1080p output, and true text-to-video generation, according to xAI's own announcement.

xAI logoxAI
The upgrade covers xAI's Grok Imagine Video 1.5 model.

Earlier versions of Grok Imagine Video 1.5 could only animate an existing image, and topped out at 720p resolution, as The Decoder reported when the model first shipped. The new release changes both constraints, and adds a reference system that lets creators lock specific characters, objects, or locations across a video.

What Changed in the Grok Imagine Video 1.5 Update

According to xAI and reporting from TestingCatalog and Blockchain.News, the upgrade bundles four capabilities:

  • Text-to-video: users can describe a scene in a prompt and generate video directly, without first generating or uploading a starting image.
  • Multi-reference support: up to seven locked reference elements, such as faces, locations, or objects, for consistency across a generation.
  • Native synchronized audio: clips can include music, sound effects, and lip-synced dialogue.
  • Native 1080p resolution, available alongside video extension and reference-guided generation.
1080pNative resolution
Up to 7Locked reference elements
3Input modes (text, image, voice)
4Live surfaces (web, iOS, Android, API)

Text-to-video and native 1080p are generally available on grok.com/imagine and in the Grok apps for iOS and Android. Image references, text-to-video, and native 1080p are also live in the xAI API through the grok-imagine-video-1.5 model, per xAI's developer documentation.

From Image-to-Video to True Text-to-Video

The shift from image-to-video to text-to-video is a meaningful step in how the model works, not just a convenience feature. Earlier Grok Imagine video generation required an image as a starting point, meaning a creator's workflow ran through the image model first. Text-to-video removes that dependency: a prompt alone is enough to produce a moving scene, putting Grok Imagine Video 1.5 in the same category as other prompt-to-clip generators competing across the AI video landscape in 2026.

That matters for how the tool gets used. Image-to-video is well suited to animating a single asset, a product photo, a portrait, a piece of concept art. Text-to-video is better suited to generating an entire scene from a description, which is closer to how a director briefs a shot than how a photo editor briefs a retouch.

Multi-Reference Control: The Real Unlock

Resolution and prompt flexibility are useful, but the most consequential change for actual production use is reference-conditioned generation. Up to seven locked elements, such as a character's face, a specific location, or a recurring object, can now be held constant across a generation, per xAI's announcement.

“

Character consistency, not resolution, is what turns AI video from a one-off demo into something a creator can build a sequence around.

This is the problem that has limited AI video's usefulness for anything beyond single clips. A generator that produces a stunning six-second shot but cannot reproduce the same face, outfit, or set in the next shot cannot support a story, an ad campaign, or a series of any kind. Reference-conditioned generation, first popularized by image models and now extending into video, is the mechanism that closes that gap. Seven simultaneous locked elements is a meaningfully higher ceiling than earlier single-reference or two-reference systems, giving creators room to hold a cast of characters and a setting steady at once.

Native Audio and Lip-Sync Collapse the Workflow

Grok Imagine Video 1.5 also supports native synchronized audio, including music, sound effects, and lip-synced dialogue, alongside video extension and reference-guided generation, according to xAI. Historically, generating a video clip and then sourcing or editing matching audio and dialogue sync has been a separate step, often requiring a different tool entirely. Folding audio generation, including lip sync, into the same model that generates the visual clip removes a tool switch that has been a persistent friction point in AI-assisted video production.

An Nvidia data-center GPU accelerator, the class of hardware that powers large-scale AI training and inference clusters
An Nvidia data-center GPU accelerator, representative of the class of hardware used across large AI training and inference clusters. xAI has not disclosed the specific GPU models behind every Grok Imagine Video 1.5 inference cluster. Photo: Nvidia, via Wikimedia Commons, CC BY-SA 4.0.

Why API Availability Changes the Calculus

Image references, text-to-video, and native 1080p being live in the xAI API, through the grok-imagine-video-1.5 model, per xAI's developer documentation, matters beyond the consumer Grok app. A model that only exists inside a mobile app is a destination; a model with an API is a building block. Developers can now wire reference-conditioned, native-audio 1080p video generation into their own products, whether that is a marketing automation tool, a game asset pipeline, or an app most people will never know runs on Grok underneath. That is the same pattern that turned earlier text and image models from consumer curiosities into infrastructure other companies build on.

The Provenance Question

Higher-fidelity, reference-locked, prompt-driven video generation raises the same concerns that have followed every step-change in generative video quality: easier misuse for impersonation and deepfakes alongside the legitimate creative-productivity gains. Regulators are actively responding to this category of risk. California's SB 942, the AI Transparency Act, became operative on August 2, 2026, requiring covered generative AI providers with over one million monthly California users to offer manifest disclosures and embed machine-readable latent provenance data in AI-generated image, video, and audio content, according to coverage of the law's compliance requirements. The operative date was deliberately pushed from January 2026 to August 2026 to align with the EU AI Act's Article 50 provenance-marking timeline. Neither rule targets Grok Imagine specifically, but both reflect the same underlying tension: tools that make consistent, synchronized, high-resolution AI video easier to produce also make it easier to misuse, and provenance labeling is becoming the regulatory answer on both sides of the Atlantic.

Where This Leaves Creators

For creators and developers evaluating AI video tools in 2026, no single model wins on every dimension, and Grok Imagine Video 1.5 is one option among several fast-moving video generators rather than a settled default. That is part of why model-agnostic platforms have become useful: rather than committing to one generative video engine, a platform like Metir gives creators access to multiple AI models so they can route a given shot, whether it needs reference consistency, native lip-synced audio, or the fastest turnaround, to whichever model handles it best, without locking a workflow to a single vendor.

Grok Imagine Video 1.5's expansion says less about any one feature and more about direction: AI video is moving from single disposable clips toward tools that can hold a character, a voice, and a setting steady across a sequence, generate straight from a prompt with no starting image required, and plug into other software through an API rather than staying locked inside a single consumer app.

Sources:

  • Imagine Video 1.5 with References | xAI
  • Grok Imagine Video 1.5 | xAI
  • xAI adds character references and 1080p to Imagine Video 1.5 | TestingCatalog
  • xAI updates Grok Imagine to 1.5 with image-to-video generation at 720p resolution | The Decoder
  • xAI Expands Grok Imagine Video 1.5 with Text, Image, and Voice Inputs | Blockchain.News
  • Grok Imagine Video 1.5 Preview | xAI Developer Documentation
  • California AI Transparency Act (SB 942): 2026 Compliance Guide

Image credits

Header image: Elon Musk, founder and CTO of xAI, speaking on stage at a Tesla Fremont factory event in July 2026. Photo by Steve Jurvetson, via Wikimedia Commons, CC BY 4.0. The photo depicts Musk at an unrelated 2026 event, not the Grok Imagine product. In-body image: an Nvidia Tesla A100 data-center GPU, official Nvidia product photo via Wikimedia Commons, CC BY-SA 4.0, used to illustrate the class of hardware behind large-scale AI training and inference rather than any specific xAI cluster.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Blog
  • Enterprise

Company

  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational