
Image: METAL
Summary
- xAI has launched a new image generation model, Imagine Image 2.0, on Grok as a "Quality Mode"
- It ranked second overall in both the image editing and text-to-image categories on the Arena benchmark, trailing OpenAI's GPT-Image-2 by a small margin
- It comes with editing tools such as Magic Wand, Multi-Ref Editing, and Smart Resize, along with pre-configured templates
Launch overview
xAI has unveiled a new image generation model, Imagine Image 2.0, on the Grok platform. The model is available as a new "Quality Mode" on grok.com/imagine and in Grok's iOS and Android apps, with API access to follow soon, xAI said. The company explained that the model was designed to follow detailed instructions accurately, keep typography and layout clean even in complex images, and maintain consistency across multiple generations.
Arena benchmark results
As of August 7, 2026, on the Arena leaderboard, the faster "low" version of Imagine 2.0 ranked second overall in both the image editing and text-to-image categories. In the Image Edit Arena, it scored an Elo of 1,439, narrowly trailing OpenAI's GPT-Image-2, which scored 1,463. In the Text-to-Image Arena, it scored 1,320, also behind GPT-Image-2's 1,380.
| Model | Image Edit Arena Elo | Text-to-Image Arena Elo |
|---|---|---|
| GPT-Image-2 (OpenAI) | 1,463 bar:100 | 1,380 bar:100 |
| Imagine 2.0 low (xAI) | 1,439 bar:98 | 1,320 bar:96 |
Reve 2.1, Meta's Muse-Image, Alibaba's Qwen-Image-3.0-Pro, Google's Gemini, and ByteDance's SeedDream all ranked lower than Imagine 2.0 in both categories. xAI said the new model outperformed its predecessor's "Quality" version by a wide margin.

New editing tools
Imagine Image 2.0 adds several editing features optimized for iterative work. "Magic Wand" allows edits to only a selected area, and a segmentation feature supports precise region selection. A background removal tool can extract a subject onto a transparent background.
"Multi-Ref Editing" combines up to five input images into a single output. "Smart Resize" converts an existing image to a desired aspect ratio, with the model automatically filling in the added space.
A related METAL LAB article covering the trajectory of OpenAI's image generation models also offers insight into recent shifts in the image generation competitive landscape.

Templates and video pre-production features
xAI also introduced templates that pre-configure commonly used image workflows. These span photo editing, product photography, marketing assets, design tools, game assets, and streaming emotes.
The company also unveiled a feature that generates characters, locations, and props separately while maintaining a consistent visual style across the overall image. xAI said it is positioning this as a stepping stone toward a full video production workflow in the future. The company explained that the feature creates a "consistent visual world" that can serve as a starting point for video production.
Outlook
The benchmark results, in which Imagine 2.0 narrowly trailed GPT-Image-2, show that competition among top-tier image generation models is narrowing to an extremely tight margin. The timing of the API release and whether the feature will expand into video generation are seen as the next points to watch.





Comments