One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

xAI unveils image generator Imagine 2.0, ranks second behind GPT-Image-2 on Arena benchmark

New version adds editing tools and templates, Elo score trails OpenAI's model by a narrow margin

파란 소파와 붉은 의자가 있는 모던한 거실 인테리어

이미지: The Decoder

Summary

  • xAI has launched a new image generation model, Imagine Image 2.0, on Grok as a "Quality Mode"
  • It ranked second overall in both the image editing and text-to-image categories on the Arena benchmark, trailing OpenAI's GPT-Image-2 by a small margin
  • It comes with editing tools such as Magic Wand, Multi-Ref Editing, and Smart Resize, along with pre-configured templates
Video from the source
출시 경로
grok.com/imagine 및 Grok iOS·Android 앱 내 '퀄리티 모드'
API 지원
곧 제공 예정(xAI 발표 기준)
Image Edit Arena Elo
Imagine 2.0(low) 1,439 vs GPT-Image-2 1,463
Text-to-Image Arena Elo
Imagine 2.0(low) 1,320 vs GPT-Image-2 1,380
전체 순위
두 카테고리 모두 전체 2위, Reve 2.1·Muse-Image·Qwen-Image-3.0-Pro·Gemini·SeedDream보다 상위
편집 기능
Magic Wand, 세그멘테이션, 배경 제거, Multi-Ref Editing(최대 5장), Smart Resize
템플릿 분야
사진 편집, 제품 촬영, 마케팅 소재, 디자인 도구, 게임 에셋, 스트리밍 이모지

Launch overview

xAI has unveiled a new image generation model, Imagine Image 2.0, on the Grok platform. The model is available as a new "Quality Mode" on grok.com/imagine and in Grok's iOS and Android apps, with API access to follow soon, xAI said. The company explained that the model was designed to follow detailed instructions accurately, keep typography and layout clean even in complex images, and maintain consistency across multiple generations.

Arena benchmark results

As of August 7, 2026, on the Arena leaderboard, the faster "low" version of Imagine 2.0 ranked second overall in both the image editing and text-to-image categories. In the Image Edit Arena, it scored an Elo of 1,439, narrowly trailing OpenAI's GPT-Image-2, which scored 1,463. In the Text-to-Image Arena, it scored 1,320, also behind GPT-Image-2's 1,380.

ModelImage Edit Arena EloText-to-Image Arena Elo
GPT-Image-2 (OpenAI)1,463 bar:1001,380 bar:100
Imagine 2.0 low (xAI)1,439 bar:981,320 bar:96

Reve 2.1, Meta's Muse-Image, Alibaba's Qwen-Image-3.0-Pro, Google's Gemini, and ByteDance's SeedDream all ranked lower than Imagine 2.0 in both categories. xAI said the new model outperformed its predecessor's "Quality" version by a wide margin.

New editing tools

Imagine Image 2.0 adds several editing features optimized for iterative work. "Magic Wand" allows edits to only a selected area, and a segmentation feature supports precise region selection. A background removal tool can extract a subject onto a transparent background.

"Multi-Ref Editing" combines up to five input images into a single output. "Smart Resize" converts an existing image to a desired aspect ratio, with the model automatically filling in the added space.

A related METAL LAB article covering the trajectory of OpenAI's image generation models also offers insight into recent shifts in the image generation competitive landscape.

Templates and video pre-production features

xAI also introduced templates that pre-configure commonly used image workflows. These span photo editing, product photography, marketing assets, design tools, game assets, and streaming emotes.

The company also unveiled a feature that generates characters, locations, and props separately while maintaining a consistent visual style across the overall image. xAI said it is positioning this as a stepping stone toward a full video production workflow in the future. The company explained that the feature creates a "consistent visual world" that can serve as a starting point for video production.

Outlook

The benchmark results, in which Imagine 2.0 narrowly trailed GPT-Image-2, show that competition among top-tier image generation models is narrowing to an extremely tight margin. The timing of the API release and whether the feature will expand into video generation are seen as the next points to watch.