One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

LTX-2.5, open video generation model runs on a single consumer RTX GPU

An open-weight world model that produces a 10-second clip in 6.8 seconds, optimized for NVIDIA GPUs

이미지: METAL LAB 생성

Summary

  • LTX has released LTX-2.5, an open-weight video generation world model
  • It lowers VRAM requirements to run locally on NVIDIA RTX GPUs and DGX Spark
  • In on-premise benchmarks it generated a 10-second clip in 6.8 seconds, outpacing closed models
Introducing LTX-2.5 — 공식 발표 영상 (출처: LTX)
모델명
LTX-2.5 (개발사 LTX)
공개일
2026년 8월 11일
지원 하드웨어
NVIDIA RTX GPU, NVIDIA DGX Spark
생성 속도(온프레미스, 2x GB200)
10초 클립을 6.8초에 생성
생성 속도(LTX API)
10초 클립에 23.7초
언어 백본
Gemma 4
동시 공개
NVIDIA Nemotron 3.5 Lightning (같은 날)
통합 도구
ComfyUI 데이 원 지원, LoRA 파인튜닝 지원

The moment has arrived when making a video no longer requires a camera crew, a render farm, or a cloud bill — just a single graphics card on a desk. LTX has released LTX-2.5, an open-weight video generation world model. It arrives as social clips, ad creative, and even film previsualization shift from the cloud to local GPUs.

A single RTX card replaces a studio

LTX said the model is optimized to run locally on NVIDIA RTX GPUs and NVIDIA DGX Spark. By lowering VRAM requirements, the company says it has made it possible to run a frontier-grade world model on hardware creators already own. The biggest change is "native multi-shot" generation. It renders a sequence spanning multiple shots as a single coherent output, keeping character appearance consistent from shot to shot. This addresses the flickering and breakage that made earlier open models hard to use for campaigns, according to the report. On top of that, a more refined Gemma 4 language backbone and a new decoder that reduces artifacts in fast motion push the output closer to a level usable without post-production, the report said.

The entire workflow runs inside ComfyUI, on a single consumer RTX GPU. Brand characters or signature styles can be locked in quickly through LoRA fine-tuning — a method that keeps the large model frozen and trains only small correction layers to fit a specific style. No studio, no cloud, and no IP ever leaves the device.

Speed makes the difference

For local operation to matter, speed has to back it up. According to an image-to-video benchmark LTX released, generation times for a 10-second clip are as follows.

Model/MethodGeneration TimeRelative Speed
LTX-2.5 (on-premise, 2x GB200)6.8s2
LTX-2.5 (LTX API)23.7s6
Omni Flash / Grok 1.5 / Veo 3.152–70s15
Seedance 2.0196s49
FLUX 3259s65
Seedance 2.5317s80
Kling 3.0 Pro398s100

On an on-premise basis, generation time is shorter than the clip's own runtime. That works out to about 7.6 times faster than the fastest closed alternative and roughly 58 times faster than the slowest. A speed gap of this size makes it genuinely feasible to batch out dozens of versions overnight and pick the best one the next morning. The report's core argument is that ad creative typically fatigues within 7 to 10 days — not for lack of ideas, but for lack of time and budget to produce enough of them.

The concept of a world model

While a large language model is trained to predict the next word, a world model is trained to predict the next moment. It generates an environment and simulates how objects and people move within it, letting users act inside that space. This structure is used not only in film, advertising, and games, but also in simulating robots moving through warehouses or factories. NVIDIA's open world model Cosmos 3, released on August 8, pointed in the same direction by combining vision reasoning with action prediction. LTX-2.5 differs slightly in that it compresses an actual production-scale video pipeline down to the scale of a single local GPU.

NVIDIA's August local AI series

LTX-2.5 is one pillar of NVIDIA's month-long "Local AI" series running throughout August. Released the same day, the open model Nemotron 3.5 Lightning is a 30B-parameter MoE (mixture-of-experts) model aimed at always-on agents, alongside the open-source library NeMo Switchyard, which routes each workflow step to the optimal model. The shared message is that open models are expanding hardware choice across RTX PCs, workstations, data centers, and the cloud.

What actually changes

The practical takeaway is that a solo creator or a small team can now match the output of a full-time production crew in a single day. With no per-clip charges or credit consumption, a GPU can churn out multiple versions overnight, and all that's left in the morning is to open a folder and pick one. Testing ten hooks and localizing for five markets is no longer a matter of booking a studio — it's a job for a single graphics card on a desk.