One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

V-RAE: Rethinking Video Latent Spaces for Generation

arXiv:2608.135562026-08-12

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents

Authors · Minghui Guo

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB