每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

V-RAE: Rethinking Video Latent Spaces for Generation

arXiv:2608.135562026-08-12

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents

作者 · Minghui Guo

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道