매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

V-RAE: Rethinking Video Latent Spaces for Generation

arXiv:2608.135562026-08-12

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents

저자 · Minghui Guo

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사