매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

arXiv:2608.174022026-08-17

Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture-of-Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP-style vision encoders remains underexplored at State-of-the-Art (SOTA) levels. In this work, we systematically study MoE designs for vision encoder scaling

저자 · Bonan Zhang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사