每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

arXiv:2608.174022026-08-17

Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture-of-Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP-style vision encoders remains underexplored at State-of-the-Art (SOTA) levels. In this work, we systematically study MoE designs for vision encoder scaling

作者 · Bonan Zhang

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道