매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

arXiv:2608.161572026-08-16

Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads cont

저자 · Shuo Yang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사