One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

arXiv:2608.161572026-08-16

Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads cont

Authors · Shuo Yang

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB