每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

arXiv:2608.161572026-08-16

Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads cont

作者 · Shuo Yang

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道