매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

arXiv:2608.123852026-08-14

arXiv:2608.12385v1 Announce Type: new Abstract: As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training cost. The two inference phases stress hardware differently: prompt prefill is parallel and typically compute-bound, whereas autoregressive decode is sequential and often memory-bandwidth-bound. Conventional width or depth scaling increases both costs together because every added layer is evaluated in both phases.

저자 · Liming Liu, Mingze Wang, Tuo Zhao

arXiv에서 원문 보기