Apple scales a diffusion-style language model to 1.7 billion parameters
After training on 2.1 trillion tokens, self-distillation produces sentences in as few as 4 steps, testing an alternative to autoregressive models
10
One email each morning — yesterday's AI, sortedGet it in your inbox
Tag
After training on 2.1 trillion tokens, self-distillation produces sentences in as few as 4 steps, testing an alternative to autoregressive models
That's the last story.