METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Apple scales up a diffusion-style language model to 1.7 billion parameters

Trained on 2.1 trillion tokens, then self-distilled to generate sentences in as few as 4 steps, testing an alternative to autoregression

Apple scales up a diffusion-style language model to 1.7 billion parameters

Image: METAL

Summary

  • Apple Machine Learning Research has published a paper on a 1.7-billion-parameter flow-matching language model trained on 2.1 trillion tokens
  • A self-distilled version of the model, called Categorical Flow Map, generates text in as few as 4 inference steps
  • Previous research had only been validated at scales below 1 billion parameters, leaving scalability an open question — this marks the first large-scale validation

What was announced

A paper titled "Scaling Categorical Flow Maps," published by Apple Machine Learning Research on August 7, 2026, presents results from scaling up a language model built not on autoregression but on flow matching. The team trained a 1.7-billion-parameter base model on 2.1 trillion tokens, then applied self-distillation — a method where a model learns a faster version of itself using its own outputs as a teacher — to produce a compressed model called Categorical Flow Map (CFM). According to the paper, this CFM generates diverse, high-quality text in as few as 4 inference steps while maintaining token entropy — a measure of sentence diversity — close to that of real data. The researchers also proposed a likelihood-calculation method that lets the model's performance be converted into scores on standard language model benchmarks, and reported that the resulting scores fall within a range comparable to existing discrete diffusion approaches.

Why it matters

Most language models in wide use today, such as the GPT family, are autoregressive: they predict text token by token, in order, from left to right. Diffusion and flow-matching models, by contrast, first gained traction in image generation, where generation starts from noise and refines multiple points simultaneously to produce a final result. Applying this approach to text allows multiple tokens to be refined in parallel rather than sequentially, which in theory enables faster generation in fewer steps. The catch is that text, unlike images, consists of discrete words and tokens rather than continuous values. Recent research has tried to bridge this gap by using flow matching to connect one-hot encoded data distributions with Gaussian noise, but until now this had only been validated at scales below 1 billion parameters, leaving open the question of whether the same results would hold at larger scale. This paper provides the first answer to that question at a scale of 1.7 billion parameters and 2.1 trillion tokens, marking a next step for flow-matching research. Notably, flow-matching architectures have already gained a foothold in image generation — Black Forest Labs released FLUX.2-dev, a 32-billion-parameter rectified flow transformer, this past August.

Black Forest Labs releases 32-billion-parameter image model FLUX.2-dev

What changes as a result

This announcement is a research paper, not a commercial product. Still, it adds to the evidence that flow matching could work as an alternative architecture even at large scale, in a language model market where autoregressive models have effectively become the standard. The researchers also shared practical guidance on how to weight loss functions and design time schedules, which is likely to serve as a direct reference for other teams attempting this approach going forward. Whether 4-step generation can deliver the speed and quality needed for practical deployment in real services remains to be confirmed through further benchmarking and validation.

Comments