METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Apple identifies cause of 'outlier token' problem in diffusion transformers

Apple found that unusually large-value tokens forming inside image-generation AI degrade image quality, and proposed a fix

Apple identifies cause of 'outlier token' problem in diffusion transformers

Image: METAL

Summary

  • Apple ML Research published a paper in August analyzing the 'outlier token' problem inside Diffusion Transformers (DiT)
  • The team found that simply masking outlier tokens doesn't work, and traced the root cause to corrupted local patch information
  • To address this, they proposed a 'Dual-Stage Registers (DSR)' technique and validated it on ImageNet and large-scale text-to-image generation

What was published

Apple ML Research released a paper in August that directly tackles the 'outlier token' phenomenon occurring inside Diffusion Transformers (DiT) used for image generation. The team confirmed that this phenomenon appears both in the pretrained Vision Transformer (ViT) encoder and in the DiT itself within a Representation Autoencoder (RAE)-based DiT pipeline. The tokens were especially pronounced in the middle layers inside the DiT.

An interesting finding was that simply masking or removing high-value outlier tokens did not improve performance. This suggests the problem isn't caused by a handful of extreme values, but by the fact that the local patch information at those positions is itself corrupted. To address this, the team proposed 'Dual-Stage Registers (DSR)': using learned registers where available, falling back to registers applied iteratively at inference time when they are not, and attaching a separate diffusion register to the denoiser (the noise-removal network).

Why 'outlier tokens' matter

Vision transformers cut images into small pieces (patches) and treat each one as a 'token,' much like language models process sentences word by word. As training progresses, however, it has been reported that some tokens develop unusually large values (high norm) while carrying almost none of the actual image information for that position, yet still drawing excessive attention from other tokens. In ViT research, a known remedy has been to set aside dedicated 'register' slots to absorb these tokens, but how this phenomenon operates within generative DiT models had not been properly clarified until now.

DiT is the architecture that recent image-generation models such as Stable Diffusion 3 and the FLUX series have adopted in place of the traditional U-Net structure. Apple has previously pursued research along these lines, including comparing the performance of diffusion-based versus autoregressive approaches and scaling a diffusion language model up to 1.7 billion parameters. This new paper can be seen as an extension of that work, pinpointing the root cause of subtle defects that erode image quality.

So what changes

The outlier token problem has long been cited as one cause of the occasional localized blotches or distorted patches seen in generated images. Because Apple's research addresses this at the level of training architecture rather than as a stopgap fix, it could inform future design guidelines for improving the stability of DiT-based image and video generation models. That said, these are research-paper-level results, and how they might be incorporated into actual commercial models remains to be seen in follow-up work.

Comments