METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Google Research Unveils a Control Framework for Diffusion Models

Google Research unveiled Diffusion Controller, a control framework that keeps the original image generation model frozen and uses a small side network to correct the direction of generation. A gray-box version with 12 million parameters achieved higher preference win rates than LoRA.

Google Research Unveils a Control Framework for Diffusion Models

Image: METAL

Summary

  • Google Research on September 29 unveiled Diffusion Controller, which recasts the image generation process of diffusion models as a continuous control problem.
  • A lightweight side network that receives only intermediate outputs corrects the direction at every step while the original model stays fixed, so it can be attached even to closed models whose weights are not released.
  • In Stable Diffusion v1.4 experiments, the gray-box method reached SFT and RWL win rates of 0.667 and 0.682, ahead of LoRA's 0.577 and 0.611, and a white-box variant recorded 0.935 under PPO.

Google Research on September 29 unveiled Diffusion Controller, a control framework that steers text-to-image diffusion models in a desired direction. At its core, the massive original model is left frozen while a small side network attached to it finely corrects the direction at every step of image creation. Chih-wei Hsu and Moonkyung Ryu, software engineers at Google Research who wrote the blog post, likened the side network to a motorcycle's "steering damper" and said it "seamlessly attaches to even access-restricted, closed-source models." In other words, adjustments aligned with user intent are possible without modifying the original model.

The problem it tackles is the tug-of-war between prompt fidelity and image quality. The blog uses a lizard wearing sunglasses as an example. The model might produce a realistic lizard without the sunglasses, or, when forced to include them, distort the lizard's face. Until now, developers have separately used guidance techniques that adjust the prompt's influence during generation and methods that retrain the model with lightweight adapters such as LoRA, reward-weighted regression or policy gradients. Hsu and Ryu diagnosed that "these tools have historically been treated as distinct and unrelated fixes" and that the field has lacked a single mathematical language to unify and analyze them.

Diffusion Controller rewrites the entire process of gradually removing noise to complete a picture as a single continuous control problem. According to the 37-page paper, which METAL reviewed in the original, the researchers cast reverse diffusion sampling as stochastic control within a linearly-solvable Markov decision process (LS-MDP) and defined control as reweighting the transition probabilities of the pretrained model. Two training methods follow from these optimality conditions. One is a PPO-style policy gradient method that caps the size of changes with a clipping rule, and the other is a reward-weighted loss that gives weight to generation paths that produced good results. The paper was posted to arXiv on March 7, with seven authors including Tong Yang, a Carnegie Mellon University student who interned at Google Research, Yuejie Chi of Yale University, and Bo Dai of Google DeepMind.

같은 프롬프트로 만든 이미지 비교. 정장 입고 시가 문 검은 고양이, 스파게티 먹는 파랑어치, 선글라스 쓴 도마뱀을 사전학습 모델과 LoRA, Diffusion Controller가 각각 그린 결과다

The same theory also determines the model structure. Because the optimal score function splits into a fixed pretrained baseline plus a small control correction, the conclusion is that the original model can be frozen and only the correction trained. The side network takes as input only the intermediate denoising outputs that the original model exposes. That is why the researchers call it a gray-box setting. It can be attached as long as intermediate outputs are visible, even without white-box access to inspect internal weights. The side network described in the paper is a lightweight UNet operating at a 64×64 latent size, and the best-performing configuration has about 11.81 million trainable parameters.

The experiments used Stable Diffusion v1.4 as the base model across three training regimes: supervised fine-tuning (SFT), reward-weighted loss (RWL) and PPO. The metric was the win rate against the original model as judged by HPS-v2, which predicts human preferences. The gray-box Diffusion Controller, with 12 million parameters, reached win rates of 0.667 in SFT and 0.682 in RWL. LoRA, which uses 17 million parameters and requires access to the model's internals, reached only 0.577 and 0.611 under the same conditions. A comparison version with a simplified side network scored 0.566 and 0.506, effectively no different from the original. In PPO, LoRA's 0.905 beat the gray-box method's 0.696, but the white-box variant Diffusion Controller-J, which trains LoRA and the side network jointly, rose to 0.935. This is the 90% win rate the blog mentioned.

Human evaluation was added as well. The researchers showed more than 50 raters sets of images made from 50 HPS-v2 prompts, generating four images per prompt, and had them pick the better set. The paper noted that the raters were paid contractors earning above the living wage in their country of employment. According to the blog, Diffusion Controller received the best results for both subjective quality and prompt matching on complex prompts with multiple attributes. On the reward-weighted loss side, the paper said it achieved a higher win rate than the existing DPOK method with a quarter of the samples. At inference time, the strength of control can be raised or lowered by changing a single guidance strength value.

SFT·RWL·PPO 세 학습 방식에서 학습 단계별 HPS-v2 승률 곡선. SFT와 RWL에서는 Diffusion Controller 계열이 LoRA 위에 있고, PPO에서는 LoRA와 흰 상자 변형이 0.9 안팎까지 오른다

From an AI engineer's standpoint, the target of this research is closed models. Hsu and Ryu explained in the blog that "the world's best image-generation models are often corporate secrets," and that the side network lets engineers adjust locked models without touching the underlying code. This means that even if a model provider does not hand over weights, a business structure becomes possible in which customers separately train a small side network on their own preference data, as long as intermediate outputs are exposed. That the experiments were limited to a single model, Stable Diffusion v1.4, also means the effect on the latest large models still needs to be confirmed separately. Google Research named personalization tools, safety mechanisms to reduce harmful image generation, and control of video models as next steps. METAL reported that Google Research unveiled a framework for long-form AI video generation, and whether this control approach moves into that video generation line is the next point to watch.

Comments