LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
在正式微调前先偷看几步训练的梯度,让LoRA的初始化更聪明
LoRA能用很少的显存微调大模型,但效果一直不如全参数微调。这篇论文提出LoRA-GA2,先用一个轻量探针跑几步早期训练获取梯度信息,再据此决定每层该分配多少低秩容量、以及如何初始化LoRA权重。在GLUE、GSM8K和HumanEval等基准上,该方法持续优于现有的LoRA变体。
他们做了什么
- 以往基于梯度的LoRA方法只看训练最开始那一瞬间的梯度(单步梯度),无法反映真实训练过程中优化路径的变化。
- 另一类方法在每一步训练中都尝试对齐梯度,虽然更精确,但需要改动优化器本身,大幅增加显存占用和训练时间。
- LoRA-GA2改用一个轻量优化器AdaLomo临时跑几步,把梯度累积保存在CPU上,再恢复原始预训练权重,用这份累积信号来分配各层的秩并初始化LoRA矩阵,这一过程不额外占用GPU显存,只增加很小的时间开销。
- 在秩分配上,论文把敏感度分数(该层对任务有多重要)和有效秩分数(该层梯度方向有多分散)相乘使用,避免只依赖单一指标造成的容量浪费。
- 在T5-Base的GLUE基准(秩为8)上,LoRA-GA2平均分比最强基线高0.66分;在Llama3.1-8B-Base上,GSM8K比最强基线高1.03分,HumanEval高0.87分;在CLIP-ViT-B/16图像分类任务上也比最强基线高0.63分。

| Method | MNLI | SST-2 | CoLA | QNLI | MRPC | Average |
|---|---|---|---|---|---|---|
| Full | 86.33±0.00 | 94.75±0.21 | 80.70±0.24 | 93.19±0.22 | 84.56±0.73 | 87.91 |
| LoRA (13) | 85.30±0.04 | 94.04±0.11 | 69.35±0.05 | 93.19±0.22 | 84.56±0.73 | 85.29 |
| Convergence Optimization Methods for LoRA | ||||||
| rsLoRA (15) | 85.73±0.10 | 94.19±0.23 | 72.32±1.12 | 93.12±0.09 | 52.86±2.27 | 79.64 |
| DoRA (19) | 85.67±0.09 | 94.04±0.53 | 72.04±0.94 | 93.04±0.06 | 68.08±0.51 | 82.57 |
| LoRA+ (9) | 85.81±0.09 | 93.85±0.24 | 77.53±0.20 | 93.14±0.03 | 74.43±1.39 | 84.95 |
| Initialization Optimization Methods for LoRA | ||||||
| PiSSA (22) | 85.75±0.07 | 94.07±0.06 | 74.27±0.39 | 93.15±0.14 | 76.31±0.51 | 84.71 |
| LoRA-GA (30) | 85.70±0.09 | 94.11±0.18 | 80.57±0.20 | 93.18±0.06 | 85.29±0.24 | 87.77 |
| Adaptive Methods for LoRA | ||||||
| AdaLoRA (37) | 85.45±0.11 | 93.69±0.20 | 69.16±0.24 | 91.66±0.05 | 68.14±0.28 | 81.62 |
| RaLoRA (34) | 85.76±0.03 | 94.22±0.29 | 78.11±0.45 | 93.36±0.14 | 84.74±0.27 | 87.24 |
| GoRA (10) | 85.91±0.22 | 94.68±0.43 | 79.86±0.35 | 93.27±0.08 | 86.10±0.20 | 87.96 |
| LoRA-GA2 (Ours) | 85.91±0.01 | 94.72±0.37 | 82.39±0.24 | 93.19±0.05 | 86.88±0.14 | 88.62 |

| Method | GSM8K | HumanEval |
|---|---|---|
| Full | 73.69±0.28 | 51.63±1.27 |
| LoRA (13) | 67.78±1.25 | 43.09±0.35 |
| rsLoRA (15) | 68.36±0.74 | 45.78±2.80 |
| DoRA (19) | 69.17±1.00 | 43.70±1.54 |
| LoRA+ (9) | 71.29±0.93 | 44.51±2.11 |
| OLoRA (2) | 68.54±0.42 | 43.29±2.44 |
| PiSSA (22) | 68.56±1.03 | 44.10±1.54 |
| LoRA-GA (30) | 71.39±0.90 | 43.29±0.61 |
| AdaLoRA (37) | 70.63±0.77 | 41.46±3.66 |
| RaLoRA (34) | 72.25±0.59 | 48.78±1.61 |
| GoRA (10) | 72.91±0.76 | 48.98±2.14 |
| LoRA-GA2 (Ours) | 73.94±0.48 | 49.85±0.33 |

| Method | Cars | DTD | EuroSAT | GTSRB | RESISC45 | SUN397 | SVHN | Average |
|---|---|---|---|---|---|---|---|---|
| Zero-shot | 63.75 | 44.39 | 42.22 | 35.22 | 56.46 | 62.56 | 15.53 | 45.73 |
| LoRA (13) | 82.31±0.08 | 76.97±0.51 | 98.38±0.20 | 97.10±0.06 | 94.99±0.11 | 77.19±0.19 | 96.62±0.06 | 89.08±0.10 |
| MELoRA (26) | 82.65±0.38 | 75.16±0.59 | 98.64±0.05 | 98.88±0.05 | 95.78±0.16 | 74.69±0.22 | 96.95±0.09 | 88.96±0.15 |
| MoRA (14) | 84.61±0.21 | 77.34±0.14 | 98.65±0.16 | 98.68±0.18 | 96.33±0.19 | 78.12±0.06 | 97.17±0.15 | 90.13±0.16 |
| AdaLoRA (37) | 73.58±0.09 | 73.79±0.48 | 96.96±0.12 | 58.87±0.38 | 89.07±0.60 | 72.00±0.10 | 94.26±0.13 | 79.79±0.27 |
| DoRA (19) | 82.44±0.26 | 76.86±0.84 | 98.43±0.17 | 97.25±0.12 | 95.10±0.16 | 77.30±0.17 | 96.63±0.04 | 89.14±0.07 |
| rsLoRA (15) | 83.94±0.22 | 77.64±0.33 | 98.51±0.17 | 98.69±0.17 | 95.90±0.20 | 77.96±0.21 | 96.94±0.06 | 89.94±0.06 |
| LoRA+ (9) | 86.61±0.36 | 73.33±1.30 | 98.54±0.14 | 98.99±0.20 | 96.06±0.38 | 76.80±0.34 | 96.98±0.08 | 89.62±0.19 |
| PiSSA (22) | 83.36±0.38 | 77.38±0.57 | 98.54±0.09 | 98.32±0.09 | 95.92±0.40 | 77.46±0.13 | 97.00±0.09 | 89.71±0.25 |
| OLoRA (2) | 83.85±0.13 | 78.60±0.25 | 98.62±0.03 | 98.49±0.14 | 96.01±0.28 | 77.30±0.08 | 97.15±0.14 | 90.00±0.15 |
| RaLoRA (34) | 86.63±0.30 | 77.75±0.20 | 98.66±0.27 | 98.98±0.11 | 96.62±0.28 | 77.86±0.05 | 97.24±0.11 | 90.53±0.03 |
| LoRA-GA2 (Ours) | 87.82±0.13 | 79.49±0.11 | 98.83±0.11 | 99.01±0.08 | 96.48±0.24 | 79.05±0.14 | 97.42±0.07 | 91.16±0.05 |

| Variant | GSM8K | HumanEval |
|---|---|---|
| One-step gradient | 72.13±0.34 | 49.65±1.01 |
| No SVD init. | 68.92±0.12 | 48.37±2.50 |
| No rank alloc. | 72.40±1.32 | 48.88±0.63 |
| Sensitivity only | 72.98±0.65 | 48.17±2.28 |
| Effective rank only | 73.51±0.51 | 48.78±1.00 |
| Full LoRA-GA2 | 73.94±0.48 | 49.85±0.33 |

为什么重要
LoRA已经是广泛使用的微调方法,这类改进意味着在同样的计算资源下能得到更好的模型效果。由于几乎不增加显存、时间开销也很小,这种方法很容易直接接入现有的LoRA训练流程。
本文术语
- LoRA(低秩适配) · 不直接修改大模型全部权重,只训练两个小矩阵来完成微调,从而节省显存的方法
- 单步/多步梯度 · 梯度指示模型该往哪个方向调整才能降低损失;单步只在训练最开始测一次,多步则在训练早期连续多个步骤中观察
- SVD(奇异值分解)初始化 · 把累积梯度分解成几个主要方向和强度,并让LoRA权重的初始值对齐这些方向
- 有效秩 · 衡量某一层梯度分布在多少个独立方向上的指标
- 敏感度 · 衡量某一层权重变化对最终损失影响大小的分数,反映该层对任务的重要程度
论文原文摘要(英文)
Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent studies have sought to narrow this gap by employing one-step gradient approximations of pretrained weights to align LoRA updates with the principal directions or intrinsic dimensionalities of full fine-tuning updates. Nevertheless, these approaches fail to capture the full dynamics of the gradients. In this paper, we propose LoRA-GA$^2$, an effective fine-tuning algorithm that fully leverages multi-step gradient information. Specifically, we introduce a lightweight probe for multi-step gradients of pretrained weights that incurs no additional GPU memory cost and only marginal time overhead. We further employ a spectrum-aware, importance-based rank allocation and optimal initialization derived from multi-step gradients. Extensive experimental results demonstrate that LoRA-GA$^2$ consistently outperforms existing LoRA variants while preserving the efficiency advantages of vanilla LoRA. For instance, LoRA-GA$^2$ surpasses the leading baseline by an average of 0.66 points on the GLUE benchmark, and outperforms the strongest baseline by 1.03 points on GSM8K and 0.87 points on HumanEval, respectively.
在 arXiv 阅读最新论文
- Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages让看图AI遵守隐藏的系统规则会明显拖累准确率,用户一旦故意要求它违规,情况会更糟
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction文本信息缺失或损坏时,这个AI不靠一次性猜测,而是反复修正猜测结果,从而更准确地判断情绪
- Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder首个用俄语提问就能搜索1C企业软件代码的公开基准和专用AI模型问世
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
- Reliable Financial Named Entity Recognition under Domain ShiftAI在正式文件里学到的自信,一到推特上就变得不可信
- When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models只插一句和图片无关的话,多模态AI的判断就会按固定规律偏移
METAL LAB 最新报道
图片来源: Haonan He et al., arXiv:2608.19800, CC BY 4.0