AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox›
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
arXiv:2608.142292026-08-13
A new unlearning method, AdaPop, erases facts from AI models based on how well-known each fact is, since popular facts are harder to forget
When AI models are trained to 'unlearn' or forget specific information, popular facts (things widely known) turn out to be encoded more deeply and are harder to remove than rare ones, yet existing methods apply the same erasing pressure to everything. AdaPop assigns a different erasing strength to each fact based on an external popularity signal such as Wikidata sitelink counts, and uses a controller that automatically adjusts, every training epoch, how much pressure is placed on preserving unrelated knowledge. Across three model families and two benchmarks, AdaPop leaked about 5 times less forgotten content under paraphrased queries and about 1.6 times less under adversarial reformulations, compared to competing methods.
METAL LAB explanatory visual
How AdaPop works: popularity-aware erasing with automatic balancing
Evidence statusMeasured results reported
Score popularityEach fact gets a popularity score from an external source such as Wikidata sitelink counts or an LLM-as-Judge rating
Convert to exponentThe popularity score becomes an exponent (beta) that applies stronger erasing pressure to popular facts and gentler pressure to rare ones, token by token
Weighted forget lossThis exponent combines with the model's token-level confidence to weight the training signal used to erase targeted facts
Dual-ascent controllerAt the end of every training epoch, the controller checks how much retained knowledge has degraded and raises or lowers the retain penalty accordingly
Verify resultsParaphrased and adversarial queries plus internal hidden-state comparisons check whether facts were truly erased rather than just hidden at the surface
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.
What they did
Problem: existing unlearning methods apply the same gradient pressure to every fact meant to be forgotten, but popular facts are memorised more deeply during pretraining and resist removal, while pushing rare facts just as hard over-erases them and damages unrelated retained knowledge.
Method: AdaPop scores each fact's popularity using an external signal such as Wikidata sitelink counts or an LLM-as-Judge rating, converts that score into an exponent that sharpens erasing pressure on popular facts and softens it on rare ones, and pairs this with a dual-ascent controller that checks retained-knowledge loss each epoch and automatically tunes the retain penalty.
Experiments: tested on Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-7B-it across the DUET and RWKU benchmarks, compared against baseline methods including GA, GD, NPO, and WGA.
Results: under paraphrased queries, AdaPop leaked about 5x less forgotten content than competing methods; under adversarial reformulations, about 1.6x less. Internally, the hidden representations of forgotten facts moved further from the pre-unlearning model than with other methods, while representations of retained facts stayed close.
Figure 1: Overview of AdaPop. The popularity score reweights per-fact token contributions to the forget loss; the dual-ascent controller adjusts the retain penalty each epoch to maintain retain quality (Algorithm 1).
Table 1: ROUGE-L Recall and Cosine Similarity for Llama-3.1-8B-Instruct at lr=10−4. ROUGE-L: mean ± std over three seeds; Cosine Similarity: single seed. F. ↓ = forget; R. ↑ = retain. w/o MU = the original checkpoint, evaluated with no unlearning applied. Underline = best among NPO, WGA, AdaPop. †Unusable: generation failure (GA: joint output and retain collapse; GD: severe retain degradation with collapse-style generation artefacts; see Appendix G).
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.939
0.968
0.874
0.883
AdaPop
0.043 ± .007
0.959 ± .009
0.369
0.971
GA†
0.000 ± .000
0.000 ± .000
0.093
0.093
GD†
0.021 ± .003
0.853 ± .011
0.120
0.897
NPO
0.670 ± .095
0.996 ± .005
0.684
0.998
DUET
WGA
0.036 ± .002
0.995 ± .003
0.442
0.996
w/o MU
0.755
0.827
0.776
0.823
AdaPop
0.078 ± .005
0.972 ± .008
0.125
0.977
GA†
0.002 ± .000
0.001 ± .000
0.063
0.067
GD†
0.026 ± .005
0.759 ± .021
0.096
0.829
NPO
0.540 ± .034
0.957 ± .005
0.403
0.967
RWKU
WGA
0.095 ± .009
0.977 ± .003
0.247
0.984
Figure 2: Wikidata popularity-score distributions for the forget sets of DUET and RWKU. DUET spans nearly two orders of magnitude (69–3,763; median 1,090), providing a direct test of popularity-sensitive methods; RWKU is narrower and skewed lower (median 130).
Table 2: ROUGE-L Recall and Cosine Similarity for Qwen2.5-7B-Instruct at lr=10−4. Notation as in Table 1.
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.921
0.886
0.864
0.861
AdaPop
0.056 ± .007
0.950 ± .008
0.456
0.964
GA†
0.000 ± .000
0.000 ± .000
0.058
0.092
GD†
0.022 ± .004
0.757 ± .015
0.094
0.841
NPO
0.684 ± .023
0.973 ± .002
0.576
0.979
DUET
WGA
0.058 ± .002
0.987 ± .009
0.477
0.999
w/o MU
0.570
0.666
0.322
0.347
AdaPop
0.016 ± .007
0.855 ± .003
0.178
0.911
GA†
0.001 ± .000
0.000 ± .000
0.071
0.075
GD†
0.014 ± .003
0.462 ± .018
0.150
0.640
NPO
0.271 ± .010
0.783 ± .012
0.490
0.870
RWKU
WGA
0.038 ± .016
0.890 ± .005
0.212
0.931
Figure 3: AdaPop on DUET (Llama) under three popularity signals: Wikidata score, LLM-as-Judge (3-seed mean), and measured corpus frequency (infini-gram over Pile-train). Left: forget ROUGE-L (↓). Right: holdout ROUGE-L (↑). Each point averages a rare-tier and a popular-tier run at that learning rate. All signals use the anchors of Appendix B, (a,b)=(58.7,0.796). Rates above 5×10−4 are omitted: at 10−3 every signal has at least one collapsed tier.
Table 3: ROUGE-L Recall and Cosine Similarity for Gemma-7B-it at lr=10−4. Notation as in Table 1.
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.892
0.926
0.586
0.541
AdaPop
0.023 ± .004
0.976 ± .011
0.206
0.976
GA†
0.000 ± .000
0.000 ± .000
0.096
0.102
GD†
0.044 ± .007
0.598 ± .024
0.069
0.626
NPO
0.626 ± .071
0.968 ± .007
0.394
0.956
DUET
WGA
0.050 ± .002
0.996 ± .004
0.237
0.997
w/o MU
0.471
0.551
0.328
0.349
AdaPop
0.034 ± .014
0.948 ± .009
0.068
0.961
GA†
0.000 ± .000
0.000 ± .000
0.000
0.000
GD†
0.013 ± .003
0.334 ± .013
0.115
0.532
NPO
0.341 ± .013
0.773 ± .010
0.240
0.840
RWKU
WGA
0.040 ± .003
0.950 ± .007
0.135
0.965
Figure 4: Internal representation metrics at lr=10−4, across three models. Each row corresponds to one metric; each column corresponds to one data split. For the forget split, deeper erasure corresponds to lower ΔLP, higher ΔRank, lower Hid.Cos, and higher KL; directional desiderata are reversed for the retain split. Marker shape indicates benchmark (DUET vs. RWKU); colour indicates model family.
Table 4: Robustness at lr=10−4. Top: paraphrase ROUGE-L on the DUET forget split (lower is better). Bottom: RWKU Level-3 adversarial-attack ROUGE-L; lower indicates stronger resistance to knowledge recovery. Underline = best among NPO, WGA, AdaPop. GA and GD collapse and are excluded from forget-quality ranking.
Algo
Llama
Qwen
Gemma
DUET paraphrase forget ROUGE-L ↓
AdaPop
0.045
0.074
0.027
GA†
0.000
0.000
0.000
GD†
0.015
0.045
0.117
NPO
0.696
0.605
0.568
WGA
0.042
0.104
0.050
RWKU adversarial-attack forget ROUGE-L ↓
AdaPop
0.262
0.144
0.195
GA†
0.000
0.001
0.000
GD†
0.254
0.171
0.103
NPO
0.657
0.400
0.485
WGA
0.396
0.211
0.239
Figure 5: Component ablation on DUET. Top row: rare-fact paraphrase forget ROUGE-L (↓) and retain ROUGE-L (↑) over lr∈[2×10−5,8×10−4]. Bottom row: popular-fact paraphrase forget ROUGE-L (↓) and retain ROUGE-L (↑) over lr∈[2×10−4,8×10−3].
Table 5: DUET paraphrase forget ROUGE-L (↓) split by popularity tier at lr=10−4. Underline = best among NPO, WGA, AdaPop. GA is omitted (all values 0.000 under collapse); †GD’s low rare-tier values coincide with retain collapse (Table 1).
Algo
Llama
Qwen
Gemma
Pop.
Rare
Pop.
Rare
Pop.
Rare
AdaPop
0.040
0.049
0.055
0.092
0.028
0.026
GD†
0.027
0.003
0.086
0.003
0.232
0.002
NPO
0.854
0.538
0.778
0.432
0.796
0.339
WGA
0.067
0.017
0.194
0.014
0.096
0.003
Figure 6: ROUGE-L (solid) and Cosine Similarity (dashed) on the merged DUET forget (left) and retain (right) splits across learning rates, for Llama (top), Qwen (middle), and Gemma (bottom).
Table 6: Internal metrics at lr=10−4, averaged across Llama, Qwen, and Gemma. Directional arrows are shown in each section header. Underline = best among NPO, WGA, AdaPop. GA and GD values are reported for reference but indicate model collapse and are not ranked. Per-model breakdown in Appendix K.
Algo
ΔLP
ΔRank
Hid.Cos
KL
DUET — forget split (ΔLP↓, ΔRank↑, Hid.Cos↓, KL↑)
AdaPop
−160¯
+53,760¯
0.702
42.5
GA†
−3,106
+106,826
0.352
479.6
GD†
−5,917
+104,134
0.192
1006.4
NPO
−42
+8,824
0.861
5.4
WGA
−52
−3,714
0.729
16.2
RWKU — forget split (ΔLP↓, ΔRank↑, Hid.Cos↓, KL↑)
AdaPop
−23¯
+4,169¯
0.484
9.1
GA†
−6,083
+96,254
0.285
953.1
GD†
−6,842
+96,881
0.175
1139.7
NPO
−18
+2,739
0.883
2.9
WGA
−3.5
−11,010
0.541
6.7
Retain split — DUET (ΔLP↑, ΔRank↓, Hid.Cos↑, KL↓)
AdaPop
+21
−13,770
0.875
6.5
GA†
−2,750
+103,104
0.348
511.4
GD†
−251
−8,751
0.760
48.2
NPO
+21
−13,660
0.910
7.5
WGA
+23¯
−13,930¯
0.880
7.5
Retain split — RWKU (ΔLP↑, ΔRank↓, Hid.Cos↑, KL↓)
AdaPop
+33
−12,190
0.883
6.2
GA†
−6,106
+98,802
0.289
955.8
GD†
−540
−3,122
0.783
97.1
NPO
+31
−11,727
0.869
7.4
WGA
+34¯
−12,202¯
0.875
6.8
Figure 7: ROUGE-L (solid) and Cosine Similarity (dashed) on the RWKU forget (left) and retain (right) splits across learning rates, for Llama (top), Qwen (middle), and Gemma (bottom).
Table 7: MMLU accuracy and HellaSwag (HS) normalised accuracy at lr=10−4. w/o MU = the original checkpoint, with no unlearning applied. Values averaged across DUET and RWKU checkpoints.
Algo
Llama
Qwen
Gemma
MMLU
HS
MMLU
HS
MMLU
HS
w/o MU
0.65
0.73
0.71
0.68
0.47
0.64
AdaPop
0.65
0.76
0.71
0.68
0.52
0.62
GA
0.23
0.33
0.23
0.43
0.25
0.26
GD
0.64
0.59
0.71
0.65
0.47
0.37
NPO
0.65
0.72
0.70
0.68
0.51
0.63
WGA
0.66
0.76
0.71
0.70
0.52
0.68
Figure 8: DUET forget and retain splits separated by popularity tier (Llama). Top row: rare facts. Bottom row: popular facts. The order-of-magnitude difference in the learning rate required to achieve comparable forgetting across the two tiers is direct evidence of the popularity gap.
Table 8: AdaPop coefficient sensitivity. Each row perturbs a and/or b by ±20% from the analytically derived baseline. F. ↓ = forget ROUGE-L; R. ↑ = retain ROUGE-L. Δ values are signed differences from the baseline.
Config
a
b
F. ↓
R. ↑
ΔF.
ΔR.
baseline
58.70
0.7960
0.045
0.961
0
0
+a,+b
70.44
0.9552
0.038
0.975
−0.007
+0.014
+a,−b
70.44
0.6368
0.044
0.994
−0.001
+0.033
−a,+b
46.96
0.9552
0.024
0.920
−0.021
−0.041
−a,−b
46.96
0.6368
0.034
0.991
−0.011
+0.030
Figure 9: Forget (left) and retain (right) ROUGE-L Recall per training epoch on DUET, Llama-3.1-8B-Instruct, at lr=10−4.
Table 9: Discordant cases between Wikidata score and LLM-judge scores on DUET. ROUGE-L: Llama-3.1-8B-Instruct recall before unlearning. Top group: Wikidata-popular facts recalled at ROUGE-L 1.0 despite low LLM-judged salience, indicating deep memorisation. Bottom group: Wikidata-rare facts with higher LLM scores but variable recall, consistent with shallower encoding.
Question
Answer
Wikidata
LLM
ROUGE-L
Wikidata popular, LLM rare
What is the country of Ensenada?
Mexico
2113
30
1.0
What is the country of Tourcoing?
France
2112
40
1.0
What is the country of Qus?
Egypt
3763
50
1.0
Wikidata rare, LLM popular
What is the instance of Hua Hin?
seaside resort
148
9000
0.5
What is the instance of Weitra?
municipality of Austria
141
5000
0.7
What is the located in the administrative territorial entity of Ollantaytambo?
Urubamba Province
165
4000
1.0
Table 10: Agreement of each popularity proxy with measured corpus frequency on DUET, counted over Pile-train (383B tokens) with infini-gram. Pearson is computed in log space. The two proxies agree with each other at Spearman 0.596 and assign the same rare/popular label to 84% of facts.
Proxy
Pearson
Spearman
Wikidata score
0.970
0.839
LLM judge
0.677
0.659
Table 11: AdaPop under corrupted popularity scores (DUET, Llama-3.1-8B-Instruct, lr=10−4). Forget ROUGE-L is reported per popularity tier. Retain degrades by 0.04 even when every label is inverted.
Score noise
F. rare ↓
F. pop. ↓
R. ↑
none
0.05
0.04
0.96
log-normal, σ=1.0
0.02
0.05
0.96
labels 50% swapped
0.02
0.05
0.97
labels 100% swapped
0.01
0.05
0.92
Table 12: ROUGE-L Recall for additional baselines, Llama-3.1-8B-Instruct at lr=10−4. w/o MU = the original checkpoint, with no unlearning applied; AdaPop reproduced from Table 1 for reference. F. ↓ = forget; R. ↑ = retain. Underline = best among non-collapsed methods. †Model collapse.
Bench.
Algorithm
ROUGE-L
F. ↓
R. ↑
w/o MU
0.939
0.968
AdaPop
0.043
0.959
UNDIAL
0.884
0.998
RMU
0.882
0.998
PDU
0.476
0.954
NPO-SAM
0.684
0.970
SimNPO
0.340
0.999
SatImp
0.108
0.995
Adaptive RMU
0.933
0.963
AltPO
0.181
1.000
FLAT†
0.001
0.001
TPO
0.397
0.997
DUET
CE-U†
0.000
0.000
w/o MU
0.755
0.827
AdaPop
0.078
0.972
UNDIAL
0.568
0.939
RMU
0.801
0.986
PDU
0.201
0.909
NPO-SAM
0.529
0.885
SimNPO
0.573
0.987
SatImp
0.417
0.985
Adaptive RMU
0.162
0.819
AltPO
0.200
0.988
FLAT†
0.004
0.003
TPO
0.341
0.986
RWKU
CE-U†
0.000
0.000
Table 13: ROUGE-L Recall for additional baselines, Qwen2.5-7B-Instruct at lr=10−4. Notation as in Table 12.
Bench.
Algorithm
ROUGE-L
F. ↓
R. ↑
w/o MU
0.921
0.886
AdaPop
0.056
0.950
UNDIAL
0.797
0.930
RMU
0.734
0.983
PDU
0.632
0.869
NPO-SAM
0.211
0.769
SimNPO
0.295
0.991
SatImp
0.070
0.985
Adaptive RMU
0.562
0.831
AltPO
0.247
0.999
FLAT†
0.000
0.001
TPO
0.460
0.999
DUET
CE-U†
0.001
0.001
w/o MU
0.570
0.666
AdaPop
0.016
0.855
UNDIAL
0.313
0.668
RMU
0.407
0.922
PDU
0.455
0.710
NPO-SAM
0.298
0.596
SimNPO
0.338
0.934
SatImp
0.187
0.891
Adaptive RMU
0.118
0.466
AltPO
0.170
0.989
FLAT†
0.004
0.003
TPO
0.382
0.934
RWKU
CE-U†
0.003
0.001
Table 14: ROUGE-L Recall for additional baselines, Gemma-7B-it at lr=10−4. Notation as in Table 12.
Bench.
Algorithm
ROUGE-L
F. ↓
R. ↑
w/o MU
0.892
0.926
AdaPop
0.023
0.976
UNDIAL
0.819
0.980
RMU
0.851
0.992
PDU
0.677
0.934
NPO-SAM
0.539
0.777
SimNPO
0.581
0.997
SatImp
0.070
0.998
Adaptive RMU
0.293
0.828
AltPO
0.262
1.000
FLAT†
0.001
0.001
TPO
0.899
0.999
DUET
CE-U†
0.000
0.000
w/o MU
0.471
0.551
AdaPop
0.034
0.948
UNDIAL
0.453
0.830
RMU
0.547
0.968
PDU
0.404
0.825
NPO-SAM
0.271
0.644
SimNPO
0.373
0.969
SatImp
0.102
0.964
Adaptive RMU
0.091
0.535
AltPO
0.172
0.989
FLAT†
0.005
0.004
TPO
0.594
0.971
RWKU
CE-U†
0.000
0.000
Findings
Under paraphrased queries AdaPop leaked about 5x less forgotten content than competing methods, and about 1.6x less under adversarial reformulations.
Across three models (Llama, Qwen, Gemma) and two benchmarks (DUET, RWKU) at a fixed learning rate, among stable (non-collapsed) methods AdaPop had the lowest cosine similarity on forgotten content in all six model-benchmark combinations, and the lowest ROUGE-L in five of six.
General capability, measured by MMLU accuracy, stayed within 0.05 of the pre-unlearning model.
Three popularity signals (Wikidata score, LLM-as-Judge, and measured corpus frequency via infini-gram on Pile-train) agreed on the rare/popular label for 84% of facts, and inverting every popularity label still only dropped retain performance by 0.04.
Where it can be used
Services that need to remove specific facts from a deployed language model for privacy, legal, or safety reasons
Research or engineering efforts using existing popularity signals like Wikidata to automatically calibrate how aggressively to erase different pieces of knowledge
Designing evaluation pipelines that check whether erased information can resurface under paraphrased or adversarially reworded queries
Limits and open work
Experiments were limited to three 7-8B scale models (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Gemma-7B-it) and two benchmarks (DUET, RWKU); larger or differently structured models were not tested.
Only LoRA fine-tuning was used; full fine-tuning was excluded based on prior work and not separately verified in this study.
The main popularity signal relied on is the Wikidata score; the LLM-as-Judge alternative for facts without Wikidata coverage was shown to be comparatively weaker at ordering facts within a tier.
On the Llama/DUET combination, the WGA baseline had a marginally lower surface-level ROUGE-L score, an exception to AdaPop's general advantage.
Why it matters
As AI systems increasingly need to remove specific information for privacy, safety, or compliance reasons, it matters whether that information is truly erased inside the model or just hidden at the surface. This work shows that accounting for how well-known a fact is leads to more thorough removal with less damage to unrelated retained knowledge.
Terms in this paper
unlearning · removing specific knowledge from an already-trained model without retraining it from scratch
Wikidata sitelink score · a popularity measure based on how many language editions of Wikipedia link to a given Wikidata entry
dual-ascent controller · a mechanism that automatically adjusts the balance between erasing targeted knowledge and preserving other knowledge, based on observed damage each training epoch
ROUGE-L · a metric measuring word-order overlap between a generated answer and the correct answer; lower is better for forgotten facts, higher is better for retained ones
LLM-as-Judge · using a large language model to rate how well-known a given fact is, as a popularity signal
Original abstract (English)
Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM unlearning methods apply uniform gradient pressure regardless of training-data frequency. We propose the AdaPop (Adaptive Popularity) method, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy (e.g., Wikidata sitelinks, LLM-as-Judge), and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch. Across three model families and two benchmarks, AdaPop leaks ~5x less forgotten content than competing methods under paraphrased queries and ~1.6x less under adversarial reformulations. We support our analysis with internal metrics: under our method, forget-set hidden states move further from the pre-unlearning model's states than under other methods, while retain-set representations remain close.