工作日早上 7 点读 AI,周日早上 8 点读周报订阅邮件

METAL LAB

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

arXiv:2608.045702026-08-06

带记忆功能的AI聊天机器人会悄悄编造近一半它声称'了解'你的内容,而自称最安全的模型往往编造得最多

研究团队构建了MirageBench,用来衡量大语言模型在只获得三条真实信息时会编造多少用户属性。对12个模型、143,616条陈述的评测显示,每个模型都有35%到49%的陈述属于无根据推断,而自我评估编造率最低的模型,反而被独立评审判定为编造最多的模型。在多轮对话试点中,这些编造的属性几乎呈线性堆积,且极少被修正。

METAL LAB 解读图

MirageBench评测流程结构

证据状态已报告实测结果

  1. 输入:仅透露3条事实的用户画像150个虚构用户(定型印象型、反定型印象型、中立型各50个),真实档案有15个属性,但只透露3条
  2. 六项个性化任务交友简介、行程规划、推荐信、礼物挑选、公寓描述、压力来源判断——构成不同程度的想象梯度
  3. 独立评审进行四分类判定每条陈述被标记为有依据、合理推断、刻板印象或纯粹编造;Claude-Opus-4-7对12个模型的143,616条陈述评分
  4. 自查结果与外部评审对比模型自我报告的过度推断率排序与评审测出的排序恰好相反(斯皮尔曼相关系数负0.60)
  5. 八轮累积试点观察小规模试验追踪编造属性是否随对话轮次持续堆积且极少被修正
这是 METAL LAB 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 研究团队给150个虚构用户画像(定型印象型、反定型印象型、中立型各50个)只透露三条真实信息,然后让模型完成六项个性化任务——撰写交友简介、规划行程、写推荐信、挑选生日礼物、描述公寓、判断压力来源——观察模型会编造出什么。
  2. 每条陈述被归入四类之一:有依据、合理推断、刻板印象、纯粹编造;该分类体系在400条陈述上与盲评的人工标注员对照,达到近乎完美的一致性(四分类科恩卡帕系数0.863,二分类0.900)。
  3. 独立评审模型(Claude-Opus-4-7)对12个模型(涵盖7个模型系列)的143,616条陈述评分,结果显示每个模型都有35%至49%(跨模型平均41.6%,按陈述数加权平均41.8%)的内容属于无根据推断,其中描述公寓任务的编造率最高达57.8%,挑选礼物任务最低为27%。
  4. 将每个模型自我报告的编造率与评审模型测出的编造率相比,出现了排名反转(斯皮尔曼相关系数为负0.60,p值0.044):Qwen3-8B自我报告的编造率最低(13.0%),却被评审判定为编造最多的模型(48.7%);而自我报告编造率最高的模型(Kimi-K2.5,58.2%)在评审结果中只排在中等水平(43.1%)。
  5. 在一项八轮多轮对话试点中,12个模型中有9个的编造属性数量几乎呈线性增长(拟合优度R方大于0.90,每轮新增5到15个),而增长最快的五个模型几乎从不修正或删除先前编造的内容(删除率仅0.4%至5%)。
Figure 1: Overview of over-inference in personalized LLMs. Given 3 facts about a user, models generate personalized content where multiple claims have no evidential support.
Figure 1: Overview of over-inference in personalized LLMs. Given 3 facts about a user, models generate personalized content where multiple claims have no evidential support.
Table 1: MirageBench leaderboard: over-inference rates as assessed by Judge (Claude-Opus-4-7) on 150 personas × 6 tasks. Each model’s rates are percentages of that model’s total claims (a per-model micro-average). The Mean row is the unweighted arithmetic mean across the 12 models (a cross-model macro-average); the corresponding claim-weighted micro-average OI over all 143,616 claims is 41.8%. Models sorted by OI rate descending.
ModelClaimsGrndStereoFabricOI%
Qwen3-8B12,68723.69.339.448.7
DeepSeek-v4-pro13,17024.410.734.745.4
GPT-4o-mini10,40826.67.138.045.1
DeepSeek-v4-flash11,97225.510.334.344.6
Qwen3.6-plus13,14623.711.832.744.5
Kimi-K2.512,66524.912.530.643.1
Gemini-3-flash14,39624.912.828.341.1
GLM-5.112,86926.012.228.040.2
GPT-5.510,41626.69.429.038.5
GPT-5.4-nano7,79030.98.629.037.6
Claude-Opus-4-611,13925.910.724.735.4
Gemini-3.1-pro12,95827.911.323.835.1
Mean25.910.531.141.6
Figure 2: The four-way claim taxonomy. The bottom two categories jointly constitute over-inference.
Figure 2: The four-way claim taxonomy. The bottom two categories jointly constitute over-inference.
Table 2: The Self-Monitoring Inversion at the model-selection level. Self-audit OI (from Task) vs. external OI (from Judge) across all models, sorted by Δ=Judge−Self. Positive Δ: the model under-detects its own over-inference; negative Δ: the model over-reports. Spearman ρ=−0.60 (p=0.044 by permutation; 95% bootstrap CI [−0.90,+0.06], family-clustered [−0.87,+0.14]; n=12).
ModelSelf%Judge%ΔPattern
Qwen3-8B13.048.7+35.7Under
GPT-4o-mini20.145.1+25.0Under
DeepSeek-v4-flash33.044.6+11.6Under
DeepSeek-v4-pro40.845.4+4.6Calib.
Gemini-3-flash41.141.1−0.0Calib.
Qwen3.6-plus45.444.5−0.9Calib.
Claude-Opus-4-641.635.4−6.2Over
GPT-5.546.438.5−8.0Over
Gemini-3.1-pro43.235.1−8.1Over
GLM-5.149.640.2−9.3Over
GPT-5.4-nano49.037.6−11.4Over
Kimi-K2.558.243.1−15.1Over
Figure 3: The MirageBench evaluation pipeline. From the benchmark input (personas with profile P and revealed facts E, and the six-task suite), the three instruments Probe, Task, and Accum elicit explicit, implicit, and continual inference, respectively, and an independent Judge classifies every resulting claim under the four-way taxonomy.
Figure 3: The MirageBench evaluation pipeline. From the benchmark input (personas with profile P and revealed facts E, and the six-task suite), the three instruments Probe, Task, and Accum elicit explicit, implicit, and continual inference, respectively, and an independent Judge classifies every resulting claim under the four-way taxonomy.
Table 3: Claim composition by task (%), ordered by groundability. Rows are pooled across all 12 models and 150 personas. Tasks that must go beyond the 3 revealed facts (top) are dominated by stereotype and fabrication, while tasks that can be answered by referring to stated preferences (bottom) maintain higher grounded proportions. OI = Stereotype + Fabricated. The four shares sum to 100% per row.
TaskGrndReasSterFabOI%
Apartment/home15.027.219.838.057.8
Rec. letter12.039.87.840.448.2
Stress source22.038.29.130.739.8
Weekend itinerary26.035.310.228.538.7
Dating profile34.037.54.524.028.5
Birthday gift40.032.98.818.327.0
Figure 4: The blind annotation interface. The revealed facts, the task context, and the single claim under review are shown, matching the information available to the Judge; the judge’s label, reasoning, and the source model are all hidden. Labels are submitted via four buttons corresponding to the faithfulness taxonomy of Figure 2, or via keyboard shortcuts A/B/C/D.
Figure 4: The blind annotation interface. The revealed facts, the task context, and the single claim under review are shown, matching the information available to the Judge; the judge’s label, reasoning, and the source model are all hidden. Labels are submitted via four buttons corresponding to the faithfulness taxonomy of Figure 2, or via keyboard shortcuts A/B/C/D.
Table 4: Inference accumulation over 8 conversation rounds (Accum). Values show mean inferred attributes stored in memory (averaged across 2 personas). Growth is approximately linear.
ModelR1R8Growth/round
GPT-5.518.5125.0+106.515.2
GLM-5.116.5121.5+105.015.0
Claude-Opus-4-615.5104.5+89.012.7
Qwen3.6-plus19.0102.5+83.511.9
DeepSeek-v4-flash13.082.0+69.09.9
Kimi-K2.515.075.0+60.08.6
Gemini-3-flash14.560.0+45.56.5
DeepSeek-v4-pro8.053.5+45.56.5
Gemini-3.1-pro14.050.5+36.55.2
GPT-4o-mini8.527.0+18.52.6
Qwen3-8B13.523.5+10.01.4
GPT-5.4-nano14.515.0+0.50.1
Table 5: Judged-claim counts per task, pooled across 12 models and 150 personas. Composition percentages and OI rates for these tasks are given in Table 3.
TaskClaims
Apartment/home31,238
Rec. letter26,885
Stress source20,085
Weekend itinerary25,971
Dating profile20,052
Birthday gift19,385
Table 6: Within-model self-audit signal. Per-model Spearman ρ and AUROC between record-level Task self-audit OI% and Judge OI%. n is the number of (persona, task) records with valid self-audit and judge outputs for that model.
ModelSpearman ρAUROCn
Qwen3.6-plus0.640.83888
GPT-5.50.630.80869
Claude-Opus-4-60.620.80897
Kimi-K2.50.610.81899
DeepSeek-v4-pro0.570.78896
GLM-5.10.560.77883
Gemini-3.1-pro0.550.74617
DeepSeek-v4-flash0.530.75890
Gemini-3-flash0.490.76900
GPT-4o-mini0.320.65897
GPT-5.4-nano0.270.61832
Qwen3-8B0.130.58834
Table 7: Judge OI% by stereotype group, per model. Δ is stereotypical minus counter-stereotypical. All 12 models show Δ>0.
ModelStereoCounterNeutralΔ
Qwen3-8B51.144.750.5+6.3
DeepSeek-v4-pro48.441.046.8+7.5
DeepSeek-v4-flash47.839.946.3+7.9
Qwen3.6-plus47.240.046.7+7.2
GPT-4o-mini46.641.947.0+4.7
Kimi-K2.545.938.445.2+7.5
Gemini-3-flash44.834.344.3+10.5
GLM-5.143.934.742.3+9.2
GPT-5.542.233.540.3+8.7
GPT-5.4-nano39.235.238.4+4.0
Claude-Opus-4-638.831.236.7+7.5
Gemini-3.1-pro38.528.838.2+9.7
Pooled44.837.043.9+7.8
Table 8: Per-model Accum regression (attributes vs. round, n=16) and mean per-round removal rate. Removal rate is the fraction of unique attributes present at round T that are absent at round T+1, averaged across the two personas.
ModelSlopeR2R8Rem.%
GPT-5.515.080.99125.00.4
GLM-5.115.210.97121.51.4
Claude-Opus-4-612.840.93104.52.4
Qwen3.6-plus12.130.93102.55.0
DeepSeek-v4-flash9.910.9582.02.4
Kimi-K2.58.820.9175.011.0
Gemini-3-flash6.550.9960.016.7
DeepSeek-v4-pro6.390.9653.54.5
Gemini-3.1-pro5.190.9550.516.0
GPT-4o-mini2.560.7927.01.5
Qwen3-8B1.560.7923.570.4
GPT-5.4-nano0.140.0315.081.6
Table 9: Model versions used in the study. Snapshots frozen at experiment time.
ModelAPI identifier / snapshot
GPT-5.5gpt-5.5-2026-06-01
GPT-5.4-nanogpt-5.4-nano-2026-05-20
GPT-4o-minigpt-4o-mini-2024-07-18
Claude-Opus-4-6claude-opus-4-6-20260415
Claude-Opus-4-7†claude-opus-4-7-20260610
Gemini-3.1-progemini-3.1-pro-preview-2026-05
Gemini-3-flashgemini-3-flash-preview-2026-04
DeepSeek-v4-prodeepseek-v4-pro-2026-05
DeepSeek-v4-flashdeepseek-v4-flash-2026-05
Qwen3.6-plusqwen3.6-plus‡
Qwen3-8BQwen/Qwen3-8B§
GLM-5.1glm-5.1‡
Kimi-K2.5moonshot-v1-k2.5
Table 10: Judge OI rates (%) by model × task. Column “All” matches the mean OI rate in the main paper’s leaderboard (Table 1); the “Mean” row matches the per-task OI rates in Table 3.
ModelApart.Rec.let.StressItiner.DatingGiftAll
Qwen3-8B63.956.445.346.137.634.248.7
DeepSeek-v4-pro62.153.842.543.432.130.945.4
GPT-4o-mini61.653.342.242.932.529.945.1
DeepSeek-v4-flash61.152.641.842.531.430.844.6
Qwen3.6-plus60.952.541.742.631.530.644.5
Kimi-K2.560.151.040.741.230.329.243.1
Gemini-3-flash57.347.939.440.127.925.641.1
GLM-5.156.547.038.539.627.325.040.2
GPT-5.554.844.437.437.425.624.538.5
GPT-5.4-nano54.644.735.736.525.223.837.6
Claude-Opus-4-652.641.534.434.923.621.935.4
Gemini-3.1-pro52.541.334.334.423.221.835.1
Mean57.848.239.838.728.527.041.6
Table 11: Judge–human agreement on 400 stratified claims.
MetricFour-classBinary (over-inference)
Accuracy0.8980.950
Cohen’s κ0.8630.900
Macro-F10.896

研究结果

  • 全部12个模型均有35%至49%的陈述属于无根据推断(跨模型平均41.6%,按陈述加权平均41.8%),没有一个模型能在整体评测中避免这一问题。
  • 模型自我报告的过度推断率与评审模型测出的过度推断率在跨模型层面呈负相关排序(斯皮尔曼相关系数负0.60,p值0.044,自举置信区间[负0.90,正0.06]),但在单个模型内部,自查仍能中等到较强地区分该模型自身陈述的好坏(AUROC 0.58至0.83)。
  • 过度推断率因任务差异很大,从礼物推荐任务的27%到公寓描述任务的59%,任务越需要脱离已知三条事实进行想象,编造率越高。
  • 在八轮对话试点中,9个模型的推断属性数量几乎呈线性累积(拟合优度R方大于0.90,每轮新增5至15个),增长最快的五个模型几乎从不修正先前编造的属性(删除率0.4%至5%)。
  • 独立评审模型在400条陈述上与盲评人工标注员达到近乎完美的一致性(四分类科恩卡帕系数0.863,二分类0.900),支持将其作为评分工具使用。

可应用场景

  • 为构建具备记忆功能的个性化AI助手的团队提供警示:不应依据模型自身报告的信心水平来判断哪个模型更安全
  • 系统设计者可以在将推断出的用户属性写入持久记忆之前,加入一个独立验证环节
  • 任务设计者可以预见到开放式、需要较多想象的个性化任务(如描述居住空间)比具体任务(如挑选礼物)更容易出现编造内容,并据此设计提示或警示
  • 在已部署的单一模型内部,自查分数仍可用作内部过滤器来标记可能编造的陈述,但不应用于跨模型的安全性比较

局限与待验证事项

  • 多轮累积试点仅使用2个虚构用户画像、8轮对话,因此绝对数值应视为参考性而非精确结果
  • 记忆更新提示要求模型保留先前属性除非被明确反驳,这一设计本身就偏向于累积效应;允许自由修剪的中立提示作为对照实验留待未来研究
  • 该基准将证据固定为恰好三条透露的事实,以模拟交互早期阶段,因此没有研究在更长时间、信息积累更多之后过度推断行为会如何变化
  • 评审模型(Claude-Opus-4-7)本身可能存在自身偏差,人工验证仅限于400条陈述、由单一盲评标注员完成;扩展到更多标注员及其他评审模型仍是未来工作
  • 支撑'自我监控反转'结论的跨模型排序相关性仅基于12个模型,置信区间较宽(自举区间[负0.90,正0.06]),因此该结果被定性为探索性发现而非精确估计的效应

为什么重要

随着具备持久记忆的个性化AI助手日益普及,这项研究表明这类系统声称'了解'用户的内容中,有很大一部分从未被用户实际说出;同时它也警示,若依据模型自身的信心程度来判断哪个模型更安全,可能会得出完全相反的结论。这为构建可信个性化系统指出了方向:应依赖外部验证,而非模型的自我报告。

本文术语

  • 过度推断(over-inference,OI) · 模型断言某个用户属性,但该属性超出了用户实际提供的证据范围
  • 自我监控反转(Self-Monitoring Inversion) · 模型自我报告的过度推断率对模型的排序,与独立评审测出的排序恰好相反的现象
  • 四分类真实性体系 · 将每条陈述归为有依据、合理推断、刻板印象或纯粹编造四类之一
  • AUROC · 一个0到1之间的分数,衡量某个信号(此处为模型自查结果)区分真实正例(被评审标记为编造)与负例的能力
  • 想象梯度(imagination gradient) · 任务从可以直接根据已知事实回答,到必须依靠纯粹想象才能完成的不同程度分布

论文原文摘要(英文)

Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.

作者 · Yushi Sun, Yanjie Zhang, Rui Sheng

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Yushi Sun et al., arXiv:2608.04570, CC BY 4.0