Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems
要测试访谈式对话系统需要大量不同性格的虚拟用户,这项研究用大语言模型自动生成这些虚拟用户人设
询问旅行计划或甜点喜好等信息的访谈类对话系统,用真人测试成本很高。这篇论文只需几个人工写的示例人设,就能让大语言模型自动批量生成风格各异的虚拟用户人设,再用这些人设驱动模拟用户进行对话测试。实验显示,这样生成的模拟对话比只用固定人设时更加多样化。
他们做了什么
- 方法在两个日语访谈对话系统上做了测试,一个询问旅行相关信息,一个询问甜点偏好
- 从10个人工编写的种子人设出发,用少样本上下文学习的方式让GPT-4o为每种条件生成100个新人设
- 生成人设时还指定了两种与沟通风格相关的性格特质:拟人化程度(把系统当物品还是当人对待)和表达详略程度(说话是绕弯子还是直接)
- 用生成的人设驱动基于GPT-4o的模拟用户,与基于GPT-4o-mini的访谈系统对话,再用发言长度差异、词汇型符比等指标衡量多样性
- 仅靠大语言模型生成新人设就已经提升了内容多样性(实词型符比在旅行领域从.106升到.122,在甜点领域从.109升到.133),而加入详略程度这一特质后,发言长度的标准差也明显提升(旅行领域从7.0升到18.2,甜点领域从8.0升到17.7),说明文体多样性也增加了
| Condition | Personality | #Dialogues | Ave. utterance length (S.D.) | Ave. utterance | length (S.D.) | Total words | Total | words | Unique words | Unique | words | Unique bigrams | Unique | bigrams | TTR |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ave. utterance | |||||||||||||||
| length (S.D.) | |||||||||||||||
| Total | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| bigrams | |||||||||||||||
| BL | 100 | 28.2 | (7.7) | 42,260 | 15,747 | 31,620 | .373 | ||||||||
| noPT | 100 | 28.5 | (7.0) | 42,784 | 16,362 | 32,730 | .382 | ||||||||
| APM | All | 100 | 28.4 | (7.2) | 42,563 | 16,089 | 32,334 | .378 | |||||||
| High | 50 | 30.1 | (7.6) | 22,547 | 8,406 | 17,098 | .373 | ||||||||
| Low | 50 | 26.7 | (6.3) | 20,016 | 7,683 | 15,236 | .384 | ||||||||
| EL | All | 100 | 36.0 | (18.2) | 53,951 | 18,556 | 39,010 | .344 | |||||||
| High | 50 | 50.7 | (13.9) | 38,001 | 12,175 | 26,883 | .320 | ||||||||
| Low | 50 | 21.3 | (5.7) | 15,950 | 6,381 | 12,127 | .400 | ||||||||
| APM+EL | All | 100 | 31.4 | (13.4) | 47,111 | 16,856 | 34,839 | .358 | |||||||
| High+High | 25 | 46.2 | (12.3) | 17,308 | 5,624 | 12,309 | .325 | ||||||||
| High+Low | 25 | 23.6 | (4.8) | 8,836 | 3,451 | 6,787 | .391 | ||||||||
| Low+High | 25 | 35.2 | (9.9) | 13,184 | 4,651 | 9,796 | .353 | ||||||||
| Low+Low | 25 | 20.8 | (5.9) | 7,783 | 3,130 | 5,947 | .402 |
| Condition | Personality | #Dialogues | Ave. utterance length (S.D.) | Ave. utterance | length (S.D.) | Total words | Total | words | Unique words | Unique | words | Unique bigrams | Unique | bigrams | TTR |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ave. utterance | |||||||||||||||
| length (S.D.) | |||||||||||||||
| Total | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| bigrams | |||||||||||||||
| BL | 100 | 25.9 | (7.8) | 25,152 | 11,138 | 20,534 | .443 | ||||||||
| noPT | 100 | 25.6 | (8.0) | 25,509 | 11,105 | 20,634 | .435 | ||||||||
| APM | All | 100 | 25.6 | (8.2) | 25,680 | 11,124 | 20,862 | .433 | |||||||
| High | 50 | 26.5 | (8.2) | 13,269 | 5,723 | 10,824 | .431 | ||||||||
| Low | 50 | 24.6 | (8.0) | 12,411 | 5,401 | 10,038 | .435 | ||||||||
| EL | All | 100 | 33.9 | (17.7) | 34,020 | 13,214 | 26,117 | .388 | |||||||
| High | 50 | 47.5 | (14.8) | 23,755 | 8,578 | 17,808 | .361 | ||||||||
| Low | 50 | 20.3 | (6.1) | 10,265 | 4,636 | 8,309 | .452 | ||||||||
| APM+EL | All | 100 | 28.1 | (12.2) | 28,067 | 11,549 | 22,182 | .411 | |||||||
| High+High | 25 | 39.4 | (13.1) | 9,857 | 3,743 | 7,643 | .380 | ||||||||
| High+Low | 25 | 22.2 | (6.7) | 5,560 | 2,458 | 4,533 | .442 | ||||||||
| Low+High | 25 | 30.6 | (10.5) | 7,642 | 3,087 | 6,015 | .404 | ||||||||
| Low+Low | 25 | 20.0 | (5.9) | 5,008 | 2,261 | 3,991 | .451 |
为什么重要
开发者不必招募真人测试者,就能用这种方法对访谈对话系统进行覆盖多种用户行为的压力测试,降低开发成本和人力投入。人设越多样,发现系统未曾预料到的问题的概率就越高。
本文术语
- 人设(persona) · 赋予虚拟用户的性格、偏好和说话风格等信息
- 访谈对话系统 · 通过提问从用户那里收集信息的对话式人工智能系统
- 用户模拟器 · 代替真人与对话系统进行交互的虚拟对话对象
- 型符比(TTR) · 不同词数量占总词数的比例,用来衡量词汇多样性的指标
- 少样本上下文学习 · 在提示中给模型看几个示例,让它据此生成风格相似的新内容
论文原文摘要(英文)
This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort and cost. Therefore, testing with user simulators can be beneficial. Since most conventional user simulators have been primarily designed for training task-oriented dialogue systems, little attention has been paid to the personas of the simulated users. During development, testing interview dialogue systems requires simulating a wide range of user behaviors, but manually creating a large number of personas is labor-intensive. We propose a method that automatically generates personas for user simulators using a large language model. Furthermore, by assigning personality traits related to communication styles when generating personas, we aim to increase the diversity of communication styles in the user simulator. Experimental results show that the proposed method enables the user simulator to generate utterances with greater variation.
在 arXiv 阅读最新论文
- LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment在正式微调前先偷看几步训练的梯度,让LoRA的初始化更聪明
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction文本信息缺失或损坏时,这个AI不靠一次性猜测,而是反复修正猜测结果,从而更准确地判断情绪
- Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured Partitioning别再机械切分时间序列,按语义把它切成有意义的块
- Reliable Financial Named Entity Recognition under Domain ShiftAI在正式文件里学到的自信,一到推特上就变得不可信
- Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis让AI分析脑影像数据时,把“为什么这个结论可信”也一并记录下来
- GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-Hailing滴滴把打车派单从预测-计算-匹配三段式流程改成一次生成完成,线上效果提升明显
- Auditing Cross-Lingual Fairness in Language Model Watermarking本该识别AI生成文本的水印技术在非英语语言中表现明显更差,而且这种差距按语系而非单个语言呈现
METAL LAB 最新报道
图片来源: Mikio Nakano et al., arXiv:2608.19549, arxiv-nonexclusive