Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems
인터뷰형 대화 시스템을 테스트하려면 다양한 가상 사용자가 필요한데, LLM으로 그런 가짜 사용자 성격을 자동으로 만들어냈다
여행 정보나 취향을 묻는 인터뷰 대화 시스템은 사람이 직접 테스트하려면 시간과 비용이 많이 든다. 이 논문은 소수의 예시 페르소나만 주면 LLM이 대화 스타일이 서로 다른 가상 사용자 페르소나를 대량으로 만들어내는 방법을 제안한다. 실험 결과 이렇게 만든 가상 사용자들이 내는 발화가 기존 방식보다 더 다양해졌다.
무엇을 했나
- 여행 정보를 묻는 시스템과 디저트 취향을 묻는 시스템, 두 개의 일본어 인터뷰 대화 시스템에 방법을 적용했다.
- 사람이 직접 만든 페르소나 10개를 예시로 주고 GPT-4o에게 조건당 100개의 새 페르소나를 만들게 했다.
- 페르소나를 만들 때 시스템을 물건처럼 대하는지 사람처럼 대하는지(의인화 정도), 말을 에둘러 하는지 직설적으로 하는지(정교함 정도) 두 가지 성격 특성을 함께 지정했다.
- 생성된 페르소나로 GPT-4o 기반 가상 사용자가 GPT-4o-mini 기반 인터뷰 시스템과 대화하게 하고, 발화 길이·어휘 다양성 등 여러 지표로 다양성을 측정했다.
- 페르소나만 LLM으로 새로 생성해도 내용 다양성이 늘었고(내용어 타입-토큰 비율이 여행 도메인 .106에서 .122로, 디저트 도메인 .109에서 .133으로 상승), 정교함 특성을 지정하면 발화 길이의 편차가 커져(표준편차가 여행 도메인 7.0에서 18.2로, 디저트 도메인 8.0에서 17.7로 상승) 문체 다양성도 늘었다.
| Condition | Personality | #Dialogues | Ave. utterance length (S.D.) | Ave. utterance | length (S.D.) | Total words | Total | words | Unique words | Unique | words | Unique bigrams | Unique | bigrams | TTR |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ave. utterance | |||||||||||||||
| length (S.D.) | |||||||||||||||
| Total | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| bigrams | |||||||||||||||
| BL | 100 | 28.2 | (7.7) | 42,260 | 15,747 | 31,620 | .373 | ||||||||
| noPT | 100 | 28.5 | (7.0) | 42,784 | 16,362 | 32,730 | .382 | ||||||||
| APM | All | 100 | 28.4 | (7.2) | 42,563 | 16,089 | 32,334 | .378 | |||||||
| High | 50 | 30.1 | (7.6) | 22,547 | 8,406 | 17,098 | .373 | ||||||||
| Low | 50 | 26.7 | (6.3) | 20,016 | 7,683 | 15,236 | .384 | ||||||||
| EL | All | 100 | 36.0 | (18.2) | 53,951 | 18,556 | 39,010 | .344 | |||||||
| High | 50 | 50.7 | (13.9) | 38,001 | 12,175 | 26,883 | .320 | ||||||||
| Low | 50 | 21.3 | (5.7) | 15,950 | 6,381 | 12,127 | .400 | ||||||||
| APM+EL | All | 100 | 31.4 | (13.4) | 47,111 | 16,856 | 34,839 | .358 | |||||||
| High+High | 25 | 46.2 | (12.3) | 17,308 | 5,624 | 12,309 | .325 | ||||||||
| High+Low | 25 | 23.6 | (4.8) | 8,836 | 3,451 | 6,787 | .391 | ||||||||
| Low+High | 25 | 35.2 | (9.9) | 13,184 | 4,651 | 9,796 | .353 | ||||||||
| Low+Low | 25 | 20.8 | (5.9) | 7,783 | 3,130 | 5,947 | .402 |
| Condition | Personality | #Dialogues | Ave. utterance length (S.D.) | Ave. utterance | length (S.D.) | Total words | Total | words | Unique words | Unique | words | Unique bigrams | Unique | bigrams | TTR |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ave. utterance | |||||||||||||||
| length (S.D.) | |||||||||||||||
| Total | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| bigrams | |||||||||||||||
| BL | 100 | 25.9 | (7.8) | 25,152 | 11,138 | 20,534 | .443 | ||||||||
| noPT | 100 | 25.6 | (8.0) | 25,509 | 11,105 | 20,634 | .435 | ||||||||
| APM | All | 100 | 25.6 | (8.2) | 25,680 | 11,124 | 20,862 | .433 | |||||||
| High | 50 | 26.5 | (8.2) | 13,269 | 5,723 | 10,824 | .431 | ||||||||
| Low | 50 | 24.6 | (8.0) | 12,411 | 5,401 | 10,038 | .435 | ||||||||
| EL | All | 100 | 33.9 | (17.7) | 34,020 | 13,214 | 26,117 | .388 | |||||||
| High | 50 | 47.5 | (14.8) | 23,755 | 8,578 | 17,808 | .361 | ||||||||
| Low | 50 | 20.3 | (6.1) | 10,265 | 4,636 | 8,309 | .452 | ||||||||
| APM+EL | All | 100 | 28.1 | (12.2) | 28,067 | 11,549 | 22,182 | .411 | |||||||
| High+High | 25 | 39.4 | (13.1) | 9,857 | 3,743 | 7,643 | .380 | ||||||||
| High+Low | 25 | 22.2 | (6.7) | 5,560 | 2,458 | 4,533 | .442 | ||||||||
| Low+High | 25 | 30.6 | (10.5) | 7,642 | 3,087 | 6,015 | .404 | ||||||||
| Low+Low | 25 | 20.0 | (5.9) | 5,008 | 2,261 | 3,991 | .451 |
왜 중요한가
실제 사용자를 모아 인터뷰 시스템을 테스트하는 대신 이런 방법으로 사람 노동을 줄이면서도 다양한 상황을 미리 시험해 볼 수 있다. 다양한 성격의 가짜 사용자를 많이 만들수록 시스템이 미처 예상하지 못한 문제를 발견할 확률이 높아진다.
이 논문의 용어
- 페르소나 · 가상 사용자에게 부여하는 성격, 취향, 말투 등의 프로필
- 인터뷰 대화 시스템 · 여러 사용자에게 질문을 던져 정보를 수집하는 대화형 AI 시스템
- 사용자 시뮬레이터 · 실제 사람 대신 대화 시스템과 상호작용하도록 만든 가상의 대화 상대
- 타입-토큰 비율(TTR) · 전체 단어 수 대비 서로 다른 단어 수의 비율로, 어휘가 얼마나 다양한지 보여주는 지표
- few-shot 예시 · 모델에게 몇 개의 예시를 보여줘서 비슷한 형식으로 새로운 결과를 만들게 하는 방법
논문 원문 초록 (영문)
This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort and cost. Therefore, testing with user simulators can be beneficial. Since most conventional user simulators have been primarily designed for training task-oriented dialogue systems, little attention has been paid to the personas of the simulated users. During development, testing interview dialogue systems requires simulating a wide range of user behaviors, but manually creating a large number of personas is labor-intensive. We propose a method that automatically generates personas for user simulators using a large language model. Furthermore, by assigning personality traits related to communication styles when generating personas, we aim to increase the diversity of communication styles in the user simulator. Experimental results show that the proposed method enables the user simulator to generate utterances with greater variation.
arXiv에서 원문 보기최신 논문
- Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms데이터 플랫폼 변경도 코드처럼 '설계도 조각'을 붙여서 검토하면 어떨까: 실험 설계 논문
- Are LLMs becoming similarly creative? Evidence from three years of models최신 AI 챗봇일수록 서로 비슷한 답을 내놓는다는 3년치 조사 결과
- Auditing Cross-Lingual Fairness in Language Model WatermarkingAI 생성 텍스트를 잡아내는 워터마크 기술이 영어 아닌 언어에서는 훨씬 부실하게 작동하고, 그 격차는 개별 언어가 아니라 언어 계열 단위로 나타난다
- TESTNAV: Pareto-Guided Search for Compositional Robustness TestingAI 모델을 여러 손상이 겹친 입력으로 시험할 때, 굳이 다 테스트하지 않고도 '진짜 위험한 실패'만 골라내는 탐색법
- Optimal Skill Selection for LLM Agents with Provable Bicriteria GuaranteesAI 에이전트에게 어떤 '스킬 문서'를 몇 개나 줘야 잘 작동하는지, 수학적으로 최적해를 보장하며 골라주는 방법
- Reliable Financial Named Entity Recognition under Domain Shift금융 AI가 서류체 문장에서 배운 자신감은 트위터로 가면 거짓말이 된다
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction텍스트가 빠지거나 망가져도, AI가 그 자리를 대신할 '가짜 텍스트'를 한 번에 만들지 않고 여러 번 고쳐가며 감정을 더 정확히 읽어낸다
METAL LAB 최신 기사
그림 출처: Mikio Nakano et al., arXiv:2608.19549, arxiv-nonexclusive