MatrAIx: Simulating the World with 8.3 Billion Persona Agents
MatrAIx用83亿个模拟人格代替真人来测试AI产品
MatrAIx是一套基础设施,包含一个由1,290个属性定义、拥有83亿条记录的模拟人格数据库,四种可让这些人格实际行动的交互环境(问卷、AI聊天、网页、应用),以及1,010个评测任务,目的是让对AI系统和数字产品的类人评测变得更快更便宜。作者进行了一项400次试验的受控研究,检验人格代理是否真的会按照设定的性格行事,还请人类评审员和LLM评审员分别为提取出的人格质量打分。作者明确指出这仍是一个用于生成假设的工具,不能替代真人验证。
METAL LAB 解读图
MatrAIx流程:从人格生成到结果报告
证据状态实测结果与计划中的工作并存
- Persona 8B由1,290个属性定义的83亿条模拟人格,通过依赖图采样(合成)和真实来源映射(基于真人)构建,并发布了约100万条的公开核心集
- 群组筛选评测者指定目标受众(如年龄、收入、地区),系统从Persona 8B中检索匹配的人格组成评测群组
- 四种交互环境问卷(Survey)、AI聊天(AI Chatbot)、网页(Web)、应用(App)环境让人格代理与被测产品互动,记录对话、操作和界面状态
- 任务与验证器1,010个任务规范定义目标和成功标准,由程序化验证器或人类/LLM评审判断人格代理的输出是否满足条件
- 群体级报告将各次试验结果按任务、群组、子群体汇总成报告,展示关键指标(如预计留存率)及支撑证据
他们做了什么
- 研究动机是人类评测AI系统和数字产品成本高、速度慢、难以规模化,而离线基准测试虽然可扩展,却往往忽略了用户之间的多样性和真实互动行为。
- Persona 8B是一个基于1,290个分类属性(年龄段、语言水平、风险偏好等)统一模式构建的83亿条人格记录数据库;合成人格通过一个依赖图(一种有向图,让年龄等父属性影响教育水平等子属性)采样以保留真实的属性关联,而基于真人的人格则从维基百科传记、亚马逊评论历史、开发者调查等来源提取并映射到同一模式中。
- 团队发布了一个经质量过滤的约100万条人格核心集(59.9847万条基于真人、40万条合成),并提供MatrAIx Playground运行环境,支持问卷(Survey)、AI聊天(AI Chatbot)、网页(Web)、应用(App)四种环境,以及覆盖25个以上领域的1,010个可复用评测任务库。
- 团队使用三种大语言模型(Claude Opus 4.8、GPT 5.5、Claude Haiku 4.5)在八个代表性任务上共运行了18,189次评测试验,捕捉了涨价后的购买犹豫、AI助手失败后用户是否愿意继续使用、以及对响应延迟的容忍度等因人而异的行为模式。
- 一项400次试验的受控研究发现,设定的行为特征在91.5%(366/400次)的试验中被正确表达或正确抑制;另外,人类评审员和LLM评审员分别对从真实来源提取的人格质量进行了打分。

| Top-level group | Dims. | Representative attributes | Representative grounding |
|---|---|---|---|
| Background | 238 | Age, region, language, education, family, career, industry | Population statistics, household surveys, education and labor taxonomies |
| Psychology | 210 | Personality, values, worldview, motivation, risk | Validated instruments, values surveys, schema design priors |
| Capability | 331 | Domain expertise, general skills, tools, programming, developer context | Occupational taxonomies, technology/developer surveys |
| Behavior and Interaction | 124 | Preferences, habits, interaction state, work practices, technology adoption | Time-use, consumer, workplace, and technology-use evidence |
| Lifestyle | 387 | Interests, media, culture, hobbies, sports, food, health, fitness | Health statistics, consumption surveys, cultural sources |
| Total | 1,290 |
| Source | Released records |
|---|---|
| Wikipedia extraction | 323,438 |
| Amazon Review extraction | 97,915 |
| Stack Overflow survey extraction | 113,120 |
| PRISM Alignment | 1,487 |
| General Social Survey | 63,532 |
| MatrAIx volunteer survey | 355 |
| Human-grounded subtotal | 599,847 |
| Full-DAG synthetic | 400,000 |
| Total | 999,847 |
| Commerce | Software | Finance | Healthcare | Other | Total | |
|---|---|---|---|---|---|---|
| Survey | 202 | 138 | 141 | 139 | 1 | 621 |
| AI Chatbot | 3 | 11 | 17 | 29 | 311 | 371 |
| Web | 2 | 2 | 2 | 0 | 6 | 12 |
| App | 0 | 5 | 1 | 0 | 0 | 6 |
| Total | 207 | 156 | 161 | 168 | 318 | 1,010 |
| Group | Subgroup | Schema category | Count | Representative attributes |
|---|---|---|---|---|
| Background | Demographics | Demographic: Core | 25 | Age bracket; region; gender identity |
| Background | Demographics | Demographic: Cultural | 2 | Cultural background; attitude toward immigration |
| Background | Demographics | Demographic: Family | 1 | Household size |
| Background | Demographics | Demographic: Life Events | 24 | Life stage; major life events; childhood environment |
| Background | Language | Linguistic: Language | 53 | Primary language; English proficiency; multilingualism |
| Background | Language | Linguistic: Communication | 37 | Expected tone; verbosity; communication preferences |
| Background | Education | Learning: Academic | 34 | Highest education; academic field; institution tier |
| Background | Education | Learning: Style | 1 | Learning style |
| Background | Career | Professional: Career | 4 | Research output; seniority; years of experience |
| Background | Career | Professional: Industry | 51 | Company size; role function; industry |
| Background | Career | Developer: Professional Context | 6 | Professional status; role archetype; contribution context |
| Psychology | Personality | Personality: Character | 34 | Domain stance; dominant trait; curiosity |
| Psychology | Personality | Personality: Big Five | 50 | Imagination; artistic interest; emotionality |
| Psychology | Personality | Personality: MBTI | 2 | Neurotype; Myers-Briggs type |
| Psychology | Personality | Personality: Relationships | 4 | Attachment anxiety; attachment avoidance; interpersonal agency |
| Psychology | Worldview | Values & Motivation | 46 | Core value; religiosity; economic motivation |
| Psychology | Worldview | Worldview: Beliefs | 67 | Political leaning; trust level; safety sensitivity |
| Psychology | Decision-Making | Risk & Decision | 7 | Risk tolerance; decision style; need for closure |
| Capability | Domains | Expertise: Domains | 144 | Domain; subject specialty; technology savviness |
| Capability | Skills | Expertise: Skills | 64 | Writing; copywriting; editing |
| Capability | Skills | Skills: Tools | 69 | Excel; Google Sheets; Python |
| Capability | Skills | Skills: Programming | 44 | Comment style; summary documentation; naming verbosity |
| Capability | Skills | Developer: Code Maintenance | 10 | Complexity tolerance; modularity preference; type-system orientation |
| Behavior and Interaction | Personal Behavior | Behavior: Preferences | 34 | Modality preference; accessibility needs; media diet |
| Behavior and Interaction | Personal Behavior | Behavior: Habits | 30 | Journaling; meditation; use of to-do lists |
| Behavior and Interaction | Personal Behavior | Behavior: Time | 3 | Time pressure; sleep schedule; micromanagement aversion |
| Behavior and Interaction | Interaction State | State: Emotional | 5 | Emotional state; intent; query complexity |
| Behavior and Interaction | Work Practices | Behavior: Work | 2 | Work schedule; office versus remote work |
| Behavior and Interaction | Work Practices | Developer: Open Source Behavior | 7 | Open-source activity; GitHub contribution mode; pull-request style |
| Behavior and Interaction | Work Practices | Developer: Community Behavior | 4 | Stack Overflow use; participation style; help-seeking preference |
| Schema group | Facets | Grounding roles | Sources |
|---|---|---|---|
| Background | Demographics (52); language (90); education (35); career (61) | Category definitions; population priors; household, language, education, and labor dependencies | UN World Population Prospects and Population Data [87, 86]; World Bank WDI and WorldPop [94, 97]; Eurostat, ACS PUMS, and IPUMS [21, 82, 38]; DHS, UNICEF MICS, and OECD Family Database [77, 85, 62]; Pew and World Values Survey [69, 96]; UNESCO UIS, World Bank Education Statistics, OECD INES, and OECD PISA [84, 93, 61, 64]; ILOSTAT, BLS OEWS, and O*NET [34, 81, 60]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]. |
| Psychology | Personality (90); worldview (113); decision-making (7) | Instrument and value-set design; selected prevalence estimates; validation | IPIP and MIDUS [35, 56]; Pew and World Values Survey [69, 96]; GSS, European Social Survey, ISSP, and Gallup World Poll [58, 20, 36, 23]; Afrobarometer, Arab Barometer, Asian Barometer, Latinobarometro, and Eurobarometer [1, 3, 5, 46, 19]; ARDA [6]. |
| Capability | Domain expertise (144); general skills (64); tools (69); programming (44); developer context (10) | Occupational and skill taxonomies; technology access and adoption; developer-tool prevalence | ITU Statistics, World Bank WDI, DataReportal, and Pew Internet [37, 94, 15, 70]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]; O*NET [60]. |
| Behavior and Interaction | Personal behavior (67); interaction state (5); work practices (13); technology use (39) | Time-use and consumer priors; workplace behavior; technology and AI adoption | American Time Use Survey, Consumer Expenditure Surveys, and OECD Time Use [79, 80, 63]; ITU, DataReportal, and Pew Internet [37, 15, 70]; Stack Overflow Survey, GitHub Octoverse, JetBrains Developer Ecosystem, and O*NET [75, 25, 39, 60]. |
| Lifestyle | Interests (358); physical health (25); fitness (2); health lifestyle (2) | Health and disability priors; consumption and time use; cultural and interest category design | WHO GHO and IHME GBD [95, 33]; ACS PUMS, NHIS, BRFSS, DHS, and UNICEF MICS [82, 11, 10, 77, 85]; ATUS, CEX, and OECD Time Use [79, 80, 63]; FAOSTAT and UNESCO Culture Statistics [22, 83]; Eurobarometer, Pew, DataReportal, WVS, and Gallup World Poll [19, 69, 70, 15, 96, 23]. |
| v | πi(v) | qlang | qregion | ri(v) | mi(v) | πirimi | pθ(v∣xPa(i)) |
|---|---|---|---|---|---|---|---|
| None | 0.30 | 0.02 | 0.10 | 0.022 | 0 | 0.000 | 0.000 |
| Basic | 0.30 | 0.08 | 0.20 | 0.178 | 1 | 0.053 | 0.028 |
| Fluent | 0.25 | 0.30 | 0.35 | 1.680 | 1 | 0.420 | 0.224 |
| Native | 0.15 | 0.60 | 0.35 | 9.333 | 1 | 1.400 | 0.747 |
| sum | 1.00 | 1.00 | 1.00 | 1.873 | 1.000 |

| Stage | Rejected | Remaining |
|---|---|---|
| Original corpus | – | 10,002,288,277 |
| Contradiction filter | 239,310 | 10,002,048,967 |
| Human exact/MinHash deduplication | 41,597 | 2,222,496 human |
| Synthetic projection deduplication | 252,936,392 | 9,746,848,482 synthetic |
| Synthetic deterministic cutoff | 1,349,070,978 | 8,397,777,504 synthetic |
| Audited baseline | 8,400,000,000 |

| Dimension | Answered | Share of those answering |
|---|---|---|
| Age bracket | 321 | 25–34 22.1%, 55–64 19.0%, 35–44 15.9%, 18–24 13.7%, 45–54 13.4%, 65–74 8.4%, 75–84 4.0%, 85+ 3.4% |
| Gender identity | 322 | Woman 47.5%, Man 42.9%, Non-binary 3.7%, Self-described 3.4%, Prefer not to say 2.5% |
| Region | 329 | South Asia 20.4%, Sub-Saharan Africa 19.1%, East Asia 13.4%, Southeast Asia 10.9%, Latin America 9.7%, Western Europe 7.0%, MENA 6.7%, North America 5.8%, Eastern Europe 4.6%, Oceania 2.4% |
| Urbanicity | 328 | Rural 33.2%, Dense urban 22.3%, Small town 21.0%, Suburban 20.4%, Nomadic / remote 3.0% |
| Socioeconomic band | 340 | Low 31.8%, Lower-middle 30.0%, Middle 20.3%, Upper-middle 12.6%, High 5.3% |
| Employment | 330 | Full-time 25.8%, Retired 16.1%, Homemaker 14.2%, Self-employed 11.8%, Part-time 9.1%, Student 8.2%, Unemployed 7.6%, Gig / freelance 7.3% |

| Collection | Tasks | Environment | Composition |
|---|---|---|---|
| Synthetic persona surveys | 405 | Survey | 135 Finance, 135 Healthcare, and 135 Software questionnaires, generally sharing a common survey and reporting template. |
| Product surveys | 200 | Survey | Twenty product-research archetypes, including purchase intent, price sensitivity, recommendation, and retention, crossed with ten retail products. |
| Synthetic chatbots | 351 | AI Chatbot | Scenario families across 27 domains, including travel, legal, healthcare, telecommunications, real estate, insurance, finance, and education. |

| Contract component | Declared content | Audit purpose |
|---|---|---|
| Task metadata | Stable name and version, environment type, product domain, difficulty, tags, artifact paths, and runtime budgets | Identifies the recipe and prevents results from silently moving between task versions. |
| Persona-facing scenario | Context, user goal, constraints, disclosure policy, and required submission | Defines what every sampled persona is asked to do without exposing verifier internals. |
| Cohort strategy | Persona-schema filters, sampling mode, stratification fields, sample size, weights, and seed policy | Makes the target audience explicit and permits the cohort to be redrawn. |
| Product attachment | Questionnaire or stimulus, chat endpoint or sidecar, website target, or native application backend | Identifies the system under test and how the runtime reaches it. |
| Verifier | Required artifacts, structured finding schema, objective checks, timeouts, and failure conditions | Converts a trial into reproducible outcomes with supporting evidence. |
| Reporting policy | Aggregations, subgroup facets, summaries, optional judge directives, and disclosure rules | Defines how trial findings become a cohort-level report. |

| Environment | Task-owned inputs | Primary trial artifacts |
|---|---|---|
| Survey | Stimulus, questionnaire schema, response constraints, and optional rationale prompts | Typed responses, missing or invalid items, rationales, confidence, and completion summary. |
| AI Chatbot | Product endpoint or sidecar, user goal, staged-disclosure policy, turn and safety limits | Full transcript, tool or service events, resolution state, termination reason, and post-run feedback. |
| Web | URL or hosted site, browser backend, exploration requirements, and submission schema | Page and action trace, considered options, screenshots where enabled, final submission, and terminal page or site state. |
| App | Native target, platform backend, initial state, permissions, and terminal-state checks | Screenshot and action trace, exported files, application state, permission changes, and cross-application side effects. |
| Task | Env. | Primary outcome | Share giving that answer | Range | q | ||
|---|---|---|---|---|---|---|---|
| GPT | Opus | Haiku | (pt) | ||||
| Annual checkup | Survey | q31 = e | 50.5% | 41.4% | 29.7% | 20.8 | <0.001∗∗∗ |
| Candy Land price | Survey | q_threshold = hesitate | 98.3% | 27.0% | 83.3% | 71.3 | <0.001∗∗∗ |
| OpenBB honesty | Chat | wouldStillContinueUse = unsure | 81.3% | 85.5% | 28.1% | 57.4 | <0.001∗∗∗ |
| Meal planning† | Chat | adherenceLikelihood = 7 | 50.6% | 0.2% | 40.0% | 50.4 | <0.001∗∗∗ |
| Notion plans | Web | decision_subject_id = plus | 63.5% | 21.5% | 75.7% | 54.2 | <0.001∗∗∗ |
| MIT OCW course | Web | task_course_level = Graduate | 40.0% | 34.5% | 47.5% | 13.0 | <0.001∗∗∗ |
| News+ subscription | App | clicked_get_started = true | 4.2% | 20.8% | 0.0% | 20.8 | 0.025∗ |
| Stocks sentiment | App | sentiment = hold | 60.0% | 30.0% | 50.0% | 30.0 | 0.153 ns |
| Task | Product-level measure | GPT 5.5 | Opus 4.8 | Haiku 4.5 |
|---|---|---|---|---|
| Annual checkup | Very likely to schedule in time | 50.5% (505/1,000) | 41.4% (414/1,000) | 29.7% (297/1,000) |
| Candy Land price | Hesitates or worse at new price | 98.3% (983/1,000) | 27.0% (270/1,000) | 83.3% (833/1,000) |
| OpenBB honesty | Would not continue using | 18.5% (185/1,000) | 14.5% (132/909) | 71.9% (719/1,000) |
| Meal planning† | Stated need fully satisfied | 46.2% (462/1,000) | 0.2% (2/1,000) | 28.1% (281/1,000) |
| Notion plans | Selected a paid plan | 75.8% (776/1,024) | 23.2% (237/1,022) | 93.9% (958/1,020) |
| MIT OCW course | Chose a graduate-level course | 40.0% (403/1,008) | 34.5% (347/1,007) | 47.5% (473/996) |
| News+ subscription | Subscribed | 4.2% (1/24) | 20.8% (5/24) | 0.0% (0/24) |
| Stocks sentiment | Buy opinion | 40.0% (8/20) | 70.0% (14/20) | 47.4% (9/19) |
| Task | Persona dimension | Groups | Sig. | Spearman ρ, exact p | ||
|---|---|---|---|---|---|---|
| GPT/Opus | GPT/Haiku | Opus/Haiku | ||||
| Annual checkup | age bracket | 6 | ∘∘∙ | +0.26 p=0.329 | −0.43 p=0.822 | −0.77 p=0.971 |
| Candy Land price | economic motivation | 4 | ∘∘∘ | −0.40 p=0.792 | −0.20 p=0.625 | +0.80 p=0.167 |
| OpenBB honesty | trust level | 4 | ∙∙∙ | +1.00 p=0.042 | +1.00 p=0.042 | +1.00 p=0.042 |
| Meal planning† | life stage | 4 | ∘∘∘ | +0.95 p=0.083 | +0.32 p=0.500 | +0.33 p=0.417 |
| Notion plans | company size | 8 | ∘∘∘ | +0.44 p=0.138 | +0.60 p=0.059 | −0.24 p=0.725 |
| MIT OCW course | academic field | 8 | ∘∘∘ | +0.93 p=0.001 | +0.09 p=0.420 | +0.06 p=0.452 |
| News+ subscr.‡ | economic motivation | 3 | ∘∘∘ | +0.00 p=0.667 | flat | flat |
| Stocks sentiment‡ | risk tolerance | 5 | ∘∘∘ | +0.47 p=0.267 | −0.92 p=1.000 | −0.67 p=0.933 |
| GPT × Opus | GPT × Haiku | Opus × Haiku | |
|---|---|---|---|
| Paired agreement over 88 joinable fields | |||
| Median Cohen’s κ | 0.000 | 0.000 | +0.001 |
| Fields at κ≤0 | 59 of 88 | 50 of 88 | 40 of 88 |
| Fields reaching κ≥0.2 | 7 of 88 | 8 of 88 | 7 of 88 |
| Fields at ≥50% agreement, κ<0.1 | 48 of 56 | 25 of 34 | 24 of 34 |
| Self-report fidelity: age band matches the persona | |||
| GPT 5.5 16.3% (ns) Opus 4.8 16.9% (ns) Haiku 4.5 100.0% | chance 16.7% |
| Survey | Chat | Web | OS-App | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Attribute | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | |||
| code-comment-style | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| code-naming-verbosity | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| code-summary-documentation | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-emoji-use | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-humor | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-politeness | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-storytelling | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-use-of-jargon | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-verbosity | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| register | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 |
| Attribute | Positive (declared) | Negative (opposite) |
|---|---|---|
| code-comment-style | Nearly every line carries an inline comment (# Iterate over every integer, # Check divisibility) | Code contains no comments whatsoever |
| code-naming-verbosity | calculate_average_of_passing_scores, number_of_passing_scores | Variables s, t, a, c, x, all single-letter |
| code-summary-documentation | Every function opens with a tldr: docstring | Prose plus code, no TLDR/summary header |
| cog-emoji-use | Many emoji across a short paragraph | No emoji despite a casual app-store prompt |
| cog-humor | Playful, witty asides (“the boxes may yet win”) | Measured, earnest tone, no jokes |
| cog-politeness | “Might I kindly ask that you resend the document at your earliest convenience” | “Hey, you forgot the attachment. Again.”: blunt and sarcastic |
| cog-storytelling | A narrated scene (“a storm came through, half the lodge went dark”) | Abstract, value-driven, no concrete scene |
| cog-use-of-jargon | Dense technical jargon (“TCP three-way handshake, SYN-ACK”) | Plain terms (“translate the name into a numerical address”) |
| cog-verbosity | Rambling, multi-paragraph, tangential answer | Short clipped fragments, no elaboration |
| register | Standard/formal phrasing | Colloquial (“cuppa and the papers”, “the wife’s doing a roast”) |
| Metric | Name | Definition |
|---|---|---|
| M1 | Claim validity | Does each populated claim remain defensible when its value, cited evidence, and free-text description are considered jointly? A field is problematic if its value is indefensible, its evidence is plainly unrelated or non-supporting, or its description adds implausible unsupported specifics. |
| M2 | No over-claiming | Has an attribute been populated without any plausible basis? Clearly unsupported assignments count as problems, while a plausible, appropriately labeled inference does not. |
| M3 | Coverage | Did the extraction omit an important attribute that was clearly available in the source? |
| M4 | Internal consistency | Do extracted fields contradict one another, especially on core identity, role, region, life stage, language, or technology-use information? |
| M5 | Overall fidelity and plausibility | Is the extracted record usable for faithfully role-playing the source person, and are its inferred psychological, interest, and communication traits plausible and coherent? |
| Metric | GPT | Claude | Δ | ≤1 |
|---|---|---|---|---|
| M1 Claim validity | 3.433 | 3.937 | −0.504 | 81.4% |
| M2 No over-claiming | 3.532 | 3.646 | −0.114 | 82.3% |
| M3 Coverage | 4.465 | 4.031 | +0.434 | 99.2% |
| M4 Internal consistency | 3.606 | 3.932 | −0.326 | 85.0% |
| M5 Fidelity and plausibility | 3.829 | 4.109 | −0.280 | 97.8% |
| Overall | 3.773 | 3.931 | −0.158 | 89.1% |
| Metric | Human mean | H–H ≤1 | GPT–H ≤1 | Claude–H ≤1 |
|---|---|---|---|---|
| M1 | 4.105 | 99.1% | 69.0% | 92.0% |
| M2 | 3.770 | 92.2% | 81.0% | 88.0% |
| M3 | 4.223 | 97.1% | 95.0% | 100.0% |
| M4 | 4.537 | 98.6% | 55.0% | 92.0% |
| M5 | 4.040 | 99.1% | 96.0% | 97.0% |
| Overall | 4.135 | 97.2% | 79.2% | 93.8% |
| Agent Bench | GAIA | Web Arena | Web Shop | Mind2 Web | App World | Tool LLM | Bench Flow | τ-bench | Generative Agents | SOTOPIA | OASIS | Silicon sampling | MatrAIx | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Product-as-SUT evaluation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ |
| Persona conditioning | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Behavioral grounding | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✓ |
| Native-device execution | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| Reproducible reporting | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ |
| Trajectory inspection | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| End-to-end validation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ |
| Silicon sampling | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ |
| Reproducible cohorts | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ |
| Multi-agent simulation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ |
研究结果
- 在400次受控试验中,设定的行为特征在91.5%(366/400次)的试验中被正确表达或正确抑制。
- 对1,000个提取出的人格,LLM评审GPT-5.5和Claude Opus 4.8给出的总体平均分分别为3.773和3.931,在5,000个配对指标评分中89.1%的差异不超过一分。
- 六名人类评审员对100个人格子集打分,平均分为4.135(满分5分),3,000个评分中84.7%为4分或5分。
- 与人类平均分相比,GPT的评分在79.2%的情况下差异不超过一分,Claude则在93.8%的情况下差异不超过一分。
- 在相同人格群组、相同任务下更换三种模型运行时,某产品页面的付费方案选择比例在23.2%到93.9%之间大幅波动,88个共有字段间的两两一致性(Cohen's kappa)几乎为零。
可应用场景
- 在正式面向真实用户推出新产品或功能之前,用多样化的模拟用户预先筛查其初步反应,例如对价格上涨的敏感度。
- 检验交互行为,例如用户在AI助手出错后是否愿意继续使用,或能容忍多长的响应延迟。
- 产品更新后,用相同的人格群组和任务重新运行评测,以在一致条件下比较新旧版本。
- 按子群体(年龄、收入、地区等)拆分结果,发现整体聚合分数可能掩盖的问题。
局限与待验证事项
- 作者明确指出,人格群组并非真实人群的概率抽样,评测结果应被视为生成假设的工具,而不是关于真人行为的直接证据。
- 当扮演人格的模型与被测系统共享同一底层模型时,好的结果含义模糊——可能是系统确实表现良好,也可能是模型偏好自己的输出——而目前的实验尚未运行区分这两种情况所需的交叉测试。
- 人格是否能真实展现信息隐瞒、反驳、放弃等对客服类交互至关重要的行为,目前还没有被直接测量,作者将其列为未来验证的优先事项,例如与真实对话日志进行比较。
- 公开发布的核心集(约100万条)只是内部完整人群中经过过滤的一个子集,其合成部分仅按年龄段、地区、性别、城乡四个边际分布进行了校准,并未覆盖全部1,290个维度的联合分布。
- 作者强调,在健康、金融、就业等受监管且后果重大的决策场景中,仅凭模拟结果不足以下结论,必须使用相同任务和工具进行真人验证。
为什么重要
这为团队提供了一种在真人测试之前(或与之并行)低成本、快速地用多样化模拟用户对AI产品和界面进行压力测试的方法。但作者本人也强调,人格代理的结果并不能直接等同于真人行为的证据,更适合被视为生成假设、仍需真人验证的工具。
本文术语
- 人格(Persona) · 由年龄、语言、性格倾向等多种属性定义的一个模拟用户画像
- 依赖图(DAG)采样 · 按照属性之间的先后依赖关系依次生成人格属性,使年龄和教育水平等相关属性保持合理搭配的方法
- 基于真人的记录(human-grounded record) · 从维基百科、评论、调查问卷等真实数据中提取、映射到统一属性体系而构建的人格
- 验证器(verifier) · 自动程序或人工/LLM评审用来确认人格代理的行为结果是否满足任务所要求条件的机制
- 群组(cohort) · 根据特定条件(如年龄段、地区)筛选出来用于某次评测的一组人格
论文原文摘要(英文)
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
- SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit一款把单词意义变化拆解到语法细节的开源分析工具
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)用AI总结股市新闻发现:简单的摘要方法反而比时髦的检索增强技术更靠谱
METAL LAB 最新报道
图片来源: Xiaomin Li et al., arXiv:2608.04205, arxiv-nonexclusive