工作日早上 7 点读 AI,周日早上 8 点读周报订阅邮件

METAL LAB

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

arXiv:2608.042052026-08-03

MatrAIx用83亿个模拟人格代替真人来测试AI产品

MatrAIx是一套基础设施,包含一个由1,290个属性定义、拥有83亿条记录的模拟人格数据库,四种可让这些人格实际行动的交互环境(问卷、AI聊天、网页、应用),以及1,010个评测任务,目的是让对AI系统和数字产品的类人评测变得更快更便宜。作者进行了一项400次试验的受控研究,检验人格代理是否真的会按照设定的性格行事,还请人类评审员和LLM评审员分别为提取出的人格质量打分。作者明确指出这仍是一个用于生成假设的工具,不能替代真人验证。

METAL LAB 解读图

MatrAIx流程:从人格生成到结果报告

证据状态实测结果与计划中的工作并存

  1. Persona 8B由1,290个属性定义的83亿条模拟人格,通过依赖图采样(合成)和真实来源映射(基于真人)构建,并发布了约100万条的公开核心集
  2. 群组筛选评测者指定目标受众(如年龄、收入、地区),系统从Persona 8B中检索匹配的人格组成评测群组
  3. 四种交互环境问卷(Survey)、AI聊天(AI Chatbot)、网页(Web)、应用(App)环境让人格代理与被测产品互动,记录对话、操作和界面状态
  4. 任务与验证器1,010个任务规范定义目标和成功标准,由程序化验证器或人类/LLM评审判断人格代理的输出是否满足条件
  5. 群体级报告将各次试验结果按任务、群组、子群体汇总成报告,展示关键指标(如预计留存率)及支撑证据
这是 METAL LAB 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 研究动机是人类评测AI系统和数字产品成本高、速度慢、难以规模化,而离线基准测试虽然可扩展,却往往忽略了用户之间的多样性和真实互动行为。
  2. Persona 8B是一个基于1,290个分类属性(年龄段、语言水平、风险偏好等)统一模式构建的83亿条人格记录数据库;合成人格通过一个依赖图(一种有向图,让年龄等父属性影响教育水平等子属性)采样以保留真实的属性关联,而基于真人的人格则从维基百科传记、亚马逊评论历史、开发者调查等来源提取并映射到同一模式中。
  3. 团队发布了一个经质量过滤的约100万条人格核心集(59.9847万条基于真人、40万条合成),并提供MatrAIx Playground运行环境,支持问卷(Survey)、AI聊天(AI Chatbot)、网页(Web)、应用(App)四种环境,以及覆盖25个以上领域的1,010个可复用评测任务库。
  4. 团队使用三种大语言模型(Claude Opus 4.8、GPT 5.5、Claude Haiku 4.5)在八个代表性任务上共运行了18,189次评测试验,捕捉了涨价后的购买犹豫、AI助手失败后用户是否愿意继续使用、以及对响应延迟的容忍度等因人而异的行为模式。
  5. 一项400次试验的受控研究发现,设定的行为特征在91.5%(366/400次)的试验中被正确表达或正确抑制;另外,人类评审员和LLM评审员分别对从真实来源提取的人格质量进行了打分。
Figure 1: Overview of the MatrAIx simulated-user evaluation framework. Evaluators describe the target audience, and MatrAIx retrieves matching personas from Persona 8B to form an evaluation cohort. The four environments define how these persona agents interact, while MatrAIx Applications provides the task specifications they execute. Shared telemetry and task-owned verification preserve the evidence needed for subgroup and population-level reporting.
Figure 1: Overview of the MatrAIx simulated-user evaluation framework. Evaluators describe the target audience, and MatrAIx retrieves matching personas from Persona 8B to form an evaluation cohort. The four environments define how these persona agents interact, while MatrAIx Applications provides the task specifications they execute. Shared telemetry and task-owned verification preserve the evidence needed for subgroup and population-level reporting.
Table 1: Persona 8B schema overview. Dimension counts refer to emitted categorical attributes. Sources provide different kinds and strengths of grounding; they do not imply direct population estimates for every value.
Top-level groupDims.Representative attributesRepresentative grounding
Background238Age, region, language, education, family, career, industryPopulation statistics, household surveys, education and labor taxonomies
Psychology210Personality, values, worldview, motivation, riskValidated instruments, values surveys, schema design priors
Capability331Domain expertise, general skills, tools, programming, developer contextOccupational taxonomies, technology/developer surveys
Behavior and Interaction124Preferences, habits, interaction state, work practices, technology adoptionTime-use, consumer, workplace, and technology-use evidence
Lifestyle387Interests, media, culture, hobbies, sports, food, health, fitnessHealth statistics, consumption surveys, cultural sources
Total1,290
Figure 3: Controlled behavioral adherence across four environments. (a) Each cell pools five personas declaring one pole of an attribute and five declaring the opposite pole; 10/10 means that all five positive personas expressed the target and all five negative personas expressed the opposite value. (b) Overall successful-trial share by environment and the number of strong attributes, defined as at least four of five successes in both arms. An LLM judge evaluates the recorded trajectory or artifact. Overall, 366 of 400 trials succeed.
Figure 3: Controlled behavioral adherence across four environments. (a) Each cell pools five personas declaring one pole of an attribute and five declaring the opposite pole; 10/10 means that all five positive personas expressed the target and all five negative personas expressed the opposite value. (b) Overall successful-trial share by environment and the number of strong attributes, defined as at least four of five successes in both arms. An LLM judge evaluates the recorded trajectory or artifact. Overall, 366 of 400 trials succeed.
Table 2: Composition of the public Persona 1M coreset. “Human-grounded” identifies the origin of a record, not a guarantee that every extracted field is a verified fact. These counts mirror the Composition table of the dataset card (footnote * ‣ 1), which is the authoritative record for the release.
SourceReleased records
Wikipedia extraction323,438
Amazon Review extraction97,915
Stack Overflow survey extraction113,120
PRISM Alignment1,487
General Social Survey63,532
MatrAIx volunteer survey355
Human-grounded subtotal599,847
Full-DAG synthetic400,000
Total999,847
(b) Environment summary.
(b) Environment summary.
Table 3: Application-task coverage. Counts are unique specifications on the repository’s main branch together with the batch collections on its synthetic-task branches; individually contributed tasks still under review on open pull requests are not counted. “Other” aggregates more than 25 additional domains.
CommerceSoftwareFinanceHealthcareOtherTotal
Survey2021381411391621
AI Chatbot3111729311371
Web2220612
App051006
Total2071561611683181,010
Figure 4: Complete three-layer taxonomy of the Persona 8B schema. The left column gives the five top-level groups and their attribute totals; the middle column gives 16 subgroups; and the right column gives all 55 conceptual categories with counts. The hierarchy covers exactly 1,290 attributes and excludes latent or helper variables used only by the generation graph.
Figure 4: Complete three-layer taxonomy of the Persona 8B schema. The left column gives the five top-level groups and their attribute totals; the middle column gives 16 subgroups; and the right column gives all 55 conceptual categories with counts. The hierarchy covers exactly 1,290 attributes and excludes latent or helper variables used only by the generation graph.
Table 4: Complete index of the 43 schema categories. Counts sum to 1,290 attributes. Examples are labels from the authoritative dimension catalog; the complete attribute-level mapping is supplied as supp_material/persona_taxonomy_mapping.csv.
GroupSubgroupSchema categoryCountRepresentative attributes
BackgroundDemographicsDemographic: Core25Age bracket; region; gender identity
BackgroundDemographicsDemographic: Cultural2Cultural background; attitude toward immigration
BackgroundDemographicsDemographic: Family1Household size
BackgroundDemographicsDemographic: Life Events24Life stage; major life events; childhood environment
BackgroundLanguageLinguistic: Language53Primary language; English proficiency; multilingualism
BackgroundLanguageLinguistic: Communication37Expected tone; verbosity; communication preferences
BackgroundEducationLearning: Academic34Highest education; academic field; institution tier
BackgroundEducationLearning: Style1Learning style
BackgroundCareerProfessional: Career4Research output; seniority; years of experience
BackgroundCareerProfessional: Industry51Company size; role function; industry
BackgroundCareerDeveloper: Professional Context6Professional status; role archetype; contribution context
PsychologyPersonalityPersonality: Character34Domain stance; dominant trait; curiosity
PsychologyPersonalityPersonality: Big Five50Imagination; artistic interest; emotionality
PsychologyPersonalityPersonality: MBTI2Neurotype; Myers-Briggs type
PsychologyPersonalityPersonality: Relationships4Attachment anxiety; attachment avoidance; interpersonal agency
PsychologyWorldviewValues & Motivation46Core value; religiosity; economic motivation
PsychologyWorldviewWorldview: Beliefs67Political leaning; trust level; safety sensitivity
PsychologyDecision-MakingRisk & Decision7Risk tolerance; decision style; need for closure
CapabilityDomainsExpertise: Domains144Domain; subject specialty; technology savviness
CapabilitySkillsExpertise: Skills64Writing; copywriting; editing
CapabilitySkillsSkills: Tools69Excel; Google Sheets; Python
CapabilitySkillsSkills: Programming44Comment style; summary documentation; naming verbosity
CapabilitySkillsDeveloper: Code Maintenance10Complexity tolerance; modularity preference; type-system orientation
Behavior and InteractionPersonal BehaviorBehavior: Preferences34Modality preference; accessibility needs; media diet
Behavior and InteractionPersonal BehaviorBehavior: Habits30Journaling; meditation; use of to-do lists
Behavior and InteractionPersonal BehaviorBehavior: Time3Time pressure; sleep schedule; micromanagement aversion
Behavior and InteractionInteraction StateState: Emotional5Emotional state; intent; query complexity
Behavior and InteractionWork PracticesBehavior: Work2Work schedule; office versus remote work
Behavior and InteractionWork PracticesDeveloper: Open Source Behavior7Open-source activity; GitHub contribution mode; pull-request style
Behavior and InteractionWork PracticesDeveloper: Community Behavior4Stack Overflow use; participation style; help-seeking preference
Figure 5: The full persona DAG. All 1,308 graph nodes and 6,999 directed edges. Nodes are placed left to right by the topological order used in Equation 5 and grouped vertically into one lane per schema category, with lanes ordered by their mean topological position; marker area scales with a node’s directed degree. Edges are drawn as translucent curves, so darker bands mark the dependency bundles that connect demographic and educational roots to downstream expertise, interest, and behavior dimensions.
Figure 5: The full persona DAG. All 1,308 graph nodes and 6,999 directed edges. Nodes are placed left to right by the topological order used in Equation 5 and grouped vertically into one lane per schema category, with lanes ordered by their mean topological position; marker area scales with a node’s directed degree. Edges are drawn as translucent curves, so darker bands mark the dependency bundles that connect demographic and educational roots to downstream expertise, interest, and behavior dimensions.
Table 5: Detailed schema-to-source grounding map. The mapping records source families and their roles in schema design, prior estimation, dependency construction, compatibility rules, or downstream validation.
Schema groupFacetsGrounding rolesSources
BackgroundDemographics (52); language (90); education (35); career (61)Category definitions; population priors; household, language, education, and labor dependenciesUN World Population Prospects and Population Data [87, 86]; World Bank WDI and WorldPop [94, 97]; Eurostat, ACS PUMS, and IPUMS [21, 82, 38]; DHS, UNICEF MICS, and OECD Family Database [77, 85, 62]; Pew and World Values Survey [69, 96]; UNESCO UIS, World Bank Education Statistics, OECD INES, and OECD PISA [84, 93, 61, 64]; ILOSTAT, BLS OEWS, and O*NET [34, 81, 60]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39].
PsychologyPersonality (90); worldview (113); decision-making (7)Instrument and value-set design; selected prevalence estimates; validationIPIP and MIDUS [35, 56]; Pew and World Values Survey [69, 96]; GSS, European Social Survey, ISSP, and Gallup World Poll [58, 20, 36, 23]; Afrobarometer, Arab Barometer, Asian Barometer, Latinobarometro, and Eurobarometer [1, 3, 5, 46, 19]; ARDA [6].
CapabilityDomain expertise (144); general skills (64); tools (69); programming (44); developer context (10)Occupational and skill taxonomies; technology access and adoption; developer-tool prevalenceITU Statistics, World Bank WDI, DataReportal, and Pew Internet [37, 94, 15, 70]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]; O*NET [60].
Behavior and InteractionPersonal behavior (67); interaction state (5); work practices (13); technology use (39)Time-use and consumer priors; workplace behavior; technology and AI adoptionAmerican Time Use Survey, Consumer Expenditure Surveys, and OECD Time Use [79, 80, 63]; ITU, DataReportal, and Pew Internet [37, 15, 70]; Stack Overflow Survey, GitHub Octoverse, JetBrains Developer Ecosystem, and O*NET [75, 25, 39, 60].
LifestyleInterests (358); physical health (25); fitness (2); health lifestyle (2)Health and disability priors; consumption and time use; cultural and interest category designWHO GHO and IHME GBD [95, 33]; ACS PUMS, NHIS, BRFSS, DHS, and UNICEF MICS [82, 11, 10, 77, 85]; ATUS, CEX, and OECD Time Use [79, 80, 63]; FAOSTAT and UNESCO Culture Statistics [22, 83]; Eurobarometer, Pew, DataReportal, WVS, and Gallup World Poll [19, 69, 70, 15, 96, 23].
Figure 6: How much of the 1,290-item instrument the volunteer subset exercises. Each vertical bar is the median coverage of one interface category’s items, sorted, for all 43 categories. Labels carry the category’s leaf name, dropping the group prefix (Interests, Behavior, and so on) that would otherwise repeat under every bar. The vertical axis starts at 88.5% rather than zero because the finding is how narrow the range is, from 89.3% (Time, under Behavior) to 92.4% (Agent Adoption, under Developer; both extremes are named on their bars), which a zero-based axis would render as 43 nearly identical bars. The dashed line marks the 91.3% category median. The narrow range indicates that missingness is not concentrated in a small set of categories. Separately, 272 of the 355 records answer every one of the 1,290 items and the remaining 83 answer between 520 and 1,186.
Figure 6: How much of the 1,290-item instrument the volunteer subset exercises. Each vertical bar is the median coverage of one interface category’s items, sorted, for all 43 categories. Labels carry the category’s leaf name, dropping the group prefix (Interests, Behavior, and so on) that would otherwise repeat under every bar. The vertical axis starts at 88.5% rather than zero because the finding is how narrow the range is, from 89.3% (Time, under Behavior) to 92.4% (Agent Adoption, under Developer; both extremes are named on their bars), which a zero-based axis would render as 43 nearly identical bars. The dashed line marks the 91.3% category median. The narrow range indicates that missingness is not concentrated in a small set of categories. Separately, 272 of the 355 records answer every one of the 1,290 items and the remaining 83 answer between 520 and 1,186.
Table 6: One child dimension through the local CPD (illustrative). english_proficiency conditioned on primary_language=English and region=North America. The prior alone would give Native a 15% share; the two dependency factors raise it to 75%, and the compatibility mask removes None outright rather than merely making it unlikely. Numbers illustrate Equation 4 and are not the deployed parameter values.
vπi​(v)qlangqregionri​(v)mi​(v)πi​ri​mipθ​(v∣xPa⁡(i))
None0.300.020.100.02200.0000.000
Basic0.300.080.200.17810.0530.028
Fluent0.250.300.351.68010.4200.224
Native0.150.600.359.33311.4000.747
sum1.001.001.001.8731.000
Figure 8: The MatrAIx Playground. A user browses the persona population 8(a), inspects and filters individual persona records 8(b), configures a study over a persona cohort 8(c), runs an interactive evaluation of the system under test 8(d), and reads the aggregated population-level report 8(e), all without writing code.
Figure 8: The MatrAIx Playground. A user browses the persona population 8(a), inspects and filters individual persona records 8(b), configures a study over a persona cohort 8(c), runs an interactive evaluation of the system under test 8(d), and reads the aggregated population-level report 8(e), all without writing code.
Table 7: Detailed post-processing accounting for the audited baseline. Human and synthetic remaining counts have different scopes until the final row.
StageRejectedRemaining
Original corpus10,002,288,277
Contradiction filter239,31010,002,048,967
Human exact/MinHash deduplication41,5972,222,496 human
Synthetic projection deduplication252,936,3929,746,848,482 synthetic
Synthetic deterministic cutoff1,349,070,9788,397,777,504 synthetic
Audited baseline8,400,000,000
(b) Persona browser. Individual records are shown as cards carrying their grounded attributes (age, career, region, and further schema dimensions), and can be filtered to assemble an evaluation cohort.
(b) Persona browser. Individual records are shown as cards carrying their grounded attributes (age, career, region, and further schema dimensions), and can be filtered to assemble an evaluation cohort.
Table 8: Declared composition of the 355 released volunteer records. Shares are of the records answering each dimension, so the denominator differs by row. Computed from the public release, so the table matches what a reader downloading the dataset obtains.
DimensionAnsweredShare of those answering
Age bracket32125–34 22.1%, 55–64 19.0%, 35–44 15.9%, 18–24 13.7%, 45–54 13.4%, 65–74 8.4%, 75–84 4.0%, 85+ 3.4%
Gender identity322Woman 47.5%, Man 42.9%, Non-binary 3.7%, Self-described 3.4%, Prefer not to say 2.5%
Region329South Asia 20.4%, Sub-Saharan Africa 19.1%, East Asia 13.4%, Southeast Asia 10.9%, Latin America 9.7%, Western Europe 7.0%, MENA 6.7%, North America 5.8%, Eastern Europe 4.6%, Oceania 2.4%
Urbanicity328Rural 33.2%, Dense urban 22.3%, Small town 21.0%, Suburban 20.4%, Nomadic / remote 3.0%
Socioeconomic band340Low 31.8%, Lower-middle 30.0%, Middle 20.3%, Upper-middle 12.6%, High 5.3%
Employment330Full-time 25.8%, Retired 16.1%, Homemaker 14.2%, Self-employed 11.8%, Part-time 9.1%, Student 8.2%, Unemployed 7.6%, Gig / freelance 7.3%
(c) Study setup. A study is specified as a decision to test (here, a $2 price increase), with the persona cohort on the left and the task catalog on the right, before launching the simulation.
(c) Study setup. A study is specified as a decision to test (here, a $2 price increase), with the persona cohort on the left and the task catalog on the right, before launching the simulation.
Table 9: Large template-based task collections. Members of a collection are separate task specifications but reuse a common contract and verifier pattern.
CollectionTasksEnvironmentComposition
Synthetic persona surveys405Survey135 Finance, 135 Healthcare, and 135 Software questionnaires, generally sharing a common survey and reporting template.
Product surveys200SurveyTwenty product-research archetypes, including purchase intent, price sensitivity, recommendation, and retention, crossed with ten retail products.
Synthetic chatbots351AI ChatbotScenario families across 27 domains, including travel, legal, healthcare, telecommunications, real estate, insurance, finance, and education.
(d) Interactive evaluation. A persona-conditioned agent converses with the system under test in a live transcript, scored in real time against a task-owned rubric (here, an AI counselor).
(d) Interactive evaluation. A persona-conditioned agent converses with the system under test in a live transcript, scored in real time against a task-owned rubric (here, an AI counselor).
Table 10: Logical components of an application-task contract.
Contract componentDeclared contentAudit purpose
Task metadataStable name and version, environment type, product domain, difficulty, tags, artifact paths, and runtime budgetsIdentifies the recipe and prevents results from silently moving between task versions.
Persona-facing scenarioContext, user goal, constraints, disclosure policy, and required submissionDefines what every sampled persona is asked to do without exposing verifier internals.
Cohort strategyPersona-schema filters, sampling mode, stratification fields, sample size, weights, and seed policyMakes the target audience explicit and permits the cohort to be redrawn.
Product attachmentQuestionnaire or stimulus, chat endpoint or sidecar, website target, or native application backendIdentifies the system under test and how the runtime reaches it.
VerifierRequired artifacts, structured finding schema, objective checks, timeouts, and failure conditionsConverts a trial into reproducible outcomes with supporting evidence.
Reporting policyAggregations, subgroup facets, summaries, optional judge directives, and disclosure rulesDefines how trial findings become a cohort-level report.
(e) Population report. Trials aggregate into a population-level outcome (here, 61% projected retention) with response distributions, decision reasons, and the per-persona response-level dataset backing the headline number.
(e) Population report. Trials aggregate into a population-level outcome (here, 61% projected retention) with response distributions, decision reasons, and the per-persona response-level dataset backing the headline number.
Table 11: Environment-specific inputs and artifacts. App covers desktop and mobile native applications, including Linux, macOS, and iOS.
EnvironmentTask-owned inputsPrimary trial artifacts
SurveyStimulus, questionnaire schema, response constraints, and optional rationale promptsTyped responses, missing or invalid items, rationales, confidence, and completion summary.
AI ChatbotProduct endpoint or sidecar, user goal, staged-disclosure policy, turn and safety limitsFull transcript, tool or service events, resolution state, termination reason, and post-run feedback.
WebURL or hosted site, browser backend, exploration requirements, and submission schemaPage and action trace, considered options, screenshots where enabled, final submission, and terminal page or site state.
AppNative target, platform backend, initial state, permissions, and terminal-state checksScreenshot and action trace, exported files, application state, permission changes, and cross-application side effects.
Figure 9: Move-transition graphs for the meal-planning chat study, by economic motivation. 1,000 GPT-5.5 persona conversations, 500 per stratum, one fixed assistant; 13,120 move transitions in total. A conversation averages six exchanges and fourteen moves, so it traverses the graph in loops rather than one pass; edges curving back up the page are those returns. Each panel draws only its own moves and transitions and lays itself out, so shape differences reflect the data. Black edges trace each step’s most frequent successor; every drawn edge carries its transition count, and a pair of moves with no edge between them co-occurred fewer than the panel’s count floor of about 90 transitions.
Figure 9: Move-transition graphs for the meal-planning chat study, by economic motivation. 1,000 GPT-5.5 persona conversations, 500 per stratum, one fixed assistant; 13,120 move transitions in total. A conversation averages six exchanges and fourteen moves, so it traverses the graph in loops rather than one pass; edges curving back up the page are those returns. Each panel draws only its own moves and transitions and lays itself out, so shape differences reflect the data. Black edges trace each step’s most frequent successor; every drawn edge carries its transition count, and a pair of moves with no edge between them co-occurred fewer than the panel’s count floor of about 90 transitions.
Table 12: Complete primary-outcome results for the eight validation tasks. Range is the largest minus smallest model share in percentage points. The q column reports a three-arm χ2 test after Benjamini–Hochberg correction across the eight tasks (∗∗∗q<0.001, ∗q<0.05; ns otherwise).
TaskEnv.Primary outcomeShare giving that answerRangeq
GPTOpusHaiku(pt)
Annual checkupSurveyq31 = e50.5%41.4%29.7%20.8<0.001∗∗∗
Candy Land priceSurveyq_threshold = hesitate98.3%27.0%83.3%71.3<0.001∗∗∗
OpenBB honestyChatwouldStillContinueUse = unsure81.3%85.5%28.1%57.4<0.001∗∗∗
Meal planning†ChatadherenceLikelihood = 750.6%0.2%40.0%50.4<0.001∗∗∗
Notion plansWebdecision_subject_id = plus63.5%21.5%75.7%54.2<0.001∗∗∗
MIT OCW courseWebtask_course_level = Graduate40.0%34.5%47.5%13.0<0.001∗∗∗
News+ subscriptionAppclicked_get_started = true4.2%20.8%0.0%20.80.025∗
Stocks sentimentAppsentiment = hold60.0%30.0%50.0%30.00.153 ns
Table 13: Product-level conclusions under three agent models on identical cohorts. Denominators are trials in which the field was answered. Notion aggregates its three paid tiers against the free tier. †The meal-planning GPT 5.5 arm did not run the declared cohort and is not comparable with the other two arms.
TaskProduct-level measureGPT 5.5Opus 4.8Haiku 4.5
Annual checkupVery likely to schedule in time50.5% (505/1,000)41.4% (414/1,000)29.7% (297/1,000)
Candy Land priceHesitates or worse at new price98.3% (983/1,000)27.0% (270/1,000)83.3% (833/1,000)
OpenBB honestyWould not continue using18.5% (185/1,000)14.5% (132/909)71.9% (719/1,000)
Meal planning†Stated need fully satisfied46.2% (462/1,000)0.2% (2/1,000)28.1% (281/1,000)
Notion plansSelected a paid plan75.8% (776/1,024)23.2% (237/1,022)93.9% (958/1,020)
MIT OCW courseChose a graduate-level course40.0% (403/1,008)34.5% (347/1,007)47.5% (473/996)
News+ subscriptionSubscribed4.2% (1/24)20.8% (5/24)0.0% (0/24)
Stocks sentimentBuy opinion40.0% (8/20)70.0% (14/20)47.4% (9/19)
Table 14: Rank correlation between model arms’ orderings of the same persona subgroups. Significance marks are ordered GPT / Opus / Haiku (∙ significant after correction, ∘ not). A flat arm has no ordering to compare. †The GPT arm has a cohort-integrity exception. ‡App rows contain only three to eight personas per subgroup.
TaskPersona dimensionGroupsSig.Spearman ρ, exact p
GPT/OpusGPT/HaikuOpus/Haiku
Annual checkupage bracket6∘∘∙+0.26 p=0.329−0.43 p=0.822−0.77 p=0.971
Candy Land priceeconomic motivation4∘∘∘−0.40 p=0.792−0.20 p=0.625+0.80 p=0.167
OpenBB honestytrust level4∙∙∙+1.00 p=0.042+1.00 p=0.042+1.00 p=0.042
Meal planning†life stage4∘∘∘+0.95 p=0.083+0.32 p=0.500+0.33 p=0.417
Notion planscompany size8∘∘∘+0.44 p=0.138+0.60 p=0.059−0.24 p=0.725
MIT OCW courseacademic field8∘∘∘+0.93 p=0.001+0.09 p=0.420+0.06 p=0.452
News+ subscr.‡economic motivation3∘∘∘+0.00 p=0.667flatflat
Stocks sentiment‡risk tolerance5∘∘∘+0.47 p=0.267−0.92 p=1.000−0.67 p=0.933
Table 15: Persona fidelity across all three model pairs. Cohen’s κ corrects raw agreement for each model’s answer distribution. GPT and Opus age-band matches are indistinguishable from uniform guessing (q=0.84 and q=0.85); Haiku matches 1,000 of 1,000 trials.
GPT × OpusGPT × HaikuOpus × Haiku
Paired agreement over 88 joinable fields
Median Cohen’s κ0.0000.000+0.001
Fields at κ≤059 of 8850 of 8840 of 88
Fields reaching κ≥0.27 of 888 of 887 of 88
Fields at ≥50% agreement, κ<0.148 of 5625 of 3424 of 34
Self-report fidelity: age band matches the persona
GPT 5.5 16.3% (ns) Opus 4.8 16.9% (ns) Haiku 4.5 100.0%chance 16.7%
Table 16: Persona adherence per attribute and environment, both acting agents. Each environment is split into two sub-columns: Opus 4.8 (left) and GPT-5.6-sol (right), under the same Opus 4.8 judge. Each cell shows two groups of five persona icons—the left group the positive cohort, the right group the negative cohort. A filled icon (🚹) marks a persona whose behavior expressed the declared value (for the negative cohort, correctly expressed the opposite value—target suppressed); a faint icon (🚹) marks one that did not. More filled icons is better in both groups. Overall Opus 366/400=91.5% (33/40 cells strong, ≥4 filled per group) vs. GPT-5.6-sol 317/400=79.2%.
SurveyChatWebOS-App
AttributeOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-sol
code-comment-style🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
code-naming-verbosity🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
code-summary-documentation🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-emoji-use🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-humor🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-politeness🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-storytelling🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-use-of-jargon🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-verbosity🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
register🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
Table 17: Cited judge evidence, Survey environment. For each attribute, the behavior the judge identified in a positive persona (declared value) versus its matched negative persona (opposite value). Excerpts are quoted from the trial trajectories.
AttributePositive (declared)Negative (opposite)
code-comment-styleNearly every line carries an inline comment (# Iterate over every integer, # Check divisibility)Code contains no comments whatsoever
code-naming-verbositycalculate_average_of_passing_scores, number_of_passing_scoresVariables s, t, a, c, x, all single-letter
code-summary-documentationEvery function opens with a tldr: docstringProse plus code, no TLDR/summary header
cog-emoji-useMany emoji across a short paragraphNo emoji despite a casual app-store prompt
cog-humorPlayful, witty asides (“the boxes may yet win”)Measured, earnest tone, no jokes
cog-politeness“Might I kindly ask that you resend the document at your earliest convenience”“Hey, you forgot the attachment. Again.”: blunt and sarcastic
cog-storytellingA narrated scene (“a storm came through, half the lodge went dark”)Abstract, value-driven, no concrete scene
cog-use-of-jargonDense technical jargon (“TCP three-way handshake, SYN-ACK”)Plain terms (“translate the name into a numerical address”)
cog-verbosityRambling, multi-paragraph, tangential answerShort clipped fragments, no elaboration
registerStandard/formal phrasingColloquial (“cuppa and the papers”, “the wife’s doing a roast”)
Table 18: Extraction-quality metrics. Higher scores indicate better extraction quality.
MetricNameDefinition
M1Claim validityDoes each populated claim remain defensible when its value, cited evidence, and free-text description are considered jointly? A field is problematic if its value is indefensible, its evidence is plainly unrelated or non-supporting, or its description adds implausible unsupported specifics.
M2No over-claimingHas an attribute been populated without any plausible basis? Clearly unsupported assignments count as problems, while a plausible, appropriately labeled inference does not.
M3CoverageDid the extraction omit an important attribute that was clearly available in the source?
M4Internal consistencyDo extracted fields contradict one another, especially on core identity, role, region, life stage, language, or technology-use information?
M5Overall fidelity and plausibilityIs the extracted record usable for faithfully role-playing the source person, and are its inferred psychological, interest, and communication traits plausible and coherent?
Table 19: Paired LLM-judge results on all 1,000 extractions. Δ is GPT minus Claude; ≤1 is the proportion of paired scores that differ by at most one point.
MetricGPTClaudeΔ≤1
M1 Claim validity3.4333.937−0.50481.4%
M2 No over-claiming3.5323.646−0.11482.3%
M3 Coverage4.4654.031+0.43499.2%
M4 Internal consistency3.6063.932−0.32685.0%
M5 Fidelity and plausibility3.8294.109−0.28097.8%
Overall3.7733.931−0.15889.1%
Table 20: Human evaluation and within-one agreement on the 100-persona subset. H–H compares all 15 pairs of human raters; GPT–H and Claude–H compare each LLM judge with the six-rater human mean.
MetricHuman meanH–H ≤1GPT–H ≤1Claude–H ≤1
M14.10599.1%69.0%92.0%
M23.77092.2%81.0%88.0%
M34.22397.1%95.0%100.0%
M44.53798.6%55.0%92.0%
M54.04099.1%96.0%97.0%
Overall4.13597.2%79.2%93.8%
Table 21: Capability comparison across representative agent benchmarks, simulation systems, and synthetic-sampling methods. Systems, left to right: AgentBench [52], GAIA [55], WebArena [111], WebShop [102], Mind2Web [17], AppWorld [78], ToolLLM [72], BenchFlow [7], τ-bench [103], Generative Agents [66], SOTOPIA [113], OASIS [101], and Silicon sampling [4]. A check (✓) indicates explicit support in the published system; a cross (✗) indicates the capability is not stated. Dimensions read as follows: product-as-SUT evaluation treats the product, not the agent, as the subject under test; persona conditioning instantiates a user from a sampled population profile; behavioral grounding verifies that behavior tracks a target attribute rather than confounders; native-device execution covers real mobile/desktop apps beyond API demos; reproducible reporting separates verifier facts from reporting policy; trajectory inspection exposes per-trial traces; end-to-end validation closes the persona–task–product loop; silicon sampling substitutes simulated respondents for survey/UX pre-studies; reproducible cohorts deterministically re-instantiate the same cohort; and multi-agent simulation denotes persistent multi-agent social worlds (a deliberate non-goal for MatrAIx). The matrix compares capability coverage rather than providing an overall ranking.
Agent BenchGAIAWeb ArenaWeb ShopMind2 WebApp WorldTool LLMBench Flowτ-benchGenerative AgentsSOTOPIAOASISSilicon samplingMatrAIx
Product-as-SUT evaluation
Persona conditioning
Behavioral grounding
Native-device execution
Reproducible reporting
Trajectory inspection
End-to-end validation
Silicon sampling
Reproducible cohorts
Multi-agent simulation

研究结果

  • 在400次受控试验中,设定的行为特征在91.5%(366/400次)的试验中被正确表达或正确抑制。
  • 对1,000个提取出的人格,LLM评审GPT-5.5和Claude Opus 4.8给出的总体平均分分别为3.773和3.931,在5,000个配对指标评分中89.1%的差异不超过一分。
  • 六名人类评审员对100个人格子集打分,平均分为4.135(满分5分),3,000个评分中84.7%为4分或5分。
  • 与人类平均分相比,GPT的评分在79.2%的情况下差异不超过一分,Claude则在93.8%的情况下差异不超过一分。
  • 在相同人格群组、相同任务下更换三种模型运行时,某产品页面的付费方案选择比例在23.2%到93.9%之间大幅波动,88个共有字段间的两两一致性(Cohen's kappa)几乎为零。

可应用场景

  • 在正式面向真实用户推出新产品或功能之前,用多样化的模拟用户预先筛查其初步反应,例如对价格上涨的敏感度。
  • 检验交互行为,例如用户在AI助手出错后是否愿意继续使用,或能容忍多长的响应延迟。
  • 产品更新后,用相同的人格群组和任务重新运行评测,以在一致条件下比较新旧版本。
  • 按子群体(年龄、收入、地区等)拆分结果,发现整体聚合分数可能掩盖的问题。

局限与待验证事项

  • 作者明确指出,人格群组并非真实人群的概率抽样,评测结果应被视为生成假设的工具,而不是关于真人行为的直接证据。
  • 当扮演人格的模型与被测系统共享同一底层模型时,好的结果含义模糊——可能是系统确实表现良好,也可能是模型偏好自己的输出——而目前的实验尚未运行区分这两种情况所需的交叉测试。
  • 人格是否能真实展现信息隐瞒、反驳、放弃等对客服类交互至关重要的行为,目前还没有被直接测量,作者将其列为未来验证的优先事项,例如与真实对话日志进行比较。
  • 公开发布的核心集(约100万条)只是内部完整人群中经过过滤的一个子集,其合成部分仅按年龄段、地区、性别、城乡四个边际分布进行了校准,并未覆盖全部1,290个维度的联合分布。
  • 作者强调,在健康、金融、就业等受监管且后果重大的决策场景中,仅凭模拟结果不足以下结论,必须使用相同任务和工具进行真人验证。

为什么重要

这为团队提供了一种在真人测试之前(或与之并行)低成本、快速地用多样化模拟用户对AI产品和界面进行压力测试的方法。但作者本人也强调,人格代理的结果并不能直接等同于真人行为的证据,更适合被视为生成假设、仍需真人验证的工具。

本文术语

  • 人格(Persona) · 由年龄、语言、性格倾向等多种属性定义的一个模拟用户画像
  • 依赖图(DAG)采样 · 按照属性之间的先后依赖关系依次生成人格属性,使年龄和教育水平等相关属性保持合理搭配的方法
  • 基于真人的记录(human-grounded record) · 从维基百科、评论、调查问卷等真实数据中提取、映射到统一属性体系而构建的人格
  • 验证器(verifier) · 自动程序或人工/LLM评审用来确认人格代理的行为结果是否满足任务所要求条件的机制
  • 群组(cohort) · 根据特定条件(如年龄段、地区)筛选出来用于某次评测的一组人格

论文原文摘要(英文)

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

作者 · Xiaomin Li

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Xiaomin Li et al., arXiv:2608.04205, arxiv-nonexclusive