MatrAIx: Simulating the World with 8.3 Billion Persona Agents
MatrAIx lets you test AI products and apps against 8.3 billion simulated personas instead of real users
MatrAIx is an infrastructure that combines a database of 8.3 billion simulated personas defined by 1,290 attributes, four interactive environments (Survey, AI Chatbot, Web, App) where those personas can act, and 1,010 evaluation tasks, all aimed at making human-style testing of AI systems and digital products faster and cheaper. The authors ran a 400-trial controlled study to check whether persona agents actually follow their assigned traits, and separately had human and LLM judges rate the quality of personas extracted from real sources. The authors are explicit that this remains a hypothesis-generating tool, not a substitute for real human validation.
METAL LAB explanatory visual
The MatrAIx pipeline: from persona generation to reported results
Evidence statusMeasured results and planned work
- Persona 8B8.3 billion simulated personas defined by 1,290 attributes, built via dependency-graph sampling (synthetic) and mapping from real sources (human-grounded); a ~1 million-persona coreset is publicly released
- Cohort selectionEvaluators specify a target audience (e.g., age, income, region), and matching personas are retrieved from Persona 8B to form an evaluation cohort
- Four environmentsSurvey, AI Chatbot, Web, and App environments let persona agents interact with the product under test, recording conversations, actions, and screen states
- Tasks and verifiers1,010 task specifications define goals and success criteria; programmatic verifiers or human/LLM judges check whether persona-agent outputs met those criteria
- Population-level reportIndividual trial results are aggregated by task, cohort, and subgroup into a report showing headline metrics (e.g., projected retention rate) alongside the underlying evidence
What they did
- The motivation is that human evaluation of AI systems and digital products is costly and slow to scale, while offline benchmarks scale well but ignore how different users actually behave and interact.
- Persona 8B is a database of 8.3 billion persona records built on a shared schema of 1,290 categorical attributes (age bracket, language proficiency, risk tolerance, etc.); synthetic personas are sampled from a dependency graph (a directed graph where 'parent' attributes like age influence 'child' attributes like education) that preserves realistic correlations, while human-grounded personas are extracted from sources like Wikipedia biographies, Amazon review histories, and developer surveys into the same schema.
- A public coreset of about 1 million quality-filtered personas (599,847 human-grounded and 400,000 synthetic) was released, alongside the MatrAIx Playground that runs persona agents in four environments—Survey, AI Chatbot, Web, and App—and a library of 1,010 reusable evaluation tasks across more than 25 domains.
- The team ran 18,189 evaluation trials across eight representative tasks using three LLMs (Claude Opus 4.8, GPT 5.5, Claude Haiku 4.5), capturing behaviors like hesitation after a price increase, willingness to keep using an assistant after it fails, and tolerance for response latency.
- A 400-trial controlled study found that assigned behavioral traits were correctly expressed or suppressed in 91.5% (366/400) of trials, and separate human and LLM judges scored the quality of personas extracted from real-world sources.

| Top-level group | Dims. | Representative attributes | Representative grounding |
|---|---|---|---|
| Background | 238 | Age, region, language, education, family, career, industry | Population statistics, household surveys, education and labor taxonomies |
| Psychology | 210 | Personality, values, worldview, motivation, risk | Validated instruments, values surveys, schema design priors |
| Capability | 331 | Domain expertise, general skills, tools, programming, developer context | Occupational taxonomies, technology/developer surveys |
| Behavior and Interaction | 124 | Preferences, habits, interaction state, work practices, technology adoption | Time-use, consumer, workplace, and technology-use evidence |
| Lifestyle | 387 | Interests, media, culture, hobbies, sports, food, health, fitness | Health statistics, consumption surveys, cultural sources |
| Total | 1,290 |
| Source | Released records |
|---|---|
| Wikipedia extraction | 323,438 |
| Amazon Review extraction | 97,915 |
| Stack Overflow survey extraction | 113,120 |
| PRISM Alignment | 1,487 |
| General Social Survey | 63,532 |
| MatrAIx volunteer survey | 355 |
| Human-grounded subtotal | 599,847 |
| Full-DAG synthetic | 400,000 |
| Total | 999,847 |
| Commerce | Software | Finance | Healthcare | Other | Total | |
|---|---|---|---|---|---|---|
| Survey | 202 | 138 | 141 | 139 | 1 | 621 |
| AI Chatbot | 3 | 11 | 17 | 29 | 311 | 371 |
| Web | 2 | 2 | 2 | 0 | 6 | 12 |
| App | 0 | 5 | 1 | 0 | 0 | 6 |
| Total | 207 | 156 | 161 | 168 | 318 | 1,010 |
| Group | Subgroup | Schema category | Count | Representative attributes |
|---|---|---|---|---|
| Background | Demographics | Demographic: Core | 25 | Age bracket; region; gender identity |
| Background | Demographics | Demographic: Cultural | 2 | Cultural background; attitude toward immigration |
| Background | Demographics | Demographic: Family | 1 | Household size |
| Background | Demographics | Demographic: Life Events | 24 | Life stage; major life events; childhood environment |
| Background | Language | Linguistic: Language | 53 | Primary language; English proficiency; multilingualism |
| Background | Language | Linguistic: Communication | 37 | Expected tone; verbosity; communication preferences |
| Background | Education | Learning: Academic | 34 | Highest education; academic field; institution tier |
| Background | Education | Learning: Style | 1 | Learning style |
| Background | Career | Professional: Career | 4 | Research output; seniority; years of experience |
| Background | Career | Professional: Industry | 51 | Company size; role function; industry |
| Background | Career | Developer: Professional Context | 6 | Professional status; role archetype; contribution context |
| Psychology | Personality | Personality: Character | 34 | Domain stance; dominant trait; curiosity |
| Psychology | Personality | Personality: Big Five | 50 | Imagination; artistic interest; emotionality |
| Psychology | Personality | Personality: MBTI | 2 | Neurotype; Myers-Briggs type |
| Psychology | Personality | Personality: Relationships | 4 | Attachment anxiety; attachment avoidance; interpersonal agency |
| Psychology | Worldview | Values & Motivation | 46 | Core value; religiosity; economic motivation |
| Psychology | Worldview | Worldview: Beliefs | 67 | Political leaning; trust level; safety sensitivity |
| Psychology | Decision-Making | Risk & Decision | 7 | Risk tolerance; decision style; need for closure |
| Capability | Domains | Expertise: Domains | 144 | Domain; subject specialty; technology savviness |
| Capability | Skills | Expertise: Skills | 64 | Writing; copywriting; editing |
| Capability | Skills | Skills: Tools | 69 | Excel; Google Sheets; Python |
| Capability | Skills | Skills: Programming | 44 | Comment style; summary documentation; naming verbosity |
| Capability | Skills | Developer: Code Maintenance | 10 | Complexity tolerance; modularity preference; type-system orientation |
| Behavior and Interaction | Personal Behavior | Behavior: Preferences | 34 | Modality preference; accessibility needs; media diet |
| Behavior and Interaction | Personal Behavior | Behavior: Habits | 30 | Journaling; meditation; use of to-do lists |
| Behavior and Interaction | Personal Behavior | Behavior: Time | 3 | Time pressure; sleep schedule; micromanagement aversion |
| Behavior and Interaction | Interaction State | State: Emotional | 5 | Emotional state; intent; query complexity |
| Behavior and Interaction | Work Practices | Behavior: Work | 2 | Work schedule; office versus remote work |
| Behavior and Interaction | Work Practices | Developer: Open Source Behavior | 7 | Open-source activity; GitHub contribution mode; pull-request style |
| Behavior and Interaction | Work Practices | Developer: Community Behavior | 4 | Stack Overflow use; participation style; help-seeking preference |
| Schema group | Facets | Grounding roles | Sources |
|---|---|---|---|
| Background | Demographics (52); language (90); education (35); career (61) | Category definitions; population priors; household, language, education, and labor dependencies | UN World Population Prospects and Population Data [87, 86]; World Bank WDI and WorldPop [94, 97]; Eurostat, ACS PUMS, and IPUMS [21, 82, 38]; DHS, UNICEF MICS, and OECD Family Database [77, 85, 62]; Pew and World Values Survey [69, 96]; UNESCO UIS, World Bank Education Statistics, OECD INES, and OECD PISA [84, 93, 61, 64]; ILOSTAT, BLS OEWS, and O*NET [34, 81, 60]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]. |
| Psychology | Personality (90); worldview (113); decision-making (7) | Instrument and value-set design; selected prevalence estimates; validation | IPIP and MIDUS [35, 56]; Pew and World Values Survey [69, 96]; GSS, European Social Survey, ISSP, and Gallup World Poll [58, 20, 36, 23]; Afrobarometer, Arab Barometer, Asian Barometer, Latinobarometro, and Eurobarometer [1, 3, 5, 46, 19]; ARDA [6]. |
| Capability | Domain expertise (144); general skills (64); tools (69); programming (44); developer context (10) | Occupational and skill taxonomies; technology access and adoption; developer-tool prevalence | ITU Statistics, World Bank WDI, DataReportal, and Pew Internet [37, 94, 15, 70]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]; O*NET [60]. |
| Behavior and Interaction | Personal behavior (67); interaction state (5); work practices (13); technology use (39) | Time-use and consumer priors; workplace behavior; technology and AI adoption | American Time Use Survey, Consumer Expenditure Surveys, and OECD Time Use [79, 80, 63]; ITU, DataReportal, and Pew Internet [37, 15, 70]; Stack Overflow Survey, GitHub Octoverse, JetBrains Developer Ecosystem, and O*NET [75, 25, 39, 60]. |
| Lifestyle | Interests (358); physical health (25); fitness (2); health lifestyle (2) | Health and disability priors; consumption and time use; cultural and interest category design | WHO GHO and IHME GBD [95, 33]; ACS PUMS, NHIS, BRFSS, DHS, and UNICEF MICS [82, 11, 10, 77, 85]; ATUS, CEX, and OECD Time Use [79, 80, 63]; FAOSTAT and UNESCO Culture Statistics [22, 83]; Eurobarometer, Pew, DataReportal, WVS, and Gallup World Poll [19, 69, 70, 15, 96, 23]. |
| v | πi(v) | qlang | qregion | ri(v) | mi(v) | πirimi | pθ(v∣xPa(i)) |
|---|---|---|---|---|---|---|---|
| None | 0.30 | 0.02 | 0.10 | 0.022 | 0 | 0.000 | 0.000 |
| Basic | 0.30 | 0.08 | 0.20 | 0.178 | 1 | 0.053 | 0.028 |
| Fluent | 0.25 | 0.30 | 0.35 | 1.680 | 1 | 0.420 | 0.224 |
| Native | 0.15 | 0.60 | 0.35 | 9.333 | 1 | 1.400 | 0.747 |
| sum | 1.00 | 1.00 | 1.00 | 1.873 | 1.000 |

| Stage | Rejected | Remaining |
|---|---|---|
| Original corpus | – | 10,002,288,277 |
| Contradiction filter | 239,310 | 10,002,048,967 |
| Human exact/MinHash deduplication | 41,597 | 2,222,496 human |
| Synthetic projection deduplication | 252,936,392 | 9,746,848,482 synthetic |
| Synthetic deterministic cutoff | 1,349,070,978 | 8,397,777,504 synthetic |
| Audited baseline | 8,400,000,000 |

| Dimension | Answered | Share of those answering |
|---|---|---|
| Age bracket | 321 | 25–34 22.1%, 55–64 19.0%, 35–44 15.9%, 18–24 13.7%, 45–54 13.4%, 65–74 8.4%, 75–84 4.0%, 85+ 3.4% |
| Gender identity | 322 | Woman 47.5%, Man 42.9%, Non-binary 3.7%, Self-described 3.4%, Prefer not to say 2.5% |
| Region | 329 | South Asia 20.4%, Sub-Saharan Africa 19.1%, East Asia 13.4%, Southeast Asia 10.9%, Latin America 9.7%, Western Europe 7.0%, MENA 6.7%, North America 5.8%, Eastern Europe 4.6%, Oceania 2.4% |
| Urbanicity | 328 | Rural 33.2%, Dense urban 22.3%, Small town 21.0%, Suburban 20.4%, Nomadic / remote 3.0% |
| Socioeconomic band | 340 | Low 31.8%, Lower-middle 30.0%, Middle 20.3%, Upper-middle 12.6%, High 5.3% |
| Employment | 330 | Full-time 25.8%, Retired 16.1%, Homemaker 14.2%, Self-employed 11.8%, Part-time 9.1%, Student 8.2%, Unemployed 7.6%, Gig / freelance 7.3% |

| Collection | Tasks | Environment | Composition |
|---|---|---|---|
| Synthetic persona surveys | 405 | Survey | 135 Finance, 135 Healthcare, and 135 Software questionnaires, generally sharing a common survey and reporting template. |
| Product surveys | 200 | Survey | Twenty product-research archetypes, including purchase intent, price sensitivity, recommendation, and retention, crossed with ten retail products. |
| Synthetic chatbots | 351 | AI Chatbot | Scenario families across 27 domains, including travel, legal, healthcare, telecommunications, real estate, insurance, finance, and education. |

| Contract component | Declared content | Audit purpose |
|---|---|---|
| Task metadata | Stable name and version, environment type, product domain, difficulty, tags, artifact paths, and runtime budgets | Identifies the recipe and prevents results from silently moving between task versions. |
| Persona-facing scenario | Context, user goal, constraints, disclosure policy, and required submission | Defines what every sampled persona is asked to do without exposing verifier internals. |
| Cohort strategy | Persona-schema filters, sampling mode, stratification fields, sample size, weights, and seed policy | Makes the target audience explicit and permits the cohort to be redrawn. |
| Product attachment | Questionnaire or stimulus, chat endpoint or sidecar, website target, or native application backend | Identifies the system under test and how the runtime reaches it. |
| Verifier | Required artifacts, structured finding schema, objective checks, timeouts, and failure conditions | Converts a trial into reproducible outcomes with supporting evidence. |
| Reporting policy | Aggregations, subgroup facets, summaries, optional judge directives, and disclosure rules | Defines how trial findings become a cohort-level report. |

| Environment | Task-owned inputs | Primary trial artifacts |
|---|---|---|
| Survey | Stimulus, questionnaire schema, response constraints, and optional rationale prompts | Typed responses, missing or invalid items, rationales, confidence, and completion summary. |
| AI Chatbot | Product endpoint or sidecar, user goal, staged-disclosure policy, turn and safety limits | Full transcript, tool or service events, resolution state, termination reason, and post-run feedback. |
| Web | URL or hosted site, browser backend, exploration requirements, and submission schema | Page and action trace, considered options, screenshots where enabled, final submission, and terminal page or site state. |
| App | Native target, platform backend, initial state, permissions, and terminal-state checks | Screenshot and action trace, exported files, application state, permission changes, and cross-application side effects. |
| Task | Env. | Primary outcome | Share giving that answer | Range | q | ||
|---|---|---|---|---|---|---|---|
| GPT | Opus | Haiku | (pt) | ||||
| Annual checkup | Survey | q31 = e | 50.5% | 41.4% | 29.7% | 20.8 | <0.001∗∗∗ |
| Candy Land price | Survey | q_threshold = hesitate | 98.3% | 27.0% | 83.3% | 71.3 | <0.001∗∗∗ |
| OpenBB honesty | Chat | wouldStillContinueUse = unsure | 81.3% | 85.5% | 28.1% | 57.4 | <0.001∗∗∗ |
| Meal planning† | Chat | adherenceLikelihood = 7 | 50.6% | 0.2% | 40.0% | 50.4 | <0.001∗∗∗ |
| Notion plans | Web | decision_subject_id = plus | 63.5% | 21.5% | 75.7% | 54.2 | <0.001∗∗∗ |
| MIT OCW course | Web | task_course_level = Graduate | 40.0% | 34.5% | 47.5% | 13.0 | <0.001∗∗∗ |
| News+ subscription | App | clicked_get_started = true | 4.2% | 20.8% | 0.0% | 20.8 | 0.025∗ |
| Stocks sentiment | App | sentiment = hold | 60.0% | 30.0% | 50.0% | 30.0 | 0.153 ns |
| Task | Product-level measure | GPT 5.5 | Opus 4.8 | Haiku 4.5 |
|---|---|---|---|---|
| Annual checkup | Very likely to schedule in time | 50.5% (505/1,000) | 41.4% (414/1,000) | 29.7% (297/1,000) |
| Candy Land price | Hesitates or worse at new price | 98.3% (983/1,000) | 27.0% (270/1,000) | 83.3% (833/1,000) |
| OpenBB honesty | Would not continue using | 18.5% (185/1,000) | 14.5% (132/909) | 71.9% (719/1,000) |
| Meal planning† | Stated need fully satisfied | 46.2% (462/1,000) | 0.2% (2/1,000) | 28.1% (281/1,000) |
| Notion plans | Selected a paid plan | 75.8% (776/1,024) | 23.2% (237/1,022) | 93.9% (958/1,020) |
| MIT OCW course | Chose a graduate-level course | 40.0% (403/1,008) | 34.5% (347/1,007) | 47.5% (473/996) |
| News+ subscription | Subscribed | 4.2% (1/24) | 20.8% (5/24) | 0.0% (0/24) |
| Stocks sentiment | Buy opinion | 40.0% (8/20) | 70.0% (14/20) | 47.4% (9/19) |
| Task | Persona dimension | Groups | Sig. | Spearman ρ, exact p | ||
|---|---|---|---|---|---|---|
| GPT/Opus | GPT/Haiku | Opus/Haiku | ||||
| Annual checkup | age bracket | 6 | ∘∘∙ | +0.26 p=0.329 | −0.43 p=0.822 | −0.77 p=0.971 |
| Candy Land price | economic motivation | 4 | ∘∘∘ | −0.40 p=0.792 | −0.20 p=0.625 | +0.80 p=0.167 |
| OpenBB honesty | trust level | 4 | ∙∙∙ | +1.00 p=0.042 | +1.00 p=0.042 | +1.00 p=0.042 |
| Meal planning† | life stage | 4 | ∘∘∘ | +0.95 p=0.083 | +0.32 p=0.500 | +0.33 p=0.417 |
| Notion plans | company size | 8 | ∘∘∘ | +0.44 p=0.138 | +0.60 p=0.059 | −0.24 p=0.725 |
| MIT OCW course | academic field | 8 | ∘∘∘ | +0.93 p=0.001 | +0.09 p=0.420 | +0.06 p=0.452 |
| News+ subscr.‡ | economic motivation | 3 | ∘∘∘ | +0.00 p=0.667 | flat | flat |
| Stocks sentiment‡ | risk tolerance | 5 | ∘∘∘ | +0.47 p=0.267 | −0.92 p=1.000 | −0.67 p=0.933 |
| GPT × Opus | GPT × Haiku | Opus × Haiku | |
|---|---|---|---|
| Paired agreement over 88 joinable fields | |||
| Median Cohen’s κ | 0.000 | 0.000 | +0.001 |
| Fields at κ≤0 | 59 of 88 | 50 of 88 | 40 of 88 |
| Fields reaching κ≥0.2 | 7 of 88 | 8 of 88 | 7 of 88 |
| Fields at ≥50% agreement, κ<0.1 | 48 of 56 | 25 of 34 | 24 of 34 |
| Self-report fidelity: age band matches the persona | |||
| GPT 5.5 16.3% (ns) Opus 4.8 16.9% (ns) Haiku 4.5 100.0% | chance 16.7% |
| Survey | Chat | Web | OS-App | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Attribute | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | Opus 4.8 | GPT-5.6-sol | |||
| code-comment-style | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| code-naming-verbosity | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| code-summary-documentation | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-emoji-use | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-humor | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-politeness | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-storytelling | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-use-of-jargon | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| cog-verbosity | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | |||
| register | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 | 🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹 |
| Attribute | Positive (declared) | Negative (opposite) |
|---|---|---|
| code-comment-style | Nearly every line carries an inline comment (# Iterate over every integer, # Check divisibility) | Code contains no comments whatsoever |
| code-naming-verbosity | calculate_average_of_passing_scores, number_of_passing_scores | Variables s, t, a, c, x, all single-letter |
| code-summary-documentation | Every function opens with a tldr: docstring | Prose plus code, no TLDR/summary header |
| cog-emoji-use | Many emoji across a short paragraph | No emoji despite a casual app-store prompt |
| cog-humor | Playful, witty asides (“the boxes may yet win”) | Measured, earnest tone, no jokes |
| cog-politeness | “Might I kindly ask that you resend the document at your earliest convenience” | “Hey, you forgot the attachment. Again.”: blunt and sarcastic |
| cog-storytelling | A narrated scene (“a storm came through, half the lodge went dark”) | Abstract, value-driven, no concrete scene |
| cog-use-of-jargon | Dense technical jargon (“TCP three-way handshake, SYN-ACK”) | Plain terms (“translate the name into a numerical address”) |
| cog-verbosity | Rambling, multi-paragraph, tangential answer | Short clipped fragments, no elaboration |
| register | Standard/formal phrasing | Colloquial (“cuppa and the papers”, “the wife’s doing a roast”) |
| Metric | Name | Definition |
|---|---|---|
| M1 | Claim validity | Does each populated claim remain defensible when its value, cited evidence, and free-text description are considered jointly? A field is problematic if its value is indefensible, its evidence is plainly unrelated or non-supporting, or its description adds implausible unsupported specifics. |
| M2 | No over-claiming | Has an attribute been populated without any plausible basis? Clearly unsupported assignments count as problems, while a plausible, appropriately labeled inference does not. |
| M3 | Coverage | Did the extraction omit an important attribute that was clearly available in the source? |
| M4 | Internal consistency | Do extracted fields contradict one another, especially on core identity, role, region, life stage, language, or technology-use information? |
| M5 | Overall fidelity and plausibility | Is the extracted record usable for faithfully role-playing the source person, and are its inferred psychological, interest, and communication traits plausible and coherent? |
| Metric | GPT | Claude | Δ | ≤1 |
|---|---|---|---|---|
| M1 Claim validity | 3.433 | 3.937 | −0.504 | 81.4% |
| M2 No over-claiming | 3.532 | 3.646 | −0.114 | 82.3% |
| M3 Coverage | 4.465 | 4.031 | +0.434 | 99.2% |
| M4 Internal consistency | 3.606 | 3.932 | −0.326 | 85.0% |
| M5 Fidelity and plausibility | 3.829 | 4.109 | −0.280 | 97.8% |
| Overall | 3.773 | 3.931 | −0.158 | 89.1% |
| Metric | Human mean | H–H ≤1 | GPT–H ≤1 | Claude–H ≤1 |
|---|---|---|---|---|
| M1 | 4.105 | 99.1% | 69.0% | 92.0% |
| M2 | 3.770 | 92.2% | 81.0% | 88.0% |
| M3 | 4.223 | 97.1% | 95.0% | 100.0% |
| M4 | 4.537 | 98.6% | 55.0% | 92.0% |
| M5 | 4.040 | 99.1% | 96.0% | 97.0% |
| Overall | 4.135 | 97.2% | 79.2% | 93.8% |
| Agent Bench | GAIA | Web Arena | Web Shop | Mind2 Web | App World | Tool LLM | Bench Flow | τ-bench | Generative Agents | SOTOPIA | OASIS | Silicon sampling | MatrAIx | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Product-as-SUT evaluation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ |
| Persona conditioning | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Behavioral grounding | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✓ |
| Native-device execution | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| Reproducible reporting | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ |
| Trajectory inspection | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| End-to-end validation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ |
| Silicon sampling | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ |
| Reproducible cohorts | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ |
| Multi-agent simulation | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ |
Findings
- In the 400-trial controlled study, the assigned behavior was expressed or correctly suppressed in 91.5% of trials (366/400).
- Across 1,000 extracted personas, LLM judges GPT-5.5 and Claude Opus 4.8 gave overall mean scores of 3.773 and 3.931 respectively, and 89.1% of 5,000 paired metric scores differed by at most one point.
- Six human raters scoring a 100-persona subset gave a mean of 4.135 out of 5, with 84.7% of 3,000 scores rated 4 or 5.
- Compared to the human mean, GPT scores were within one point in 79.2% of cases and Claude scores in 93.8% of cases.
- When the same cohort was run through three different models on identical tasks, one product's paid-plan selection share ranged from 23.2% to 93.9%, and pairwise agreement (Cohen's kappa) across 88 shared fields was close to zero.
Where it can be used
- Screening how diverse user groups might react to a new product or feature (e.g., a price increase) before running costly human studies.
- Checking interactive behaviors such as whether users keep using a chatbot after it fails, or how much response delay they tolerate.
- Re-running the same persona cohort and tasks after a product update to compare versions under consistent conditions.
- Breaking down results by subgroup (age, income, region, etc.) to spot problems that overall aggregate scores might hide.
Limits and open work
- The authors explicitly state that persona cohorts are not probability samples of real populations, and that results should be treated as hypothesis-generating rather than direct evidence about real human behavior.
- When the model playing the persona shares the same backbone as the system being evaluated, favorable results become ambiguous (self-preference bias vs. genuine satisfaction), and the current experiments did not run the crossing test needed to separate this effect.
- Whether personas realistically show behaviors like withholding information, pushing back, or abandoning a task—key for realistic support-style interactions—has not been directly measured yet; the authors list comparison against real conversation logs as a priority for future work.
- The public coreset (about 1 million personas) is a filtered subset of the full internal population, and its synthetic portion is calibrated to only four demographic marginals (age, region, gender, urbanicity), not the full 1,290-dimensional joint distribution.
- The authors caution that for consequential decisions in regulated areas like health, finance, or employment, simulated results are not sufficient and require human validation using the same task and instrument.
Why it matters
This gives teams a way to cheaply and quickly stress-test AI products and interfaces against a wide, diverse range of simulated user backgrounds before—or alongside—expensive human studies. But the authors themselves caution that persona-agent results are not direct evidence of real human behavior, so this is best read as a way to generate hypotheses that still need human confirmation.
Terms in this paper
- Persona · A simulated user profile defined by attributes such as age, language, and personality traits.
- DAG (dependency graph) sampling · Generating persona attributes one at a time in an order where earlier ('parent') attributes shape later ('child') ones, so correlated traits like age and education level stay realistic.
- Human-grounded record · A persona built from real-world data such as Wikipedia biographies, reviews, or survey responses, mapped into the same attribute schema.
- Verifier · An automated check or human/LLM judge that confirms whether a persona agent's output met a task's required conditions.
- Cohort · A selected group of personas (e.g., matching certain age or region criteria) used for a specific evaluation run.
Original abstract (English)
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI assistants would rather double-check facts than ask you a question, even when asking is the right call
- SynFlow: A Multidimensional Diachronic Semantic Analysis ToolkitAn open-source tool that breaks down how a word's meaning changed, not just that it changed
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)Testing AI summaries of stock news, the simple approach beat the trendy retrieval-based one
Latest from METAL LAB
- Sakana AI Signs Deal With Japan's Defense Ministry for Intelligence Analysis AI Trial
- Hermes Agent builds its own skills the more you use it
- Is training AI on copyrighted books legal? Courts are still fighting it out
- Chinese gray market sells Anthropic Claude tokens at 10% of list price
- Even the Best AI Runaway Response Plan Among Five Major Labs Scores Only 3
Figures: Xiaomin Li et al., arXiv:2608.04205, arxiv-nonexclusive