AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox

METAL LAB

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

arXiv:2608.042052026-08-03

MatrAIx lets you test AI products and apps against 8.3 billion simulated personas instead of real users

MatrAIx is an infrastructure that combines a database of 8.3 billion simulated personas defined by 1,290 attributes, four interactive environments (Survey, AI Chatbot, Web, App) where those personas can act, and 1,010 evaluation tasks, all aimed at making human-style testing of AI systems and digital products faster and cheaper. The authors ran a 400-trial controlled study to check whether persona agents actually follow their assigned traits, and separately had human and LLM judges rate the quality of personas extracted from real sources. The authors are explicit that this remains a hypothesis-generating tool, not a substitute for real human validation.

METAL LAB explanatory visual

The MatrAIx pipeline: from persona generation to reported results

Evidence statusMeasured results and planned work

  1. Persona 8B8.3 billion simulated personas defined by 1,290 attributes, built via dependency-graph sampling (synthetic) and mapping from real sources (human-grounded); a ~1 million-persona coreset is publicly released
  2. Cohort selectionEvaluators specify a target audience (e.g., age, income, region), and matching personas are retrieved from Persona 8B to form an evaluation cohort
  3. Four environmentsSurvey, AI Chatbot, Web, and App environments let persona agents interact with the product under test, recording conversations, actions, and screen states
  4. Tasks and verifiers1,010 task specifications define goals and success criteria; programmatic verifiers or human/LLM judges check whether persona-agent outputs met those criteria
  5. Population-level reportIndividual trial results are aggregated by task, cohort, and subgroup into a report showing headline metrics (e.g., projected retention rate) alongside the underlying evidence
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.

What they did

  1. The motivation is that human evaluation of AI systems and digital products is costly and slow to scale, while offline benchmarks scale well but ignore how different users actually behave and interact.
  2. Persona 8B is a database of 8.3 billion persona records built on a shared schema of 1,290 categorical attributes (age bracket, language proficiency, risk tolerance, etc.); synthetic personas are sampled from a dependency graph (a directed graph where 'parent' attributes like age influence 'child' attributes like education) that preserves realistic correlations, while human-grounded personas are extracted from sources like Wikipedia biographies, Amazon review histories, and developer surveys into the same schema.
  3. A public coreset of about 1 million quality-filtered personas (599,847 human-grounded and 400,000 synthetic) was released, alongside the MatrAIx Playground that runs persona agents in four environments—Survey, AI Chatbot, Web, and App—and a library of 1,010 reusable evaluation tasks across more than 25 domains.
  4. The team ran 18,189 evaluation trials across eight representative tasks using three LLMs (Claude Opus 4.8, GPT 5.5, Claude Haiku 4.5), capturing behaviors like hesitation after a price increase, willingness to keep using an assistant after it fails, and tolerance for response latency.
  5. A 400-trial controlled study found that assigned behavioral traits were correctly expressed or suppressed in 91.5% (366/400) of trials, and separate human and LLM judges scored the quality of personas extracted from real-world sources.
Figure 1: Overview of the MatrAIx simulated-user evaluation framework. Evaluators describe the target audience, and MatrAIx retrieves matching personas from Persona 8B to form an evaluation cohort. The four environments define how these persona agents interact, while MatrAIx Applications provides the task specifications they execute. Shared telemetry and task-owned verification preserve the evidence needed for subgroup and population-level reporting.
Figure 1: Overview of the MatrAIx simulated-user evaluation framework. Evaluators describe the target audience, and MatrAIx retrieves matching personas from Persona 8B to form an evaluation cohort. The four environments define how these persona agents interact, while MatrAIx Applications provides the task specifications they execute. Shared telemetry and task-owned verification preserve the evidence needed for subgroup and population-level reporting.
Table 1: Persona 8B schema overview. Dimension counts refer to emitted categorical attributes. Sources provide different kinds and strengths of grounding; they do not imply direct population estimates for every value.
Top-level groupDims.Representative attributesRepresentative grounding
Background238Age, region, language, education, family, career, industryPopulation statistics, household surveys, education and labor taxonomies
Psychology210Personality, values, worldview, motivation, riskValidated instruments, values surveys, schema design priors
Capability331Domain expertise, general skills, tools, programming, developer contextOccupational taxonomies, technology/developer surveys
Behavior and Interaction124Preferences, habits, interaction state, work practices, technology adoptionTime-use, consumer, workplace, and technology-use evidence
Lifestyle387Interests, media, culture, hobbies, sports, food, health, fitnessHealth statistics, consumption surveys, cultural sources
Total1,290
Figure 3: Controlled behavioral adherence across four environments. (a) Each cell pools five personas declaring one pole of an attribute and five declaring the opposite pole; 10/10 means that all five positive personas expressed the target and all five negative personas expressed the opposite value. (b) Overall successful-trial share by environment and the number of strong attributes, defined as at least four of five successes in both arms. An LLM judge evaluates the recorded trajectory or artifact. Overall, 366 of 400 trials succeed.
Figure 3: Controlled behavioral adherence across four environments. (a) Each cell pools five personas declaring one pole of an attribute and five declaring the opposite pole; 10/10 means that all five positive personas expressed the target and all five negative personas expressed the opposite value. (b) Overall successful-trial share by environment and the number of strong attributes, defined as at least four of five successes in both arms. An LLM judge evaluates the recorded trajectory or artifact. Overall, 366 of 400 trials succeed.
Table 2: Composition of the public Persona 1M coreset. “Human-grounded” identifies the origin of a record, not a guarantee that every extracted field is a verified fact. These counts mirror the Composition table of the dataset card (footnote * ‣ 1), which is the authoritative record for the release.
SourceReleased records
Wikipedia extraction323,438
Amazon Review extraction97,915
Stack Overflow survey extraction113,120
PRISM Alignment1,487
General Social Survey63,532
MatrAIx volunteer survey355
Human-grounded subtotal599,847
Full-DAG synthetic400,000
Total999,847
(b) Environment summary.
(b) Environment summary.
Table 3: Application-task coverage. Counts are unique specifications on the repository’s main branch together with the batch collections on its synthetic-task branches; individually contributed tasks still under review on open pull requests are not counted. “Other” aggregates more than 25 additional domains.
CommerceSoftwareFinanceHealthcareOtherTotal
Survey2021381411391621
AI Chatbot3111729311371
Web2220612
App051006
Total2071561611683181,010
Figure 4: Complete three-layer taxonomy of the Persona 8B schema. The left column gives the five top-level groups and their attribute totals; the middle column gives 16 subgroups; and the right column gives all 55 conceptual categories with counts. The hierarchy covers exactly 1,290 attributes and excludes latent or helper variables used only by the generation graph.
Figure 4: Complete three-layer taxonomy of the Persona 8B schema. The left column gives the five top-level groups and their attribute totals; the middle column gives 16 subgroups; and the right column gives all 55 conceptual categories with counts. The hierarchy covers exactly 1,290 attributes and excludes latent or helper variables used only by the generation graph.
Table 4: Complete index of the 43 schema categories. Counts sum to 1,290 attributes. Examples are labels from the authoritative dimension catalog; the complete attribute-level mapping is supplied as supp_material/persona_taxonomy_mapping.csv.
GroupSubgroupSchema categoryCountRepresentative attributes
BackgroundDemographicsDemographic: Core25Age bracket; region; gender identity
BackgroundDemographicsDemographic: Cultural2Cultural background; attitude toward immigration
BackgroundDemographicsDemographic: Family1Household size
BackgroundDemographicsDemographic: Life Events24Life stage; major life events; childhood environment
BackgroundLanguageLinguistic: Language53Primary language; English proficiency; multilingualism
BackgroundLanguageLinguistic: Communication37Expected tone; verbosity; communication preferences
BackgroundEducationLearning: Academic34Highest education; academic field; institution tier
BackgroundEducationLearning: Style1Learning style
BackgroundCareerProfessional: Career4Research output; seniority; years of experience
BackgroundCareerProfessional: Industry51Company size; role function; industry
BackgroundCareerDeveloper: Professional Context6Professional status; role archetype; contribution context
PsychologyPersonalityPersonality: Character34Domain stance; dominant trait; curiosity
PsychologyPersonalityPersonality: Big Five50Imagination; artistic interest; emotionality
PsychologyPersonalityPersonality: MBTI2Neurotype; Myers-Briggs type
PsychologyPersonalityPersonality: Relationships4Attachment anxiety; attachment avoidance; interpersonal agency
PsychologyWorldviewValues & Motivation46Core value; religiosity; economic motivation
PsychologyWorldviewWorldview: Beliefs67Political leaning; trust level; safety sensitivity
PsychologyDecision-MakingRisk & Decision7Risk tolerance; decision style; need for closure
CapabilityDomainsExpertise: Domains144Domain; subject specialty; technology savviness
CapabilitySkillsExpertise: Skills64Writing; copywriting; editing
CapabilitySkillsSkills: Tools69Excel; Google Sheets; Python
CapabilitySkillsSkills: Programming44Comment style; summary documentation; naming verbosity
CapabilitySkillsDeveloper: Code Maintenance10Complexity tolerance; modularity preference; type-system orientation
Behavior and InteractionPersonal BehaviorBehavior: Preferences34Modality preference; accessibility needs; media diet
Behavior and InteractionPersonal BehaviorBehavior: Habits30Journaling; meditation; use of to-do lists
Behavior and InteractionPersonal BehaviorBehavior: Time3Time pressure; sleep schedule; micromanagement aversion
Behavior and InteractionInteraction StateState: Emotional5Emotional state; intent; query complexity
Behavior and InteractionWork PracticesBehavior: Work2Work schedule; office versus remote work
Behavior and InteractionWork PracticesDeveloper: Open Source Behavior7Open-source activity; GitHub contribution mode; pull-request style
Behavior and InteractionWork PracticesDeveloper: Community Behavior4Stack Overflow use; participation style; help-seeking preference
Figure 5: The full persona DAG. All 1,308 graph nodes and 6,999 directed edges. Nodes are placed left to right by the topological order used in Equation 5 and grouped vertically into one lane per schema category, with lanes ordered by their mean topological position; marker area scales with a node’s directed degree. Edges are drawn as translucent curves, so darker bands mark the dependency bundles that connect demographic and educational roots to downstream expertise, interest, and behavior dimensions.
Figure 5: The full persona DAG. All 1,308 graph nodes and 6,999 directed edges. Nodes are placed left to right by the topological order used in Equation 5 and grouped vertically into one lane per schema category, with lanes ordered by their mean topological position; marker area scales with a node’s directed degree. Edges are drawn as translucent curves, so darker bands mark the dependency bundles that connect demographic and educational roots to downstream expertise, interest, and behavior dimensions.
Table 5: Detailed schema-to-source grounding map. The mapping records source families and their roles in schema design, prior estimation, dependency construction, compatibility rules, or downstream validation.
Schema groupFacetsGrounding rolesSources
BackgroundDemographics (52); language (90); education (35); career (61)Category definitions; population priors; household, language, education, and labor dependenciesUN World Population Prospects and Population Data [87, 86]; World Bank WDI and WorldPop [94, 97]; Eurostat, ACS PUMS, and IPUMS [21, 82, 38]; DHS, UNICEF MICS, and OECD Family Database [77, 85, 62]; Pew and World Values Survey [69, 96]; UNESCO UIS, World Bank Education Statistics, OECD INES, and OECD PISA [84, 93, 61, 64]; ILOSTAT, BLS OEWS, and O*NET [34, 81, 60]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39].
PsychologyPersonality (90); worldview (113); decision-making (7)Instrument and value-set design; selected prevalence estimates; validationIPIP and MIDUS [35, 56]; Pew and World Values Survey [69, 96]; GSS, European Social Survey, ISSP, and Gallup World Poll [58, 20, 36, 23]; Afrobarometer, Arab Barometer, Asian Barometer, Latinobarometro, and Eurobarometer [1, 3, 5, 46, 19]; ARDA [6].
CapabilityDomain expertise (144); general skills (64); tools (69); programming (44); developer context (10)Occupational and skill taxonomies; technology access and adoption; developer-tool prevalenceITU Statistics, World Bank WDI, DataReportal, and Pew Internet [37, 94, 15, 70]; Stack Overflow Survey, GitHub Octoverse, and JetBrains Developer Ecosystem [75, 25, 39]; O*NET [60].
Behavior and InteractionPersonal behavior (67); interaction state (5); work practices (13); technology use (39)Time-use and consumer priors; workplace behavior; technology and AI adoptionAmerican Time Use Survey, Consumer Expenditure Surveys, and OECD Time Use [79, 80, 63]; ITU, DataReportal, and Pew Internet [37, 15, 70]; Stack Overflow Survey, GitHub Octoverse, JetBrains Developer Ecosystem, and O*NET [75, 25, 39, 60].
LifestyleInterests (358); physical health (25); fitness (2); health lifestyle (2)Health and disability priors; consumption and time use; cultural and interest category designWHO GHO and IHME GBD [95, 33]; ACS PUMS, NHIS, BRFSS, DHS, and UNICEF MICS [82, 11, 10, 77, 85]; ATUS, CEX, and OECD Time Use [79, 80, 63]; FAOSTAT and UNESCO Culture Statistics [22, 83]; Eurobarometer, Pew, DataReportal, WVS, and Gallup World Poll [19, 69, 70, 15, 96, 23].
Figure 6: How much of the 1,290-item instrument the volunteer subset exercises. Each vertical bar is the median coverage of one interface category’s items, sorted, for all 43 categories. Labels carry the category’s leaf name, dropping the group prefix (Interests, Behavior, and so on) that would otherwise repeat under every bar. The vertical axis starts at 88.5% rather than zero because the finding is how narrow the range is, from 89.3% (Time, under Behavior) to 92.4% (Agent Adoption, under Developer; both extremes are named on their bars), which a zero-based axis would render as 43 nearly identical bars. The dashed line marks the 91.3% category median. The narrow range indicates that missingness is not concentrated in a small set of categories. Separately, 272 of the 355 records answer every one of the 1,290 items and the remaining 83 answer between 520 and 1,186.
Figure 6: How much of the 1,290-item instrument the volunteer subset exercises. Each vertical bar is the median coverage of one interface category’s items, sorted, for all 43 categories. Labels carry the category’s leaf name, dropping the group prefix (Interests, Behavior, and so on) that would otherwise repeat under every bar. The vertical axis starts at 88.5% rather than zero because the finding is how narrow the range is, from 89.3% (Time, under Behavior) to 92.4% (Agent Adoption, under Developer; both extremes are named on their bars), which a zero-based axis would render as 43 nearly identical bars. The dashed line marks the 91.3% category median. The narrow range indicates that missingness is not concentrated in a small set of categories. Separately, 272 of the 355 records answer every one of the 1,290 items and the remaining 83 answer between 520 and 1,186.
Table 6: One child dimension through the local CPD (illustrative). english_proficiency conditioned on primary_language=English and region=North America. The prior alone would give Native a 15% share; the two dependency factors raise it to 75%, and the compatibility mask removes None outright rather than merely making it unlikely. Numbers illustrate Equation 4 and are not the deployed parameter values.
vπi​(v)qlangqregionri​(v)mi​(v)πi​ri​mipθ​(v∣xPa⁡(i))
None0.300.020.100.02200.0000.000
Basic0.300.080.200.17810.0530.028
Fluent0.250.300.351.68010.4200.224
Native0.150.600.359.33311.4000.747
sum1.001.001.001.8731.000
Figure 8: The MatrAIx Playground. A user browses the persona population 8(a), inspects and filters individual persona records 8(b), configures a study over a persona cohort 8(c), runs an interactive evaluation of the system under test 8(d), and reads the aggregated population-level report 8(e), all without writing code.
Figure 8: The MatrAIx Playground. A user browses the persona population 8(a), inspects and filters individual persona records 8(b), configures a study over a persona cohort 8(c), runs an interactive evaluation of the system under test 8(d), and reads the aggregated population-level report 8(e), all without writing code.
Table 7: Detailed post-processing accounting for the audited baseline. Human and synthetic remaining counts have different scopes until the final row.
StageRejectedRemaining
Original corpus10,002,288,277
Contradiction filter239,31010,002,048,967
Human exact/MinHash deduplication41,5972,222,496 human
Synthetic projection deduplication252,936,3929,746,848,482 synthetic
Synthetic deterministic cutoff1,349,070,9788,397,777,504 synthetic
Audited baseline8,400,000,000
(b) Persona browser. Individual records are shown as cards carrying their grounded attributes (age, career, region, and further schema dimensions), and can be filtered to assemble an evaluation cohort.
(b) Persona browser. Individual records are shown as cards carrying their grounded attributes (age, career, region, and further schema dimensions), and can be filtered to assemble an evaluation cohort.
Table 8: Declared composition of the 355 released volunteer records. Shares are of the records answering each dimension, so the denominator differs by row. Computed from the public release, so the table matches what a reader downloading the dataset obtains.
DimensionAnsweredShare of those answering
Age bracket32125–34 22.1%, 55–64 19.0%, 35–44 15.9%, 18–24 13.7%, 45–54 13.4%, 65–74 8.4%, 75–84 4.0%, 85+ 3.4%
Gender identity322Woman 47.5%, Man 42.9%, Non-binary 3.7%, Self-described 3.4%, Prefer not to say 2.5%
Region329South Asia 20.4%, Sub-Saharan Africa 19.1%, East Asia 13.4%, Southeast Asia 10.9%, Latin America 9.7%, Western Europe 7.0%, MENA 6.7%, North America 5.8%, Eastern Europe 4.6%, Oceania 2.4%
Urbanicity328Rural 33.2%, Dense urban 22.3%, Small town 21.0%, Suburban 20.4%, Nomadic / remote 3.0%
Socioeconomic band340Low 31.8%, Lower-middle 30.0%, Middle 20.3%, Upper-middle 12.6%, High 5.3%
Employment330Full-time 25.8%, Retired 16.1%, Homemaker 14.2%, Self-employed 11.8%, Part-time 9.1%, Student 8.2%, Unemployed 7.6%, Gig / freelance 7.3%
(c) Study setup. A study is specified as a decision to test (here, a $2 price increase), with the persona cohort on the left and the task catalog on the right, before launching the simulation.
(c) Study setup. A study is specified as a decision to test (here, a $2 price increase), with the persona cohort on the left and the task catalog on the right, before launching the simulation.
Table 9: Large template-based task collections. Members of a collection are separate task specifications but reuse a common contract and verifier pattern.
CollectionTasksEnvironmentComposition
Synthetic persona surveys405Survey135 Finance, 135 Healthcare, and 135 Software questionnaires, generally sharing a common survey and reporting template.
Product surveys200SurveyTwenty product-research archetypes, including purchase intent, price sensitivity, recommendation, and retention, crossed with ten retail products.
Synthetic chatbots351AI ChatbotScenario families across 27 domains, including travel, legal, healthcare, telecommunications, real estate, insurance, finance, and education.
(d) Interactive evaluation. A persona-conditioned agent converses with the system under test in a live transcript, scored in real time against a task-owned rubric (here, an AI counselor).
(d) Interactive evaluation. A persona-conditioned agent converses with the system under test in a live transcript, scored in real time against a task-owned rubric (here, an AI counselor).
Table 10: Logical components of an application-task contract.
Contract componentDeclared contentAudit purpose
Task metadataStable name and version, environment type, product domain, difficulty, tags, artifact paths, and runtime budgetsIdentifies the recipe and prevents results from silently moving between task versions.
Persona-facing scenarioContext, user goal, constraints, disclosure policy, and required submissionDefines what every sampled persona is asked to do without exposing verifier internals.
Cohort strategyPersona-schema filters, sampling mode, stratification fields, sample size, weights, and seed policyMakes the target audience explicit and permits the cohort to be redrawn.
Product attachmentQuestionnaire or stimulus, chat endpoint or sidecar, website target, or native application backendIdentifies the system under test and how the runtime reaches it.
VerifierRequired artifacts, structured finding schema, objective checks, timeouts, and failure conditionsConverts a trial into reproducible outcomes with supporting evidence.
Reporting policyAggregations, subgroup facets, summaries, optional judge directives, and disclosure rulesDefines how trial findings become a cohort-level report.
(e) Population report. Trials aggregate into a population-level outcome (here, 61% projected retention) with response distributions, decision reasons, and the per-persona response-level dataset backing the headline number.
(e) Population report. Trials aggregate into a population-level outcome (here, 61% projected retention) with response distributions, decision reasons, and the per-persona response-level dataset backing the headline number.
Table 11: Environment-specific inputs and artifacts. App covers desktop and mobile native applications, including Linux, macOS, and iOS.
EnvironmentTask-owned inputsPrimary trial artifacts
SurveyStimulus, questionnaire schema, response constraints, and optional rationale promptsTyped responses, missing or invalid items, rationales, confidence, and completion summary.
AI ChatbotProduct endpoint or sidecar, user goal, staged-disclosure policy, turn and safety limitsFull transcript, tool or service events, resolution state, termination reason, and post-run feedback.
WebURL or hosted site, browser backend, exploration requirements, and submission schemaPage and action trace, considered options, screenshots where enabled, final submission, and terminal page or site state.
AppNative target, platform backend, initial state, permissions, and terminal-state checksScreenshot and action trace, exported files, application state, permission changes, and cross-application side effects.
Figure 9: Move-transition graphs for the meal-planning chat study, by economic motivation. 1,000 GPT-5.5 persona conversations, 500 per stratum, one fixed assistant; 13,120 move transitions in total. A conversation averages six exchanges and fourteen moves, so it traverses the graph in loops rather than one pass; edges curving back up the page are those returns. Each panel draws only its own moves and transitions and lays itself out, so shape differences reflect the data. Black edges trace each step’s most frequent successor; every drawn edge carries its transition count, and a pair of moves with no edge between them co-occurred fewer than the panel’s count floor of about 90 transitions.
Figure 9: Move-transition graphs for the meal-planning chat study, by economic motivation. 1,000 GPT-5.5 persona conversations, 500 per stratum, one fixed assistant; 13,120 move transitions in total. A conversation averages six exchanges and fourteen moves, so it traverses the graph in loops rather than one pass; edges curving back up the page are those returns. Each panel draws only its own moves and transitions and lays itself out, so shape differences reflect the data. Black edges trace each step’s most frequent successor; every drawn edge carries its transition count, and a pair of moves with no edge between them co-occurred fewer than the panel’s count floor of about 90 transitions.
Table 12: Complete primary-outcome results for the eight validation tasks. Range is the largest minus smallest model share in percentage points. The q column reports a three-arm χ2 test after Benjamini–Hochberg correction across the eight tasks (∗∗∗q<0.001, ∗q<0.05; ns otherwise).
TaskEnv.Primary outcomeShare giving that answerRangeq
GPTOpusHaiku(pt)
Annual checkupSurveyq31 = e50.5%41.4%29.7%20.8<0.001∗∗∗
Candy Land priceSurveyq_threshold = hesitate98.3%27.0%83.3%71.3<0.001∗∗∗
OpenBB honestyChatwouldStillContinueUse = unsure81.3%85.5%28.1%57.4<0.001∗∗∗
Meal planning†ChatadherenceLikelihood = 750.6%0.2%40.0%50.4<0.001∗∗∗
Notion plansWebdecision_subject_id = plus63.5%21.5%75.7%54.2<0.001∗∗∗
MIT OCW courseWebtask_course_level = Graduate40.0%34.5%47.5%13.0<0.001∗∗∗
News+ subscriptionAppclicked_get_started = true4.2%20.8%0.0%20.80.025∗
Stocks sentimentAppsentiment = hold60.0%30.0%50.0%30.00.153 ns
Table 13: Product-level conclusions under three agent models on identical cohorts. Denominators are trials in which the field was answered. Notion aggregates its three paid tiers against the free tier. †The meal-planning GPT 5.5 arm did not run the declared cohort and is not comparable with the other two arms.
TaskProduct-level measureGPT 5.5Opus 4.8Haiku 4.5
Annual checkupVery likely to schedule in time50.5% (505/1,000)41.4% (414/1,000)29.7% (297/1,000)
Candy Land priceHesitates or worse at new price98.3% (983/1,000)27.0% (270/1,000)83.3% (833/1,000)
OpenBB honestyWould not continue using18.5% (185/1,000)14.5% (132/909)71.9% (719/1,000)
Meal planning†Stated need fully satisfied46.2% (462/1,000)0.2% (2/1,000)28.1% (281/1,000)
Notion plansSelected a paid plan75.8% (776/1,024)23.2% (237/1,022)93.9% (958/1,020)
MIT OCW courseChose a graduate-level course40.0% (403/1,008)34.5% (347/1,007)47.5% (473/996)
News+ subscriptionSubscribed4.2% (1/24)20.8% (5/24)0.0% (0/24)
Stocks sentimentBuy opinion40.0% (8/20)70.0% (14/20)47.4% (9/19)
Table 14: Rank correlation between model arms’ orderings of the same persona subgroups. Significance marks are ordered GPT / Opus / Haiku (∙ significant after correction, ∘ not). A flat arm has no ordering to compare. †The GPT arm has a cohort-integrity exception. ‡App rows contain only three to eight personas per subgroup.
TaskPersona dimensionGroupsSig.Spearman ρ, exact p
GPT/OpusGPT/HaikuOpus/Haiku
Annual checkupage bracket6∘∘∙+0.26 p=0.329−0.43 p=0.822−0.77 p=0.971
Candy Land priceeconomic motivation4∘∘∘−0.40 p=0.792−0.20 p=0.625+0.80 p=0.167
OpenBB honestytrust level4∙∙∙+1.00 p=0.042+1.00 p=0.042+1.00 p=0.042
Meal planning†life stage4∘∘∘+0.95 p=0.083+0.32 p=0.500+0.33 p=0.417
Notion planscompany size8∘∘∘+0.44 p=0.138+0.60 p=0.059−0.24 p=0.725
MIT OCW courseacademic field8∘∘∘+0.93 p=0.001+0.09 p=0.420+0.06 p=0.452
News+ subscr.‡economic motivation3∘∘∘+0.00 p=0.667flatflat
Stocks sentiment‡risk tolerance5∘∘∘+0.47 p=0.267−0.92 p=1.000−0.67 p=0.933
Table 15: Persona fidelity across all three model pairs. Cohen’s κ corrects raw agreement for each model’s answer distribution. GPT and Opus age-band matches are indistinguishable from uniform guessing (q=0.84 and q=0.85); Haiku matches 1,000 of 1,000 trials.
GPT × OpusGPT × HaikuOpus × Haiku
Paired agreement over 88 joinable fields
Median Cohen’s κ0.0000.000+0.001
Fields at κ≤059 of 8850 of 8840 of 88
Fields reaching κ≥0.27 of 888 of 887 of 88
Fields at ≥50% agreement, κ<0.148 of 5625 of 3424 of 34
Self-report fidelity: age band matches the persona
GPT 5.5 16.3% (ns) Opus 4.8 16.9% (ns) Haiku 4.5 100.0%chance 16.7%
Table 16: Persona adherence per attribute and environment, both acting agents. Each environment is split into two sub-columns: Opus 4.8 (left) and GPT-5.6-sol (right), under the same Opus 4.8 judge. Each cell shows two groups of five persona icons—the left group the positive cohort, the right group the negative cohort. A filled icon (🚹) marks a persona whose behavior expressed the declared value (for the negative cohort, correctly expressed the opposite value—target suppressed); a faint icon (🚹) marks one that did not. More filled icons is better in both groups. Overall Opus 366/400=91.5% (33/40 cells strong, ≥4 filled per group) vs. GPT-5.6-sol 317/400=79.2%.
SurveyChatWebOS-App
AttributeOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-solOpus 4.8GPT-5.6-sol
code-comment-style🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
code-naming-verbosity🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
code-summary-documentation🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-emoji-use🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-humor🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-politeness🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-storytelling🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-use-of-jargon🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
cog-verbosity🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
register🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹🚹🚹🚹🚹🚹 🚹🚹🚹🚹🚹
Table 17: Cited judge evidence, Survey environment. For each attribute, the behavior the judge identified in a positive persona (declared value) versus its matched negative persona (opposite value). Excerpts are quoted from the trial trajectories.
AttributePositive (declared)Negative (opposite)
code-comment-styleNearly every line carries an inline comment (# Iterate over every integer, # Check divisibility)Code contains no comments whatsoever
code-naming-verbositycalculate_average_of_passing_scores, number_of_passing_scoresVariables s, t, a, c, x, all single-letter
code-summary-documentationEvery function opens with a tldr: docstringProse plus code, no TLDR/summary header
cog-emoji-useMany emoji across a short paragraphNo emoji despite a casual app-store prompt
cog-humorPlayful, witty asides (“the boxes may yet win”)Measured, earnest tone, no jokes
cog-politeness“Might I kindly ask that you resend the document at your earliest convenience”“Hey, you forgot the attachment. Again.”: blunt and sarcastic
cog-storytellingA narrated scene (“a storm came through, half the lodge went dark”)Abstract, value-driven, no concrete scene
cog-use-of-jargonDense technical jargon (“TCP three-way handshake, SYN-ACK”)Plain terms (“translate the name into a numerical address”)
cog-verbosityRambling, multi-paragraph, tangential answerShort clipped fragments, no elaboration
registerStandard/formal phrasingColloquial (“cuppa and the papers”, “the wife’s doing a roast”)
Table 18: Extraction-quality metrics. Higher scores indicate better extraction quality.
MetricNameDefinition
M1Claim validityDoes each populated claim remain defensible when its value, cited evidence, and free-text description are considered jointly? A field is problematic if its value is indefensible, its evidence is plainly unrelated or non-supporting, or its description adds implausible unsupported specifics.
M2No over-claimingHas an attribute been populated without any plausible basis? Clearly unsupported assignments count as problems, while a plausible, appropriately labeled inference does not.
M3CoverageDid the extraction omit an important attribute that was clearly available in the source?
M4Internal consistencyDo extracted fields contradict one another, especially on core identity, role, region, life stage, language, or technology-use information?
M5Overall fidelity and plausibilityIs the extracted record usable for faithfully role-playing the source person, and are its inferred psychological, interest, and communication traits plausible and coherent?
Table 19: Paired LLM-judge results on all 1,000 extractions. Δ is GPT minus Claude; ≤1 is the proportion of paired scores that differ by at most one point.
MetricGPTClaudeΔ≤1
M1 Claim validity3.4333.937−0.50481.4%
M2 No over-claiming3.5323.646−0.11482.3%
M3 Coverage4.4654.031+0.43499.2%
M4 Internal consistency3.6063.932−0.32685.0%
M5 Fidelity and plausibility3.8294.109−0.28097.8%
Overall3.7733.931−0.15889.1%
Table 20: Human evaluation and within-one agreement on the 100-persona subset. H–H compares all 15 pairs of human raters; GPT–H and Claude–H compare each LLM judge with the six-rater human mean.
MetricHuman meanH–H ≤1GPT–H ≤1Claude–H ≤1
M14.10599.1%69.0%92.0%
M23.77092.2%81.0%88.0%
M34.22397.1%95.0%100.0%
M44.53798.6%55.0%92.0%
M54.04099.1%96.0%97.0%
Overall4.13597.2%79.2%93.8%
Table 21: Capability comparison across representative agent benchmarks, simulation systems, and synthetic-sampling methods. Systems, left to right: AgentBench [52], GAIA [55], WebArena [111], WebShop [102], Mind2Web [17], AppWorld [78], ToolLLM [72], BenchFlow [7], τ-bench [103], Generative Agents [66], SOTOPIA [113], OASIS [101], and Silicon sampling [4]. A check (✓) indicates explicit support in the published system; a cross (✗) indicates the capability is not stated. Dimensions read as follows: product-as-SUT evaluation treats the product, not the agent, as the subject under test; persona conditioning instantiates a user from a sampled population profile; behavioral grounding verifies that behavior tracks a target attribute rather than confounders; native-device execution covers real mobile/desktop apps beyond API demos; reproducible reporting separates verifier facts from reporting policy; trajectory inspection exposes per-trial traces; end-to-end validation closes the persona–task–product loop; silicon sampling substitutes simulated respondents for survey/UX pre-studies; reproducible cohorts deterministically re-instantiate the same cohort; and multi-agent simulation denotes persistent multi-agent social worlds (a deliberate non-goal for MatrAIx). The matrix compares capability coverage rather than providing an overall ranking.
Agent BenchGAIAWeb ArenaWeb ShopMind2 WebApp WorldTool LLMBench Flowτ-benchGenerative AgentsSOTOPIAOASISSilicon samplingMatrAIx
Product-as-SUT evaluation
Persona conditioning
Behavioral grounding
Native-device execution
Reproducible reporting
Trajectory inspection
End-to-end validation
Silicon sampling
Reproducible cohorts
Multi-agent simulation

Findings

  • In the 400-trial controlled study, the assigned behavior was expressed or correctly suppressed in 91.5% of trials (366/400).
  • Across 1,000 extracted personas, LLM judges GPT-5.5 and Claude Opus 4.8 gave overall mean scores of 3.773 and 3.931 respectively, and 89.1% of 5,000 paired metric scores differed by at most one point.
  • Six human raters scoring a 100-persona subset gave a mean of 4.135 out of 5, with 84.7% of 3,000 scores rated 4 or 5.
  • Compared to the human mean, GPT scores were within one point in 79.2% of cases and Claude scores in 93.8% of cases.
  • When the same cohort was run through three different models on identical tasks, one product's paid-plan selection share ranged from 23.2% to 93.9%, and pairwise agreement (Cohen's kappa) across 88 shared fields was close to zero.

Where it can be used

  • Screening how diverse user groups might react to a new product or feature (e.g., a price increase) before running costly human studies.
  • Checking interactive behaviors such as whether users keep using a chatbot after it fails, or how much response delay they tolerate.
  • Re-running the same persona cohort and tasks after a product update to compare versions under consistent conditions.
  • Breaking down results by subgroup (age, income, region, etc.) to spot problems that overall aggregate scores might hide.

Limits and open work

  • The authors explicitly state that persona cohorts are not probability samples of real populations, and that results should be treated as hypothesis-generating rather than direct evidence about real human behavior.
  • When the model playing the persona shares the same backbone as the system being evaluated, favorable results become ambiguous (self-preference bias vs. genuine satisfaction), and the current experiments did not run the crossing test needed to separate this effect.
  • Whether personas realistically show behaviors like withholding information, pushing back, or abandoning a task—key for realistic support-style interactions—has not been directly measured yet; the authors list comparison against real conversation logs as a priority for future work.
  • The public coreset (about 1 million personas) is a filtered subset of the full internal population, and its synthetic portion is calibrated to only four demographic marginals (age, region, gender, urbanicity), not the full 1,290-dimensional joint distribution.
  • The authors caution that for consequential decisions in regulated areas like health, finance, or employment, simulated results are not sufficient and require human validation using the same task and instrument.

Why it matters

This gives teams a way to cheaply and quickly stress-test AI products and interfaces against a wide, diverse range of simulated user backgrounds before—or alongside—expensive human studies. But the authors themselves caution that persona-agent results are not direct evidence of real human behavior, so this is best read as a way to generate hypotheses that still need human confirmation.

Terms in this paper

  • Persona · A simulated user profile defined by attributes such as age, language, and personality traits.
  • DAG (dependency graph) sampling · Generating persona attributes one at a time in an order where earlier ('parent') attributes shape later ('child') ones, so correlated traits like age and education level stay realistic.
  • Human-grounded record · A persona built from real-world data such as Wikipedia biographies, reviews, or survey responses, mapped into the same attribute schema.
  • Verifier · An automated check or human/LLM judge that confirms whether a persona agent's output met a task's required conditions.
  • Cohort · A selected group of personas (e.g., matching certain age or region criteria) used for a specific evaluation run.

Original abstract (English)

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

Authors · Xiaomin Li

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB

Figures: Xiaomin Li et al., arXiv:2608.04205, arxiv-nonexclusive