PersonaOS

Methodology

What grounds a persona, how it would be validated, what can go wrong, and how the field compares

Grounding layers

Four sources, each labelled on the attribute, with coverage as counts.

LayerSourceChipCoverage
1 · OccupationO*NET 30.1 (U.S. Dept. of Labor) + BLS OES May 2025Imported
2 · Title spacePersonaOS taxonomy: industry → sub-area → role family → base role × specialization × seniority, expanded across firm size and segmentImported
3 · BehaviourFixed rules over layers 1–2: voice by seniority, modulators by firm size, frustrations and work context, constraints by buying role, trait sketch by a documented O*NET mappingGeneratedEvery persona; labelled on the chip; never a probability
4 · HumanGrounding rounds: a real person in the role answers a five-question instrument for one product; the curator applies proposals as version 2Human-answered0 rounds · 0 applied · 0 attributes in this browser

Validation protocol

What is measured now and what needs real people. No result here is simulated to look validated.

CheckMeasuresMethodStatus
Test–retest baselineHow consistently a real person answers the same instrument twice.Grounding round answered twice, two weeks apart; agreement is the ceiling every persona metric is read against.Needs human data
Hold-out agreementWhether a persona’s answer matches a held-out human answer for the same role and product.For each grounding round, one question is withheld from the materials review; the version-1 persona answers it; agreement is counted per audience and reported with an exact interval.Needs human data
Distributional matchWhether a panel’s choice counts match human counts for the same stimulus.Panel next-step counts against a human sample’s counts; distance reported, never a single accuracy number.Needs human data
Subgroup parityWhether agreement holds across seniority, industry and firm size, not only on average.The hold-out metric sliced by the audience facets; a gap is reported as counts per slice.Needs human data
Provenance on every attributeThat no attribute exists without a source.Origin and evidence reference on each attribute; provenance counts on every persona page.Live in this demo
Traceable reactionsThat every sentence a persona produces points at the attribute it came from.Reaction lines carry attribute chips; panels count which attributes shaped the sample.Live in this demo
No self-graded confidenceThat the system never scores its own certainty.Confidence is null on every attribute; rollups are counts, rates are computed on read.Live in this demo
Occupational anchoringThat role-level facts come from incumbent surveys and labor statistics, not from a model.Every base role maps to an O*NET-SOC occupation; work context, tasks, knowledge, skills, interests, values, styles, employment and pay are loaded from the public release with their domain source.Live in this demo

Known distortions and the design answer

The failure modes the literature reports for LLM personas, and what in PersonaOS addresses each.

DistortionWhat it looks likeDesign answer
Flattened varianceLLM personas converge on the average answer, so a sample of 500 behaves like 5.Variation is structural, not prompted: 122,596 titles × ~25 firm contexts × an expertise score assigned per row; panels report the spread as counts.
Demographic stereotypingPrompted demographics produce caricatures; accuracy drops for under-represented groups.Personas are occupational, and the occupation layer is survey data from incumbents in that occupation, not a description of them.
Sycophancy and position biasModels agree with the stimulus and prefer the first option.Reactions are composed from attributes and constraints, so a Junior cannot sign and a Budget Holder cannot approve above the team line, whatever the stimulus says.
Hallucinated specificsInvented tools, numbers and quotes that read as evidence.Evidence by reference: a quote exists only if a person said it in a grounding round; O*NET and BLS values carry their release and code.
Self-graded confidenceA confidence number the model made up is read as calibration.Confidence is never self-graded; groundedness is shown as counts by provenance, and calibration is a protocol that needs human data.
Static populationA persona set that never learns from the product it is used on.Grounding rounds attach human evidence per persona and product; the public row is never mutated, the grounded version is a layer, and an applied round is spent.

Positioning

Vendor claims as stated on their public sites, September 2026. Links open the source.

VendorPopulation basisScaleGrounding sourceProvenance per attributePublished validationB2B structureActs in product
PersonaOS (Marketrix)Occupational title space modeled on O*NET/SOC, expanded across firm size and segment3,000,000 personas · 122,596 titles · 194 O*NET occupationsO*NET incumbent surveys + BLS OES per role; grounding rounds per persona and productOrigin and evidence reference on every attribute; counts, never self-gradedProtocol published here; hold-out agreement needs customer rounds (not claimed)Native: seniority, buying role, firm size, segment, expertise tierPersonas drive real products in Marketrix simulations, studies and QA flows
SimileDigital twins of real people from two-hour interviews; behavioral foundation model1,000-person study; model trained on 2.9M responses from 210 experiments (as stated)Recorded interviews, daily studies, client dataPer-result predicted accuracy tag (as stated)85% of human self-retest accuracy on GSS (2024 paper); weekly evaluations (as stated)Consumer and population focusSimulation of responses; not stated to operate products
AaruSimulated populations acting under hypothetical conditions“Scale of the real world”; 40,000 simulated investors tracker (as stated)Constructed worlds; client researchNot stated0.90 median correlation on an EY wealth study (as stated)Consumer, investor and civic populationsAgents act in constructed worlds
Artificial SocietiesSocieties of personas built from surveys, focus groups and interviews3M+ personas (as stated)Human statements and behavioursNot stated86% distribution accuracy across 1,000 panels (as stated)Audience and communications focusSimulates opinion formation in networks
Synthetic UsersOCEAN-profiled AI participants for interviews10–12 participants per qualitative study; hundreds for quantitative (as stated)Optional RAG over customer transcripts and ticketsNot stated85–92% parity with organic research (as stated)Any audience by descriptionInterviews only
EvidenzaSynthetic B2B buyers and executives by product categoryThousands of synthetic customers per study (as stated)Category-based generationNot stated88% accuracy in 100+ validations; 0.81 correlation with Salesforce (as stated)B2B executive focusSurveys and interviews
MatrAIxPersona agents for evaluating AI products8.3 billion persona agents, 1,290 attributes each (as stated)Not statedNot statedNot statedEvaluation of chatbots, prototypes and appsRuns agents through prototypes and workflows