Four sources, each labelled on the attribute, with coverage as counts.
| Layer | Source | Chip | Coverage |
|---|---|---|---|
| 1 · Occupation | O*NET 30.1 (U.S. Dept. of Labor) + BLS OES May 2025 | Imported | … |
| 2 · Title space | PersonaOS taxonomy: industry → sub-area → role family → base role × specialization × seniority, expanded across firm size and segment | Imported | … |
| 3 · Behaviour | Fixed rules over layers 1–2: voice by seniority, modulators by firm size, frustrations and work context, constraints by buying role, trait sketch by a documented O*NET mapping | Generated | Every persona; labelled on the chip; never a probability |
| 4 · Human | Grounding rounds: a real person in the role answers a five-question instrument for one product; the curator applies proposals as version 2 | Human-answered | 0 rounds · 0 applied · 0 attributes in this browser |
What is measured now and what needs real people. No result here is simulated to look validated.
| Check | Measures | Method | Status |
|---|---|---|---|
| Test–retest baseline | How consistently a real person answers the same instrument twice. | Grounding round answered twice, two weeks apart; agreement is the ceiling every persona metric is read against. | Needs human data |
| Hold-out agreement | Whether a persona’s answer matches a held-out human answer for the same role and product. | For each grounding round, one question is withheld from the materials review; the version-1 persona answers it; agreement is counted per audience and reported with an exact interval. | Needs human data |
| Distributional match | Whether a panel’s choice counts match human counts for the same stimulus. | Panel next-step counts against a human sample’s counts; distance reported, never a single accuracy number. | Needs human data |
| Subgroup parity | Whether agreement holds across seniority, industry and firm size, not only on average. | The hold-out metric sliced by the audience facets; a gap is reported as counts per slice. | Needs human data |
| Provenance on every attribute | That no attribute exists without a source. | Origin and evidence reference on each attribute; provenance counts on every persona page. | Live in this demo |
| Traceable reactions | That every sentence a persona produces points at the attribute it came from. | Reaction lines carry attribute chips; panels count which attributes shaped the sample. | Live in this demo |
| No self-graded confidence | That the system never scores its own certainty. | Confidence is null on every attribute; rollups are counts, rates are computed on read. | Live in this demo |
| Occupational anchoring | That role-level facts come from incumbent surveys and labor statistics, not from a model. | Every base role maps to an O*NET-SOC occupation; work context, tasks, knowledge, skills, interests, values, styles, employment and pay are loaded from the public release with their domain source. | Live in this demo |
The failure modes the literature reports for LLM personas, and what in PersonaOS addresses each.
| Distortion | What it looks like | Design answer |
|---|---|---|
| Flattened variance | LLM personas converge on the average answer, so a sample of 500 behaves like 5. | Variation is structural, not prompted: 122,596 titles × ~25 firm contexts × an expertise score assigned per row; panels report the spread as counts. |
| Demographic stereotyping | Prompted demographics produce caricatures; accuracy drops for under-represented groups. | Personas are occupational, and the occupation layer is survey data from incumbents in that occupation, not a description of them. |
| Sycophancy and position bias | Models agree with the stimulus and prefer the first option. | Reactions are composed from attributes and constraints, so a Junior cannot sign and a Budget Holder cannot approve above the team line, whatever the stimulus says. |
| Hallucinated specifics | Invented tools, numbers and quotes that read as evidence. | Evidence by reference: a quote exists only if a person said it in a grounding round; O*NET and BLS values carry their release and code. |
| Self-graded confidence | A confidence number the model made up is read as calibration. | Confidence is never self-graded; groundedness is shown as counts by provenance, and calibration is a protocol that needs human data. |
| Static population | A persona set that never learns from the product it is used on. | Grounding rounds attach human evidence per persona and product; the public row is never mutated, the grounded version is a layer, and an applied round is spent. |
Vendor claims as stated on their public sites, September 2026. Links open the source.
| Vendor | Population basis | Scale | Grounding source | Provenance per attribute | Published validation | B2B structure | Acts in product |
|---|---|---|---|---|---|---|---|
| PersonaOS (Marketrix) | Occupational title space modeled on O*NET/SOC, expanded across firm size and segment | 3,000,000 personas · 122,596 titles · 194 O*NET occupations | O*NET incumbent surveys + BLS OES per role; grounding rounds per persona and product | Origin and evidence reference on every attribute; counts, never self-graded | Protocol published here; hold-out agreement needs customer rounds (not claimed) | Native: seniority, buying role, firm size, segment, expertise tier | Personas drive real products in Marketrix simulations, studies and QA flows |
| Simile | Digital twins of real people from two-hour interviews; behavioral foundation model | 1,000-person study; model trained on 2.9M responses from 210 experiments (as stated) | Recorded interviews, daily studies, client data | Per-result predicted accuracy tag (as stated) | 85% of human self-retest accuracy on GSS (2024 paper); weekly evaluations (as stated) | Consumer and population focus | Simulation of responses; not stated to operate products |
| Aaru | Simulated populations acting under hypothetical conditions | “Scale of the real world”; 40,000 simulated investors tracker (as stated) | Constructed worlds; client research | Not stated | 0.90 median correlation on an EY wealth study (as stated) | Consumer, investor and civic populations | Agents act in constructed worlds |
| Artificial Societies | Societies of personas built from surveys, focus groups and interviews | 3M+ personas (as stated) | Human statements and behaviours | Not stated | 86% distribution accuracy across 1,000 panels (as stated) | Audience and communications focus | Simulates opinion formation in networks |
| Synthetic Users | OCEAN-profiled AI participants for interviews | 10–12 participants per qualitative study; hundreds for quantitative (as stated) | Optional RAG over customer transcripts and tickets | Not stated | 85–92% parity with organic research (as stated) | Any audience by description | Interviews only |
| Evidenza | Synthetic B2B buyers and executives by product category | Thousands of synthetic customers per study (as stated) | Category-based generation | Not stated | 88% accuracy in 100+ validations; 0.81 correlation with Salesforce (as stated) | B2B executive focus | Surveys and interviews |
| MatrAIx | Persona agents for evaluating AI products | 8.3 billion persona agents, 1,290 attributes each (as stated) | Not stated | Not stated | Not stated | Evaluation of chatbots, prototypes and apps | Runs agents through prototypes and workflows |