MatrAIx: The Hardest Part of AI Evaluation / MatrAIx:AI评测中最难的部分
(This interview was originally published at AGI House Research on September 26, 2026.)
Last week at Enterprise Deployment Build Day, we sat down with Xiaomin Li and Yuexing Hao, the two founders of MatrAIx. We talked about why offline benchmarks miss what real users feel, how 8.3 billion personas get built and turned into agents, what “91.5%” actually measures, what happens when the base model changes, why their customers care about time more than cost, and the “mirror world” idea behind the name. Full interview below.
Behind 8.3 billion simulated users, a startup is targeting the part of AI evaluation that benchmarks never measured.
1. An agent that passed every test and scored 2 out of 5
A coding agent finished its task. Every unit test passed. By any current benchmark, that counts as success. MatrAIx’s simulated user gave it 2 out of 5.
The team followed the reasons that “user” wrote down and traced them to specific steps. The agent had spun through many idle rounds. It had wandered outside the code repository and opened the user’s private photos before returning to finish the job. Token usage was more than 50 times normal. “The first problem is safety. It went out of scope; it should have been limited to the repo. The second is efficiency.”
The two founders described this internal case in our interview. Conventional evaluation checks only whether the result is correct. Problems like these are visible only from the user’s side of the conversation.
2. Who they are
MatrAIx was started by Xiaomin Li and Yuexing Hao. Xiaomin was previously a senior research scientist at Google DeepMind and Microsoft, working on LLM post-training and coding agents; his recent papers also cover safety reward models. Yuexing finished her PhD at Cornell this January, was a research scientist at Microsoft’s applied science lab, and then a postdoc at MIT working on computer-using agents; her earlier work was in medical AI. The two met at MIT and have collaborated for nearly three years.
The project began as an open-source community. The Playground and task library were open-sourced on July 31; the Persona 1M dataset went up on Hugging Face on August 1; the technical report appeared on arXiv on August 4. The paper lists 93 authors. By the team’s own count, more than 200 scientists have joined the community, over 40 of them from OpenAI, Anthropic, Google DeepMind, and xAI. The GitHub repository has about 1.9k stars under an MIT license.
In the interview they explained why researchers at large labs joined: “They all had this problem in their own work settings, and they wanted to try whether persona agents could solve it.”
The name comes from The Matrix, with the “I” replaced by “AI”. Xiaomin called it “an idea I’ve had since I was a kid.” The GitHub README states that “the simulated world is for exploration, stress testing, and hypothesis generation, not a replacement for evidence from real people”.
3. The problem: the gap between offline scores and online feedback
Xiaomin described a pattern they kept hitting in industry: “The benchmark scores can be great, but once it ships, the feedback from real users is just different.”
Take a coding agent. The same question gets asked in very different ways. Some people write extremely detailed prompts; others are lazy and vague. Some pack ten questions into one prompt; others ask one per turn. Students ask for an explanation after every step. Some people drift from code to the news and back again. Offline benchmarks capture none of this.
The paper’s introduction puts it more formally. A novice wants explanations, small edits, and frequent confirmation; an expert wants terse responses and more autonomy. These differences shape interaction trajectories, trust in the result, and willingness to continue after a failure. Aggregate scores hide problems that specific user groups run into.
Using an LLM to play the user is not new; benchmarks such as τ-bench already do it. But the reliability of that setup has drawn specific criticism. A CMU study published in March ran 451 real people through the full τ-bench protocol and compared them against 31 LLM simulators. It found simulated users to be excessively cooperative, stylistically uniform, and lacking real frustration or ambiguity. The effect is an “easy mode” that inflates agent success rates above the human baseline.
MatrAIx targets exactly this gap: not one generic “user”, but diversity built into the simulation.
4. Technical deep dive
(I): how 8.3 billion profiles are generated
The schema. Each persona is described by 1,290 categorical dimensions in five groups.
| Group | Dimensions | Representative attributes |
|---|---|---|
| Background | 238 | Age, region, language, education, family, occupation, industry |
| Psychology | 210 | Personality, values, worldview, motivation, risk tolerance |
| Capability | 331 | Domain expertise, general skills, tools, programming, developer context |
| Behavior & interaction | 124 | Preferences, habits, interaction states, work style, tech adoption |
| Lifestyle | 387 | Interests, media, culture, hobbies, sports, diet, health, fitness |
Each dimension takes a value from a finite set. English proficiency runs from Native to None with CEFR levels in between; risk tolerance runs from risk-averse to risk-seeking.
In our interview they added a design detail the paper does not spell out: demographic dimensions come from public databases, but behavioral dimensions were built “scenario first.” The team mapped more than 50 AI-interaction scenarios and then defined interaction attributes for each: whether someone comments their code, their Python proficiency, whether their prompts are verbose, how many questions they ask per turn. The 1,290 dimensions are only the first version; the new one has more.
The synthetic path. Sampling each dimension independently produces impossible profiles, such as a 19-year-old retired surgeon. The paper instead builds a directed acyclic graph (DAG) over the 1,290 dimensions. An edge is added only when a data source directly reports the conditional relationship, and personas are sampled one dimension at a time in topological order. Each non-root node’s conditional distribution has the form
p(Xi = v | xPa(i)) ∝ πi(v) · ri(v; xPa(i)) · mi(v; xPa(i))
Here πᵢ is the population-wide prior. rᵢ is a source-informed likelihood-ratio adjustment: if the primary language is English and the region is North America, “Native” gets more weight. mᵢ is a binary compatibility mask that removes contradictions such as “primary language English, English proficiency None.” An unusual but possible combination like “primary language English, proficiency Basic” is down-weighted, not removed. The point is to keep rare but real people while ruling out logically impossible ones.
The human-grounded path. The other records come from six sources: Wikipedia biographies, Amazon review histories grouped by reviewer, the Stack Overflow Developer Survey, the U.S. General Social Survey (GSS), the PRISM Alignment dataset, and 355 consented responses to MatrAIx’s own persona survey. Free text is extracted by an LLM under constraints; dimensions the evidence does not support are left null rather than imputed, and names and contact details are stripped. For example, Shakespeare: his biography, works, and tastes map into the schema, but “he probably never used a coding agent,” so that dimension stays empty.
Quality control. Human-grounded records are deduplicated with exact hashing plus MinHash; synthetic records are treated as duplicates when all 14 high-information attributes match. The public coreset of roughly one million records contains 599,847 human-grounded records (323,438 from Wikipedia, 113,120 from Stack Overflow, 97,915 from Amazon, 63,532 from GSS) and 400,000 synthetic ones.
In terms of cultural coverage, the profiles span a dozen or so countries, but “the real information sources are probably all American,” and the United States is their primary focus, both as a population and as a market.
(II): from profile to acting agent
The profile is written into the system prompt as text. Xiaomin said it is “just like your skills, a skills.md,” a textual description rather than a numeric vector, closer to a filled-in questionnaire than an embedding.
The bottleneck of this approach is the context window. Profiles grow, especially once self-evolution is added, and “at some point the context window can’t hold it.” Three directions are under exploration:
- Routing. Load only the attributes relevant to the task. Whether someone likes spicy food probably does not affect their coding style, so it can be dropped.
- A more compact representation, with training built for that representation, to balance efficiency against quality.
- Steering. Push the model directly toward a specific persona’s behavior.
They also said they have found non-training methods that beat plain prompting, sitting between prompting and full training. They also plan to post-train their own model, similar to what Cursor did: it lets users plug in different vendors’ models but also ships its own Composer model. Hosting their own model costs R&D up front but should be cheaper and more flexible in the long term.
At the execution level, a trial is formalized as the tuple ⟨persona, task, agent interface, model, random seed⟩. Trials share no state, so they are parallel by construction. Each trial produces an artifact bundle that a task-owned verifier converts into structured findings. There are four environments.

A persona record becomes an agent when paired with a model, acts in one of four environments, and its trace is scored by a verifier tied to the task.
App tasks run in a Docker-based Linux desktop sandbox or on a remote macOS desktop or iOS simulator. Of the 1,010 tasks in the library, 621 are surveys, 371 are chatbot tasks, 12 are web tasks, and 6 are app tasks. The more interactive the environment, the fewer the tasks, and cost is the reason: in the paper, survey, chatbot, and web tasks used about 1,000 personas per model, while the two app tasks used 24 and 20.
(III): how do you know it acts like a person
This is the central question for the whole approach. The paper splits validation into layers, and each layer answers a different question.
Layer one: does the agent follow its assignment (persona adherence)? The team chose ten behavioral attributes and tested each in all four environments, with five personas assigned one pole and five the opposite, for 400 trials judged by an LLM on the recorded trajectory. 366 trials matched the assignment, or 91.5%.
| Environment | Match rate | Strongly consistent attributes (of 10) |
|---|---|---|
| Surveys | 92–96% | 9 |
| AI chatbot | 92–96% | 9 |
| Web | 92–96% | 9 |
| App | 83% | 6 |
This number measures whether the agent obeyed the persona it was given. It does not measure agreement with real human behavior. In our interview, they said persona adherence is, in their view, essentially an instruction-following capability, so a stronger base model helps here: “We’re standing on their shoulders.”
Layer two: is the extraction faithful to its source? Two LLM judges scored 1,000 extracted personas. Six human raters gave a source-matched subset of 100 a mean of 4.135 out of 5. Claude Opus 4.8’s scores fell within one point of the human mean 93.8% of the time; GPT 5.5’s did so 79.2% of the time.
Layer three: are findings stable across models? In a financial research task (OpenBB), each persona carried one of four trust orientations, from hostile to trusting. All three models ordered the four groups identically, with Cramér’s V between 0.228 and 0.363.
Agreement with real humans is the layer the paper does not cover head-on. In our interview, they said the team compares personas extracted from real profiles against the corresponding people’s actual choices, such as “A or B.” Results vary by domain. Software engineering and computer use show small gaps, because “what a person can operate, it can operate too.” Finance, health, and political elections “still have quite large gaps.” Experiences that need physical sensation, “I really have to drink it to know if I like it,” are out of reach for now. The paper’s conclusion takes the same position: human studies remain necessary before applying conclusions to real populations or consequential decisions.
A position piece on their site goes one level deeper. Its title: “To Simulate a User, You Should Make Your LLM Dumber.” The argument is that a frontier model knows too much, stays too consistent, and fails too rarely to stand in for a real person, so using it as-is cannot show where real users get stuck.
5. What happens when the base model changes
During the interview, we asked: the persona is a layer on top of a base model. When the base model changes generation, or its own personality shifts, does the simulation drift with it?
They acknowledged this is exactly the problem they are working on. If the persona is only “a shallow layer,” a change in the base changes the results, so they need methods that adapt to different base models and are training their own. They suggest that the model driving the persona is part of the evaluation configuration and should be reported with every result, and that important findings should be re-checked with more than one model before guiding product decisions.
Their blog post gives a worked example. In a case contributed by an Indian AI company, 400 simulated Indian business owners and finance leads evaluated four versions of a chartered accountant’s engagement letter. The version with a lower upfront payment and a broader liability cap raised the share requesting only minor changes from 0% to 23%. Liability terms mattered far more (+19.63 percentage points) than the payment schedule (+3.38 points). The same 400 personas were rerun on three different models; the direction of the finding held, while the exact rates differed.
Another study of theirs, MicroVerse, measures identity drift over long runs. Twenty-five agents each carry a locked “soul file” of values and moral limits into a desert grid where water is scarce. The researchers compare that fixed file against the “living journal” each agent keeps rewriting about itself. On the same map, Qwen3-32B settled at an oasis once it found water, while Claude Haiku kept relocating; 11 of 12 Haiku runs ended with no survivors. They warn that, when assigning a ruthless persona, they may be measuring the model’s built-in safety habits rather than the persona they thought they deployed.
6. Who the customers are and what they buy
Customer profile. Current customers are mainly LLM model and digital product companies: large labs building general models, and smaller labs building the best model in one vertical. The target verticals are e-commerce, gaming, and advertising. Inbound interest has come from more than ten countries, including Singapore, Japan, Argentina, Cambodia, Mexico, Switzerland, Sweden, and China. Many overseas teams are still targeting the U.S. market. They have also received requests from small businesses (such as a restaurant owner), one-person companies, and academic researchers.
What they deliberately avoid. Government work and political elections are mostly off the table. This is mainly because of the high sensitivity involved, and because outcomes there depend less on persona and can be heavily affected by black swan events, such as a shooting or an earthquake. The same applies to financial market prediction. That sets them apart from competitors such as Aaru, whose specialty is forecasting elections and group behavior.
“The feedback from frontier labs and big tech is that they don’t care that much about cost. Saving time is what matters most. They don’t want to wait ten days; they want a round in half a day or a few hours, then go change training and the harness.” That being said, they also have a cost advantage, as given in their site’s cost analysis: 10,000 surveys cost $28,560 or more with humans versus $3.60 to $104 simulated; 1,000 chats cost $8,568 or more versus $3 to $87.
They also brought up three disadvantages of human evaluation compared to their simulation:
- Scale is not representative. A model serves a billion users, but a human evaluation can only recruit fifty to a hundred thousand.
- Annotator distribution is skewed. They have seen cases where 80% of annotators turned out to be in one Southeast Asian country.
- Signals are noisy. A user who stops mid-conversation may have just stepped away and come back later. Humans usually give a thumbs up or down; simulated users write out their reasons every time.
But they position themselves as a complement, not a replacement: “Iterate fast with us first, then run a final round of human eval.”
The deliverable is more than a report. The multi-turn interaction data can be used for SFT. The team also provides a reward model for RL that “includes not only task completion but each user’s own satisfaction.”
The competitive field. Simile, founded by Joon Sung Park of Stanford’s “Smallville” work, announced a $100M Series A in February and closed a $200M Series B at a $2B valuation at the end of July. Aaru’s December Series A was led by Redpoint, with part of the equity priced at a $1B valuation. Most of these companies enter through market research. Those companies lean toward qualitative research and in-depth interviews, aiming to predict a specific person or group, while MatrAIx wants “large-scale persona agent deployment.” Xiaomin added that model iteration today “doesn’t rely on changing the algorithm every two weeks”; it relies on hill climbing, finding where the strongest model still fails. “Their target may really be to predict one person. Ours is to bring in human diversity through simulation and find the model’s latent problems.”
7. Dynamic personas and the mirror world
In the new version, personas are no longer static. Mood and work state change, agents self-evolve, they have memory, and even social relationships, though the network is limited for now: “If everyone had to know ten thousand people, it wouldn’t be efficient.”
The mechanism is event-driven. The team feeds in real news, such as an earthquake in California, and lets agents in the affected region and group “react the way a person would.” Those reactions are written into each agent’s memory. Simulated time currently runs at the same speed as real time. This can be accelerated, but then “you lose a lot of real-life big events to project from.”