Unlisted research program map

The research questions behind AnthroSim.

The program tests selected claims against ANES 2020, 500 recorded DeliData discussions, and 369,215 Wikipedia Articles-for-Deletion debates.

The program contains 8 numbered studies in 5 accessible drafts, plus 2 exploratory pilots with no draft yet. Review scores come from the automated Stanford Agentic Reviewer at paperreview.ai. Human peer review is pending for every draft.

These studies measure specific political and reasoning tasks. Use the results to plan real research; they do not predict what a particular customer, voter, patient, or juror will do.

Draft available Result in draft Partial result in draft Exploratory, no draft

Numbered studies

1
Study 1 · IE-ASIM-2026-01

Simulating a population

Draft available
Question

When personas are built from survey data, do their answers reproduce the population distribution, and can changes to that population be audited?

Answer so far

On 74 ANES demographic cells, calibrated personas had about one-third the vote-distribution error of naive repeated prompting and preserved within-cell variation. Direct readout, verbalized sampling, a calibrated label, and a training-data lookup matched or beat the persona arm on these coarse static cells. Steering was tested for ideology and religious attendance; it worked on DeepSeek and only partly on GPT-4o-mini.

Why it matters

Simpler methods are stronger baselines for static shares in well-surveyed groups. Persona value must be measured on individual variation and downstream interaction, with calibration and steering audits published alongside the result.

Automated review, paperreview.ai

6/7, 7/7, 5/7 · latest assessment: justify publication

2
Study 2 · IE-ASIM-2026-02/03

Group discussion

Draft available
Question

Can simulated discussion reproduce which people change their answers in real groups?

Answer so far

In a 50-group Wason evaluation, the original engine scored worse than no conversation. A private step that reasoned from each participant's seeded misconception reduced composite error by 45% on DeepSeek. Three-seed transfer tests found the opposite effect on Qwen 9B and an inconclusive result on Llama 3.1 8B.

Why it matters

A conversation mechanism can improve one model and harm another. Revalidate it whenever the model or engine changes.

Automated review, paperreview.ai

4/7, 4/7, 5/7, 7/7 · latest assessment: recommend acceptance

3
Study 3 · IE-ASIM-2026-02/03

Model size

Draft available
Question

Does a larger model produce more human-like deliberation?

Answer so far

On the Wason task, solo drift toward the textbook answer rose across nine Qwen sizes, although the series was not monotonic. The 0.8B model's wrong-to-correct rate was 0.044, compared with 0.573 for DeepSeek. Under generated deliberation, error was statistically tied from 4B upward. Both member fidelity and change dynamics followed the thinking model; the speaking model had a smaller effect.

Why it matters

Model choice trades member-level reading against task-specific drift. Select the thinking model from measured requirements. The single-run split suggests a cheaper speaking model may preserve most measured behavior, but this needs interval-bearing replication.

Automated review, paperreview.ai

4/7, 4/7, 5/7, 7/7 · latest assessment: recommend acceptance

4
Study 4 · IE-ASIM-2026-04

Evidence in the room

Draft available
Question

When sources enter a discussion, do simulated voters change at the human rate and for supported reasons?

Answer so far

Without evidence arrival, generated agents changed at 0.09 times the real rate. On 48 committed voters in content-bearing debates, an external belief state plus a provenance gate matched the gross change magnitude, 0.3125; the instability-adjusted result was marginal. Across six seeds per evidence-bearing cell, evidence helped at the level of passing the preregistered band, while rankings by evidence volume and model size remained unresolved. In a 39-turn audit, 19 of 34 source-claiming turns, 56%, were judged unsupported by their excerpt.

Why it matters

Evidence must reach the room, and source claims need a content check. This study does not rank evidence volume or model size.

Automated review, paperreview.ai

5/7, 4/7, 4/7, 5/7, 5/7 · latest assessment: recommend acceptance, or a strong borderline in favor

5
Study 5 · IE-ASIM-2026-05/06

Who sees the votes

Draft available
Question

Does seeing the room's votes change what a simulated group decides?

Answer so far

On one model and task, an anonymous running count increased within-group agreement by 17.8 percentage points. Adding names produced no detected additional effect. One specified false tally produced no detected change in the share ending on its target relative to the true-tally arm. On within-group agreement, it produced no detected difference from the private control and was 16.6 points below the true tally. Simulated groups reached about half the convergence observed in the real groups.

Why it matters

Vote visibility is a causal configuration choice. The experiment detected an informational channel and left identity-based social pressure unresolved for this task.

Automated review, paperreview.ai

6/7, 5/7, 6/7 · latest assessment: recommend acceptance

6
Study 6 · IE-ASIM-2026-05/06

Does it read like a person?

Draft available
Question

Can visible AI-writing tells be removed, and does removing them improve behavioral fidelity?

Answer so far

A speech checklist removed 69% of the measured writing tells, with no detected behavioral change. An instruction aimed at hedging in private thought barely changed the visible tells and improved composite fidelity by 0.282. In a 34,000-message census, tell rates varied sharply by model.

Why it matters

Transcript style and behavioral fidelity require separate measurements. Cleaner prose does not establish a more faithful simulated decision.

Automated review, paperreview.ai

6/7, 5/7, 6/7 · latest assessment: recommend acceptance

7
Study 7 · IE-ASIM-2026-07/08

Documented individuals

Draft available
Question

Do measured facts about a person improve prediction of that person's answer?

Answer so far

Across three models, one predictive political attribute raised balanced accuracy from 0.60-0.64 to 0.86-0.88. Traits sampled from the person's demographics scored 0.49-0.52. Role-playing the person and asking about them directly were statistically equivalent, agreeing on 92-96% of individuals.

Why it matters

Measured provenance carries useful signal. A structured profile provides an audit trail; the persona framing added no measured accuracy in this experiment.

Automated review, paperreview.ai

6/7, 7/7, 7/7 · latest assessment: recommend acceptance

8
Study 8 · IE-ASIM-2026-07/08

Audiences no survey measured

Draft available
Question

Can deeper demographic descriptions reach a population that existing surveys missed?

Answer so far

Across specifications containing one to five coarse demographic attributes, direct prompting remained the most accurate tested arm, and the training-data lookup beat the fusion generator at every depth. Even the deepest survey cells still had training support. The support gap appeared when a variable was missing. That result comes from masking an observed variable; a genuinely never-collected variable remains untested.

Why it matters

A defensible unsupported-audience claim must identify the missing variable, show that it carries signal, and report the survey's own uncertainty floor.

Automated review, paperreview.ai

6/7, 7/7, 7/7 · latest assessment: recommend acceptance

9
Study 9 · IE-ASIM-2026-09-PILOT

Steering a group by choosing who is in it

Exploratory
Question

Given an outcome we name in advance, can a search over who sits in a simulated room find the composition most likely to produce it, and does that composition survive being run again?

Answer so far

Composition steers a room through what members are told they believe, not through who they are told they are. On a transit measure where ANES 2020 measures the real demographic gaps (age 18.9 points, income 14.7, sex 8.0), the simulation reproduced none of them: age moved 2.6 points, income and sex moved the wrong way, and every real gap fell outside the simulated interval. A member told what they think moved the outcome 90.6 points. The personas are not being ignored, they are audible in the speech at 0.93 to 0.99 classifier AUC; the model voices the demographic and then does not reason from it. Separately, a searched best mix that looked worth 37.5 points over an ordinary panel was worth 10 on fresh panels, too small for this pilot to distinguish from zero.

Why it matters

A synthetic panel assembled from census demographics will not carry the real demographic gradient on a values question, and the error is directional rather than random, so it can be checked in advance against any survey covering the outcome. Use measured attitudes as the input and publish the survey anchor beside the result.

Read the study → No draft yet
10
Study 10 · IE-ASIM-2026-10-PILOT

Do simulated people mean what they say?

Exploratory
Question

A simulated participant can speak fluently and still argue one way while voting another, or say things that contradict the person they were told to be. Neither shows up in an outcome distribution. Can a larger model audit a smaller one turn by turn and catch it?

Answer so far

Yes, and the failure rate is high. Auditing 528 simulated board members with a larger judge, about a third fail: one in nine argues one way and votes the other, and one in four says something contradicting the background they were given. Auditing the transcript alone catches only half of that, because persona contradictions are invisible without knowing who the member was supposed to be. Repairing the discussion prompt first took agreement between a member's words and their own vote from 0.605 to 0.885, so the residual is not simply a prompting artefact.

Why it matters

This is the check that decides whether a simulated panel can be trusted at the level of an individual participant rather than an average. It is also the natural place to spend a larger model: generate cheaply, adjudicate well.

Read the study → No draft yet
Cross-cutting measurements

Questions answered inside the drafts

Do simulated people all sound alike?

Result in draft
Question

How do simulated speakers differ from real group members, and what does orchestration improve?

Answer so far

Standalone agent arms used 0.33-0.44 vocabulary breadth versus 0.71 for people and repeated bigrams 0.23-0.36 versus 0.07. The orchestrated engine matched real rooms on cross-speaker overlap and did not separate statistically on vocabulary breadth, although it still repeated more. A classifier identified most simulated authors more easily than real authors.

Why it matters

The measured defect is narrow vocabulary and repetition. Attribution already distinguishes the agents easily.

Read the result →

Which models can run the engine?

Partial result in draft
Question

Does the private-thought mechanism transfer across model families?

Answer so far

Across three seeds per arm, the loop improved DeepSeek, worsened Qwen 9B, and was inconclusive on Llama 3.1 8B. Smaller Qwen results were also inconclusive. The tested families provide no reliable predictor of transfer.

Why it matters

Run the same calibration test on every model intended for production and repeat it after material engine changes.

Read the result →