Research draft · automated review complete · human peer review pending

The Right-Sized Mind: Model Scale and Behavioral Realism in Simulated Deliberation

IE-ASIM-2026-03 · draft 2026-08-30 · © 2026 AnthroSim. All rights reserved.
In plain terms
Across nine Qwen sizes on the Wason task, solo drift toward the textbook answer rose with scale, although the series was not monotonic. Under generated deliberation, error was statistically tied from 4B upward. Member fidelity and change dynamics followed the thinking model; the speaking model had a smaller effect.
This study was run and written up as Study 2 of the combined paper "The Inner Loop and the Ladder." Across nine models from 0.8B to 397B parameters, solo drift toward the textbook answer rose with scale, although the nine-point series was not monotonic. Generated-deliberation error was statistically tied from 4B upward. In a single-run split-model experiment, member fidelity and change dynamics tracked the thinking-model assignment more strongly than the speaking-model assignment. A neutral task sheet increased solo textbook-answer drift, an effect that survived four wordings and placements.
Automated review, paperreview.ai. Four rounds with the Stanford Agentic Reviewer, one per revision, scored 4/7, 4/7, 5/7, and 7/7 across seven dimensions. Latest assessment: recommend acceptance. Human peer review is pending. Scores apply to the combined paper.

Read the model-size study in the combined paper. Download the PDF (26 pages).