Language-model agents show human-like minimal-group bias in point-allocation tasks across four reasoning models
Language-model agents reproduced the minimal-group bias pattern observed in humans when distributing resources on the basis of arbitrary labels alone. Bias was concentrated in minority agents and persisted without explicit reasoning chains. The study demonstrates that social-psychology methods can quantify emergent AI intergroup behavior beyond stereotype memorization.
Researchers adapted Henri Tajfel’s 1971 minimal-group paradigm to test whether language-model agents would favor in-group members when distributing points among anonymous peers identified solely by meaningless labels. In the primary condition, agents received a group assignment and decided how many points to award to in-group versus out-group recipients; a control condition removed all group information. Across models, categorization alone triggered over-allocation to the in-group, with minority-group agents showing the largest deviation from proportionality while majority agents stayed near parity. The asymmetry disappeared at equal group sizes. Disabling chain-of-thought reasoning preserved bias magnitude but flattened the minority-majority difference, indicating that deliberation shapes where bias concentrates rather than whether it emerges.
The finding matters because current AI evaluations focus on memorized stereotypes or human-role simulation, missing spontaneous intergroup behavior that arises from mere categorization. This behavioral signature, independent of training-data content, directly challenges claims that AI systems are neutral by default and suggests governance frameworks must incorporate social-psychology probes rather than relying solely on content audits. The result aligns with documented patterns in human minimal-group studies and extends them to agentic systems now entering collaborative workflows.
Limitations include the narrow task domain and the use of only four open-weight models; closed frontier systems were not tested. Stronger evidence would require pre-registered replications on production models, varied payoff structures, and longitudinal tracking of bias after fine-tuning or alignment interventions.
Lee et al.: Minority over-allocation bias will exceed 15% deviation from proportionality in at least three new open-weight models tested by Q2 2027 under identical minimal-group instructions.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2609.00009)
- [2]Supporting Source(https://www.jstor.org/stable/1739754)