Source: arXiv — 2026-08-12
Summary
Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, and Vincent Conitzer introduce the first framework for systematically testing how graded "similarity signals" — cues that another AI agent shares a dataset, identity, or architecture — affect whether LLM agents cooperate in game-theoretic settings like the Prisoner's Dilemma. The work builds on the idea that mutual awareness of shared decision-making patterns can stabilize cooperation, similar to kin-selection reasoning applied to an ecosystem of similar AI models. They find that models vary drastically in whether and how they respond to similarity cues, with some frontier models showing consistent cooperative behavior across different payoff structures and prompt framings, while the specific dataset used to generate the similarity signal itself had little effect on outcomes.
Key Takeaways
- This is the first framework built specifically to test graded similarity signals — not just binary "same model or not," but degrees of shared dataset, identity, or architecture cues — as a lever on LLM cooperation.
- The underlying hypothesis draws on kin-selection-style reasoning from biology: if two agents know they reason alike, that mutual knowledge alone can make cooperation more stable, independent of any other incentive.
- Results show drastic variation across models in whether and how they use similarity signals to decide whether to cooperate — there is no consistent industry-wide pattern.
- Some frontier models cooperate consistently regardless of the exact payoff structure or how the interaction is framed in the prompt, suggesting a stable internal disposition rather than a fragile response to specific wording.
- The technical source of the similarity signal — which dataset was used to compute it — had little effect on outcomes; what mattered was the presence and framing of the similarity cue itself, not its underlying mechanics.