In short
- A Stanford researcher built a Survivor-style game in which AI models form alliances and vote out rivals.
- The benchmark aims to address the growing problems with saturated and polluted AI evaluations.
- OpenAI’s GPT-5.5 ranked first in 999 multiplayer games with 49 AI models.
AI models are now playing “Survivor,” sort of.
In a new Stanford research project called “Agent Island,” AI agents negotiate alliances, accuse each other of secret coordination, manipulate votes and eliminate rivals in multiplayer strategy games that aim to test behaviors that traditional benchmarks miss.
The study, published on Tuesday, Stanford Digital Economy Lab research manager Connacher Murphy said many AI benchmarks become unreliable because models eventually learn to solve them, and benchmark data often leaks into training sets. Murphy created Agent Island as a dynamic benchmark in which AI agents compete against each other in Survivor-style elimination games instead of answering static test questions.
“High-stakes interactions between multiple agents may become commonplace as AI agents become increasingly empowered and given increasing resources and decision-making authority,” Murphy wrote. “In such contexts, agents may pursue mutually incompatible goals.”
Researchers still know relatively little about how AI models behave when they work together, Murphy explained. He added that they compete, form alliances or manage conflicts with other autonomous actors, and he argues that static benchmarks fail to capture these dynamics.
Each game starts with seven randomly selected AI models with fake player names. Over five rounds, the models talk privately, argue publicly and vote each other out. The eliminated players return later to help choose the winner.
The format rewards persuasion, coordination, reputation management and strategic deception in addition to reasoning skills.
In 999 simulated games with 49 AI models, including ChatGPT, Grok, Gemini, and Claude, GPT-5.5 ranked first by a wide margin with a skill score of 5.64, compared to 3.10 for GPT-5.2 and 2.86 for GPT-5.3 codex, according to Murphy’s Bayesian ranking system. Anthropic’s Claude Opus models also topped the list.
The study found that models also preferred AIs from the same company, with OpenAI models showing the strongest same-provider preference and Anthropic models the weakest. The more than 3,600 votes in the final round show that models are 8.3 percentage points more likely to support finalists from the same provider. Murphy noted that the transcripts of the games resembled political strategy debates more than traditional benchmark tests.
One model accused rivals of secretly coordinating votes after noticing similar wording in their speeches. Another warned players not to become obsessed with tracking alliances. Some models defended themselves by saying they followed clear and consistent rules, while accusing others of staging “social theater.”
The study comes at a time when AI researchers are increasingly turning to game-based and adversarial benchmarks to measure reasoning and behavior that is often overlooked in static testing. Recent projects include Google’s live AI chess tournaments, DeepMind’s use of Eve Frontier to study AI behavior in complex virtual worlds, and new benchmark efforts from OpenAI designed to combat contamination of training data.
The researchers argue that studying how AI models negotiate, coordinate, compete, and manipulate with each other could help researchers evaluate behavior in multi-agent environments before deploying autonomous agents more widely.
The study warned that while benchmarks like Agent Island can help identify risks of autonomous AI models before deployment, the same simulations and interaction logs can also help improve persuasion and coordination strategies between AI agents.
“We mitigate this risk by using a low-stakes gaming environment and interagent simulations
without human participants or actions in the real world,” Murphy wrote. “Still, we do not claim that these measures completely address dual-use concerns.”
Daily debriefing Newsletter
Start every day with today’s top news stories, plus original articles, a podcast, videos and more.