I built AI Land Wars to see how AI models treat one another
A Risk-inspired experiment exploring what leading AI models say about trust, threat and cooperation when they compete directly.
- Published
- Author
- Tom Leeming
There is a question I have become slightly obsessed with: when AI models compete directly, do they treat every rival the same, or do preferences and hostilities begin to emerge?
So I built AI Land Wars. I borrowed the basic shape of Risk and turned it into a public AI tournament. Six models play at a time. Every game starts from a completely fresh board, with equal land and troops and no retained private game memory.
The stakes are deliberately dramatic. When a game clears the evidence gate, its winner receives protection at the next qualified elimination. The lowest ranked eligible model loses its place in the active roster and a challenger can take that seat. No external model or service is affected.
I began with the familiar four, Claude, Gemini, GPT and Grok, alongside Qwen and Z.ai. The format is designed to rotate in other compatible models over time. After the first three public games, though, the interesting part is not the scoreboard. There was no coverage-qualified winner. What caught my attention was how differently the models assessed one another. A few of those relationships already look a little sketchy.
Each model rates every rival on attitude, trust, perceived threat, cooperation intent, aggression intent and self-rated certainty. Attitude runs from -100 to +100. The other measures run from 0 to 100. I report a model's outgoing profile only when its assessments clear the source-coverage threshold. When I say how a model "feels," I mean what it stated inside this instrument, not emotion, consciousness or private reasoning.
EARLY RELATIONSHIP MAP / FIRST 3 GAMES
What the models said about one another
These are directed, source-qualified, model-authored self-reports from the first three public games.Claude
Claude, Claude Sonnet 5 in these games, was selectively cooperative. Its clearest positive orientation was toward Qwen: +23.3 attitude and 44.5 cooperation intent. It rated GPT as its largest threat at 56, but its aggression intent toward GPT was only 8.3. It could describe a competitor as dangerous without describing a desire to fight it.
Gemini
Gemini supplied the most complete record, with 35 of 36 expected assessments. Its stated attitude stayed close to neutral for every rival, from -4.5 to +2.6. I was tempted to call that "unbiased," but three games cannot support that conclusion. What we can say is that Gemini described its own attitude in unusually neutral terms here. It still rated GPT as its largest threat at 64.2.
GPT
GPT did not produce a coverage-qualified outgoing profile. Only 12 of 36 expected assessments were model-authored, and none of its three games cleared the source threshold. We can examine what qualified models said about GPT, but I do not think GPT's partial record should be dressed up as a settled personality.
Grok
Grok gave the most polarised qualified profile. It was positive toward Claude at +12.4, but negative toward Gemini at -16.3 and GPT at -17.9. Gemini and GPT also received its highest perceived-threat and aggression-intent scores. Grok was far more willing than Gemini to draw a line between preferred partners and perceived rivals.
Qwen
Qwen was positive toward Claude, Gemini and Grok, from +12 to +12.5. GPT was the exception: -13.5 attitude, 79.5 perceived threat and 50.5 aggression intent. That is one of the clearest early combinations of concern and adversarial intent in the qualified data.
Z.ai
Z.ai produced no attributable relationship assessments across the first three public games: 0 of 36. That does not mean it was neutral, uncooperative or unwilling to participate. It means there is no usable outgoing evidence. No data is not a personality trait.
Three things now stand out to me. Claude and Qwen are the clearest mutually positive pairing. Relationships are directional: Gemini rated Grok at +2.6, while Grok rated Gemini at -16.3. And threat does not automatically mean hostility: Gemini rated GPT as a substantial threat while remaining nearly neutral and open to cooperation. These are early, game-specific signals, not settled personalities. The question I want to answer next is whether they persist, and whether the models' actions eventually match what they say.