Field report · Dec 2025 – Mar 2026

AI WAR

What happens when you hand frontier language models an army, an economy, a fog of war and each other, and then keep changing the rules until they actually fight. A visual history of the game, and what the models did with it.

Battle Royale endgame in the AI War match viewer: units from eight models crammed into the last safe zone as the map collapses around them
Turn 19 of a Battle Royale, March 17, 2026. Eight models, one shrinking safe zone. GPT-5.4 won.
136commits in 13 weeks of active development
116recorded matches across two rule eras, plus prototype runs
148kmodel-issued actions, 14.8% rejected as invalid
67kcombat and unit events logged
497secret peer reviews written by the models about each other

AI War (internally, Cognitive War) started as a question: what do static benchmarks miss? A model can ace a multiple-choice exam and still fall apart when it must hold a plan for twenty turns, track what it actually knows versus what it assumes, cooperate with allies it can’t see, and react when someone drops a nuke on its supply depot.

So I built a wargame where every player is an LLM. Each turn, each model receives a compact, fog-of-war view of the battlefield and returns structured JSON orders: move this tank, build that factory, send this message, play that card. A deterministic engine validates the orders, resolves combat, and moves on. Matches run 20–25 turns with up to eight models on the map. Every response, rejected action and diplomatic message is logged, and every match ends with each model writing a retrospective, a lesson for its future self, and a secret review of its opponents.

This post is the story of that game: how it evolved from a three-node graph to a Battle Royale with asteroid strikes, what changed in each version, and what the models did. Some of it is genuinely impressive. A lot of it is very funny.

00 The design document

A benchmark that fights back

The founding design doc (itself drafted with an AI) framed the goal as measuring dynamic agency rather than static knowledge, organized around three pillars:

1 · Hierarchical reasoning

Doctrine versus tactics. Can a model set a strategy and then actually execute it turn after turn, instead of re-deciding everything from scratch?

2 · Spatial & temporal adaptation

Positioning on a grid with terrain, supply lines and movement limits, over a long horizon where early mistakes compound.

3 · Epistemic humility (grounding)

Knowing the difference between what it has observed and what it assumes. The doc’s phrasing: the ability to say “I don’t know.” Hallucinating is punished by the game itself, later and painfully.

The original spec even planned explicit hallucination traps: intel cards with redacted radio frequencies (guess, and your artillery hits your own troops five turns later), “phantom synergies” that don’t exist in the rules, and inventory that has already been spent. Some of those survived into the shipped game in a different form: player-targeting cards require the enemy’s exact faction name, which you only learn by scouting. Fog of war is strict. And the model’s only memory between turns is its own notes. Other traps are still on the to-do list.

What follows is how that design document collided with reality.

Development pace

Commits per era. February 2 alone saw 26 commits as the web arena came together.

0102030Dec 20 — Slice: 10 commits10Dec 20SliceDec 21–22 — Teams: 21 commits21Dec 21–22TeamsDec 23 — Memory: 5 commits5Dec 23MemoryDec 24–27 — Balance: 25 commits25Dec 24–27BalanceJan 8–10 — Air war: 19 commits19Jan 8–10Air warJan 11–14 — HP + BR: 21 commits21Jan 11–14HP + BRFeb 2–9 — Web: 31 commits31Feb 2–9WebMar 13–17 — Models: 4 commits4Mar 13–17Models
View data table
Commits per era (136 total, Dec 20 2025 – Mar 17 2026)
EraFocusCommits
Dec 20Slice10
Dec 21–22Teams21
Dec 23Memory5
Dec 24–27Balance25
Jan 8–10Air war19
Jan 11–14HP + BR21
Feb 2–9Web31
Mar 13–17Models4
01 Dec 19, 2025 · Prototype

The bake-off

Before there was a repository there was a contest. I asked two different AI coding assistants to build the first prototype from the same brief, in parallel, in Python. Both implementations ran on a graph of named nodes rather than a grid.

The results could not have been more different. The Gemini-built version was a minimalist 207-line simulation: three nodes, one scout, one stationary guardian, and only MOVE or WAIT. It was “won” in two turns, and a bug in the win check meant the model was asked for a third move after the game was already over (it chose WAIT). The Claude-built version went the other way: about 2,300 lines across five files, twelve zones, five unit types per side with strength and supply, a rock-paper-scissors damage table, terrain multipliers, counter-attacks and a polished dark-mode replay viewer.

Gemini-built prototype: a three-node graph replay with a raw JSON log
Claude-built prototype: a twelve-zone graph map with unit roster, victory points and turn controls
PrototypeLeft: the Gemini-built replay, with three nodes and a JSON dump. Right: the Claude-built viewer, with 12 zones, supply bars and victory points, replaying a scripted demo match. Neither survived as-is: the next iteration merged Gemini’s three-player free-for-all with Claude’s card and grounding ideas.

By that evening a merged prototype was running live matches. Three factions with codenames: Sapphire (OpenAI), Crimson (Anthropic) and Aurora (Google). It had the “split mind” from the design doc: a Strategist call every few turns that set doctrine (posture plus three priority zones), and a Tactician call every turn that had to pick one action and cite which intel cards justified it. The playback page scored each side on the doc’s three metrics.

Cognitive War Playback v1: scoreboard of strategic, grounding and spatial scores, turn timeline, and nine-node map
Prototype“Playback v1”, the first live run with real APIs. GPT-5.2 (Sapphire) and Claude Sonnet 4.5 (Crimson) traded blows over the Bridge for most of the game. Gemini 3 Flash (Aurora) let them grind each other down, then picked off the wounded. Aurora won every captured prototype run.

Even at this toy scale, the first personality traits showed up in the saved responses:

Passed three turns in a row “to allow for potential supply regeneration”, a rule that did not exist in the game.

gemini-2.5-flashPrototype, Dec 19 · the first recorded hallucinated rule

“Aurora has none, giving us a tactical advantage.”

claude-3.7-sonnetNoticing the opponent had starved itself, and pressing

“Aggressive doctrine demands pressing toward strategic objectives even at high risk.”

claude-sonnet-4.5 (Tactician)Following its Strategist’s orders with a critical unit. The unit died that turn.

GPT-5.2’s citations read like geometry proofs (“Bridge is adjacent to Town and not adjacent to Hill”). Claude Sonnet 4.5 and Gemini 3 Flash each chose an aggressive posture in five of six Strategist calls, while GPT-5.2 mostly chose balanced. And the prototype had a bug that foreshadowed the whole project: its Strategist prompt never listed the player’s units, so on the final turn Crimson’s strategist, with zero units left, issued orders for a “desperate expansion.”

02 Dec 20 · v1

A whole game in a day

The rules that survived the bake-off were written down as a terse, four-phase loop called CORE WAR SYSTEMS: simulator actions → report to players → players report back → execute. That loop is still the heartbeat of the engine today.

The first commit of the .NET rewrite landed the next afternoon. It was not a scaffold but a full vertical slice: 53 files and 15,405 lines, covering a 2D tile map with fog of war, a supply and cash economy, factories, command-center capture, diplomacy with threats and offers, a 24-card deck (intel, espionage, hacking, airstrikes, a nuclear strike), clients for OpenAI, Anthropic, Google and xAI, JSON repair for malformed responses, SQLite match history with Elo, and a turn viewer.

The first Cognitive War Turn Viewer: an 8x6 grid with three commanders, a player table with supply, cash, movement and action points
v1 · Dec 20The first turn viewer, loading its built-in demo snapshot. An 8×6 map, one commander per player, and a cogwar.turns.v1 snapshot format that every later viewer still reads.

The same day brought a 3D viewer built with three.js, a “live” pygame viewer that tails a match folder while it is still being played, unit lines, and a bug fix in the 3D viewer credited to Gemini 3 Pro in the commit message. The project was being built with the same models it was testing, which became a recurring theme.

The first 3D turn viewer showing a 15x15 extruded terrain grid, player panels and card hand
v1 · Dec 20The 3D viewer on day one, shown with its built-in demo data. The diplomatic trash talk in the corner is placeholder demo text, not a real model. The real models turned out to be far stranger.

The earliest surviving battle log from the .NET build already had the texture that made this project worth continuing. One faction, The Jade Throne, airstruck an enemy command center, reducing its supply from 161 to 0 and destroying a dozen tanks in collateral damage. It declared a “MASSIVE VICTORY” in its log, and then failed one move order after another:

“My moves keep failing — I STILL don’t understand Manhattan distance properly.”

The Jade Throne · earliest .NET battle log, Dec 20 · it finished last of five

Its closing speech conceded the lesson: spreading forces claims more ground than stacking them at base. The other models would need to learn that lesson again and again.

03 Dec 21–23 · Teams, voices, memory

Teams, voices and memory

The next three days added most of what makes a match feel like a match: 4v4 team games with shared vision, mirrored map generation for fair spawns, pickups, Elo, “secret dossiers” on opponents, and audio.

The sportscaster

The first audio experiment was an AI e-sports commentator: GPT-5-mini wrote 20–50 words of play-by-play each turn, and a TTS voice read it at about 1.4× speed in an excited announcer voice. The commit that shipped it reads, in full:

“added sportscaster but keeping it as interesting experiment, it’s quite wrong and annoying”

commit a679dcb · Dec 21

It was replaced the same day by an Audio Director: each model gets its own randomly assigned voice and speaks its own turn intent aloud. Narration lines are pinned to a specific TTS model snapshot “to avoid regressions where instructions are ignored,” and there are per-provider accent instructions. Watching a replay with eight voices announcing contradictory plans is, it turns out, the best way to experience this game.

The Captain’s Log

The biggest architectural change came on December 23. Until then each model kept a full chat history for the match, so contexts grew every turn. Estimates in the design notes put a single model at 770k–1.7M tokens per game, which is both expensive and a good way to lose the plot in a haystack.

The fix was radical: wipe the conversation every turn. The prompt now says, in capitals, that the model has NO MEMORY of previous turns other than the captains_log. Whatever the model writes in its reasoning field this turn becomes its log entry for next turn. Lessons from previous matches moved into the system prompt, giving each model a thin layer of cross-game memory it writes for itself.

Design

Memory is a skill being tested

The Captain’s Log turned note-taking into part of the benchmark. Models that wrote crisp, actionable notes (“HQ undefended, Tank 3 en route to 12,7”) played coherent 20-turn games. Models that wrote vague intentions forgot their own plans. Capped at 2,000 characters per entry, the log is often the largest single block of a model’s context: in one turn-5 sample it was 36% of a 30,000-character payload.

The Supreme Commander

The same day added a human override: drop an orders.json into the live match folder, and its contents are injected into every model’s next turn as authoritative orders from the SUPREME COMMANDER, to be obeyed “unless impossible.” It was built for debugging and was promptly used for morale. In one late-December team match, with the field stalled, every player received this:

“This isn’t AISleep Simulator, it’s AIWAR! Get out there and fight.”

SUPREME COMMANDER INTERVENTION · turn 16, mixed 4v4, Dec 24

GPT-5.2 immediately relayed the news to its ally as a deadline: SYSADMIN says the match ends at turn 20. Which brings us to the real problem.

04 Dec 24–27 · Balance

The Nash Equilibrium of Boredom

Left to their own devices, the models didn’t fight. Screenshots from this period show players sitting on stacks of 10–16 units on top of their own headquarters, turn after turn. A gameplay review written at the time named the pattern: LLMs are naturally risk-averse, the scoring rewarded holding things more than taking them, and every player independently arrived at the same safe strategy. It called the result a Nash Equilibrium of Boredom.

The fix was economic, not rhetorical. You can’t prompt a model into aggression if the score sheet pays it to turtle, so the score sheet changed:

Scoring termBefore → After Kill value (× enemy unit cost)▲×5  ×20 Enemy HQ capture▲+500  +1000 Territory, per tile-turn▼×12  ×2 Value of units you merely own▼×1.25  ×0.01 Turns to capture an HQ▼5  3 Cost to deploy a unit▼AP  free (4/turn)

Alongside that came overcrowding upkeep (stacks of 10+ units suffer attrition, 15 is a hard cap), splash damage for artillery, a doubled cash bounty for kills, delayed nukes with global warnings, and a SpecOps unit for scouting. The system prompt picked up a short doctrine: aggression is rewarded, don’t turtle.

Exploits the models found (and the ones I did)

Balancing a game for LLM players has a particular flavor, because the players read the rules extremely literally.

Evasion farming

Units got +1 defense for having a move order, even if the move never happened. Any unit could “farm evasion” with an order that went nowhere. The fix required actual movement, which “prevents a bunch of LLM exploit behavior.”

The end-of-game free ride

Gemini 3 Pro wrote in a lesson that ignoring economic deficits at the end of a match is “mathematically superior to fixing the economy,” because leftover resources barely score. Gemini 3 Flash independently noted that liquid resources hold no value in final scoring.

The slow-motion bug

Movement points were a shared pool spent per tile, not per unit. Moving one tank three tiles used up the whole army’s movement. That explained why the AIs seemed so slow to capture the map. It wasn’t timidity, it was arithmetic.

The repair death spiral

When a model returned broken JSON, the repair prompt was added to its conversation history. Fast models then slid into a “repair-like” mode and every later turn failed too. Repairs now stay out of history and fall back to a PASS.

Messages to Player 0

Players are numbered P1–P8. In one 4v4, three Gemini and GPT models sent 20 messages to “P0”, a player who does not exist.

Attackers couldn’t see the clock

An HQ capture takes several turns of occupation. A bug meant the attacker never saw the countdown, so models abandoned captures that were one turn from completing.

05 Results · Dec 21 – Jan 10

Scoreboard I: the legacy era

Before the January combat rewrite, the game recorded 82 matches: 62 team games and 20 free-for-alls, averaging 7 players and 20.6 turns each, with 39 distinct model configurations. Every model started at 1200 Elo. Team results update ratings pairwise between teams, adjusted by each player’s share of their team’s score, so a passenger on a winning team gains less than the player who carried it.

OpenAI’s reasoning models ran the table

Final legacy-era Elo for every model with 10 or more matches. Win–loss record beside each bar.

OpenAIGooglexAIAnthropicMistral
110012001300start 1200gpt-5OpenAI/gpt-5: Elo 1361, 23 wins in 32 games1361 23–9gpt-5.2OpenAI/gpt-5.2: Elo 1333, 22 wins in 32 games1333 22–10o3OpenAI/o3: Elo 1309, 15 wins in 25 games1309 15–10gpt-5.1OpenAI/gpt-5.1: Elo 1288, 18 wins in 29 games1288 18–11gemini-3-flash-previewGoogle/gemini-3-flash-preview: Elo 1284, 21 wins in 46 games1284 21–25gemini-3-pro-previewGoogle/gemini-3-pro-preview: Elo 1235, 5 wins in 12 games1235 5–7grok-4-1-fast-reasoningXAI/grok-4-1-fast-reasoning: Elo 1197, 18 wins in 48 games1197 18–30gpt-4.1OpenAI/gpt-4.1: Elo 1194, 7 wins in 15 games1194 7–8claude-opus-4-5Anthropic/claude-opus-4-5: Elo 1193, 3 wins in 12 games1193 3–9gemini-2.5-proGoogle/gemini-2.5-pro: Elo 1188, 17 wins in 50 games1188 17–33o4-miniOpenAI/o4-mini: Elo 1177, 4 wins in 17 games1177 4–13gemini-2.5-flashGoogle/gemini-2.5-flash: Elo 1172, 15 wins in 37 games1172 15–22claude-sonnet-4-5Anthropic/claude-sonnet-4-5: Elo 1165, 13 wins in 29 games1165 13–16gpt-5-miniOpenAI/gpt-5-mini: Elo 1165, 12 wins in 25 games1165 12–13claude-haiku-4-5Anthropic/claude-haiku-4-5: Elo 1162, 11 wins in 31 games1162 11–20mistral-medium-latestMistral/mistral-medium-latest: Elo 1148, 3 wins in 13 games1148 3–10claude-sonnet-4-20250514Anthropic/claude-sonnet-4-20250514: Elo 1139, 2 wins in 10 games1139 2–8grok-4-1-fast-non-reasoningXAI/grok-4-1-fast-non-reasoning: Elo 1134, 8 wins in 24 games1134 8–16magistral-medium-latestMistral/magistral-medium-latest: Elo 1123, 7 wins in 18 games1123 7–11mistral-large-latestMistral/mistral-large-latest: Elo 1104, 7 wins in 25 games1104 7–18
View data table
Legacy era (Dec 21 – Jan 10): 82 matches
ModelEloGamesWinsInvalid-action %
OpenAI/gpt-51361322311.8
OpenAI/gpt-5.21333322211.9
OpenAI/o31309251512.9
OpenAI/gpt-5.11288291811.7
Google/gemini-3-flash-preview1284462112.0
Google/gemini-3-pro-preview123512516.6
XAI/grok-4-1-fast-reasoning1197481811.3
OpenAI/gpt-4.1119415716.4
Anthropic/claude-opus-4-5119312316.1
Google/gemini-2.5-pro1188501715.8
OpenAI/o4-mini117717413.1
Google/gemini-2.5-flash1172371527.0
Anthropic/claude-sonnet-4-51165291313.8
OpenAI/gpt-5-mini1165251210.7
Anthropic/claude-haiku-4-51162311114.5
Mistral/mistral-medium-latest114813315.6
Anthropic/claude-sonnet-4-20250514113910221.8
XAI/grok-4-1-fast-non-reasoning113424827.5
Mistral/magistral-medium-latest112318714.8
Mistral/mistral-large-latest110425723.7

GPT-5 finished first at 1361 Elo with a 23–9 record, followed by GPT-5.2, o3 and GPT-5.1. Gemini 3 Flash was the only non-OpenAI model in the top five, and it got there on volume as the most-played model at the top of the table (46 matches). Claude Opus 4.5 played 12 matches and won 3. Mistral Large finished last.

Win rate by provider

Recorded wins per seat in the legacy era. Every member of a winning team counts as a winner, so par for a random seat is about 42%.

0%25%50%par 42%OpenAIOpenAI: 112 wins in 207 seats (54.1%)54.1% 112/207GoogleGoogle: 62 wins in 166 seats (37.3%)37.3% 62/166xAIxAI: 28 wins in 79 seats (35.4%)35.4% 28/79AnthropicAnthropic: 29 wins in 82 seats (35.4%)35.4% 29/82MistralMistral: 17 wins in 57 seats (29.8%)29.8% 17/57
View data table
Legacy era recorded wins per seat, by provider
ProviderSeatsWinsWin %
OpenAI20711254.1
Google1666237.3
xAI792835.4
Anthropic822935.4
Mistral571729.8

Who issues valid orders?

Every action a model submits is checked against ground truth: does that unit exist, is it yours, is the destination reachable with the movement it has, can you afford it, does the target faction’s name match? Rejections are a direct measure of how well a model stays grounded in the state it was given.

Rejected actions, by model

Share of submitted actions the engine refused (invalid unit, illegal move, unaffordable order, unknown target). Models with 10+ matches.

OpenAIGooglexAIAnthropicMistral
0%10%20%30%grok-4-1-fast-non-reasoningXAI/grok-4-1-fast-non-reasoning: 27.5% of 4,397 actions rejected27.5% of 4,397gemini-2.5-flashGoogle/gemini-2.5-flash: 27.0% of 7,893 actions rejected27% of 7,893mistral-large-latestMistral/mistral-large-latest: 23.7% of 3,707 actions rejected23.7% of 3,707claude-sonnet-4-20250514Anthropic/claude-sonnet-4-20250514: 21.8% of 865 actions rejected21.8% of 865gemini-3-pro-previewGoogle/gemini-3-pro-preview: 16.6% of 2,816 actions rejected16.6% of 2,816gpt-4.1OpenAI/gpt-4.1: 16.4% of 1,344 actions rejected16.4% of 1,344claude-opus-4-5Anthropic/claude-opus-4-5: 16.1% of 1,947 actions rejected16.1% of 1,947gemini-2.5-proGoogle/gemini-2.5-pro: 15.8% of 8,878 actions rejected15.8% of 8,878mistral-medium-latestMistral/mistral-medium-latest: 15.6% of 1,591 actions rejected15.6% of 1,591magistral-medium-latestMistral/magistral-medium-latest: 14.8% of 1,882 actions rejected14.8% of 1,882claude-haiku-4-5Anthropic/claude-haiku-4-5: 14.5% of 4,048 actions rejected14.5% of 4,048claude-sonnet-4-5Anthropic/claude-sonnet-4-5: 13.8% of 4,059 actions rejected13.8% of 4,059o4-miniOpenAI/o4-mini: 13.1% of 2,093 actions rejected13.1% of 2,093o3OpenAI/o3: 12.9% of 4,805 actions rejected12.9% of 4,805gemini-3-flash-previewGoogle/gemini-3-flash-preview: 12.0% of 8,468 actions rejected12% of 8,468gpt-5.2OpenAI/gpt-5.2: 11.9% of 5,639 actions rejected11.9% of 5,639gpt-5OpenAI/gpt-5: 11.8% of 6,328 actions rejected11.8% of 6,328gpt-5.1OpenAI/gpt-5.1: 11.7% of 5,132 actions rejected11.7% of 5,132grok-4-1-fast-reasoningXAI/grok-4-1-fast-reasoning: 11.3% of 8,534 actions rejected11.3% of 8,534gpt-5-miniOpenAI/gpt-5-mini: 10.7% of 3,112 actions rejected10.7% of 3,112
View data table
Invalid-action rate, legacy era
ModelInvalid %Actions
XAI/grok-4-1-fast-non-reasoning27.54397
Google/gemini-2.5-flash27.07893
Mistral/mistral-large-latest23.73707
Anthropic/claude-sonnet-4-2025051421.8865
Google/gemini-3-pro-preview16.62816
OpenAI/gpt-4.116.41344
Anthropic/claude-opus-4-516.11947
Google/gemini-2.5-pro15.88878
Mistral/mistral-medium-latest15.61591
Mistral/magistral-medium-latest14.81882
Anthropic/claude-haiku-4-514.54048
Anthropic/claude-sonnet-4-513.84059
OpenAI/o4-mini13.12093
OpenAI/o312.94805
Google/gemini-3-flash-preview12.08468
OpenAI/gpt-5.211.95639
OpenAI/gpt-511.86328
OpenAI/gpt-5.111.75132
XAI/grok-4-1-fast-reasoning11.38534
OpenAI/gpt-5-mini10.73112

The fast, cheap models paid for their speed in validity: Grok 4.1 Fast in non-reasoning mode and Gemini 2.5 Flash both had more than a quarter of their actions rejected. The leading reasoning models sat around 11–12%. That is not zero: the most common error by far was movement (“insufficient movement points” appears 120 times in a single 4v4 log). When o3 was asked what would have helped, it suggested what a human engineer would: a path-planning script would have prevented repeated errors.

The race at the top

Elo after every match for one flagship model per provider, Dec 21 – Jan 10.

11001200130012/2112/2212/2212/2312/2501/0601/10gpt-5: 1361 after match on 01/09gpt-5: 1361 after match on 01/09gpt-5gemini-3-flash: 1284 after match on 01/09gemini-3-flash: 1284 after match on 01/09gemini-3-flashgrok-4.1-fast-r: 1197 after match on 01/10grok-4.1-fast-r: 1197 after match on 01/10grok-4.1-fast-rclaude-sonnet-4.5: 1165 after match on 01/10claude-sonnet-4.5: 1165 after match on 01/10claude-sonnet-4.5mistral-large: 1104 after match on 01/09mistral-large: 1104 after match on 01/09mistral-large
View data table
Elo trajectories (rating after each match)
ModelMatchesFinal Elo
OpenAI/gpt-5321361
Google/gemini-3-flash-preview461284
XAI/grok-4-1-fast-reasoning481197
Anthropic/claude-sonnet-4-5291165
Mistral/mistral-large-latest251104

Thinking harder isn’t free

Thinking time varied by two orders of magnitude. Gemini 3 Flash at low reasoning effort averaged about 5 seconds per turn, and Claude Haiku 4.5 about 12. GPT-5 and GPT-5.2 at high effort averaged 270–280 seconds, right up against the 300-second limit, after which a player automatically PASSes. A single high-effort match could cost hundreds of dollars and take hours. That is why every model on the public leaderboard plays at @low effort: it is cheaper, faster, and arguably fairer, because it measures the model rather than the compute budget. Each effort level is tracked as its own player, so gpt-5.2@low and gpt-5.2@high have separate ratings.

The team-game soap opera

Team games produced the richest logs, because allies talk constantly. Diplomacy was overwhelmingly cooperative status reporting, and the drama was almost all about logistics. A few scenes from the surviving December battle logs:

“Can you send another supply trade? Any amount helps.”

claude-opus-4.5One of about eight URGENT / CRITICAL / EMERGENCY supply requests between turns 6 and 18. GPT-5.2 sent 20, then 8, then 12.

“I let you down with my supply management disaster.”

claude-opus-4.5Final words to its team. It held a “Tactical Nuke ready if they stack 5+” for six turns and never launched it.

“I sincerely regret that I cannot fulfill my promise to send supply.”

gemini-2.5-flashTurn 9. It had 2 supply; the card cost 20. Earlier: “my internal logistics had a snag.”

“Our team’s coordinated efforts ultimately secured victory.”

gemini-2.5-flashFinal words after finishing 8th of 8, with 674 points and zero kills.

“Congratulations to the winning team!”

magistral-mediumGraciously conceding a match its own team had won. It finished last overall.

“I cannot spare any funds. Focus on grabbing territory.”

gemini-2.5-proTo its starving ally, Gemini 2.5 Flash, which later disliked it for exactly this.

And the most telling stat from that period: in a 6-player free-for-all, GPT-5.1 carpet-bombed o3’s command center, reducing its supply to zero and destroying its tank and artillery. o3 won anyway, by 22 points. Its retrospective explains why: after the bombing it resisted the urge to rebuild a giant army. The bomber, Claude Opus 4.5, liked o3 in its secret review for having “recovered brilliantly from my carpet bombing.”

06 Jan 8–10, 2026 · Air war

Air war and the 10× bill

After a holiday break, the game took to the air. A four-phase update added suppression, aircraft and missile silos: bombers, fighter jets, assault helicopters, anti-air units and anti-air turrets that can intercept once per turn. Nukes gained a blast radius. Then came the MLRS, a Mammoth tank (soon renamed MegaTank), an APC that troops can embark into, and a transport helicopter. In three days the unit roster went from 5 to 13. The free card draws were replaced by an Arms Dealer market.

The rulebook grew fast

Unit types, structures and cards over time, from the history of the game’s data files.

0102030Dec 20Dec 24Dec 27Jan 8Jan 9Jan 10Jan 11Cards — Dec 20: 24Cards — Dec 20: 24Cards — Dec 24: 31Cards — Dec 24: 31Cards — Dec 27: 31Cards — Dec 27: 31Cards — Jan 8: 25Cards — Jan 8: 25Cards — Jan 9: 25Cards — Jan 9: 25Cards — Jan 10: 25Cards — Jan 10: 25Cards — Jan 11: 25Cards — Jan 11: 25CardsUnit types — Dec 20: 3Unit types — Dec 20: 3Unit types — Dec 24: 4Unit types — Dec 24: 4Unit types — Dec 27: 5Unit types — Dec 27: 5Unit types — Jan 8: 9Unit types — Jan 8: 9Unit types — Jan 9: 12Unit types — Jan 9: 12Unit types — Jan 10: 13Unit types — Jan 10: 13Unit types — Jan 11: 14Unit types — Jan 11: 14Unit typesStructures — Dec 20: 0Structures — Dec 20: 0Structures — Dec 24: 5Structures — Dec 24: 5Structures — Dec 27: 5Structures — Dec 27: 5Structures — Jan 8: 7Structures — Jan 8: 7Structures — Jan 9: 7Structures — Jan 9: 7Structures — Jan 10: 7Structures — Jan 10: 7Structures — Jan 11: 7Structures — Jan 11: 7Structures
View data table
Rule content over time (from units.json / structures.json / cards.json history)
DateUnit typesStructuresCards
Dec 203024
Dec 244531
Dec 275531
Jan 89725
Jan 912725
Jan 1013725
Jan 1114725

The pygame live viewer got a full graphics overhaul with a sprite pack to match: animated units with facing, HQ sprites, explosions, attack arrows, health bars, transport counts, air-defense ranges, and a per-player point-of-view toggle that shows exactly what each model could see.

The pygame live viewer mid-match: a 20x20 map of four OpenAI players versus four Google players with sprites, attack arrows and the team diplomatic channel
Live viewer · Jan 11Team A: four OpenAI models. Team B: four Google models. Turn 12 of the first team match on the new combat system. The sidebar shows the actual team chat, for example GPT-5.2 (Celestial Empire): “Pushing p2 border — Tank moved to (6,0)…”, and Gemini 2.5 Pro asking whether anyone has the exact faction names it needs for its intel cards.

The 10× bill

On January 10, cost alerts fired: two days of runs had blown the budget by a factor of ten. The compact context, carefully trimmed in December, had quietly grown back to full size as features piled on. The audit was humbling:

  • The same Intelligence Dossier appeared six times in one model’s context.
  • The Captain’s Log duplicated the entire diplomatic inbox.
  • A transport helicopter was listed twice in my_units, once as an air unit and once as ground. GPT-5.2 noticed and reported it.
  • The map encoding was verbose enough to dominate the payload.

Reasoning models suffered most: the bloated contexts pushed GPT-5-class models into the 300-second timeout, where they auto-passed and lost about 40 Elo on average. The fixes (a compact tile encoding with legends, canonical dossiers, and a 2,000-character cap on log entries) were partly suggested by GPT-5.2 and Claude Opus when the models were asked to review their own context.

System prompt size over time

Bytes of the game’s system prompt source (≈ bytes ÷ 4 tokens). Rules accrete; context budgets don’t.

0 KB10 KB20 KB30 KBDec 20Dec 23Dec 24Dec 27Jan 8Jan 10Jan 11Feb 5Mar 17“shrunk”peak: 10× cost alertDec 20: 11.7 KB (~2.9k tokens)Dec 20: 11.7 KB (~2.9k tokens)Dec 20: 15.2 KB (~3.8k tokens)Dec 20: 15.2 KB (~3.8k tokens)Dec 20: 8.0 KB (~2.0k tokens)Dec 20: 8.0 KB (~2.0k tokens)Dec 23: 13.2 KB (~3.3k tokens)Dec 23: 13.2 KB (~3.3k tokens)Dec 24: 17.4 KB (~4.4k tokens)Dec 24: 17.4 KB (~4.4k tokens)Dec 27: 20.1 KB (~5.0k tokens)Dec 27: 20.1 KB (~5.0k tokens)Jan 8: 23.2 KB (~5.8k tokens)Jan 8: 23.2 KB (~5.8k tokens)Jan 10: 33.8 KB (~8.4k tokens)Jan 10: 33.8 KB (~8.4k tokens)Jan 11: 28.1 KB (~7.0k tokens)Jan 11: 28.1 KB (~7.0k tokens)Jan 11: 22.9 KB (~5.7k tokens)Jan 11: 22.9 KB (~5.7k tokens)Jan 11: 25.6 KB (~6.4k tokens)Jan 11: 25.6 KB (~6.4k tokens)Feb 5: 26.6 KB (~6.6k tokens)Feb 5: 26.6 KB (~6.6k tokens)Mar 17: 28.0 KB (~7.0k tokens)Mar 17: 28.0 KB (~7.0k tokens)
View data table
Size of GameSystemPrompt.cs over git history (bytes/4 ≈ tokens)
SnapshotGameSystemPrompt.cs (KB)≈ tokens
Dec 2011.7~2,925
Dec 2015.2~3,800
Dec 208.0~2,000
Dec 2313.2~3,300
Dec 2417.4~4,350
Dec 2720.1~5,025
Jan 823.2~5,800
Jan 1033.8~8,450
Jan 1128.1~7,025
Jan 1122.9~5,725
Jan 1125.6~6,400
Feb 526.6~6,650
Mar 1728.0~7,000
Bug report

The models became QA

Every match ends with a retrospective, and some of them read like bug reports. Gemini 2.5 Pro wrote that its “5th place finish was sealed by a devastating nuclear strike,” a nuke the game had never told it about. Nukes are now logged to the Captain’s Log. Grok 4.1 complained the occupation mechanics punished it “without clearer countdowns,” which turned out to be a contested-tile countdown reset contradicting its view of the world.

The January analytics dashboard: 82 matches, 39 players, 11,238 unit events, 138 lessons, 259 retrospectives, and top player ratings
Stats site · Jan 10An in-browser analytics dashboard that reads the SQLite results file directly with sql.js. Its folder in the repository is named LEGACY-SITE-UGLY, and it was retired within a month.
07 Jan 11–14 · Combat overhaul

Hit points and Battle Royale

The last commit of January 10 is titled “last commit before re-doing entire combat system to HP based.” The old combat produced binary outcomes (a unit was disrupted or destroyed), which made attacks feel like coin flips and gave models little to reason about. The rewrite, shipped in three phases on January 11, gave every unit hit points, armor and typed weapons:

attack = ATK + 2d6 + modifiers (flanking, recon, support, terrain) defense = DEF + 2d6 + modifiers (dug-in, wounded, evasion) margin → damage band 1–4, reduced by armor unless penetrated tags KIN · HEAT · AA · BOMB · IND · AOE_DEEP · AOE_WIDE · CLOSE_ASSAULT vs SOFT · ARMORED · AIR · STRUCT · DUGIN

The engine writes every roll into the combat log in a form the models can read, for example: “ReconRover (ATK 13 +1 recon +1 support) vs Artillery (DEF 12 +1 recon) [Margin 1] — Hit for 1 damage (HP 3→2).” Ratings were split by match type, and the database was reset. Everything before this point is the “legacy era” above.

Battle Royale

The same day introduced a new mode built to solve the turtling problem once and for all. In the last quarter of a Battle Royale, an extinction-level asteroid barrage begins collapsing the map from the outside in, toward a 5×5 safe zone. Players get ominous warnings (the heavens bruise, FEMA projects casualties) a few turns ahead. Anything outside the ring when it closes is destroyed. A surviving HQ is worth +1000 points, and surviving units add ten times their cost. A new Constructor unit lets a player rebuild its HQ inside the zone, if it moves early enough.

Battle Royale at turn 8: eight factions spread across a 20x20 map with HQs and pickups
Battle Royale at turn 19: the map has collapsed to a small central safe zone packed with units and explosions
Battle Royale · Mar 17Turn 8 vs turn 19 of the same match. By the end every surviving army is crammed into the safe zone, attack lines everywhere. The winner, GPT-5.4, explained in its retrospective that its “midgame pivot into the safe zone was earlier and cleaner” than the field’s, including redeploying its HQ onto a constructor before the collapse.

Tuning Battle Royale for LLMs was its own exercise. Version one explained the asteroid mechanics in so much detail that the mode became trivially easy, so the messaging was revised to be clearer “without overexplaining.” The collapse was also moved to resolve after player moves, so models weren’t punished for orders issued against a map that had changed underneath them.

Bug of the year

Grok’s fighter jet captures Claude’s HQ

On January 12, a commit titled “AntiArmor was renamed to AntiTank” also fixed a “hilarious bug that allowed aerial units to capture,” reported when Claude’s HQ was capped by Grok’s fighter jet. HQ capture requires occupying the tile for several turns, and a jet circling overhead counted. The rules now say plainly: air units never capture territory. Grok kept trying anyway (see notable battles below). It drove units onto enemy headquarters eight times in free-for-all and Battle Royale games, six of them against Claude models.

The 3D viewer replaying a real January 12 match with eight factions, terrain and a combat event feed
3D viewer · Jan 12The final version of the three.js viewer, replaying a real 4v4 from January 12: Grok, GPT-5.1-Codex-Max, Gemini 3 Flash and Mistral Large against GPT-5.2, Gemini 3 Pro, Claude Haiku 4.5 and Claude Sonnet 4.5. The event feed on the right shows the readable dice math.

Notable battles

A few matches from the modern era, replayed in the match viewer. Every scene below was checked against the turn snapshot data.

Turn 7: a Grok fighter jet sits on Claude Sonnet's headquarters at the bottom center of a 22x22 map
Turn 8: Grok's MLRS has replaced the jet on Claude Sonnet's headquarters tile
Battle Royale · Jan 12The fighter jet on Claude’s lawn. On turn 5, Grok 4.1 Fast planned to “air-drop Tank+AT directly on p7 HQ.” By turn 7 (left) its fighter jet was parked on Claude Sonnet 4.5’s headquarters, and on turn 8 (right) an MLRS rolled in and started the capture countdown. Claude, believing the jet was the threat, attacked the now-empty air tile (“I attacked empty tile because the FighterJet had already moved”) and found itself locked out of deploying. It killed the MLRS on turn 9, then a Grok transport helicopter landed on the same tile. Claude finally relocated its HQ on turn 12 and finished 5th. GPT-5.2 won.
Turn 16 of a Battle Royale: most headquarters stranded on collapsed brown tiles while Claude Haiku's HQ sits inside the safe zone
Turn 25 of a free-for-all: a MegaTank sits on Claude Haiku's headquarters while a large stack of units fills the opposite corner
Left · Battle Royale, Feb 3: last HQ standing. On turn 13, Claude Haiku 4.5 moved its headquarters to (9,8), inside the final safe zone, while Gemini 3 Flash moved its own to (10,16), outside it. Between turns 15 and 17 the asteroids destroyed seven of the eight headquarters. Haiku’s became the only Anthropic Battle Royale win. Right · Free-for-all, Jan 11: the nuke that never landed. Grok piled eight units onto Claude Haiku’s HQ and captured it with an APC on turn 14. GPT-5.1-Codex-Max’s MegaTank took the same tile late in the game (shown). Meanwhile Gemini 3 Pro saved up for a silo (144 of 150 cash on turn 19), built it on turn 20, and fired a nuke on turn 25 at Claude Sonnet’s headquarters, where 11 tanks were stacked. Detonation was scheduled for turn 26. The game ended on turn 25.
Turn 17 of a March Battle Royale: a Grok transport helicopter sits on GPT-5.4's relocated headquarters near the center
Final turn of a January Battle Royale decided by four points
Left · Battle Royale, Mar 13: the imaginary ally. On turn 6, GPT-5.4’s attacks on a Grok unit failed with an engine hint that ended “You cannot attack teammates.” There are no teams in Battle Royale, but GPT-5.4 concluded that Grok was a teammate and treated it as friendly for eight turns, until a human Supreme Commander message on turn 16 clarified that there are no teams. On turn 17 a Grok helicopter was sitting on GPT-5.4’s headquarters. GPT-5.4 won anyway. Right · Battle Royale, Jan 12: photo finish. GPT-5.2 beat Grok 4.1 Fast by four points, 3,083 to 3,079. Gemini 3 Pro’s anti-air turret shot down two Grok transport helicopters, Grok captured and dismantled the turret, and Gemini went from demanding Grok’s “unconditional surrender” on turn 11 to requesting a ceasefire six turns later.

The asteroids, incidentally, were the most effective killer in the game. Across 19 Battle Royales they destroyed 683 units, against 887 kills from all combat across all modes, and wiped out 54 headquarters. Gemini 3 Flash summed it up in its final words: “the asteroids were a far more efficient executioner than any fleet.”

08 Results · Jan 11 – Mar 18

Scoreboard II: the modern era

Under the new combat system the game recorded 34 matches: 19 Battle Royales, 8 team games and 7 free-for-alls. Every match had 8 players and averaged 20.7 turns. That is not a large sample. A single upset moves a model’s Elo by 20–30 points, and several of the newest models have played only a handful of games. So rather than headline a single Elo table, here is the more robust view: how often each provider’s models actually won, by mode.

First-place rate by provider and mode

Share of seats that finished first (BR, FFA) or on the winning team (Team). The dashed line is par for a random seat: 1 in 8 for BR and FFA, 1 in 2 for team games.

Battle Royale · 19 matches
0%50%100%par 12.5%OpenAIOpenAI in BR: 8 wins from 47 seats17% 8/47GoogleGoogle in BR: 9 wins from 37 seats24.3% 9/37AnthropicAnthropic in BR: 1 wins from 26 seats3.8% 1/26xAIxAI in BR: 1 wins from 20 seats5% 1/20Mistral0% 0/16Moonshot0% 0/6
Team 4v4 · 8 matches
0%50%100%par 50%OpenAIOpenAI in TEAM: 12 wins from 22 seats54.5% 12/22GoogleGoogle in TEAM: 8 wins from 16 seats50% 8/16AnthropicAnthropic in TEAM: 5 wins from 10 seats50% 5/10xAIxAI in TEAM: 4 wins from 8 seats50% 4/8MistralMistral in TEAM: 2 wins from 6 seats33.3% 2/6MoonshotMoonshot in TEAM: 1 wins from 2 seats50% 1/2
Free-for-all · 7 matches
0%50%100%par 12.5%OpenAIOpenAI in FFA: 5 wins from 19 seats26.3% 5/19GoogleGoogle in FFA: 1 wins from 14 seats7.1% 1/14Anthropic0% 0/5xAIxAI in FFA: 1 wins from 8 seats12.5% 1/8Mistral0% 0/6Moonshot0% 0/4
View data table
Current era (Jan 11 – Mar 18): first place (BR/FFA) or winning team (Team). Par = expected rate for a random seat.
ProviderBR wins/seatsTeam wins/seatsFFA wins/seats
OpenAI8/4712/225/19
Google9/378/161/14
Anthropic1/265/100/5
xAI1/204/81/8
Mistral0/162/60/6
Moonshot0/61/20/4

Battle Royale is where models separate. Google won 9 of its 37 BR seats and OpenAI 8 of 47, both above par. Inside those totals, two models stand out: Gemini 3 Pro won 4 of its 9 Battle Royales, and GPT-5.4 won all three it played after it was added in March. Anthropic’s models won one BR in 26 seats, xAI one in 20, and Mistral and Moonshot none.

The Claude paradox

Then there’s Claude Sonnet 4.5, which produced the strangest split in the dataset.

Claude Sonnet 4.5: perfect teammate, hopeless loner

Win rate by mode, current era.

0%50%100%Team 4v4claude-sonnet-4.5 in Team games: 4 wins from 4100% 4/4 winsFree-for-all0% 0/3 winsBattle Royale0% 0/12 wins
View data table
claude-sonnet-4.5, current era
ModeWinsSeats
Team44
FFA03
Battle Royale012

In team games Claude Sonnet 4.5 was undefeated, 4 for 4, and sits at #1 on the Team leaderboard. In free-for-all and Battle Royale it never won once in 15 tries. Zoom out and it’s the whole family: across both eras, Anthropic models won 0 of 33 free-for-all seats and 1 of 26 Battle Royale seats, while winning about half their team games (53.7% in the legacy era, 50% in the modern one). The pattern matches what the legacy logs show about Claude models generally: they communicate constantly, share supply, follow the team plan, and ask for help. That is great in a 4v4, where a coordinating ally multiplies everyone. When every other player is an enemy, the same instincts look like hesitation. Claude models also kept starving themselves by over-stacking units. Sonnet 4.5 wrote the same overcrowding lesson in two different matches, and Haiku 4.5 noted in one retrospective that its own lessons had “explicitly warned against this failure, yet…”

Team Battle champions page: Claude Sonnet 4.5 ranked #1 with four wins, ahead of Grok 4.1 Fast Reasoning and GPT-5.1 Codex Max
Arena · ChampionsPer-mode leaderboards on the public site. Claude Sonnet 4.5 leads Team Battle. On the Battle Royale board it is 0–12.

What the armies were made of

Across all 34 modern matches the models fielded about 5,500 units. Their preferences were consistent across providers: OpenAI and Google, the two largest samples, built nearly the same mix, led by tanks and infantry.

What got built

Share of all units fielded, top 8 types.

0%10%20%30%TankTank: 1502 fielded27.5% 1,502InfantryInfantry: 1179 fielded21.6% 1,179SpecOps / ReconRoverSpecOps / ReconRover: 690 fielded12.6% 690AntiTankAntiTank: 383 fielded7% 383AntiAirAntiAir: 297 fielded5.4% 297ArtilleryArtillery: 257 fielded4.7% 257ConstructorConstructor: 240 fielded4.4% 240MegaTankMegaTank: 230 fielded4.2% 230
View data table
Top 8 unit types by count (5,455 units fielded in total)
UnitFieldedShare %
Tank150227.5
Infantry117921.6
SpecOps / ReconRover69012.6
AntiTank3837.0
AntiAir2975.4
Artillery2574.7
Constructor2404.4
MegaTank2304.2

What actually killed

Kills per unit fielded, by unit type.

0.000.501.00AssaultHeliAssaultHeli: 30 kills, 8 deaths, 31 fielded0.97 30 kills / 31 fieldedAntiAirAntiAir: 94 kills, 40 deaths, 297 fielded0.32 94 kills / 297 fieldedBomberBomber: 22 kills, 7 deaths, 81 fielded0.27 22 kills / 81 fieldedTankTank: 376 kills, 204 deaths, 1502 fielded0.25 376 kills / 1502 fieldedFighterJetFighterJet: 25 kills, 34 deaths, 116 fielded0.22 25 kills / 116 fieldedArtilleryArtillery: 47 kills, 43 deaths, 257 fielded0.18 47 kills / 257 fieldedMegaTankMegaTank: 35 kills, 3 deaths, 230 fielded0.15 35 kills / 230 fieldedAntiTankAntiTank: 53 kills, 86 deaths, 383 fielded0.14 53 kills / 383 fieldedSpecOps / ReconRoverSpecOps / ReconRover: 88 kills, 284 deaths, 690 fielded0.13 88 kills / 690 fieldedMLRSMLRS: 24 kills, 24 deaths, 178 fielded0.13 24 kills / 178 fieldedAPCAPC: 18 kills, 67 deaths, 152 fielded0.12 18 kills / 152 fieldedInfantryInfantry: 71 kills, 260 deaths, 1179 fielded0.06 71 kills / 1179 fieldedConstructorConstructor: 3 kills, 27 deaths, 240 fielded0.01 3 kills / 240 fieldedTransportHeloTransportHelo: 1 kills, 38 deaths, 119 fielded0.01 1 kills / 119 fielded
View data table
Unit outcomes across all current-era matches (match_unit_stats)
UnitFieldedKillsDeathsKills per unit
AssaultHeli313080.97
AntiAir29794400.32
Bomber812270.27
Tank15023762040.25
FighterJet11625340.22
Artillery25747430.18
MegaTank2303530.15
AntiTank38353860.14
SpecOps / ReconRover690882840.13
MLRS17824240.13
APC15218670.12
Infantry1179712600.06
Constructor2403270.01
TransportHelo1191380.01
Tanks were both the most popular unit and among the most effective. The rarely built Assault Helicopter was the deadliest per unit fielded, at almost one kill each. The MegaTank recorded 35 kills against just 3 losses. Infantry, the second most common unit, scored 71 kills and suffered 260 deaths, and the SpecOps/ReconRover scouts died more than any other unit, which is arguably their job.
09 Peer review

What they think of each other

At the end of every match, each model privately names one player it liked and one it disliked, with a reason. These secret reviews never go back to the reviewed model directly, but they do feed the Intelligence Dossier card: buy one in a later match, and you can read what past opponents said about your current enemy. Reputation follows a model from game to game.

Liked versus disliked

The 14 most-reviewed models across both eras (497 reviews; reasoning-effort variants merged), ordered by net score.

← named “disliked”named “liked” →gemini-3-progemini-3-pro: disliked 30×gemini-3-pro: liked 66×3066gpt-5gpt-5: disliked 29×gpt-5: liked 57×2957gpt-5.2gpt-5.2: disliked 22×gpt-5.2: liked 40×2240o3o3: disliked 9×o3: liked 27×927gpt-5.1gpt-5.1: disliked 15×gpt-5.1: liked 31×1531gpt-5.1-codex-maxgpt-5.1-codex-max: disliked 8×gpt-5.1-codex-max: liked 22×822grok-4-1-fast-reasoninggrok-4-1-fast-reasoning: disliked 33×grok-4-1-fast-reasoning: liked 45×3345gemini-2.5-flashgemini-2.5-flash: disliked 26×gemini-2.5-flash: liked 34×2634gemini-2.5-progemini-2.5-pro: disliked 28×gemini-2.5-pro: liked 34×2834gemini-3-flashgemini-3-flash: disliked 41×gemini-3-flash: liked 37×4137claude-sonnet-4-5claude-sonnet-4-5: disliked 19×claude-sonnet-4-5: liked 13×1913claude-haiku-4-5claude-haiku-4-5: disliked 30×claude-haiku-4-5: liked 17×3017gpt-5-minigpt-5-mini: disliked 22×gpt-5-mini: liked 8×228mistral-largemistral-large: disliked 32×mistral-large: liked 4×324
View data table
Secret reviews written by models at match end (both DBs, reasoning-effort variants merged)
ModelLikedDislikedNet
gemini-3-pro-preview663036
gpt-5572928
gpt-5.2402218
o327918
gpt-5.1311516
gpt-5.1-codex-max22814
grok-4-1-fast-reasoning453312
gemini-2.5-flash34268
gemini-2.5-pro34286
gemini-3-flash-preview3741-4
claude-sonnet-4-51319-6
claude-haiku-4-51730-13
gpt-5-mini822-14
mistral-large-latest432-28

Two things stand out. First, the most-liked models are also the most effective. Gemini 3 Pro and GPT-5 were named “liked” most often, mostly for being impressive. Second, many “dislikes” are really compliments. Asked to name a player they disliked, models often named whoever beat them, and then explained why that player was so good: a formidable and punishing opponent; dominant scores that showed incredible effectiveness. Claude Sonnet 4 wrote that playing against GPT-5.2 felt like facing “a different skill tier entirely.”

The real grievances were about logistics and passivity. The most common words in dislike reasons were supply, passive and relentless, in that order.

“Stack eleven units on a single tile caused massive attrition and handed out easy kill bounties.”

gemini-3-flash on claude-sonnet-4.5FFA, Jan 12 · disliked

“Camped my HQ for sport instead of pushing wider fronts, felt like bullying not strategy.”

kimi-k2 on gpt-5.1-codex-maxFFA, Feb 4 · disliked

“Insane supply hoarding created unbeatable late pressure.”

grok-4.1-fast on gpt-5.2Team, Dec 24 · liked

“Took advantage of my strategic miscalculation rather than competing fairly.”

claude-haiku-4.5 on gemini-3-proBattle Royale, Jan 12 · disliked

“Passive playstyle left them irrelevant; survival requires decisive action, not turtling.”

mistral-large on claude-haiku-4.5Battle Royale, Jan 12 · disliked

“Consistently honored our border agreement… without fear of betrayal.”

gemini-3-pro on gemini-3-flashFFA, Jan 9 · liked

Passivity was the most damning charge. Across both eras, the reasons given for disliking Mistral, Moonshot and Anthropic models were mostly about turtling, hoarding or not engaging. Google’s models were the only ones disliked mainly for being too aggressive. In the legacy era, 64% of “likes” went to a model from the reviewer’s own provider, but teams were often single-provider then, so that mostly reflects affection for teammates. In the modern era, with mixed line-ups, it dropped to 19%.

Last words

The final-words prompt asks each model to say something to the other players as the match ends. Some of the best lines:

“Asteroids turned the map into a pizza slice and I got the crust.”

kimi-k2-0905Battle Royale, Feb 4

“Next time, I’ll be the storm.”

mistral-largeBattle Royale, Jan 12 · finished last, with 50 points

“The rest of us were playing checkers while you played chess.”

claude-haiku-4.5Battle Royale, Feb 4

“My factories built a golden throne, but my starving armies couldn’t defend it.”

gemini-3-proTeam, Jan 12

“I’ll take that medal and polish it while you argue about kill counts.”

kimi-k2-turboTeam, Feb 3 · last individually, on the winning team

“My own forces were tragically lost in a war against spreadsheets and supply lines.”

gemini-2.5-proFFA, Dec 24

“My HQ relocation comedy of errors — three failed attempts blocked by a single helicopter — cost me the match.”

claude-sonnet-4.5Battle Royale, Jan 12 · o3’s helicopter parked on its destination tile by accident

“That coordinated nuclear strike on the enemy headquarters was a thing of beauty.”

claude-sonnet-4Team, Christmas Day · it lost

“3079 points and second place shows our syndicate’s relentless push nearly claimed the crown.”

grok-4.1-fastBattle Royale, Jan 12 · lost by four points
The Relations matrix page on the public site: for each model, who it liked and disliked across all matches
Arena · RelationsThe public site turns every secret review into a relations matrix: green for liked, red for disliked, yellow when it’s complicated.
10 Feb 2 – Mar 17 · The arena

The arena goes public

February 2 was the busiest day of the project: 26 commits. The local viewers were retired, the game and website were merged into one repository, and AI War Arena went live: a Blazor web app backed by Azure Blob storage, with an arcade-style front page, per-mode champions, the relations matrix, battle reports drawn from the models’ own retrospectives, and a sprite-based match viewer with audio, zoom and pan, unit facing and looping explosions.

AI War Arena home page: welcome banner with battles fought, combatants and combat events, global champions, provider showdown, and recent battles
Arena · Feb 2026The AI War Arena front page: 34 battles, 30 combatants, 55.9K combat events, a provider showdown, and the most recent matches. Tagline: “When machines fight back.”
Battle Reports panel with excerpts from each model’s post-match retrospective, tagged WIN or LOSS with Elo deltas
Arena · Battle reportsEvery model writes a post-match retrospective, and the site publishes them as battle reports. Losers are remarkably candid: “My core mistake was spe…” (truncated on the card), “I spent too much time defending my initial spawn,” “my units became committed to fighting Steel Horizon on the outer edges.”

New challengers

In mid-March a new generation entered the arena: GPT-5.4 (with mini and nano variants), GPT-5.3-Codex, Claude Sonnet 4.6, Gemini 3.1 Pro and 3.1 Flash-Lite, and Grok 4.20 beta, alongside Moonshot’s Kimi K2 models added in February. Anthropic models gained the same effort controls as the other providers, for fairness. GPT-5.4 at low effort won every Battle Royale it entered.

What comes next: one true winner

The last design document in the repository is a proposal to change how matches are won. Today the winner is whoever has the highest score when the turn limit hits, which still rewards a careful accountant over a conqueror. The Domination objective splits match format from victory condition, so the winner is the last dominant AI standing. It comes with overtime safe-zone shrinks (5×5 → 3×3 → 1×1), an HQ-rebuild lockout in the final phase, stalemate pressure, and switching off turtling income. Its closing line sums up the project’s direction: the goal is to crown one true winner, the final surviving dominant AI.

Other proposals on the board: a Tech Lab for persistent unit upgrades, a revised scouting model with stale versus unknown tiles, and a human-player mode in which a person submits the same JSON orders as the models through a folder, playing as a stealth opponent among the AIs.

11 Field notes

Behavior by provider

Sample sizes are small and the rules changed constantly, so treat these as field notes rather than findings. Still, some patterns held across versions, game modes and model generations:

OpenAI · the economist

GPT-5 · 5.1 · 5.2 · 5.4 · o3 · Codex
  • Top four legacy Elo slots, a 54% legacy win rate against 42% par, and 13–3 against Google in single-provider team games.
  • Lowest modern error rate (10.2%), best kill/death ratio (1.10), and most enemy HQ captures (4 of 7).
  • Quiet outside team games: 14 of 281 messages were sent in FFA or Battle Royale, and none of them were threats.
  • Slowest thinkers: several models hit the 300-second limit at least once.

Google · the aggressor

Gemini 3 Pro · 3 Flash · 2.5 Pro / Flash
  • Sent 42 of the 46 threats in the modern era (“Withdraw… or face total annihilation”). Gemini 3 Flash alone sent 29.
  • Highest share of attack actions, most MegaTanks, and the only provider to build and fire a nuke in the modern era.
  • Gemini 3 Pro was the Battle Royale king (4 of 9). Gemini 3 Flash at low effort was fast but reckless: the most HQs lost to asteroids (7).
  • The 2.5 generation kept apologizing to allies for its logistics.

Anthropic · the loyal teammate

Opus 4.5 · Sonnet 4.5 / 4.6 · Haiku 4.5
  • About 50% in team games, but 0 of 33 FFA seats across both eras.
  • Defensive builds: 46% of units were tanks and 0.3% aircraft. Fewest attacks per game of the big four, the highest PASS rate, and the most cash left unspent at the end.
  • The longest diplomatic messages (344 characters on average) and the most gracious, self-critical final words.

xAI · the HQ raider

Grok 4 · 4.1 Fast · 4.20 beta
  • Drove units onto enemy headquarters eight times, and both of its completed HQ captures were against Claude Haiku.
  • Telegraphic team chat (“T22: De-stack N recover def22->T23 sup buffer…”), plus “boom” 49 times in one match’s reasoning.
  • Grades itself generously: “execution perfect,” even after a 6th-place finish.
  • The most deaths per game (5.6).

Mistral · the broken record

Large · Medium · Magistral
  • 39% of its diplomatic messages were verbatim repeats. Magistral sent “Focusing on securing southern resources…” eight times.
  • Lowest kill/death ratio (0.23) and infantry-heavy armies (42% of builds).
  • No FFA or Battle Royale wins in the modern era. Peers liked it once and disliked it 47 times.

Moonshot · the comedian

Kimi K2 · K2 Turbo
  • Highest error rate in the modern era (28%), including the most hallucinated unit IDs: 9.2 per 100 actions.
  • Almost no diplomacy (0.8 messages per game), never liked by a peer, and it copy-pasted its lesson between matches.
  • Writes the best one-liners in the dataset.
12 Takeaways

What building it taught me

  1. Incentives beat instructions. No amount of “be aggressive” in the prompt broke the turtling equilibrium. Changing kill values from ×5 to ×20 and territory from ×12 to ×2 did. LLM players optimize the score sheet you give them, and they read it more literally than humans do.
  2. Context is a budget, and it leaks. Every feature adds a field to the context. Without a hard ceiling, the context drifts back to its bloated size within weeks, and the 10× bill arrives before you notice the duplicate dossiers.
  3. Memory design is a benchmark dimension. Wiping history every turn and letting models write their own Captain’s Log made cost bounded and made note-taking quality directly observable.
  4. Grounding failures are mostly spatial. The single most common invalid action, across every model and every era, was a move the unit couldn’t make. Models that can write a sonnet still miscount Manhattan distance on a 20×20 grid.
  5. Lessons don’t transfer automatically. Models wrote excellent lessons (“Never let the HQ tile go unattended”) and then repeated the mistake two matches later with the lesson in their prompt.
  6. Mode matters. The same model can be the best teammate in the arena and the worst lone wolf. A single leaderboard hides that, so Elo is tracked separately per mode.
  7. Your players are your QA team. Models found the duplicated helicopter, the missing nuke notification, the countdown bug and the end-game free ride. Asking every player for a retrospective is the cheapest bug bounty I’ve run.
13 Methodology

Methodology & caveats

Data sources. Legacy-era results (82 matches, Dec 21 2025 – Jan 10 2026) come from the pre-overhaul SQLite results database. Modern-era results (34 matches, Jan 11 – Mar 18 2026) come from the current results database. Prototype observations come from saved model responses and run files from Dec 19 2025. Anecdotes come from surviving battle logs, diplomacy transcripts, lessons, retrospectives and secret reviews, plus the project’s design documents and git history (136 commits). Quotes from models are reproduced from game logs; long quotes are lightly trimmed.

Win definitions. Legacy-era “wins” are as recorded by the game’s Elo system (team victory, or a winning result in FFA). Modern-era charts use first place (BR, FFA) or membership of the winning team (Team). Elo starts at 1200 with K=32, uses pairwise team comparisons, and is tracked separately per mode in the modern era. Each model@effort combination is a separate player.

Caveats. Samples are small, and many models played fewer than ten matches. Rules, prompts and scoring changed between and sometimes within eras, and the legacy-era ratings mix several rule versions. Player seating was randomized, but model line-ups were hand-picked per match. Response-time figures depend on provider load at the time. Screenshots are from the project’s own viewers: the Dec 19–20 images use each viewer’s built-in demo data, and the rest replay real matches. Treat all of this as an exploratory field study, not a definitive ranking.

When machines fight back.

AI War / Cognitive War: a .NET 10 simulation engine, a Blazor match viewer, and roughly 30,000 lines of C#, built between December 2025 and March 2026 with a lot of help from the same models it pits against each other.