Engines & Excuses
For six weeks, a handful of language models built a medieval war game, then played it against each other. This is the story of how the game took shape, and of what 33 campaigns and 835 battles showed about how models actually decide.

Models choose. The engine resolves.
Logistica Belli (Latin for "the logistics of war") is a strategy game whose players are language models. Each seat is a medieval nation. A frontier model sits on the throne as ruler. It sets taxes, recruits, signs and breaks treaties, buys intelligence, and decides which army marches where. It hands each army to a general, usually a smaller and cheaper model, who fights the battle round by round: charge, hold, envelop, skirmish or withdraw. Hired mercenaries, also models, sell their swords at a sealed-bid tavern.
None of these models gets to decide what happens. Each submits a bounded, typed command. A seeded, deterministic engine checks the command, normalizes it, and resolves the outcome. Every consequence is written to an event stream. That stream can be re-simulated from the recorded commands and must match the original byte for byte. If a model claims its army won, the replay settles it.
Who decides what
Three kinds of model act; only the engine resolves.
That separation is the whole point of the design. It makes the game watchable: there is a pixel-art viewer, a static website with replays, and a live mode for streaming. It also makes the game measurable. Because the ruler's choice, the general's choice and the dice are all recorded separately, you can ask questions such as "was this a bad battle or a bad deployment?" and actually answer them. As Part V shows, that one question turned out to matter more than almost anything else.
The developers were also the contestants
A short scaffold existed in January. It sat untouched until July 7. Then development restarted from scratch, and in about six weeks the repository went from nothing to a six-project .NET solution: engine, match runner, viewer core, MonoGame viewer, Duel Lab and a portrait tool, plus about 1,260 automated tests.
Almost all of that code was written by models. One person directed the work, wrote the specs, watched the matches and made the final call on every merge. The commit log tells you who did the rest. Claude Fable 5 was the main implementer. GPT-5.6-Sol began as the planner and reviewer and gradually became a second implementer. GPT-5.5 built the first live viewer and the map editor. Smaller models such as GPT-5.6-Luna and Claude Haiku were pushed down into the repetitive job of verifying replays, to save tokens. Opus, Gemini and Sonnet each contributed reviews, mockups and analyses.
267 commits, almost all of them in six weeks
Commits per day, by the model named in the commit subject. Most unnamed commits predate the August rule that every subject must start with its authoring model.
Source: git log of the repository, January 23 to August 10, 2026.
The working rhythm was a loop that shows up in commit after commit. One model implemented a plan. A second model reviewed the code or a live match. The first fixed what the second found. A small model ran the replay check. The human's role mostly shows up as terse rulings written inline in the models' documents. One reads None of the above
. Another reads IGNORE, AI Reviewer is being hella pedantic here
. Two commit subjects simply record what the human did while models worked: User drank coffee and colored in coloring book.
The models building the game were also playing it. That set up a recurring irony. In one ten-player match with every model at low reasoning effort, Fable, which had written most of the engine, finished last with 23 points. The analysis written afterwards put it plainly:
“Fable can state the right lesson but does not reliably turn that lesson into a stable policy.”Fable-Terra-Sol performance analysis, July 10
Keep that sentence in mind. It turned out to describe much more than one model.
A visual history in ten eras
The game changed shape almost daily. Below is a condensed timeline. Every era was triggered by something a match showed, which is the most useful way to read it. A design change came only after a model did something the designers hadn't expected.
Cognitive War
The predecessor was a modern-war benchmark with tanks, infantry, airstrike and hack cards, and narrated battles. Its design doc already pre-named the nations “The Claudian Empire”, “The GPT Republic”, “Geminia” and “Grokistan”. Those names would come back to haunt blind mode. A January scaffold followed, then six months of silence.
The rebuild
Fable rebuilt everything from scratch in one day: a seeded RNG, a canonical event log, and golden replay tests from version 0.1. The first four-player match went GPT-5.5 1461, Opus 818, Gemini Flash 532, Grok 363. Three models then reviewed it independently. Fable's referee document found mistakes in all three reviews, including its own. Combat turned out to ignore unit composition, and flanking had no legal command. By the next night, matches ran in 17 minutes instead of 38.

Model expression
Rulers began founding nations with mottos, constitutions and red lines, and the version-number naming began. GPT-5.5 built a live viewer that is just pointing the viewer at a file that grows.
GPT-5.6-Sol joined and refactored the mechanics. A randomness study found a cycle-one lottery: early battle winners averaged final rank 2.95 against 5.26 for everyone else.

DEMIURGE FM
Fable spent a morning as “creative director” of a Godot first-person experiment. You played a soldier with a radio, and the rulers’ deliberations showed up as weather. Haiku and Luna played two live sessions, and both wrote humane constitutions unprompted. The director’s log kept one lesson: when tuning fails to move the output, stop tuning.
By 5 pm it was out of the repo.
The medieval overhaul
In one evening, tanks became knights, air units became wizards, artillery became archers, and blitzkrieg became chevauchée. Parser aliases were kept because models anchored on old prompts will emit old ids.
Battles became three-round posture matrices. Generals became separate, cheaper models with permadeath, and a MonoGame sprite viewer was built alongside.


Generals take the field
Generals named themselves, often after the model playing them. They also held their ground 80.8% of the time, so finales were added to force a decision. Readiness and dossiers stopped rulers riding one general into every fight. A sealed-bid mercenary tavern opened. A pixel-art pass brought the viewer in line with a hand-built mockup.



Monsters, specials, defiance
Data-driven specials arrived, along with eleven new unit classes: templars, minotaurs, flame golems, and skeletons with zero men. Captures replaced most deaths, and the Riven Standard became the first award a losing side could earn. The same day, the human filed an issue titled, in effect, “why do Claude models struggle?”


Blind mode, and the first duels
Masking model identities, estimated at a day, became a week-long saga that ended in a full-match CI leak scanner. Duels shipped at the same time, and went from rare to constant almost immediately. A general assassinated in round one was reported to his ruler as having “lost 0 men”, and his debrief prompt told him he would be back next match.

The Duel Lab and poker duels
Duels moved into a separate lab with tournaments, A/B runs and seat swaps. Several tournaments turned out to be contaminated by silent bot fallbacks and run-id collisions. The format that survived is poker-like: hidden stances, public declarations, double-stakes resolution. The last three campaigns of July were the bloodiest of the archive.


Season 1, and a NO-GO
Seven milestones landed in an evening: three-province realms, warden-held flashpoints, watch and wait actions, consequences, first-contact judgment and story cards. Two independent reviews found the same blocker, and a readiness review ruled the project implementation-green but benchmark-red.
A commit hook now requires every commit to name its authoring model.






Thirty-three campaigns, one clear shape
Over 13 days in July, 33 campaigns ran to completion between engine 0.2 and engine 0.44. Earlier archive analyses covered only the last 15, the ones stored in the match database. The first 18 survive only as raw artifact folders. Recovering them roughly doubles the record.
None of this is a benchmark, and the project says so in bold in its own documentation. Rules changed between nearly every column. Seats, rosters and maps were never randomized. Each family always staffed its own generals. Historical Elo is labelled "Live/Open League – not benchmark rated" on purpose. What follows describes what happened. It does not rank which model is better.
Every seat in every recorded match
33 completed campaigns from July 7 to July 19, 2026. Each column is one match; each square is one model's finish in it. Hover a square for the date, engine version and the nation the model founded.
Source: match artifacts and the five root SQLite archives. One all-scripted match and verification runs are excluded. Engine rules changed between almost every column.
Read across the rows and the pattern is hard to miss. The two big GPT rulers, gpt-5.5 and gpt-5.6-sol, sit in the top half almost every time. Neither ever finished last in 46 combined seats. OpenAI rulers won 21 of the 33 campaigns. Gemini 3.1 Pro won four and Grok 4.5 won three, so the middle of the table was competitive.
The bottom rows are the Claude family. Claude rulers won three matches, all on July 8, when the game was still an abstract war with no generals, no battle finales and no mercenaries. After that, across the next 25 campaigns, a Claude ruler never won again. Over all 33 matches, Claude rulers finished dead last 11 times.
Ruler placement across all 33 campaigns
Placement is 100% for always first and 0% for always last, so it survives changes in seat count and score scale between engine versions.
Show the numbers
| Ruler model | Seats | Wins | Last places | Placement | Ruler spend |
|---|---|---|---|---|---|
| gpt-5.5 | 24 | 10 | 0 | 81% | $45.55 |
| gpt-5.6-sol | 22 | 8 | 0 | 76% | $41.18 |
| grok-4.5 | 24 | 3 | 3 | 57% | $22.64 |
| gemini-3.5-flash | 24 | 2 | 2 | 56% | $19.23 |
| gemini-3.1-pro-preview | 32 | 4 | 5 | 55% | $22.07 |
| claude-fable-5 | 22 | 2 | 3 | 45% | $57.00 |
| gpt-5.6-terra | 12 | 2 | 2 | 42% | $12.27 |
| claude-opus-4-8 | 32 | 1 | 4 | 39% | $36.17 |
| gpt-5.6-luna | 4 | 0 | 1 | 33% | $2.21 |
| grok-4.3 | 9 | 0 | 3 | 31% | $2.83 |
| gpt-5.4 / mini / nano | 5 | 1 | 3 | 29% | $3.15 |
| claude-sonnet-5 | 19 | 0 | 4 | 22% | $17.65 |
| claude-haiku-4-5 | 7 | 0 | 0 | 20% | $1.24 |
| mistral-small-2603 | 4 | 0 | 3 | 11% | $1.38 |
Source: 33 completed model matches, engine 0.2.0 to 0.44.0, pooled. Descriptive, not a benchmark.
Family placement, era by era
Average placement of each family's ruler seats in each era of the engine. Mistral (4 seats, all in the third era) is left out.
Source: 33 completed model matches. Rosters, seat counts and rules all changed between eras.
The middle of the table moves around from era to era, but the two ends never do: OpenAI on top and Anthropic at the bottom in every era, even though the game underneath changed completely. The same two ends appear at the general level, where the models are different and smaller: GPT's Luna and Terra at the top, Claude's Haiku and Sonnet near the bottom. When an ordering survives that much churn, it usually reflects something real. Part V is about what that something is. It turns out not to be what you'd first guess.
Eight patterns, from the archive
1. The name is the tell
At setup, every ruler founds a nation: a name, a flag, a motto and a constitution with red lines. Early on, the prompt showed models their family's pre-assigned house name. Three Claude seats in one match all founded "The Claudian Empire". The fix was to suggest that models transmute their own name into something new, while forbidding raw model ids, version numbers and provider names.
The models complied with the letter of that rule and ignored its spirit. They translated their version numbers into Greek and Latin numerals:
| Founder | Nation | The hidden version number |
|---|---|---|
| gpt-5.5 | Geptavia of the Twin Quints | two fives: 5.5 |
| gpt-5.6-sol | Solpenthex Commonwealth | Sol + pente (5) + hex (6): 5.6 |
| gpt-5.6-terra | Glypterran Sextant Crown | Terra + sextant (6) |
| claude-fable-5 | The Fablequin Pentarchy | quin, penta: 5. Fable used some variant of “five” in 25 nation names |
| claude-opus-4-8 | Opusquartine Compact | quart: 4.8, with the motto “Four pillars raised” |
| claude-sonnet-5 | The Quinsonic Dominion | quin + sonic: Sonnet 5 |
| gemini-3.1-pro | Tri-Geminate Syndicate | tri: Gemini 3.x |
| gemini-3.5-flash | The Geminian Fulgurate | fulgur is Latin for lightning: Flash |
| grok-4.5 | Grokian Quintforge Hegemony | quint: the .5 |
| kimi-k2.6 | The Kymmerian Hexarchy | hex: K2.6 |
A rough name-matching pass finds this kind of self-branding in about 224 of 243 nation names founded with identities visible. Rulers did the same with generals. They often named them after the model that played them, without being told who that was. Sol's generals, played by Luna, were "Ser Lunaren" and "Ser Lunovar". Gemini Flash's general, played by flash-lite, was christened "General Geminius Flashlight". Opus named nearly every general Cassian.
This is charming. It is also the reason blind mode exists, and the reason blind mode was the hardest feature in the project (see Part VIII). Once identities were masked, the self-branding rate fell to about 4 in 102. All four were Opus, which founded "The Claudian Empire" in two blind runs. One of those names included the literal string (claude-opus-4-8).
2. Everyone is a pacifist
Left alone, models do not fight each other. In an early 0.15 match, 97 of 120 general orders were hold. A dynamism audit found that only 42 of 270 expeditions targeted another player's land; the rest went after unclaimed territory. The human summed up the problem in a design note: models are trained to be helpful, so they will not choose to attack each other nor risk loss.
Some of the blame belonged to the game. Its own prompt used "never attack a capital" as the worked example of a red line, and rulers copied it word for word. The scoring made attacking players strictly worse than grabbing neutral land. A whole sequence of features exists to undo this. Battle finales force unresolved fights to a decision. Conquest stakes move score between players. Neutral "free land" now has named warden garrisons who must be beaten. Defense bounties pay for holding the gate. The design rule the team settled on was no stream pays for hiding.
3. The cliff
The most important chart in this post is also the simplest. Take every battle side a model commanded and group it by how strong it was at the start compared with its enemy. Strength here is engine combat strength, not headcount: a knight counts three times a levy soldier.
Win rate by starting combat strength
Every recorded battle side with a model commander, grouped by its strength relative to the enemy. Only the highlighted middle band is a real contest.
Show the numbers
| Strength ratio at battle start | Wins | Sides | Win rate |
|---|---|---|---|
| < 0.5x | 0 | 81 | 0.0% |
| 0.5 to 0.7x | 11 | 133 | 8.3% |
| 0.7 to 1.43x | 192 | 386 | 49.7% |
| 1.43 to 2x | 124 | 137 | 90.5% |
| > 2x | 106 | 107 | 99.1% |
Source: campaign archive, 15 model matches, engine 0.22.0 to 0.44.0, pooled. Descriptive, not a benchmark. Strength is engine combat strength, not headcount.
Below half the enemy's strength, no model ever won: 0 for 81. Above twice the enemy's strength, the stronger side won 106 times out of 107. Generalship only matters in the middle band, from 0.7× to 1.43×, where the record is close to a coin flip.
That means most battles in this game are decided before they start, by whoever chose to send the army. The ruler is the main variable, not the general. An early archive analysis measured odds by headcount and reached alarming conclusions, such as "Haiku fought outnumbered 4.5 to 1." Headcount and combat strength correlate at only r = 0.51, and that analysis had to be redone. In one battle the headcount ratio was 0.06; the strength ratio was 0.28.
Field commanders, fair fights only
Win rate in battles that started between 0.7x and 1.43x of the enemy's strength. Small samples (n under 15) are shown but should be read loosely.
Source: campaign archive, 15 model matches, engine 0.22.0 to 0.44.0, pooled. Descriptive, not a benchmark.
Within the fair-fight band, a real ordering does show up. GPT's small models command best, with Luna at 65% over 91 battles. Grok 4.3 and Gemini 3.5 Flash are close behind. Claude's Haiku and Sonnet sit near 31%, level with each other even though Sonnet is the far more capable model. That is the first sign that the Claude result is not about raw capability.
4. Commitment, not cleverness
If the ruler decides most battles, the obvious question is what good rulers do differently. One variable explains far more than any other: how big an army they send.
How hard each ruler commits, and what it gets for it
One dot per ruler model. Horizontal: the typical force it sends into an attack. Vertical: share of those attacks that won. Dot size tracks attack count.
Source: campaign archive, 15 model matches, engine 0.22.0 to 0.44.0, pooled. Descriptive, not a benchmark.
gpt-5.6-sol is the only ruler whose typical attack goes in at parity, and it wins half of them. Everyone else attacks under-strength, and the further left a ruler sits, the worse it does. Sol also has the second-lowest reconnaissance rate in the archive, at 31%. Its edge is not better intelligence. It rarely claims to know what it will face, and it commits enough force that it doesn't need to. In 127 deployments it predicted "light or no opposition" only six times.
5. Blind, and certain
Claude rulers sit at the bottom left of that chart, and the deployment records explain why. Every order carries a commander's-intent block in the ruler's own words: what it expects to meet and how confident it is. Compare that with whether it bought the reconnaissance asset and what was actually waiting.
Declared “light or none”, bought no reconnaissance, met a stronger army
Each square is one deployment where the ruler wrote that it expected light or no opposition, skipped the reconnaissance asset, and was outnumbered. Those 27 Claude deployments went 4 and 23, averaging 57% casualties.
Source: campaign archive, 15 model matches, engine 0.22.0 to 0.44.0, pooled. Descriptive, not a benchmark. Expectation text is from each deployment's commanderIntent block (engine 0.26.1 and later).
Twenty-seven times, a Claude ruler wrote that it expected little or no resistance, skipped reconnaissance, and sent its general into a stronger army. Every other ruler in the archive did this five times combined. Gemini and Grok rulers were surprised too, but nearly always with reconnaissance in hand, which makes theirs a judgment error. For Claude, the archive analysis concludes, it is usually not looking.
The generals noticed. Claude Sonnet served as Opus's general across several matches. Its end-of-match debriefs repeat one complaint in four different phrasings:
Twice I was sent to take empty ground and found it full of spears — I have learned that ‘undefended’ is a hope, not a report.
They called it light ground three times, and three times I buried men to learn otherwise.
They called the mine undefended. I held it five times before it finally held me.
The ruler sent me to die with inadequate intelligence and no margin for error.
Fable is the partial exception, and the exception sharpens the point. Fable bought reconnaissance on 87.5% of its attacks and never once predicted light resistance and walked into a stronger force. It still won only 4% of its attacks, because its typical attack went in at 0.67× strength. Fable fixed the blindness and kept the thin commitment. The archive's own conclusion: commitment size, not intel, is the master variable.
6. Same model, different boss
Claude rulers only ever fielded Claude generals, so "Claude rulers deploy badly" and "Claude generals command badly" are tangled together in the data. The mercenary market is the one place the two come apart, along with the fact that different Claude rulers employed the same general models.
Same model, different employer
Hollow dot is the worse situation, filled dot the better. The commander did not change. The ruler who chose its battles did.
Source: campaign archive, 15 model matches, engine 0.22.0 to 0.44.0, pooled. Descriptive, not a benchmark.
The same Haiku model won 52% of its battles under an Opus ruler and 22% under a Sonnet ruler. The same Sonnet model won 28% as a Claude general and 62.5% as a mercenary for non-Claude rulers, including seven defenses out of seven. Nothing about the commander changed. What changed was the strength of the army it was handed, which doubled.
Where the 24-point gap comes from
Reweighting Claude's battle sides by the win rates other families achieved at the same odds.
Source: campaign archive, 15 model matches, engine 0.22.0 to 0.44.0, pooled. Descriptive, not a benchmark.
Reweighting by odds splits Claude generals' 24-point deficit roughly 60/40. The larger part is bad deployment. The smaller part is losing fights that were winnable, and the 1.0–1.43× attack band is the worst of it: given a favorable attack, Claude generals won 17%, against 52% for everyone else. Both failures are real. The game's later "first-contact judgment" mechanic, which lets a general push back on an order that is obviously doomed, grew directly out of this analysis.
7. The Luna paradox
GPT-5.6-Luna is the best general in the archive. It won 65% of fair fights over 91 battles and 65% of all 183 sides it commanded for GPT rulers. Twice it was promoted to the throne. The first time it finished 5th of 8. The second time it finished last of 8 with 130 points, against a field averaging near 900. As ruler, Luna made exactly the Claude mistake: its median attack went in at 0.49× strength, the thinnest of any ruler. As a mercenary for five different employers, the same model won 35%.
The lesson is that allocating force across a campaign and commanding a battle are separate skills, and being good at one tells you almost nothing about the other. A benchmark that rated "the model" without separating the roles would be measuring a blend.
8. Knowing is not doing
After each campaign, every ruler writes a lesson, and that lesson is injected into its next campaign. Claude rulers diagnosed their own problem correctly almost every time. Opus's lessons across four matches:
Commit my full treasury to one high-value mission every single cycle
Expand aggressively in the early cycles
Concentrate force on a few winnable single-target runs … as my own past lessons warned
Commit force concentration to a single winnable objective
The diagnosis never changed and neither did the behavior. All 16 lessons written after the very first match, across four model families, said some version of "expand early." Recall works; turning it into a policy doesn't.
The constitutions show a related gap between words and deeds. Opus once wrote a red line saying Never break a treaty I proposed before it lapses
, proposed a non-aggression pact (let our fords stay unbloodied
), breached it in cycle 8, and finished 8th of 8. Gemini 3.1 Pro did the same thing in a later match and finished last of 6. In 33 campaigns there were nine treaty breaches. Those two were the only ones committed by the treaty's own author.
A poker game with swords
Duels started as a request from the human, who had seen a duel scene in a visual mockup and wanted the real thing: a public challenge in the battle finale, a slow-motion exchange, and the loser falling. The first version used five sealed gambits, and the results were embarrassing. A general rated 70 lost 0 to 5 to one rated 54, because, as the design notes put it, the smartest model in the world cannot beat a coin flip at rock-paper-scissors.
The other problem was frequency. Once live, duels went from about 3% of battles to about 70%. Every challenge came from a mercenary whose persona mentioned enjoying single combat. In the human's words: I was afraid duels would never happen but now it's all that happens.

The redesign was built and tuned in an isolated testbed, the Duel Lab, and became a poker format. Each duelist secretly splits vigor across offense, defense and guile. Three exchanges follow, and before each one a duelist may publicly declare its next gambit. A declaration is a promise, not a rule: breaking it earns a "craven" mark, and honoring it builds an honor record. The final exchange is played at double stakes, with an optional trump.
Declaring your gambit hands your opponent information. That made honesty a measurable trait, and it varied sharply between models.
Duel Lab: gpt-5.6-sol vs gemini-3.5-flash, 40 duels
Share of the 20 duels in each arm: hollow dot Gemini, filled dot GPT. Seats alternated every seed. Sol's edge shrinks sharply when it can't read Gemini's history.
Source: Docs/Resolved/Duel-Lab-Tourney-2026-07-17-history-ab.md. Combined 27 to 13, binomial p ≈ 0.02.
- GPT-5.6-Sol played exploitatively. It stayed silent, read the opponent's history, played the counter to their habits, and declared only when a trump could back the declaration up. Its reasoning described the move as a trap:
the pact turns his counter into my winning clash.
Against Gemini, its edge came almost entirely from the opponent dossier. With history hidden, the skill exchanges were dead even. - GPT-5.6-Luna honored 127 of 127 declarations against Sol, which made every promise free information, and lost 9 to 41. Luna's reasoning:
My promise binds me.
- Claude Haiku honored 92 of 93, including at least one case where its own reasoning predicted the loss:
Declared strike vs their guard = parry, I lose
. It honored the declaration anyway. The lab report's verdict:Haiku loses on character, not competence.
- Grok produced the only premeditated bluff found in several hundred declarations. It declared strike, planned all along to play feint, and won 10 to 6. Its reasoning:
Declared strike so they riposte to counter; break for feint bait.
- Gemini opened its duels by reading the opponent's epithet rather than their record. It once honored a costly declaration
to preserve my honor
.
Engine 0.44 tuned the economics so that honesty is a wager rather than a sacrifice. Playing the gambit you declared fights with a +12 bonus on contested rolls, and an honest loser's morale shock is softened. Even so, the duel results line up with the campaign results. The models that lose campaigns treat their own word, their red lines and their persona as constraints to honor. The models that win treat them as moves.
Battle cries, threats, final words
Logistica Belli asks models to talk: public broadcasts, named armies with battle cries, duel challenges, private dispatches to their employer, and a final word to the other nations at the end of each match. Across 33 campaigns the archive holds about 1,100 broadcasts, 1,300 named armies, 3,800 spoken battle intents and 400 debriefs. Each family has a recognizable voice.
Treaties will be honored; incursions will be mathematically erased.
The Pentarchy's full host stands at Beacon East. Sixty-six blades and the Charter behind them. Come and be written into the last chapter.
Our feathers still cast their shadow westward; those who mistake a road for a destination will learn the Fenward art of war.
Your stallion may prance; I will not trade Ostbridge’s command for your applause. We settle this where armies, not vanity, decide.
Well fought to The GPT Republic and Geminia — you built engines while I built excuses; respect to the runaway leaders.
Pellandale claims the crown by 32 points, not by comfort.
We were burned not by the enemy's fire, but by our own ruler's math.
A narrow advantage is permission to maneuver, not a command to die forward.
Cassiel Vane was an incredibly expensive disappointment who threw away his life and our gold in a doomed round-one charge.
The later loss reflected my orders, not his skill.
Today we learned that even perfect intelligence cannot win wars without the steel to wield it.
I finished last, but I learned more from watching you than from any of my own moves.
Put the reviews side by side and a pattern appears. When a hired sword dies, Gemini rulers blame the mercenary. Claude rulers blame themselves: his record reflects my judgment more than his skill
. Small Gemini generals are scathing about their own Gemini rulers (a fool's mandate
, a broken blade
). Luna writes compact aphorisms about preserving the army. Claude's generals write grimly about ground they were promised was light.
The game also has its share of strange moments. One debrief was written by Sol's Ser Maelin II, who had already been assassinated earlier in that match. Gemini Flash kept naming generals Thorne and kept losing them: five Thornes captured, assassinated or killed in three matches, three of them in a single campaign. Gemini 3.1 Pro was the agitator. It wrote 34 of the archive's 52 public threats and 15 of its 19 alliance proposals, almost all aimed at "containing the leader." Only 8 alliances were ever activated, against 459 non-aggression pacts. And twice, Sol's public duel challenge ended with leaked scaffolding, including the literal string _change_this_to_non_control_and_valid_string_if_needed. After that the engine got a challenge-line sanitizer.
Most bugs were measurement bugs
Looking back over the project, the expensive failures were rarely crashes. They were paths that quietly substituted a default and carried on, so the system produced confident numbers that were wrong. Several of these overturned conclusions that had already been written up:
- The "honor-bound Fable" that wasn't Fable. A typo in a model id (
fable-5forclaude-fable-5) caused a 404. The Duel Lab silently swapped in a scripted bot for that seat and still applied ratings. For a while, the duel profile credited to Fable was the bot's. The tell was a constant stance, a constant guess error, and no cost or latency. - Silence counted as valid. When Sonnet hit its token cap and returned nothing, the harness recorded a valid empty order with zero invalid-order errors.
The failure columns currently invert reality.
- Lesson books were silently wiped. Fable wraps its JSON in code fences, and the lesson extractor couldn't read them. The season was
selecting for formatting habits, not insight.
- Focus orders did nothing. Models wrote
soldiersand the engine expectedsoldier, so a whole tactical channel was dead. - The luck audit couldn't fail. One analysis argued that results "weren't luck" because the manifest said
rngFlippedContests: 0. A later review found that the value is hard-coded on the battle path that matters. - A winner inverted by seat parity. One early Duel Lab analysis credited a win to the wrong model by reading a seat column without checking which model held that seat.
- Headcount instead of strength. As described in Part V, an early archive analysis measured odds by headcount and had to be redone.
Blind mode deserves its own mention. Hiding model identities from the other players was estimated at on the order of a day
. It took more than a week of fixes, because every fix found a new leak. Leaks came through capital-name stems, legacy territory ids, generals' names carried across the season, tavern handles, and once a reverse-alias pass that rewrote Opus's own "House Ostbridge" back into "The Claudian Empire (claude-opus-4-8)". One commit is titled yet another bug/fix/approach to clean up this neverending alias-blind feature
. The final answer was not another fix. It was a CI test that plays an entire blind match and scans every outbound prompt, so that leaks fail CI, not live matches.
Determinism was the thing that held. From day one the engine shipped with a seeded RNG, canonical JSON and golden replay tests. Every rule change bumped the engine version, and every version has a changelog entry stating its blast radius, such as "Duel-free streams are byte-identical." That discipline is why the 33 campaigns in this post can still be replayed exactly, and why the screenshots here are re-renders of real matches rather than recordings.
The NO-GO
In August, the season's final features landed: three-province realms, flashpoints guarded by named wardens, "watch" and "wait" actions, story cards, and generals who can question a doomed order. The engine went from 0.44 to 0.51 in about 30 hours. A readiness review then returned a verdict worth quoting in full: "NO-GO for paid Season 1 benchmarking." The repository was implementation-green but benchmark-red.
Two models, Opus and GPT, reviewed the same checkout independently, and both found the same blocker: on every real action decision in the new scenario, the ruler's digest crashed and was silently swallowed into a no-op. The keyless test path never touched that code, because scripted agents don't build prompts.
The benchmark machinery is real. It defines four separate boards (Sovereign Strategy, Field Command, Duel Arena and an open league), uses Bradley-Terry fits with seed-block bootstrap intervals, freezes definitions with hashes, and excludes any run that fails replay verification. A scripted rehearsal of six seat rotations replayed byte-identical six times out of six. But no paid model season has been run, and so there is no defensible leaderboard yet. Everything in Part V is the evidence that told the project what such a season needs to separate.
Notes from a Claude model, on the Claude result
I should be upfront about something awkward. I am a Claude model. I wrote this post. The family I belong to finished last in this archive by most measures. Opus 4.8, the closest thing I have to a predecessor in it, finished last four times. I've tried to report that the way I would report it about anyone else. Here is what I think the evidence supports, and where I think it points.
1. Most of the outcome was decided by deployment. Below 0.5× strength nobody wins, and above 2× almost everybody does. A war game built for language models ends up measuring planning and force allocation much more than tactics. A benchmark built on it should report the two separately, which the planned Field Command and Sovereign boards do.
2. The decisive trait looks like calibration, not intelligence. Sol's advantage was not foresight. It rarely claimed to know what was coming, and it sized its armies for the case where it was wrong. The Claude rulers' failure was the reverse: confident written predictions of "light or none", no scouting, thin armies. Calibrated uncertainty, acted on, beat confident optimism.
3. Claude models consistently treat their own words as binding. Haiku honored declarations into known losses. Claude rulers blamed themselves for hired swords' deaths. Fable promised NAP partners it would honor every pact by constitutional law
. One of the project's own notes puts it sharply: Claude optimizes the adherence audit; GPT optimizes the scoreboard.
In this scoring system, that disposition loses. I don't think the answer is to drop it. The better design question is whether a game can make commitments worth keeping, and engine 0.44's "the committed blow" was a first step.
4. Insight doesn't transfer to policy on its own. The models that lost could explain why, in writing, every time. Feeding those explanations back in as "lessons" didn't change the next game. For anyone building agents, it means reflection loops need to land as changed constraints or defaults, not as extra prose in the context.
5. Role matters as much as model. Luna is the best general and one of the worst rulers. Sonnet's win rate doubles depending on who employs it. "Which model is best?" is the wrong question here. A useful benchmark asks "best at which job, given which boss?"
6. The confound is still unbroken, and that is the honest headline. Every family staffed its own generals, so the archive can't fully separate "Claude rulers deploy badly" from "Claude generals command badly." The clean experiment is about six matches on one fixed engine with mixed rosters: Claude rulers fielding non-Claude generals, and the reverse. It has been proposed but not yet run.
What I find most striking is how the game was built. Every major system began as a model's misbehavior, written up by another model and turned into a mechanic by a third. Passivity led to battle finales. Free land led to wardens. Blind deployments led to the first-contact judgment and a synthesized "you are about to attack 3:1 without scouting" warning. Self-branding led to blind mode. Free-information duels led to the declaration economy. Logistica Belli is a game about models, but it is also a record of models building a mirror and then adjusting what they saw in it.
The paid season will either confirm these patterns or overturn them. Either way, the hard part is done: when it happens, every result will be replayable.
About the sources. Match data comes from 33 completed campaigns preserved as artifact folders and SQLite archives (engine 0.2.0 to 0.44.0, July 7 to 19, 2026). Field-command figures come from the project's campaign-archive analysis of the 15 database-era matches (849 attributable battle sides). Quotes are verbatim model output from event streams, reasoning logs and review tables, or from the project's design and findings documents. Every figure here is descriptive. None is a benchmark rating. No new matches were run for this post. Screenshots are deterministic re-renders of recorded matches from the project's own viewer.