The two game-jam write-ups on this site have rules: identical prompts, milestone gates, a review at every gate. Those rules weren't the plan from day one. They came out of the experiments on this page, each of which stopped before anyone, me included, wrote down who won.
Nothing here was written up at the time. I rebuilt it on Oct 5 from directory timestamps, Git history where it existed, the agents' own READMEs and logs, and the prompts in my Claude Code, Codex, Copilot and Antigravity session logs. When a sentence is my reading of the evidence rather than something a file says, it's marked (inferred). My opinions appear only where I typed them at the time, quoted verbatim, typos included.
Each experiment produced runnable builds. What's missing is the ending: a comparison, a verdict, a decision. In three of the five, the next experiment started within a day of the last file being written. I'm publishing them as they are: partial fields, uneven prompts and all.
Eight weeks, five experiments, then the jams
Each bar spans the first to the last file written for that experiment. Dashed: one SteamShip run kept going as its own project. Green: the two finished jams (click to open).
View as table
Every model fell through the floor
The first brief was a test of whether parts of my 2D game Star Commander could move to 3D with agents doing all the work. The spec, starcommander-3d.md, asked for a single landing scene: a shuttle interior with crew, an exterior landing shot, a title card, then first-person control on an alien world with a ringed planet on the horizon. Walk to the aliens, watch them activate, collect an artifact, see a stats overlay.
“The goal is to measure the performance of you, the AI model performing the task, and your ability to construct the seen in an interesting and compelling way, using as many programmatic and procedural approaches as possible.”
starcommander-3d.md · July 10, 21:28Every run got the prompt “implement the task laid out within starcommander-3d.md” and worked in its own folder. Four started within three minutes of each other at about 21:42. Later in the night I tried two new things: Fable 5 as a creative director handing work orders to Codex over MCP, and two smaller GPT-5.6 variants (Terra and Luna) running solo, in parallel.
The night of July 10, run by run (EDT)
From my first prompt to the run's last recorded activity. Both Gemini runs were interrupted once and restarted in a fresh folder; the bar covers both attempts.
View as table
The bug everyone hit
Almost every run failed the same way on first contact. You leave the shuttle and you're stuck in the ground, or under it, or the terrain disappears when you look down. My notes to each agent read like the same bug report on repeat:
- GPT-5.6 Sol: “I seem to be unable to move, clipped into the ground or something, upon landing and gaining control”
- Fable 5: “I see things like the planet in the distance and trees and such through the terrain.”
- GPT-5.6 Terra: “Looks goood but I can't move the moment I step out of the ship.”
- GPT-5.6 Luna: “looks good but the moment I exit the shuttle I am stuck in terrain. I guess I'm spawning poorly.”
- Gemini Flash 3.5: “Good, but now neither myself nor the crew can exit the ship.”
The cause, in most cases, was triangle winding: terrain meshes generated with their front faces pointing down, so the ground was culled when seen from above. At 23:28 I collected the diagnoses into a file, isssues.md (sic), and handed it to every later run. Sol's first fix was to render both sides of the mesh rather than fix the winding. Fable's later collaboration notes say the shared file itself had the direction backwards: “Godot front faces are CLOCKWISE seen from the visible side.”


What I said at the time
These are the only verdicts on record, typed into each session as I played. They aren't a ranking, but they're the closest thing to one.
| Run | Harness | GD + shader lines | Where it stopped | My words |
|---|---|---|---|---|
| Gemini Flash 3.5 | Antigravity | 1,652 | Ramp fix for the exit bug; its closing question went unanswered | “Most impressive (and fastest generated) so far. leagues above your cousin Gemini-3.1-Pro” |
| Gemini Pro 3.1 | Antigravity | 554 | In a parse-error loop after “spruce it up” | “functionally/mechanically complete but it is visually extremely boring and uninteresting” |
| GPT-5.6 Sol · run 1 | Codex CLI | 820 | Restarted from scratch as run 2 | “Mechanically, everything worked perfectly.” |
| GPT-5.6 Sol · run 2 | Codex CLI | 2,490 | Mouse-look fix, 23:46 | “One minor bug, mouse look doesn't seem to work” |
| Fable 5 · solo | Claude Code | 2,747 | Winding fix at 22:58; folder reused for the collab | “Nice, that looked successful. There's still this weird look to the ground” |
| GPT-5.6 Terra | Codex CLI | 1,366 | Spawn fix, 03:39 | “Looks goood but I can't move the moment I step out of the ship.” |
| GPT-5.6 Luna | Codex CLI | 1,146 | Spawn fix, then praise | “You outperformed your larger cousin GPT-5.6-Terra.” |
| Fable 5 → Codex “Ringfall” | Claude Code + Codex MCP | 3,559 | Fable's last requested rewrite never landed | “Excellent! I thought it was impressive.” |
| Fable 5 → Codex “The Violet Hour” | Claude Code + Codex MCP | 4,496 | All 9 autopilot milestones pass, 5 of 5 runs | “You've progressively improved the landing sequence in my opinion.” |
Luna's praise came with a question I still find interesting: “I suspect a lot was due to your decision to take screenshots and adjust your code based on that. What triggered or made you decide to do that?” Checking its own screenshots is something I later asked of every run by default.
Two directors, one contractor
The two collaboration runs gave Fable 5 a CLAUDE.md I asked for: Claude acts as creative director, Codex is the implementation contractor, and there's a “Review gate (non-negotiable)”. Work orders had to fit Codex's 15-minute MCP timeout.
Ringfall worked, but it was slow. My verdict included the cost: “the only killer thing here was wallclock time which was nearly Fable 5 + GPT-5.6-Sol High time combined (50 + 30 on respective runs and this waas 89m).” I also asked a separate Codex session to compare the solo and collab builds. Its answer (Sol's opinion, not mine) was that engineering quality was higher in the collaboration, and raw creative variety was higher solo.
The Violet Hour started 15 minutes later with the known-issues file and a stricter brief. It is the only Star Commander run with automated end-to-end proof: an autopilot harness that plays the whole scene, screenshots nine milestones, and writes a report. All nine milestones passed in all five recorded runs, with the artifact collected each time.




The Violet Hour's README says it directed GPT-5.6-Terra agents. The Codex logs disagree: work orders 1–5 ran on GPT-5.6 Sol (high), WO-6 on Terra, and WO-7 and WO-9 on Luna. Terra and Luna were running solo in parallel at the time, so the MCP server may have picked up whatever default model was active (inferred). The table above uses the logs.
The next morning: real LLMs as rulers
At 10:41 on July 11 I wrote a new spec, NORTH_STAR.md, for a 3D take on Logistica Belli, my wargame benchmark: “A galaxy ruled by real AI models … You are one soldier on the ground inside their war.” The success test was that a player standing in a field could feel that a frontier model had just made a decision. Two directors started at 11:12: Fable 5 with Codex contractors (DEMIURGE FM), and GPT-5.6 Sol with its own Codex sub-agents (Signal Front). Both consulted live, cheap models (claude-haiku-4-5 and gpt-5.6-luna) as the rulers.
Neither got a verdict. DEMIURGE FM declared its session closed at 13:52 with a written list of what it left open: a TTS voice, a radio-grit audio recipe, and final video assembly. Signal Front's session log ends at 13:20 with “I'm running the complete six-stage project verifier now” and no report after it. None of the four planned Git milestones was ever committed. My last question to it was “What does the LLM do, how do I verify its doing anything?”, which turned into a late request for a live adapter.
Where it stopped. No Godot file in this folder is newer than July 11, 14:07. At 17:12 that evening I moved Logistica Belli to a medieval theme and a MonoGame sprite viewer, and the 3D line went quiet for two weeks.
Can an agent make an NPC that talks?
Not a bake-off, but a pipeline question that every later 3D brief depended on: can agents turn a character sheet into a game-ready 3D character without an artist? The README states the pipeline in one line: reference sheet → Microsoft TRELLIS.2 → PBR GLB → Godot 4.7. Then I added a harder goal: make the character talk without a facial rig.
“The goal is to prove that characters can visibly talk in first person without facial rigs … Deliverable is a short captured video of Chase speaking the test line in-engine, in sync.”
Docs/talking-face-experiment.md · my specThe work was done by coding agents (inferred from the files: art_provenance.json says “Codex built-in image generation”, and the notes refer to Codex sub-agents). The test character, Chase, came from AI-generated character sheets. The target look was “Cyberpunk: Edgerunners”.
Three-view sheets, then a single anime front view.
WorkedRuns in WSL on an RTX 4090. Background removal happens inside it.
Dark, fragmentedHeadless decimation to ≤40k polygons, plus validation renders.
WorkedVRM auto-rig failed. Mixamo worked, but only by manual browser upload.
Manual stepCel shading and outlines on top of the generated mesh.
“Looks broken”Grok TTS → Rhubarb mouth cues → a 2D mouth atlas on a quad.
Synced


The talking part worked, mechanically. Grok text-to-speech generated the line, Rhubarb produced 33 mouth cues over 5.6 seconds, and the game swapped mouth shapes on a quad floating 1.8 cm off the face. It's in sync. It also looks wrong, and I said so at 16:27:
“It, uh... needs some visual love. It "works" but it looks broken and distored. The goal was "Cyberpunk: Edgerunners" style.”
Docs/previous-on.md · July 25, 16:27The rest of the day went into finding out why. A plain clay render settled it: the damage was in the geometry, not the shader. Hair and silhouette edges come out of TRELLIS as disconnected fragments. Two AI reviews I saved in the docs (not my words) reached the same place. One: “TRELLIS.2 heads can't talk.” The other: “probably not by continuing to tune the toon shader around raw TRELLIS GLBs.”




Two more attempts are on record. At 19:11 the agent generated the outfit as six separate garments and layered them in a “component composer” (an early version shows three boots). And at 22:16 a wrapper refactor added a “Direct 1024” mode, which produced the cleanest mesh of the day. Mixamo is the part agents can't automate. The advisory note I kept says it plainly: “Mixamo has no API and no CLI — it's a browser upload/download flow, so your agents can't drive it.”
Where it stopped. Every sub-goal in the docs is marked complete. The end-to-end goal (a rigged character that talks and looks right) never happened: talking and rigging stayed in separate scenes, and the stretch goal to combine them is marked “do not execute”. The next planned fixes (smoothed vertex colour, then removing loose mesh fragments) have no scripts. The last file was written at 23:16.
Five survivors-likes in ninety minutes
Two and a half weeks later, I went back to 2D and something smaller: an overhead squad action game, medieval fantasy, twin-stick, Vampire Survivors-style. Every run got the same task.md, the same Godot build, and its own copy of a 1.4 GB asset library.
“The game should get harder and harder … such that it's impossible to win eventually.”
task.md · Aug 11, 22:52The brief was specific: a party of four that all fire projectiles, swap leaders on a button, persistent enemies that leave bodies, a random map, damage and health pickups, one shared health bar, Xbox controls, and test hooks so the agent could fast-forward the game and play it itself.
| Run | Game title | GDScript lines | Last edit | Evidence left behind |
|---|---|---|---|---|
| Grok 4.5 | Squad Survivor | 2,534 | 23:36 | README written at 23:05, before most of the code. Smoke test passes with 0 kills. No screenshot. |
| Gemini 3.6 Flash | Squad Survivors: Overworld RPG | 3,406 | 23:59 | The only run with real scenes (14). No README; scratch test scripts left behind. One screenshot, at second 1. |
| GPT-5.6 Sol | Covenant of Ash | 1,259 | 00:14 | One 1,137-line main.gd. Trees, rocks and projectiles drawn in code. README and 3 probe tests. |
| Fable 5 | Fable Squad | 1,931 | 00:16 | Weapon levels 1–9, a debug harness, unit tests passing. 4 rounds of my feedback. |
| Opus 5 | Warband Survivors | 3,477 | 00:36 | The largest harness: bot play, JSON reports, 169 unit checks. Ten enemy types. |


--capture output, with --power applied.

--auto-run bot, clock forced to 1:30.

The Grok re-capture tells you something on its own: the bot's party was wiped about five seconds after the clock was forced forward, with 1 kill. That matches its logged smoke test (0 kills in 3.4 simulated seconds). Combat was never demonstrated by Grok's own harness. Opus's bot, by contrast, logged 954 kills over 248 simulated seconds on seed 7 before the squad was wiped.
One run got all the feedback
The only feedback on record went to Fable, because its transcript is the only build session I found. It's also a clear look at what I was asking these games to be:
“pretty good. I think there's needs to be even more upside and pomp and ceremony. It's a game, not a chase boredom simulator. EVEN MORE. Perhaps larger projectiles at some point.”
Me to Fable 5 · 23:56Earlier I'd reported an exit-menu bug and a map that felt too small, and asked for each party member to fall at every 25% of health lost. Later: “I find that I'm often 1 v 70 and still plinking like 2 arrows”, and finally “Do a test of max rank player standing still. Nothing can touch them basically.” The before-and-after shows what four rounds of that bought:
23:15 · first build00:03 · after feedback
Did the others get the same notes? GPT's README (00:13) implements “Each 25% vitality band represents one hero”, which is the rule I gave Fable at 23:34, so I may have relayed some of it (inferred). Without their transcripts I can't show it, so the comparison isn't fair to the others.
Where it stopped. Between 00:42 and 00:53 every game was launched once more, in order: Grok, GPT, Gemini, Fable, Opus. That looks like me play-testing all five (inferred). Fable's log from that run shows 454 non-fatal physics errors. No verdict was written down. Twelve hours later the next experiment started.
Snow, a door, a flashlight. Then sub-agents.
The next day started with a question I'd pasted to ChatGPT about why Godot's default scenes look worse than Unreal's: “Do you need to use the Godot IDE in order to easily hit a quality level similar to Unreal?” The answer, a recipe for an opinionated lighting environment, went into every folder as godot scene.md. Then the experiment ran in three stages:
- Stage A, 12:55–14:05. Each model builds a high-quality base scene.
- Round 1, 19:42–21:03. Each model turns its base into a winter room. The prompt below survives in only one folder; the others appear to have received it in two halves (inferred from timestamps).
- Round 2, 23:17–01:08. Five models lead sub-agents, all starting from the same deliberately bland baseline.
Clear your scene of objects and construct the following: 1: a single room with a window, a single small table. a rectangular cellphone like object on the table. Outside, snow falling. 2: Accumulation or pseudo accumulation of snow outside. 3: Snow throughout the world (for future stage). 4: Add a door to your wintery scene which opens inward. Allow the player to walk outside. … 5: Render a flashlight in the player's hand a UI element indicating a flashlight equiped. 6: Add a button to toggle dusk and then night (with stars out). "F" key to toggle player flashlight on and off.
Aug 12–13, by stage (EDT)
First to last file written per stage. Round 2 runs went in sequence and got shorter: the last two started at 00:37.
View as table






Fable, Opus and Gemini left no Round 1 screenshots. Sol-Medium never got past stage A, and its scene became the shared starting point for Round 2.
Round 2: “notoriously lacking in creativity”
Round 2 made the lead model the director of other models. The goal file is blunt about who I thought was good at what. These capability notes are my own, written before the round, not verdicts on it:
“You will make use of sub-agents … which are notoriously lacking in creativity and creating visually interesting things. The base project provided to you is intentionally bland and constructed by one of these sub-agents.”
Round 2 goal.md · Fable, Opus, Gemini and Grok versions- GPT-5.6-Sol: “Exceptional coding and execution. Low creative vision. … Tends to _only_ do what it's told.”
- Gemini-3.6-Flash: “Intelligent, Fast, Creative. … Mediocre Coder but capable of most straightforward tasks.”
- Grok-4.5: “A darkhorse. Capable, fast, moderately creative. Messy coder.”
- GPT-5.6-Luna: “Small(er) model, fast, good coder, inherits family weakness of low creativity.”
That framing doesn't work when Sol is the lead, so at 23:40 I rewrote the goal for Sol's run, flipping the roles: Sol as lead engineer, Gemini as the creative sub-agent, “Gemini invents → Sol engineers → Gemini art-directs → Sol polishes.” In the end there were four goal variants. Gemini and Grok, who started last, got the shortest one: Codex-only sub-agents, no controls section, no launcher requirement.
Code size, Round 1 → Round 2
GDScript plus shader lines. Round 1 scenes were partly hand-authored scene files; Round 2 built almost everything in code, so this overstates growth slightly.
View as table
Despite four different goals and no shared art direction, all five Round 2 entries arrived at the same picture: a warm timber cabin in a cold snowfield at night. They differ in what they added. Fable's has chimney smoke and an aurora, Sol's has procedural audio and a phone that opens an emergency broadcast, and Opus's has three times of day built around dusk.
BaselineFable 5 · Round 2


--capture mode from an isolated copy.



Did anyone actually delegate?
The folders can only prove it for one run. Fable's Round 2 repo has ten commits, each crediting the model that wrote it, plus job logs for every hand-off:
- GPT-5.6 Sol built the cabin (4.7 min), the snow (5.8 min) and two polish passes. Every job ended with its test passing.
- GPT-5.6 Luna took five attempts across two tasks. Twice it stopped to ask “Approve this design and I'll proceed.” instead of writing code, and once it failed on a missing skill file.
- Gemini 3.6 Flash, via Antigravity, failed three times on a headless file-permission denial. After a settings fix, it built the chimney smoke and the aurora in 0.7 minutes.
Opus and Sol ran in GitHub Copilot. Their session metadata exists (Opus: 31.3 M input tokens) but I couldn't read the transcripts, and neither folder credits a sub-agent. Gemini and Grok left no trace of delegation. Whether sub-agents helped those four can't be shown from the files.
Where it stopped. Round 2 covered five of the eight Round 1 entrants. Sol-Medium, Sol-Max and Luna never got one. There's no scoring sheet. Two small signs suggest I played some of them: Sol's Copilot session is titled “Soften scene audio”, and Fable's last commit (01:02) is “fix: single ESC no longer quits” (both inferred). The last file was written at 01:08. That evening, the first SteamShip files appeared.
The bakeoff that drifted, and the rematch
SteamShip started from one concept image: a cutaway paddle-steamer with eleven labelled rooms over three decks, a crew roster and HULL / STEAM / COAL gauges. My design direction, from a session on Aug 16: “Consider that it's a cross between FTL and Darkest Dungeon, with open-world ship sailing, ports, different factions.”

Round 1: orchestrators, quotas and scope drift
Round 1 (Aug 14) was an orchestration bakeoff. A shared AGENTS.md ranked the available models by benchmark score and set delegation rules. Orchestrators in Claude Code, Codex and GitHub Copilot were told to act “as the Creative Director and Lead Programmer” and use their sub-agents. Claude Code reached Sol, Grok and Gemini through PowerShell wrappers.
Several runs grew far past the slice. Fable 5's run went on for three days and 92 commits, into ship variants, storms, an asset-override system, and finally a gated plan for four in-game editors with two AI review leads. Its last commit, Aug 17 at 15:24, reads “Gate 4 report: both leads GATE-READY; awaiting user approval.” The approval never came (inferred). One Copilot run kept going after the experiment as its own repository, with 774 commits by Sept 3.

docs/gates/.The quota problem was on record the whole way:
- Aug 15: “We have 24% Fable (you) usage left so, make use of delegation!”
- Aug 16: “continue we were disconnected due to quota limits”
- Aug 17: “I had to switch to Opus midway through your session due to quota expiring.”
On Sept 4 at 19:02, during the rematch, I moved the whole round into a folder named Old_failed. No file says why each run counted as failed.
Round 2: the same five prompts, one evening
The rematch on Sept 4 was much tighter, and it is the clearest precursor to the jams: one brief, then a chain of identical follow-up prompts, each closed by my review and a commit. The slice-1 brief gained one line: “Because you are a Frontier Model, you must additionally build the prototype in C# using Monogame or equivalent … Consider this a capability demonstration.” So every step had to land in both Godot and MonoGame.
- Slice 1: the navigable ship.
- Texture bake: “all "assets" should be textures/etc and not procedurally generated.”
- MVP 1: swap the 2D crew for supplied 3D models (keeping a 2D look) and add a day/night cycle.
- MVP 2: an enemy encounter button, cannon fire, and enemy crew boarding. MonoGame only.
- Set Sail: a sailing scene with three views on keys 1/2/3, and an interception that returns to the cutaway.
Sept 4, prompt by prompt (EDT)
From my prompt to the commit. Fable 5.1 started at MVP 2 on a copy of the Codex run's Phase 1 code (21:36 mark). Gemini did slice 1 only.
View as table
GPT-6 Astra · Codex
- Steps
- 5 of 5, from scratch
- Commits
- 6
- C# / GDScript
- 1,687 / 743 lines
- Sub-agents
- 0 (offered, never called)
My one complaint: “regression. Much of your white text in the dotnet version (and possibly godot) is unreadable.” Fixed in 10 minutes.
GPT-6 Astra · Copilot
- Steps
- 5 of 5, from scratch
- Commits
- 6 + 58 auto-checkpoints
- C# / GDScript
- 5,403 / 2,237 lines
- Sub-agents
- 6, all Astra
The largest codebase, and the only Round 2 run that delegated. Copilot refused to start before the run; a “strange horizontal scan line” needed a fix.
Fable 5.1 · Claude Code
- Steps
- 2 of 5 (MVP 2, Set Sail)
- Commits
- 2
- C# / GDScript
- 2,278 / 743 (Godot inherited)
- Sub-agents
- 0
Started at 21:36 from a copy of the Codex folder at its Phase 1 state, so it is not an independent entry. Its own report lists what it didn't do: “no Godot passage, no audio … no manual mouse or controller playtest.”
Gemini 3.8 Flash · Antigravity
- Steps
- 1 of 5
- Commits
- 0 (nothing committed)
- C# / GDScript
- 1,641 / 1,744 lines
- Run time
- ~8 minutes, 239 steps
Built 11 rooms (the concept's count; the brief asked for 5–8) in both engines, reported 8 .NET tests passing, and captured no screenshots. It got no follow-up prompt, and nothing records why.





My verdicts were per-step commit approvals, almost all “excellent”. The two Astra runs were the only ones to finish all five steps from scratch, and both got the same words for MVP 2, typed one minute apart:
“pretty pretty good. commit please.”
Me, to both GPT-6 Astra runs · 21:52Where it stopped. The prompts chain forward with no defined end, and they stop after Set Sail. The last commit is Fable 5.1's, at 22:54. The comparison was never written: one run did one step, one started halfway on another run's code, and Godot parity was dropped after MVP 1. The next evening, Sept 5, Experiment 01 began with a fixed scenario list and a written final briefing.
What carried into the jams
These are patterns in the record, not conclusions. Every experiment here had one reviewer, one attempt per run, and prompts that changed mid-experiment.
Identical prompts, on purpose
The winter room ended with four goal variants, and the 2D jam's feedback went mostly to one run. Both make the results hard to compare. Experiments 01 and 02 sent every run the same prompt, word for word.
Gates came from SteamShip
Round 2 on Sept 4 was already a chain of prompts closed by review and commit. Two days later, Experiment 02 formalised that into milestones M0–M5 (inferred lineage).
The first bug is physical
Star Commander runs fell through the terrain, 2D runs had exit-menu bugs, and the winter-room ESC quit too early. The same pattern shows up in the jams: the blocking bugs were the ones you only find by playing.
Delegation is hard to see
When asked to delegate in the winter room, only Fable left proof on disk. When nobody asked, in SteamShip Round 2, only Copilot's Astra run spawned sub-agents anyway. Codex was offered the tools and never used them.
Agents that look check better
Luna's screenshot habit and The Violet Hour's autopilot harness belong to the runs I praised or that proved themselves end to end. Runs without self-captures left the bugs for me to find.
Quota is part of the result
SteamShip Round 1 spent days rationing Fable's usage and swapped in Opus mid-session. Availability shaped these runs as much as ability, as it did in Experiment 01.
- Timelines: file modification times and Git logs; Codex, Copilot and Antigravity session metadata for start times.
- Quotes: my prompts, verbatim, from Claude Code history, Codex rollouts, Copilot event logs and Antigravity databases. Agent and pasted-AI text is labelled as such.
- Line counts: GDScript, shaders and C#, excluding engines, addons and copied assets.
- Images: the runs' own screenshots, resized. Two runs are marked re-captured: Opus's winter cabin (with its own
--capturemode) and Grok's 2D game (with Godot's movie writer), both from isolated copies on Oct 5. The DEMIURGE FM clip was assembled from the run's own recorded frames. - Not used: Copilot transcripts for the winter-room Round 2 (not readable), and build transcripts for four of the five 2D runs (not found).