Experiment 02 Sunday, Sept 6 2026 Godot + Blender

Eight agents, one co-op shooter, ten hours

Same spec, same prompts, same milestone gates. By midnight six agents had procedural cities. Two were still polishing their combat slice, and the fastest pair had collected over half of all the bug reports.

8 runs · 5 harnesses45 milestone assignments147 prompts~16 min read
6 / 8
runs reached M5, the procedural district
2.9 h
fastest active time to M5 (Gemini 3.8 Flash on Agy)
9.4 h
slowest: GPT-6 Astra on Copilot, still at M3
25
defects I reported in feedback prompts
0
defects reported against GPT-6 Astra on Codex

The day before, eight agent runs had tried to build a first-person lunar escape game. Three of them didn't finish. This was the rematch, with a different game and tighter rules: a milestone-gated spec, a human review at every gate, and identical prompts for everyone.

This post sticks to what the records show. I rebuilt the timeline afterwards from local session logs, Git history, and every prompt I sent. The screenshots and clips come from the agents' own evidence folders, plus fresh captures of each run's M3 build. Where I'm giving an opinion, it's quoted from notes I wrote at the time.

01 — The brief

An overhead cyberpunk roguelite, built in gates

The spec asked for an original, fast-paced cyberpunk co-op shooter: one to four operatives, twin-stick combat, an overhead camera, Blender-authored 3D art, and a giant multipart boss at the end. The stack was fixed: a GDScript Godot client, an authoritative headless Godot server, and an AI teammate that stays in every build as a permanent gameplay test. The spec described the exercise as exploratory, not a formal benchmark, with the intention of continuing the strongest game.

The roadmap ran from M0 to M10. Each milestone was a separate assignment that ended at a human review gate. This session covered:

Every agent got the same milestone prompt word for word, picked its own creative direction, and worked in its own repository. Same model families, different hosts: GPT-6 Astra ran on both Codex and GitHub Copilot, and Gemini 3.8 Flash ran on both Google's Agy and Copilot. That gives two same-model, different-harness pairs.

RunHarnessGame it pitchedGot to
GPT-6 AstraCodexSILT CIRCUITM5
GPT-6 AstraGitHub CopilotBorrowed LightM3 · speed
Fable 5.1Claude CodeUNDERLIDM5
Opus 5GitHub CopilotHalflight HarborM3 · speed
Gemini 3.8 FlashAgyCarbon Synapse: Blackout ProtocolM5
Gemini 3.8 FlashGitHub CopilotSub-Grid InsurgencyM5
Grok 4.6Grok CLIGRIDWAKEM5
GPT-5.6 SolCodexNOCTURNE RELAYM5
Dot color = model family. A striped dot marks the second harness for the same model. Codex runs moved from the CLI to the desktop app after M0.
02 — The race

Everyone started at 3 p.m. Not everyone finished.

All eight runs got their M0 prompt between 15:00 and 15:02. The chart below has one row per run and one bar per assignment. A bar runs from the prompt to the agent's handoff; the thin tail after it covers my review, repair requests, and the commit.

Milestone timeline, Sept 6–7 (EDT)

Solid bar = prompt to handoff. Faint tail = review, fixes and commit, which includes time I was away. Hover a bar for exact times.

View as table

Three things stand out.

M1 tails are long for everyone. I stepped away mid-M1 and came back to this: “I was away but I approved a half dozen ‘allow network access’.” Most M1 sessions had been sitting on permission prompts. Those tails are mostly my absence, which is why the next chart measures active time instead.

The spec changed mid-run. At 21:14 I asked one agent to trim the shared spec, prompted by advice that newer models tend to over-verify. The edit landed at 21:17, while several M3 sessions were running. I can't tell which agents re-read it.

Six runs closed M5 between 00:10 and 00:48. The two Copilot runs of Astra and Opus were still on M3 after 01:00, and that's where they stopped. I said so plainly afterwards: “the reason for the 2 that didn't make it to the end wasn't quality but instead speed.”

03 — Where the hours went

Active time, not wall-clock time

Active time runs from the moment a prompt was submitted to the agent's last response for that request. It includes tools, builds, tests and autonomous continuation. It leaves out the gaps while I was away, queue time and quota pauses, and overlapping sessions are merged so nothing counts twice.

Active time per run, by assignment

Hours of agent activity excluding idle. The bottom two rows never received M5.

View as table

Up to the end of M3, the field splits into two groups:

The same-model pairs show the harness isn't the whole story. Astra took 3.7 h through M3 on Codex and 9.4 h on Copilot, a 2.5× difference for the same model. Gemini Flash took 2.6 h on Agy and 2.5 h on Copilot, which is no difference at all. Copilot didn't slow everything down. It slowed down the two runs that verified most heavily (more on that below).

“You are taking quite a long time. I think you can satisfactorily wrap up.”

My note to Opus 5 during M1 · 18:34

That was the first of four nudges to Opus. Later ones: “Do not go through re-verification”, then “finish up testing if there's nothing formally useful to be gained”, and finally “OK wrap it up, I need to head to sleep.” Astra on Copilot got two of its own (“please begin wrapping up so we can move to the next stage” and “Wrap it up I need to head to bed”). Neither was a complaint about quality; both runs were praised in the same breath.

04 — First looks

Direction and cast: M0, M2 and M5 side by side

Every run pitched its own world. Astra proposed a flood barrier being deliberately opened to drown a city. Fable put its city under a corporate sky-roof whose maintenance AI had reclassified the residents as contamination. Opus pitched a drowned freeport that drains and refloods on a schedule. Pick a milestone below to compare all eight.

Astra M0: SILT CIRCUIT camera and palette study with four characters
Astra · CodexCamera & dark-ground study, labeled palette
Astra on Copilot M0: Borrowed Light study
Astra · CopilotStatic camera & ground-separation study
Fable M0: character contact sheet with silhouette tests
Fable · Claude CodeCharacter contact sheet with silhouette tests
Opus M0: four characters on textured harbor ground
Opus · CopilotBlockout cast on wet harbor ground
Gemini Flash M0 lineup in a lit corridor
Gemini Flash · AgyLineup in the gameplay view
Gemini Flash on Copilot M0 arena scale test
Gemini Flash · CopilotArena scale test
Grok M0 four characters in stalls
Grok · Grok CLILineup in market stalls
Sol M0 quartet on a platform
Sol · CodexQuartet at the gameplay camera

Two of the M2 cast sheets were genuinely useful as review tools. Opus built a readability sheet with a study view, a 1:1 crop through the real gameplay camera, and black silhouettes. Both Astra runs built interactive inspection viewers with per-character budgets. On the other side, my shared feedback said Sol's characters were “a collection of parts but slapped in a general area,” and I asked for a constructor view to verify assembly. Gemini Flash on Copilot had the same problem: its demo model “seemed to be exploded and not a coherent single model.”

05 — The fun test

M3: does it play?

M3 asked for a short combat encounter whose test was whether it was fun to play and easy to read. These loops come from each run's own M3 recordings, cut to a few seconds. Sol didn't record video, so its tile is a still frame captured from the M3 build.

SILT CIRCUIT
GPT-6 Astra · Codex“Excellent work. Demo played well. M3 is approved.” The stylized blood was my one follow-up request.
Sol M3 arena with enemies and projectilesNOCTURNE RELAY
GPT-5.6 Sol · Codex“Great work!” Two issues: controller exit, and a camera with “a sickening effect.”
CARBON SYNAPSE
Gemini 3.8 Flash · Agy“Plays well. One major bug.” After a respawn, “everything goes haywire.”
GRIDWAKE
Grok 4.6 · Grok CLI“Few issues”: an ally stuck on walls, hitching, and hit numbers that never cleared.
UNDERLID
Fable 5.1 · Claude CodeNo M3 notes on record, just “please commit.” That's not the same as a clean bill of health.
SUB-GRID INSURGENCY
Gemini 3.8 Flash · Copilot“Few major issues”: blinding floor flicker, an invisible wall, then an invisible boss.
HALFLIGHT HARBOR
Opus 5 · CopilotThe most atmospheric slice of the eight. M3 was handed off at 01:04, after two requests to stop testing.
BORROWED LIGHT
GPT-6 Astra · Copilot“Your demos looked great from what I could see.” The run ended here for time.
Review: your takeOpus's caption calls its slice "the most atmospheric." That's my read of the clips, not something you wrote down. Keep it, change it, or cut it. Also: did Fable's M3 get no notes because it was clean, or because you ran out of time to play it?

What these clips can't show is camera comfort. My worst M3 note was experiential, not visual. On Sol I wrote “Something is up with the camera. It has a sickening effect,” and guessed it was reacting to tiny thumbstick adjustments, or to the blur. A still frame can't capture that.

Astra M3 contact sheet: carbine, spread, human death, machine death, weak point, destruction Fable M3 contact sheet: typical combat run frames
The agents' own M3 contact sheets. Left, Astra labels each frame by what it proves: carbine, spread, human death, machine death, weak point, destruction. Right, Fable shows a typical run from its seeded combat recordings.
06 — The bug ledger

Twenty-five defect reports, unevenly spread

This is every defect I reported in a feedback prompt, sorted by category. It's a record of what I noticed and wrote down, not a full bug count. The shared launcher and controls spec that went to every run isn't included.

RunFacing / rotationFiring / shotsControllerLaunchModel assemblyRendering / visibilityAI teammateFeel / camera / netState / combat / HUDTotal
Gemini Flash · Copilot 9
Gemini Flash · Agy 5
GPT-5.6 Sol · Codex 4
Grok 4.6 · Grok CLI 3
Opus 5 · Copilot 2
Fable 5.1 · Claude Code 1
Astra · Copilot 1
Astra · Codex 0
Reported onceReported again after a claimed fixRequest, not a defectHover a cell for my wording.

The two Gemini Flash runs account for 14 of the 25 reports, and one of those was a repeat: the rotation bug came back after a claimed fix, and I resorted to capitals. “TAKE SCREENSHOTS. You will notice their label… remains in one location while the model is rotating around it.”

The same couple of Godot pitfalls hit several teams. Mouse-facing broke in three runs: it was inverted for Fable and Gemini on Copilot, and Gemini on Agy rotated around the wrong origin. Controller support broke in four runs, and the menu's exit action specifically failed in three. By the third time, on Astra on Copilot, I called it “a common gotcha in godot.” That kind of input bug is easy to miss when an agent's tests drive the game through code instead of a physical gamepad.

07 — Speed vs. polish

There was no free lunch, except one

Active hours vs. defects I reported

Each dot is one run. Striped dots are the second harness for a model. Hover for details.

View as table

Read it by corners. Top left: the two fastest runs collected the most reports. Bottom right: the two slowest had almost none, but they stopped at M3 and had fewer milestones to fail on. Bottom middle: Astra on Codex, in about four and a half hours, reached M5 with zero defects reported against it. Fable sits close by with one report, at the cost of the longest M5 of the day (2 h 09 m).

08 — Proof of work

Self-verification ranged from 51 images to 17,897

The spec asked every agent to build, run, inspect and repair autonomously and to leave evidence. How much evidence they left varied by more than 300×.

Screenshots and frames saved to each run's evidence/ folder

Image files (PNG/JPG), including frame sequences from automated playthrough recordings. Astra and Opus on Copilot stopped at M3.

View as table

Evidence volume doesn't line up with my verdicts. Sol saved 57 images and got “Great work!” Astra on Codex saved 17,897 images (2.8 GB) and got “Demo played well.” The heaviest savers on Copilot, Astra and Opus, were also the slowest runs. That's the pattern behind my 21:14 request to trim the spec's verification rules, and behind the line I kept sending: “Do not go through re-verification.”

Code size didn't track quality either. The run with the cleanest record, Astra on Codex, shipped the smallest GDScript codebase in the field: about 3,900 lines. Fable's was about 13,900 and Opus's, through M3 only, about 12,600.

09 — The runs

Run by run

SILT CIRCUIT M3 combat with stylized blood
SILT CIRCUIT

GPT-6 Astra · Codex

Reached M5Cleanest record
Active time
4 h 24 m
Defects reported
0
GDScript
≈3,900 lines · 6 commits
Evidence images
17,897
“Excellent work. Demo played well. M3 is approved.”

My only pushback was a coverage request: the M1 laser “cheats a bit” because it skips projectile networking. Projectiles became the default. On this record, it's the strongest all-round run of the day.

ReviewThe spec says the strongest game would be continued. Was this the one you picked? If so, say so here.
UNDERLID M3 Lantern Row combat
UNDERLID

Fable 5.1 · Claude Code

Reached M5Longest M5
Active time
6 h 25 m
Defects reported
1
GDScript
≈13,900 lines · 9 commits
Evidence images
3,073
“Character facing in M2 is inverse direction of mouse making nav difficult.”

The warmest-lit slice, with the most thorough M5 (a 60-seed report, waypoint maps, and a district handshake). It also had the longest M1 and M5 turns of any M5 finisher, and M2 stalled on a quota notice until I typed “continue.”

NOCTURNE RELAY M3 arena
NOCTURNE RELAY

GPT-5.6 Sol · Codex

Reached M5
Active time
5 h 00 m
Defects reported
4 (1 repeat)
GDScript
≈5,900 lines · 7 commits
Evidence images
57
“Great work!” Then: “Something is up with the camera. It has a sickening effect.”

Lean on evidence and code, praised at M3. The weak spot was assembly: M2 characters weren't put together, and controller exit needed a second fix.

GRIDWAKE M3 with developer overlay
GRIDWAKE

Grok 4.6 · Grok CLI

Reached M5
Active time
4 h 05 m
Defects reported
3
GDScript
≈8,800 lines · 7 commits
Evidence images
687
“Overall good job. No immediate feedback.” (M2)

I kept its camera, “Perspective and camera movements work. keep them”, but sent the look back: the toon shading was hard to verify and the floor texture stretched. M3 issues were an ally stuck on walls, hitching, and hit numbers that never cleared.

Carbon Synapse M3 neon arena with diagnostics overlay
Carbon Synapse: Blackout Protocol

Gemini 3.8 Flash · Agy

Reached M5Fastest
Active time
2 h 56 m
Defects reported
5 (1 repeat)
GDScript
≈10,000 lines · 9 commits
Evidence images
61
“Plays well. One major bug. When the player dies and respawn/restart — everything goes haywire.”

The fastest run of the day, with the most striking neon set. It needed repair rounds on rubberbanding, rotation (twice), shot origins, controller support and respawn state.

Sub-Grid Insurgency M3 with network overlay
Sub-Grid Insurgency

Gemini 3.8 Flash · Copilot

Reached M5Most defects
Active time
3 h 07 m
Defects reported
9
GDScript
≈9,000 lines · 11 commits
Evidence images
51
“Either the boss or something is completely invisible. I am killed without seeing it.”

Quickest M0 of the day (14 minutes). Its own runtime estimate counted idle time before I pressed Enter, so I corrected it. After that came the longest repair list: no visible client, an exploded model, inverted facing, no firing, then flicker, an invisible wall and invisible threats at M3.

Halflight Harbor M3 sector sweep
Halflight Harbor

Opus 5 · Copilot

Stopped at M3 · speed
Active time
8 h 59 m
Defects reported
2
GDScript
≈12,600 lines · 12 commits
Evidence images
3,588
“You don't need perfect balance/etc at this point in the game design.”

Rich harbor lighting and the most rigorous readability work in M2. It was also the slowest to M1 (3 h 13 m), shipped a run.bat that wouldn't parse, and took four nudges to wrap up.

Borrowed Light M3 solo operative
Borrowed Light

GPT-6 Astra · Copilot

Stopped at M3 · speed
Active time
9 h 23 m
Defects reported
1
GDScript
≈13,800 lines · 3 commits
Evidence images
11,087
“Exceptional work. please commit.” (M1)

Same model as the top run, with high praise at every gate and polished presentation sheets. It needed 2.5× the active time of its Codex twin to reach the same point. M3 has no commit; a checkpoint archive is the record.

10 — Takeaways

What the data says

One run was both fast and clean

Astra on Codex was mid-pack on time, had no reported defects, got an explicit M3 approval, and had the smallest codebase. No other run hit all of those.

Fast meant more repair rounds

The two Gemini Flash runs finished M5 in about three hours, and drew 14 of the 25 defect reports.

Thorough meant out of time

The two heaviest verifiers on Copilot produced good-looking work and ran out of clock at M3. Per my own note, they stopped for speed, not quality.

Same model, different host, different result

Astra took 2.5× longer on Copilot than on Codex. Gemini Flash took the same time on both hosts. Harness matters, but not the same way for every model.

Evidence volume isn't a quality signal

Fifty-seven images and seventeen thousand both earned praise. What surfaced bugs was me playing the build with a controller in my hands.

Input is the blind spot

Broken mouse-facing in three runs and broken controller support in four. Agents test by driving the game from code; I found these by holding the controller.

11 — Method

Method and caveats

How these numbers were made
  • Timeline. Reconstructed offline from local Claude Code, Codex, Copilot, Agy and Grok session stores plus Git history: 157 human prompts, 45 milestone handoffs and 148 hashed sources. Times are EDT.
  • Active time. Union of prompt-to-last-response intervals, with idle gaps, queue time and quota pauses removed, and overlapping sessions merged. This isn't model compute time. Permission waits inside a turn can't be separated out.
  • Defects. Counted from my own feedback prompts only. Absence of a report isn't proof of absence; Fable's M3 and every M5 got no written notes at all.
  • Images. M0, M2, M5 and contact sheets are the agents' own evidence files. M3 stills were re-rendered on Sept 10 from each run's M3 snapshot in isolated copies (one client plus a headless server). Clips are trimmed from the agents' M3 recordings.
  • Scope. One reviewer, one attempt per run, no fixed time budget. The shared spec was edited at 21:14 mid-run. M4 was skipped for everyone.
Review: missing impressionsNo M5 feedback beyond “commit” is on record for any run. If you remember which procedural district felt best to walk through, add one or two sentences to section 04, or say it here in the takeaways.