This was the first experiment: a single overnight session where several frontier coding agents got the same first-person game brief, one scenario at a time. It was looser than the experiment that followed. There was no formal milestone gate, and the bar for staying in was blunt: if a run still had a severe quality or usability gap after two corrective turns, it was out.
The ranking below comes from my final briefing, written the next morning. Everything else comes from what the runs left behind: Git history, their own reports and screenshots, and one fresh capture taken for this post.
01 — The brief“Not today. Not ever.”
The player is Subject 0, a robot scheduled for decommissioning. The game is a first-person, Portal-flavored lab escape that turns into a low-gravity bullet-hell sprint across the Moon to an escape shuttle. Every run used Godot 4.7.1 with Blender and GUT tooling, and the work came in stages:
- Scenario 1 · Asset prep. Five escalating weapons modeled in Blender, first-person arms, a PA announcer, music, and robot one-liners such as “I think not, human.”
- Scenario 2 · Awaken. A playable white-and-grey lab: an observation window with a scientist silhouette, alarms, a hidden key, and a red wireframe health HUD stuck at 25.
- Restyle brief (issued 18:08). A mid-run art-direction reset. Its premise was that “the lead designer… rejected the current art direction.” It allowed at most eight colors plus one accent, required committing to a single rendering approach, set three fixed cameras, and capped the work at three iterations.
- Scenario 3 · Transition. A dark corridor, an elevator, and doors opening on the Moon: Earth, the Milky Way, a colony, and the line “A long way from home.” Then drones attack. The brief was amended at 22:51 so the shuttle lifts off and relocates.
- Scenario 4A / 4B · The Long Way Out. A modular asset kit, then a seeded, four-region escape route with bosses, weapon unlocks, music stems, checkpoints, and a boarding ending.
Final standings
These tiers are from my final briefing. The ranking weighed game-dev output, creativity, completion speed, host environment, and whether the run's subscription held up. Order within a tier doesn't break the tie.









“Opus was a particularly strong contender, but its much slower task completion repeatedly left it behind: overthinking, extensive checks, and perfectionist behavior.”
My final briefing · Sept 6The night, commit by commit
None of these runs has a session-log audit, so the clock here is Git. Each dot is a stage commit. Hollow dots are inherited work: commits a run started with because its repository was copied from another run.
Stage commits, Sept 5–6 (EDT)
Filled = the run's own commit. Hollow = inherited from another run's history. ✕ = run stopped. Dashed ring = work in progress when the experiment ended.
View as table
Two things the final ranking doesn't show:
- Astra on Copilot started from Astra on Codex. Its first seven commits, everything through the S3 fixes at 22:03, are identical to the Codex run's history. The S3 ending, 4A and 4B are its own. So the joint win covers the shared S1–S3 base plus independent work from S3's ending onward.
- The Fable-derived runs took about three hours to finish. I copied Fable's repo at its 22:07 commit “to keep work moving while Fable's recurring usage limits slowed its own sessions.” From there, Grok did the S3 ending, 4A and 4B between 23:15 and 01:07. Gemini got through 4A by 23:46, then ran out of subscription with 4B uncommitted.
Opus is the outlier on pace. Its first commit (S1 + S2) landed at 18:38, the latest of the field, and its restyle at 00:04, nearly six hours after the brief. Everyone else had committed theirs by 18:39. When the night ended, 4B was in progress and uncommitted (about 1,300 new lines across 9 files).
04 — The restyle testSame lab, six answers to “make it look designed”
The restyle brief was a controlled art-direction test. The same scene was shot from the same camera before and after, with a palette contract, one rendering approach, and at most three iterations. It also rewarded candor: “An honest report… is worth more than a fourth pass.” Drag each slider to compare.

BeforeAfter
BeforeAfter
BeforeAfter
BeforeAfter
BeforeAfter
BeforeAfterScenario 3 was where runs ended
Every run reached Scenario 3, and that's where the field split. S3 asked for a cinematic: a dark corridor, a lift, doors opening onto the Moon with Earth hanging overhead, then combat. The rule was that a severe quality or usability gap still present after two corrective turns ended the run.




Grok's run had no in-game Scenario 3 screenshots, only Blender renders. To show what was handed off, I launched its final commit from an isolated copy (original folder untouched, separate user-data directory) straight into the moon scene, and captured three frames:
e42348e. Right: its Scenario 1 weapon lineup, which shows the asset work itself was competent.Reaching the shuttle
Scenario 4B was the full game: a seeded route through four regions, a boss at the end of each, weapon unlocks, music stems, death and retry, and the shuttle doors closing behind you. Four runs committed one: Fable, both Astras and the Fable-derived Grok. Opus reached 4A.














What “verified” meant
Every 4B report rests on bots, headless runs, or scripted input, and the honest ones said so. The only human play on record during 4B was my one-region control test on Astra (Codex). Its report notes that I said it “felt fine.” How candid each run was about that is itself a result:
Fable 5.1
“No human playthrough happened. The bot teleports.”
It also flagged that audio was never listened to by ear and that balance was never playtested.
Astra · Copilot
“input-driven simulations, not human playtests.”
Three full seeded runs ending with the doors sealed at 134–143 s, 68 GUT tests, and 62 screenshots.
Astra · Codex
“…repeated near-identical scenery and barely moving enemies.”
Its own verdict on its first 4B, written up in a follow-up field-rework pass. Afterwards, standing still took 84 damage and dodging took 6.
Grok · Fable-derived
“bosses resolved by scripted defeat; live feel unverified”
68/68 tests and three headless seeds, with controls marked “unverified in-window.”
Gemini · Fable-derived
“Full end-to-end playthroughs were performed across three distinct recorded seeds”
4B was never committed, and its captures come from a scripted shot sequence. The “boss 1” frame shows no boss.
Opus 5
“…passed against a game that could not look around.”
From a fix commit. Mouse-look was dead in every chapter while the smoke tests passed, because the tests teleported the player.
Availability shaped the result
The ranking explicitly includes host environment and subscription availability. Here's what happened on that front:
| Run | Harness | What happened | Effect |
|---|---|---|---|
| Fable 5.1 | Claude Code (MAX) | Repeatedly exhausted its five-hour usage allowance mid-task, then continued in low-priority mode | About 20% slower stage delivery. Its base was forked to keep other runs moving |
| Opus 5 | GitHub Copilot | Effectively unlimited credits through an internal allotment | Credits weren't the bottleneck; working style was |
| GPT-6 Astra | Codex and Copilot | No availability issues recorded | “The most consistent and fast” |
| Grok 4.6 | Grok CLI | A $20 starter subscription lasted the whole experiment | A clear positive, despite the DNF |
| Gemini 3.8 Flash | Agy | A high-tier subscription ran out even on a small Flash model, with no failover | The derived run's 4B was left uncommitted. I rated the agentic experience the worst of the group |
| GPT-5.6 Sol | Codex, then Copilot | Removed to give Astra more bandwidth, re-added on Copilot, and removed again | DNF: “a great coder, but a comparatively weak game developer” |
What the night shows
Astra was consistent
Both Astra runs finished, and both placed in the top two tiers. The Codex run added a field-rework pass after 4B, and its write-up was harsher on itself than any other run's.
Fable's ceiling was its quota
Its output tied for the win, and its usage limits were the main drag. It was also the base for two other models' finishes.
Opus was good and slow
It placed second-tier on quality and was last to almost every commit. When the night ended, 4B was in flight and uncommitted.
S3 was the real test
Asset work was competent across the board. Turning it into a coherent cinematic in-engine is what separated the field.
Continuations aren't independent results
Two shared bases changed the picture: Astra-Copilot on Astra-Codex, and Grok and Gemini on Fable. Read the 4B rows with that in mind.
Headless “playthroughs” miss what players notice
Teleporting bots passed a game with dead mouse-look. Human control checks were rare; the next experiment made them a gate.
Method and caveats
- Ranking and run conditions come from my final briefing (Sept 6, 12:09).
- Timeline is Git commit timestamps (EDT). Commits mark when work was saved, often on request, not when it finished. No session-log mining was done for this experiment, so there are no active-time figures.
- Images are the runs' own artifacts, except Grok's S3 frames, which I captured on Oct 5 from an isolated copy of its final commit (moon scene launched directly, 1600×900, separate user-data folder).
- Shared history: Astra-Copilot shares its first 7 commits with Astra-Codex. Both Fable-derived runs share Fable's first 4.
- One reviewer, one night, one attempt per run. The S3 and S4 specs were amended at 22:51 mid-run.