Agent Game Jams
I give frontier coding agents the same game-development brief, in the same order, with the same prompts, and watch what they ship. Each run gets its own repository, its own creative direction, and a real person playing the build at the end of every stage.
These write-ups stick to the record: timestamps, Git history, the bugs I reported, and what's on screen. Some runs won and some didn't finish. Both outcomes are in here.

Subject 0: nine runs, one lunar escape
A first-person robot escape from a lab to the Moon, built overnight in Godot and Blender. Scenario 3 decided the field: two joint winners, two finishes built on another model's code, and three DNFs.

Eight agents, one co-op shooter, ten hours
An overhead cyberpunk roguelite with an authoritative server, built in milestone gates with human review at each one. Six runs reached a procedural district, two ran out of time, and the fastest pair collected over half the bug reports.
The same models, two nights running
Most model families ran in both experiments. The second brief was overhead rather than first-person, and it was gated milestone by milestone with feedback at each gate. No run was cut for quality in experiment 02.
| Model · harness | 01 · Subject 0 | 02 · Co-op shooter | Across both |
|---|---|---|---|
| GPT-6 Astra · Codex | Joint 2ndFinished 4B, then reworked it | Reached M50 defects reported · “Demo played well” | Top tier both nights; cleanest record on night two |
| GPT-6 Astra · Copilot | Joint winnerS1–S3 shared with the Codex run | Stopped at M3 · speed9.4 h active, highly praised | Same model; much slower on night two |
| Fable 5.1 · Claude Code | Joint winnerRepeatedly hit usage limits | Reached M51 defect · longest M5 | Top tier both nights, paced by its quota |
| Opus 5 · Copilot | Joint 2nd4B uncommitted at the end | Stopped at M3 · speedLongest M1 · 4 wrap-up nudges | Good work, slowest both nights |
| Gemini 3.8 Flash · Agy | DNF at S3A Fable-derived continuation placed 6th | Reached M5Fastest run · 5 defects | Fast, with lots of repair rounds |
| Gemini 3.8 Flash · Copilot | Not entered | Reached M59 defects, the most of any run | Same speed as on Agy, more bugs |
| Grok 4.6 · Grok CLI | DNF at S3A Fable-derived continuation placed 5th | Reached M53 defects · “No immediate feedback” at M2 | Much stronger on the gated brief |
| GPT-5.6 Sol · Codex | Removed“A great coder, but a comparatively weak game developer” | Reached M5“Great work!” at M3 · 4 defects | Much stronger on the gated brief |
Patterns so far
Two models held up both nights
GPT-6 Astra on Codex (joint 2nd, then the only run with zero defects reported) and Fable 5.1 (joint winner, then one defect report) were near the top both times.
Thoroughness has a clock
Opus 5 was praised and slowest both nights. Astra on Copilot went from joint winner to running out of time. Heavy self-verification was the common thread.
Gates rescue weaker starts
Grok, Gemini and Sol all failed to finish the first brief. With milestone gates and feedback at each one, all three reached the final milestone of the second.
Availability is part of the score
Usage caps, subscriptions with no failover, and harness quirks moved results as much as raw ability, especially on a single-night clock.
Agents don't hold controllers
Inverted facing, dead mouse-look, broken controller exit: the bugs that blocked play were the ones a headless test can't feel.
Caveat: n = 1
One reviewer, one attempt per run, one night each. Read these as field notes, not leaderboards.