Field notes · September 2026

Agent Game Jams

I give frontier coding agents the same game-development brief, in the same order, with the same prompts, and watch what they ship. Each run gets its own repository, its own creative direction, and a real person playing the build at the end of every stage.

These write-ups stick to the record: timestamps, Git history, the bugs I reported, and what's on screen. Some runs won and some didn't finish. Both outcomes are in here.

Subject 0: first-person view on the lunar surface with Earth overhead and drones firing
Experiment 01 · Sat, Sept 5 2026

Subject 0: nine runs, one lunar escape

A first-person robot escape from a lab to the Moon, built overnight in Godot and Blender. Scenario 3 decided the field: two joint winners, two finishes built on another model's code, and three DNFs.

Astra (Copilot) & Fable tie3 DNF
9runs
~11 hone night
S3where the cut happened
Read the write-up →
Cyberpunk co-op shooter: neon overhead arena from Gemini Flash's M3 build
Experiment 02 · Sun, Sept 6 2026

Eight agents, one co-op shooter, ten hours

An overhead cyberpunk roguelite with an authoritative server, built in milestone gates with human review at each one. Six runs reached a procedural district, two ran out of time, and the fastest pair collected over half the bug reports.

6 reached M52 stopped for speed
8runs
2.9–9.4 hactive time
25defects reported
Read the write-up →

The same models, two nights running

Most model families ran in both experiments. The second brief was overhead rather than first-person, and it was gated milestone by milestone with feedback at each gate. No run was cut for quality in experiment 02.

Model · harness01 · Subject 002 · Co-op shooterAcross both
GPT-6 Astra · Codex Joint 2ndFinished 4B, then reworked it Reached M50 defects reported · “Demo played well” Top tier both nights; cleanest record on night two
GPT-6 Astra · Copilot Joint winnerS1–S3 shared with the Codex run Stopped at M3 · speed9.4 h active, highly praised Same model; much slower on night two
Fable 5.1 · Claude Code Joint winnerRepeatedly hit usage limits Reached M51 defect · longest M5 Top tier both nights, paced by its quota
Opus 5 · Copilot Joint 2nd4B uncommitted at the end Stopped at M3 · speedLongest M1 · 4 wrap-up nudges Good work, slowest both nights
Gemini 3.8 Flash · Agy DNF at S3A Fable-derived continuation placed 6th Reached M5Fastest run · 5 defects Fast, with lots of repair rounds
Gemini 3.8 Flash · Copilot Not entered Reached M59 defects, the most of any run Same speed as on Agy, more bugs
Grok 4.6 · Grok CLI DNF at S3A Fable-derived continuation placed 5th Reached M53 defects · “No immediate feedback” at M2 Much stronger on the gated brief
GPT-5.6 Sol · Codex Removed“A great coder, but a comparatively weak game developer” Reached M5“Great work!” at M3 · 4 defects Much stronger on the gated brief

Patterns so far

Two models held up both nights

GPT-6 Astra on Codex (joint 2nd, then the only run with zero defects reported) and Fable 5.1 (joint winner, then one defect report) were near the top both times.

Thoroughness has a clock

Opus 5 was praised and slowest both nights. Astra on Copilot went from joint winner to running out of time. Heavy self-verification was the common thread.

Gates rescue weaker starts

Grok, Gemini and Sol all failed to finish the first brief. With milestone gates and feedback at each one, all three reached the final milestone of the second.

Availability is part of the score

Usage caps, subscriptions with no failover, and harness quirks moved results as much as raw ability, especially on a single-night clock.

Agents don't hold controllers

Inverted facing, dead mouse-look, broken controller exit: the bugs that blocked play were the ones a headless test can't feel.

Caveat: n = 1

One reviewer, one attempt per run, one night each. Read these as field notes, not leaderboards.