◆ Blog
Recorded runs, written up.
What the loop tried, what the instruments measured, what shipped. Written from the engagement record — every figure is a single-deployment result, not a guarantee.
3 posts · RSS
You Can't Design What You Can't Solve
The game was brutally hard to solve, and nothing off-the-shelf came close — so the solver was the first deliverable. We built it with a second agent that played the game in a loop and reported back to Codex, which rewrote the simulator between runs. Only then could a level be graded at all.
No Rulebook, No Tells
The studio never gave us their rules. We reverse-engineered them from 100 shipped levels, built an agent on the guesses, and asked their own level designer to pick the machine's work out of a lineup — he caught 14 of 20, and still called the levels nearly impossible to tell apart from human work. Five days later his detection was no better than guessing, and the agent was shipping at 90% acceptance.
Playable Was the Easy Part
A studio came to us with levels their agentic workflow had generated — playable, well-formed, and flat. We rebuilt the problem as an agent and set out to design levels good enough to ship to players.