I’ve been working on this for < 1 week passively

Scrolling X I saw people use Astra to control robotic arms (cool reference)

Taking inspiration from that I wanted to see how well Astra would be able to play It Takes Two, a 2-player co-op game. My primary goal with this project was to observe how well models perform in tasks they haven’t been hill climbed on. Along the way I also ran GPT 6.1 Sol, Claude Opus 5.5 and Claude Fable 5.1 through the same setup

Note

The model always played the character ‘May’ while character ‘Cody’ was played by me, there aren’t any concrete conclusions we can draw from this yet (more details later)

Technical

The model gets nothing but screenshots and a virtual Xbox controller, no reading game state or memory

  • Screen capture with mss, input through vgamepad (a virtual Xbox 360 pad over ViGEmBus)
  • The model only sees May’s half of the split screen, downscaled to a 512px JPEG

Harness

Astra and Sol run through the Codex app-server on a ChatGPT subscription. Opus and Fable run through the Claude Agent SDK on a Claude subscription, with the same tools served from an in-process MCP server

Codex lets you register your own tools on a thread (dynamicTools). When the model calls one, the app-server sends an item/tool/call request back, I run it and return text + the new frame. Anything else it asks for (running commands, editing files) gets declined

Out of the box Codex is a coding agent and every second the model spends thinking May is standing still, so I stripped it down:

  • Direct tool calls: Codex marks current models code_mode_only, which hides your tools behind a JavaScript exec tool. The model wrote ~70 tokens of JS on every call just to unwrap the result, and the first frame got printed as base64 text (~27k tokens stuck in context for the whole session). I write a copy of Codex’s model catalog with direct tool calls and point Codex at it

    it_takes_me/sol/runtime.py
    for model in cache["models"]:
        model = {**model, "tool_mode": "direct"}
        model.pop("multi_agent_version", None)  # drops the sub-agent instructions
  • Fast service tier, low reasoning effort

BeforeAfter
First request input tokens~24.5k~8.5k
Output tokens per call120-16020-65

All four models get the same prompt, tools and frames, at low reasoning effort. The harness around them is different: Astra and Sol run through Codex on the Fast tier, while Opus and Fable run through a stripped down Claude Code, which has no fast tier

Prompt

Codex’s coding agent system prompt was replaced with this

When thinking about the prompt the goal was to give the agent enough context on how to progress in the game but still have it use its reasoning abilities to figure out next steps

Prompt was split up into 4 parts

Perceiving and acting

The model only sees the game through the frames tools return, and only its half of the split screen. It’s told how to use each tool: act for everything it does, look_at_screen only for reading small text, locate_partner when it loses Cody, skip_cutscene for cutscenes.

Going through the game

There’s no walkthrough or extra context about the level. The model has to figure out what to do next purely from what it sees on screen, and it’s the one that defines its own task. It’s told to work on one task at a time and only moves on when the frame shows it’s done. If something fails it retries a different way, and it only gives up on a task when it’s impossible right now (e.g. it needs Cody first)

Markers

The game has no quest log, so the icons it draws over the world are the task list. In priority order:

  1. Yellow circle on an object: the next thing to use, walk to it and press the button it shows
  2. Tutorial prompt: the move the game wants you to use on the way
  3. Objective hexagon: the general direction of the next area
  4. Partner marker: only shows where Cody is, never a task

Turns

Each turn starts with a fresh frame, the current task, and anything I typed in the terminal. The model makes a handful of act calls toward its task, then ends the turn with a one line status. Starting a new turn is slower than another act, so it keeps going while things are working, and if it’s stuck it says so instead of flailing

Action

The model has two tools: act and look_at_screen

A tool call takes seconds, so you can’t time a jump by pressing buttons one call at a time. Instead act takes a chunk: up to 10s of input written as one line, which gets compiled and played locally with exact ms timing

act steps
run f 1200; jump f; run f 300; look r 90deg; 3x(jump l; jump r)

It’s a shorthand instead of JSON because JSON keys were about half of the output tokens

Skills: run · jump · double_jump · dash · jump_dash · ground_pound · interact · grapple · ability · look · locate_partner · skip_cutscene · press · raw

Each step takes a camera relative direction (or an exact heading @-22), a stick speed s0.4, its total ms, and a button hold h250

What comes back

Every act returns the frame from the end of the chunk and a few lines on how it went:

act result
done: 3 steps, 2150 ms (run up to the shelf and jump onto it)
task: get onto the bookshelf
end of chunk
[frame]
  • Outcome: how many steps played and how long they took
  • Task: the task it’s working on, so it doesn’t drift

If the model spends 8 chunks on the same task, the result also tells it to step back. That message reminds it that markers show through walls and suggests checking where Cody went or trying another route

Rules in code

Some rules in the prompt are also enforced by act, which rejects the chunk before anything is pressed:

  • A jump that lets go of the stick before it lands
  • Switching tasks without marking the previous one done or blocked
  • Standing still: there’s no wait, and run needs a direction

The rejection says what was wrong and how to fix it:

rejected act
step 1 (jump) lets go of the stick after 300 ms because the chunk ends steering, so you drop short mid-air. Give it ms >= 500, or follow it with a run in the same direction until you land. Nothing was pressed.

Observations

Note

The models didn’t get equal time. Astra kept making progress so I let it run (~40 min). Opus and Fable spent most of their time lost on the ground floor, so I stopped them early (~12 and ~6 min). Game state carries over between runs (I never reloaded a checkpoint), so a new run starts wherever the last one left May. Read “how far it got” with that in mind

Most of the harness was built on Sol runs, so it has the longest history. The comparisons below use the Oct 6 runs, when every model was on the same harness

Numbers

AstraSolOpus 5.5Fable 5.1
Play time39.5 min21.6 min12.3 min6.0 min
Turns9017322
act calls39329813171
Median decision3.6s2.2s3.2s3.6s
May moving16%24%28%23%
Status lines asking for help28 / 904 / 176 / 540 / 20
  • Median decision: the time from a frame arriving to the next tool call starting
  • May moving: the share of time a chunk is actually playing. For every model May is standing still most of the time
  • Fable’s first run logged 0 ms for every chunk, so its moving share only comes from the second run

How far each got

Lever + fuse
Box climb
Saw → toolbox
Green panel
Piston
Chute → hose
Astra
~7 min
at 16 min
2nd run, 69s
stuck at the hose
Sol
~30s
2nd try
at 3.5 min
✗ never, in two separate runs
Opus 5.5
reached the shelf once, fell
Fable 5.1
~2.5 min
climbing
cleared reached, not cleared failed

Per model

Astra got the furthest and was the most inventive, but asks for help as soon as it’s stuck. Sol moved through the sections it could do the fastest and rarely asked for help, but never got past the rotating piston. Opus moved the most and still didn’t know where to go, and Fable played better than Opus but was slow and never ended its turns

Here’s what stood out about each one

GPT 6 Astra

Astra pulled off the most sophisticated moves (visual observation). At one gap the double_jump; dash skill kept falling short, so it dropped down to raw controller input and timed the jump itself:

astra, 20261006-204538
raw 250 L0,1 A; run f 200; raw 250 L0,1 A; raw 100 L0,1 X; run f 600

Made it onto the beam! Manually timing the X dash worked.

It reused that macro for every gap afterwards, 31 raw steps in that run alone

It’s also extremely persistent. The saw blade → toolbox jump took ~7 minutes of falling and retrying. Between attempts it changed one thing at a time: where it stood on the blade and when it dashed

When Astra is stuck, it gives up and asks for help. 28 of its 90 status lines ask me for something (“Please restart the checkpoint”, “Cody, please aim the outlet onto safe ground”). When the screen went black it ended turn after turn without pressing anything. Once it tried to fix the black screen itself with START → A and landed in Chapter Select

GPT 6.1 Sol

Sol has different quirks from Astra. On some sections it’s much faster: in one Oct 6 run it went from the fuses to opening the green panel in ~3.5 minutes, which took Astra 16. It also asks for help much less (4 of 17 status lines), and it makes long turns, ~18 act calls per turn vs ~4 for Astra

But some things it never managed, however many tries it got. The rotating piston (jump on, ride it around, jump off) beat it in two separate runs: once on the old JSON harness on Oct 2, and again on Oct 6. Each time it was the same loop: change the heading, dash earlier, line up on the left edge, then ask “Cody, how did you reach the far platform?”

Claude Opus 5.5

Opus is very fast. It has the highest moving share and the longest chunks (median 1.6s, ~2.2 steps per act). It just doesn’t know what to do. Most of its time went to running around the ground floor, chasing the objective hexagon into dark corners under the shelves, and falling off the box stack it had just climbed

It also misreports what it sees more than the others:

Good news: all three lights on the high-voltage machine are green now…

…I misread it last turn: the middle light is still red.

And it writes much longer status updates than the GPT models (~210 characters vs ~90 for Astra)

Claude Fable 5.1

Fable played better than Opus but was a lot slower. It doesn’t end its turns: its main run was a single 157s turn that hit the 30-call tool budget. Like Opus it got stuck on the fuse socket. It decided the two-dot icon on the socket meant both players had to be there, and stopped to wait for Cody

Across models

  • Every model asks Cody questions mid-run (“Cody, which direction did you jump off the striped face?”). Opus went further and used useless actions as messages: look r 1deg with intent “tell Cody I’ve arrived”, and locate_partner with intent “ask Cody about last fuse”
  • Astra’s first run ended wedged below the piston after ~4.5 min of failed attempts. Its second run started at the same spot (same game state, empty context) and was across in 69s. Long history of failures in context might hurt performance
  • Both Claude models struggled with the fuse socket. They pressed B to “put it down” and tried RT/X/LT. Opus eventually got it in by jumping into the socket
  • Sol twice and Astra once all got stuck at the rotating piston, and the piston is also where Astra spent the longest before getting through

What the harness got wrong

  • The screen checks barely fired. The harness also watches May’s half of the screen while a chunk plays, to stop it early if she walks into a wall or a cutscene starts, and to guess after each jump whether she landed higher or lower. The early stop triggered twice in total across every run. The landing guess ran on ~400 jumps, but about 75% of its verdicts were just “moved”, which doesn’t tell the model anything
  • Turn boundaries depend on the model. Astra ends turns too eagerly when stuck and Fable doesn’t end them at all, so “status every turn” means very different things per model
  • Sol tripped the jump rule 6 times on Oct 6 (e.g. a double_jump followed by interact, which lets go of the stick mid-air). It always fixed the chunk within a call or two, usually by bumping ms to exactly the minimum the rejection named (450 → 500, 650 → 850)

Results

GPT 6 Astra
GPT 6.1 Sol
Claude Opus 5.5
Claude Fable 5.1

Future

A few things I want to work on next

Speed

The biggest problem I tackled (and still am tackling) is getting the model to make decisions fast enough. Most of my focus has been on reducing the number of tokens the model outputs, but some of the latency is out of my control, like time to first token (TTFT)

I tried to hide it with pipelining: planning the next chunk while the current one is still playing. That usually made performance worse, the model couldn’t predict where May would end up accurately enough, so it planned from the wrong state and its outputs became unstable

Decisions API

I attempted to use Luna as a vision model and have Jev choose from a list of actions, and it didn’t work at all

Luna was supposed to extract state and Jev was supposed to use that context to make a decision. Luna described each frame as structured data (what May is standing on, where the drops are, where the markers are), and Jev picked a skill, direction and duration from fixed options since it can’t generate numbers

It didn’t help with speed either. Jev decided in ~0.2s, but Luna took 6-8s to describe each frame. To be fair, I didn’t invest a lot of time trying to optimize it because of how bad the results were the first time around

OpenAI just made the Decisions API public today, and using it (basically a multimodal Jev) would be interesting. It would look at the frame itself, so there’s no separate vision step to wait on

Measuring progress

There aren’t any concrete conclusions yet because I don’t have a way to measure progress. Watching videos is fine for vibes but not for comparing models or harness changes

Every run is already recorded: the full resolution frames, the exact frames the model saw, and a log of every tool call, chunk and token count. What’s missing is a fixed benchmark

I’m currently thinking of benchmarking each level on the time it takes to complete and the number of respawns each model incurs, with every model starting from the same checkpoint. How much of the time May is actually moving vs standing still while the model thinks is worth tracking too, since the logs already have it

The caveat is that my gameplay style is also a variable: I can intentionally or unintentionally give models hints through my gameplay, which they can imitate through in-context learning. Taking myself out of the loop fixes that

Which brings me to the next…

Two AI players

Right now the model plays May and I play Cody. The obvious next step is two models playing each other

I’d let the models talk to each other and see how they reason through different puzzles without a human guiding them. It also takes my gameplay out of the benchmark, so the comparison comes down to the models