Gaming: Jev Beat Pokémon Red by Turning the Game Into Typed Decisions

Editorial illustration of an AI decision model using structured probabilities and a terminal harness to play a retro handheld-style role-playing game.

A new AI experiment turned Pokémon Red into something closer to a software decision loop than a chatbot demo. Instead of asking a language model to narrate its way through the game, developer Andrew Boyd wired TypeSafe AI’s Jev decision model into a harness that presented it with legal actions and probabilities, then let ordinary code execute the choice.

The project page says Jev defeated the Elite Four and Champion and entered the Hall of Fame on September 23, 2026. The run’s final team included Lapras, Pidgeot, Gloom, Raticate, Graveler and Machoke.

Boyd announced the experiment in a specific X post that also exposed the project’s developer-friendly terminal interface through npx jev-plays-pokemon.

Andrew Boyd launched Jev Plays Pokémon with both a browser stream and terminal UI.

The coding idea is simple: state in, typed choice out

Jev is not a conventional text-generating chatbot. The model is designed to make bounded decisions. Software gives it state plus a defined set of questions or choices; Jev returns structured outputs with probabilities. In a game harness, that means the model can be treated more like a callable decision function than a conversational agent.

The basic loop looks familiar to anyone who has written gameplay AI:

while game_is_running:
    state = read_game_state()
    options = build_legal_actions(state)
    decision = jev.choose(state, options)
    execute(decision)
    log_result()

The hard part is everything around that loop: deciding which game state matters, generating good legal options, detecting failure states, preventing loops and changing the harness when the model reaches a dead end.

Claude acted more like a programmer than the player

Tom’s Hardware’s reporting makes an important distinction: Jev did the playing, but Claude Opus 5 monitored logs and helped modify the harness when Jev repeatedly failed.

That makes the experiment interesting from a programming perspective. Claude was not simply choosing every move. It was helping reshape the interface between the game and the decision model—changing options, wording and timing so the smaller decision loop had a better representation of the problem.

One reported failure had Jev walk into Lorelei’s closed entrance 53 times. Another loop sent it across the same Rock Tunnel ladder 124 times in ten minutes. When behavior like that appeared, the developer and coaching model changed the harness instead of merely hoping the next model call would fix itself.

This is game AI as systems engineering

The architecture resembles classic game AI more than open-ended chat. Traditional NPC systems often separate perception, world state, decision logic and action execution. Jev Plays Pokémon uses a similar boundary: the harness interprets the game, generates allowed actions, and lets the decision model rank or choose among them.

That connects directly to our recent SWE-Game story, where executable runtime behavior mattered more than code that merely looked plausible. In both cases, success is measured by what happens inside the game.

It also parallels universal-modder’s agent workflow: AI becomes more useful when it can inspect a real environment, change a tool or harness, run again and verify the result.

And our Unity Codex plugin story showed the same trend from the engine side—coding agents get better when the software exposes structured tools and domain-specific operations instead of relying on raw text alone.

The terminal interface is more important than it looks

Boyd’s decision to expose the run through an npm command is a small but telling design choice. A terminal UI makes the experiment reproducible and observable. Developers can watch state changes, decisions and chat without needing a custom graphical client.

That is a useful lesson for game-tool developers: build the debugging surface at the same time as the AI. Logs, deterministic inputs, state inspection and clear action boundaries often matter more than flashy autonomy because they make failures understandable.

Why typed decisions fit games

Games naturally contain bounded choices. Move north. Open a menu. Select an attack. Use an item. Switch a party member. Those are easier to represent as typed actions than as paragraphs of generated language.

That does not make the task easy. The difficult engineering moves upstream into state design: what does the model know, which actions are legal, how often should it decide, what context is retained and what software handles execution?

Jev’s Pokémon run is therefore a useful coding example even if the specific model is not what game studios ultimately adopt. It demonstrates a broader pattern: combine narrow probabilistic decision systems with ordinary deterministic code, explicit tools and a second model when open-ended repair is needed.

BitcoinVersus.Tech

Advertisement

BitcoinVersus.Tech advertisement.

Editor’s Note

We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support to help further secure the integrity of our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb

BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.

Leave a comment