AI playtesting for game studios

Parallel playtesting on demand for game developers

Upload your game and agents test every edge case, finding game-breaking bugs, performance issues, softlocks, broken tutorials, and regressions.

HOW IT WORKS

1

Connect your dev build

minecraft-1.21-dev.zip

612 MB · uploaded 2 h ago

✓ build #214
2

Send agents to test your game

A goal in plain English — anything you can ask in a sentence. No scripts, no test code.

NEW TESTS

Survive the first night without dying

Craft a stone pickaxe from a fresh spawn

Reach a village before nightfall

Mine iron and smelt an ingot

+ Add test
runs · 25start · fixed seed — fresh spawn
Queue 4 tests · 100 runs
3

Agents fan out and play

Every goal becomes a test that runs in parallel on every new build — Keyer keeps score.

app.keyer.dev / minecraft

Minecraft

survival crafting · build #214 · uploaded 2 h ago

+ New test

151

runs on this build

5 / 7

goals passing

12

issues found

23 min

avg run length

GoalAgent runsGame-breaking bugsPerformance issuesWarnings
Survive the first night without dying3812View report
Craft a stone pickaxe from a fresh spawn311View report
Reach a village before nightfall22113View report
Mine iron and smelt an ingot1112View report
Craft a furnace19214View report
Trade with a villager before nightfall16115View report
Reach the Nether with iron gear14123View report

Every run: video · per-tick trace · referee verdict

4

Get the report

Bugs, crashes, performance and regressions from every suite — each one tied to a replay you can watch.

Minecraft / Reports /

Build #214

151 runscompared against build #213

5 / 7

goals passing

12

bugs found

3

new in this build

1

crash caught

58 h

of play, overnight

BUGS FOUND

High

Placed furnace vanishes after reloading the world save

Craft a furnace · 3 runs · Video · Trace

View error
High

Agent falls through an unloaded chunk at the cave entrance

Craft a furnace · 2 runs · Video · Trace

View error
Med

Villager trade window can't be reopened after closing it with Esc

Trade with a villager · 5 runs · Video · Trace

View error
Med

Furnace recipe doesn't appear until the inventory is closed and reopened

Craft a furnace · 4 runs · Video · Trace

View error
Low

Respawn screen flashes before the death animation finishes

Survive the first night · 1 runs · Video · Trace

View error

View all 12 bugs →

PERFORMANCE

Avg frame rate

141 fps

1% low

58 fps

Peak memory

3.2 GB

World load

11.9 s

⚠ +3.8 s vs #213

⚠ Leak suspect: +380 MB / hour in the 4 h soak run · trace ↗

CRASHES

Out of memory, 3 h into the soak run

1 of 151 runs · seed 8841 · Minidump · Video · Trace

Every line links to its replay — video · per-tick trace · referee verdict · JSON export

5

Relive and inspect any error

Watch what the agent played, read what it was thinking — or jump in and take the controls.

Minecraft / Craft a furnace / Run #15 /

Error · HighPlaced furnace vanishes after reloading the world save

Agent's view: the Minecraft crafting screen with cobblestone placed in the crafting gridOBJECTIVE: Craft a furnace 12:41 / 23:05+
WASD
ESPACE
Jump in here · 12:41
0:0023:05

MODEL THOUGHTS

t1284

Eight cobblestone in the inventory — that's enough. Heading back to the crafting table.

t1307

Crafting table open. I place the cobblestone in a ring and the furnace recipe lights up.

t1330

Furnace crafted and placed. Referee confirms the goal from the inventory state.

▶ Jump in at 12:41you take the controls — the game keeps going

Product dashboard, illustrative data. The demo runs are real →

RUNS ON BUILDS FROM EVERY MAJOR ENGINE

Unity logo

Unity

Unreal Engine logo

Unreal Engine

Godot logo

Godot

GameMaker logo

GameMaker

CryEngine logo

CryEngine

+ others

The coverage gap

Every build changes. Your playtests don't keep up.

Human playtests are slow and expensive.

Scheduling players, running sessions, writing up notes — by the time the report lands, the build it describes is three versions old.

Scripted bots test what you told them to.

Recorded input sequences shatter the moment a menu moves or a recipe changes. They can confirm a path still exists — they can't discover that it broke.

The bugs that matter live in the messy middle.

Not in the happy path and not in the unit tests — in the improvised, mistake-riddled way real players actually reach your content.

Capabilities

Built like an experiment, not a demo.

Pixels only

The agent gets rendered frames and the same controls a player has. No API hooks, no source access, no scripted paths. If a human can play the build, Keyer can test it.

Deterministic replays

Seeded worlds and tick-stepped simulation. A failure at tick 512 replays as a failure at tick 512 — on your machine, tomorrow, in CI.

Full evidence

Every run produces a captioned video of what the agent saw and thought, a per-tick input trace, every planning brief, and a machine-readable run report.

Private pilot

Start with one build and one question.

Can a new player reach the first crafting bench? Does the tutorial survive the UI refactor? What breaks when the recipe tree changes? Send us the build and the question — you get back the video, the trace, and the verdict.

Get a demoGeneral enquirieswe reply within a day

FAQ

Fair questions.

Do you need our source code?+

No. The agent interacts with a build the way a player does — rendered frames in, keyboard and mouse events out. The referee verifies outcomes the same way, from the game's own visible state — no SDK, no engine hook, no integration.

Is this replacing our QA team?+

No. It takes the part of QA nobody wants: running the same progression for the two-hundredth time, overnight, on every build. Humans keep the judgment-heavy work — what feels confusing, what feels unfair, what feels fun.

What kinds of games work?+

The harness is engine-agnostic — the three demo runs are a survival-crafting world, a commercial tower defense, and an arcade racer, with no game-specific code between them. It works best on games with deterministic or seedable starts and goals you can state in a sentence: 'reach the first bench', 'survive to round 60', 'finish in under a minute'.

How autonomous is it really?+

Every run on the demo page is unedited and fully autonomous: a planner writes one small brief at a time, an actor executes it from pixels, and a referee checks the result. Nobody touched the controls. Some briefs fail — those stay in the run report too — and the planner routes around them. When a run ends without a win, it's reported as exactly that.

What do we get after a run?+

A verdict verified against game state, the full video with the agent's reasoning captioned over it, the per-tick trace, every brief the planner wrote, and the facts the run added to memory. Everything is diffable between builds.