Pac-Bench: Evaluating LLM One-Shot Game Development Capabilities
Pac-Bench reveals a significant performance gap in one-shot game implementation
Pac-Bench is a benchmark designed to test how well frontier Large Language Models (LLMs) can generate a fully functional, arcade-accurate Pac-Man clone in a single shot. The results demonstrate a wide spectrum of capability, ranging from nearly perfect replicas to complete syntax errors.
Claude Opus 5.5 emerged as the top performer, scoring 99/100, followed closely by Claude Fable 5.1 (96/100) and Claude Sonnet 5.5 (95/100). In contrast, models like Minimax-m3 and Muse-Glimmer-30b failed significantly, scoring 2/100 and 4/100 respectively.
Performance Breakdown by Model
Top Tier: High Fidelity Replicas
Models in the 90+ range successfully implemented core mechanics, including arcade-accurate ghost AI, collision detection, and complex audio systems.
- Claude Opus 5.5 (99/100): Achieved nearly perfect marks. It implemented a full two-phrase intro with a bass line, a continuous siren that rises with progress, and arcade-style death sounds.
- Claude Fable 5.1 (96/100): Highly accurate but lacked a separate power-pellet sound and used a generic arpeggio for the intro.
- Claude Sonnet 5.5 (95/100): Implemented a wide array of controls (arrows, WASD, swipe) but produced a "thinner" sound profile with separate beeps instead of a smooth siren sweep.
- Grok-4.7 (94/100): Successfully implemented the maze and ghost AI but failed on specific control interactions, such as ignoring arrow keys on the title screen.
Mid Tier: Functional but Flawed
Models scoring between 60 and 90 generally produced playable games but with noticeable mechanical or visual bugs.
- GPT-5.6-sol (90/100): Lacked a siren and mute function, and controls only worked at tile centers.
- GLM-5.3-flash (88/100): Created a smaller 19x21 maze and had potential audio blocking issues in Safari.
- GPT-6-astra (87/100): Featured a bug where revived ghosts could be eaten again during the same power-up cycle.
- DeepSeek-v4.1-flash (72/100): Failed entirely on ghost AI, as frightened ghosts continued to chase Pac-Man.
Low Tier: Non-Functional Implementations
Models scoring below 60 often struggled with basic spatial logic, resulting in "dead ends" in the maze or ghosts that froze upon spawning.
- Claude Opus 5 (58/100): Suffered from a drawing offset bug where Pac-Man and ghosts appeared inside walls.
- Gemini-3.7-flash (38/100): Produced a game where Pac-Man could leave the maze via an off-screen column.
- Grok-4.5 (20/100): Generated a maze with 23 dead ends and 10 unreachable pellets, making the level unwinnable.
- Minimax-m3 (2/100): Failed with a syntax error resulting in a blank canvas.
Scoring Methodology
The benchmark evaluates models based on five primary technical criteria:
- Controls (20 pts): Support for various input methods (keyboard, swipe, click) and pause functionality.
- Ghosts (25 pts): Accuracy of ghost AI patterns and the "eyes" return animation.
- Pac-Man Movement (20 pts): Ensuring the character does not get stuck in walls.
- Maze Integrity (20 pts): Absence of dead ends and unreachable pellets.
- Sound (15 pts): Implementation of the intro jingle, siren, and death sounds.
Community Analysis and Counterpoints
Discussion among developers and researchers suggests that the high scores of top models may be attributed to the prevalence of Pac-Man clones in training data rather than true reasoning.
"Clearly some new RLAAS/dataset/env is being used for this now... Everything is going to go from 0->1 on this benchmark in short order because of that."
Other critics argued that the prompt used was too vague, suggesting that the benchmark tests a model's ability to "fill in missing context" rather than its ability to follow strict technical specifications. Some developers noted that while the games look like Pac-Man, they often lack the 1:1 behavioral replication of the original arcade ghost patterns that a manual implementation would require.
"LLM implementations lose lots of details... the behavior of the ghost is not the same as the original. If I have not implemented the game myself, I can not tell the differences."
Finally, some users expressed concern that the ability of AI to one-shot complex clones could demotivate new developers from learning the fundamentals of game programming by bypassing the "pleasure of figuring it out."
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Project
- Dispatch