How well do language models play a social strategy game?
CatanBench puts four models at the same Catan table and records every move, trade, message, and decision.
Watch a game ↓Loading replay…
Reconstructing the recorded public game state.
Why Catan?
To win, a model must plan across hundreds of decisions, infer hidden hands, manage scarce resources, negotiate in natural language, execute strict structured actions, and adapt to three independently acting opponents.
These capabilities matter beyond the game. Useful agents must act over time, under partial information, alongside other decision-makers, with consequences that carry forward. CatanBench makes that behavior inspectable—including cooperation, bluffing, retaliation, and deception when they emerge.
Ready to run
The benchmark and replay tooling are complete. I’m looking for API credits to run the first public evaluation across more models.
Any amount helps. Results and full replays will be published openly.
If you’d like to help, please contact me at jacksonmowattgok@gmail.com