Yatzi
Nordic Yatzy against a strong engine. Fifteen boxes, no American joker rules. The model chooses how to keep dice and what to score. We score every real decision against the engine oracle. The headline is how often that decision matches the oracle. See Methodology.
- Claude Fable 5Medium83.3$1.59
- Claude Sonnet 5Medium81.4$0.43
- Kimi K3Medium80.0$0.25
- Grok 4.6Medium77.8$0.09
- Claude Opus 5Medium77.2$2.42
- Grok 4.5Medium77.2$0.09
- Grok 4.6High76.5$0.09
- GPT-5.6 SolMedium75.6$0.43
- GPT-6 AstraHigh74.7$0.58
- DeepSeek V4 Flash 0731High74.4$0.05
- GPT-5.6 LunaHigh73.5$0.02
- Grok 4.5Low72.1$0.09
Quality is exact ideal %, 0–100. Protocol boardbench-0.1.
Quality vs cost
12 of 12 models
- Anthropic
- OpenAI
- SpaceXAI
- DeepSeek
- Moonshot AI
- Pareto frontier: nothing plotted is both cheaper and better
- Same model at different efforts
- Anthropic
- OpenAI
- SpaceXAI
- DeepSeek
- Moonshot AI
Horizontal axis is logarithmic.
Table
| Model | Quality | Legality | Cost | n |
|---|---|---|---|---|
| Fable 5Medium | 83.3 | 100.0 | $1.59 | 1 × 42 |
| Sonnet 5Medium | 81.4 | 99.4 | $0.43 | 3 × 118 |
| Kimi K3Medium | 80.0 | 95.7 | $0.25 | 2 × 80 |
| Grok 4.6Medium | 77.8 | 100.0 | $0.09 | 2 × 81 |
| Opus 5Medium | 77.2 | 100.0 | $2.42 | 2 × 79 |
| Grok 4.5Medium | 77.2 | 98.9 | $0.09 | 3 × 127 |
| Grok 4.6High | 76.5 | 100.0 | $0.09 | 2 × 81 |
| SolMedium | 75.6 | 100.0 | $0.43 | 2 × 82 |
| AstraHigh | 74.7 | 100.0 | $0.58 | 2 × 83 |
| DeepSeek V4 Flash 0731High | 74.4 | 100.0 | $0.05 | 2 × 78 |
| LunaHigh | 73.5 | 100.0 | $0.02 | 2 × 83 |
| Grok 4.5Low | 72.1 | 98.3 | $0.09 | 1 × 43 |
Headline is exact ideal %. n is quality matches × scored decisions.