Splendor
Two-player base Splendor against a strong engine. Gems, cards, nobles; first to the usual prestige target. The model’s buys, takes, and reserves are scored against the engine oracle. The headline is ideal-decision rate, not final prestige. See Methodology.
- DeepSeek V4 Flash 0731High96.4$0.02
- Claude Fable 5Medium96.0$1.33
- Grok 4.6High93.0$0.10
- Grok 4.5Medium92.9$0.09
- GPT-6 AstraHigh91.8$0.78
- Claude Sonnet 5Medium91.5$0.40
- Claude Opus 5Medium89.1$2.35
- Grok 4.6Medium86.5$0.09
- Grok 4.5Low86.2$0.11
- Kimi K3Medium85.0$0.25
- GPT-5.6 LunaHigh84.4$0.05
- GPT-5.6 SolMedium77.2$0.40
Quality is exact ideal %, 0–100. Protocol boardbench-0.1.
Quality vs cost
12 of 12 models
- Anthropic
- OpenAI
- SpaceXAI
- DeepSeek
- Moonshot AI
- Pareto frontier: nothing plotted is both cheaper and better
- Same model at different efforts
- Anthropic
- OpenAI
- SpaceXAI
- DeepSeek
- Moonshot AI
Horizontal axis is logarithmic.
Table
| Model | Quality | Legality | Cost | n |
|---|---|---|---|---|
| DeepSeek V4 Flash 0731High | 96.4 | 96.6 | $0.02 | 1 × 28 |
| Fable 5Medium | 96.0 | 100.0 | $1.33 | 1 × 25 |
| Grok 4.6High | 93.0 | 98.3 | $0.10 | 2 × 57 |
| Grok 4.5Medium | 92.9 | 100.0 | $0.09 | 1 × 28 |
| AstraHigh | 91.8 | 100.0 | $0.78 | 2 × 49 |
| Sonnet 5Medium | 91.5 | 90.1 | $0.40 | 2 × 59 |
| Opus 5Medium | 89.1 | 100.0 | $2.35 | 2 × 55 |
| Grok 4.6Medium | 86.5 | 100.0 | $0.09 | 2 × 52 |
| Grok 4.5Low | 86.2 | 90.6 | $0.11 | 1 × 29 |
| Kimi K3Medium | 85.0 | 96.8 | $0.25 | 2 × 60 |
| LunaHigh | 84.4 | 84.1 | $0.05 | 1 × 32 |
| SolMedium | 77.2 | 100.0 | $0.40 | 2 × 57 |
Headline is exact ideal %. n is quality matches × scored decisions.