Yatzi

Nordic Yatzy against a strong engine. Fifteen boxes, no American joker rules. The model chooses how to keep dice and what to score. We score every real decision against the engine oracle. The headline is how often that decision matches the oracle. See Methodology.

Quality is exact ideal %, 0–100. Protocol boardbench-0.1.

Quality vs cost

12 of 12 models
  • Anthropic
  • OpenAI
  • SpaceXAI
  • DeepSeek
  • Moonshot AI
75808590$0.01$0.1$1$10Claude Fable 5 Medium · 83.3% · $1.59 · Pareto frontierClaude Opus 5 Medium · 77.2% · $2.42Claude Sonnet 5 Medium · 81.4% · $0.43 · Pareto frontierDeepSeek V4 Flash 0731 High · 74.4% · $0.05 · Pareto frontierGPT-5.6 Luna High · 73.5% · $0.02 · Pareto frontierGPT-5.6 Sol Medium · 75.6% · $0.43GPT-6 Astra High · 74.7% · $0.58Grok 4.5 Medium · 77.2% · $0.09Grok 4.5 Low · 72.1% · $0.09Grok 4.6 High · 76.5% · $0.09 · Pareto frontierGrok 4.6 Medium · 77.8% · $0.09 · Pareto frontierKimi K3 Medium · 80.0% · $0.25 · Pareto frontierClaude Fable 5 MediumClaude Sonnet 5 MediumKimi K3 MediumGrok 4.6 MediumClaude Opus 5 MediumGrok 4.5 MediumGrok 4.6 HighGPT-5.6 Sol MediumGPT-6 Astra HighDeepSeek V4 Flash 0731 HighGPT-5.6 Luna HighGrok 4.5 Low
  • Pareto frontier: nothing plotted is both cheaper and better
  • Same model at different efforts
  • Anthropic
  • OpenAI
  • SpaceXAI
  • DeepSeek
  • Moonshot AI

Horizontal axis is logarithmic.

Table

ModelQualityLegalityCostn
Fable 5Medium83.3100.0$1.591 × 42
Sonnet 5Medium81.499.4$0.433 × 118
Kimi K3Medium80.095.7$0.252 × 80
Grok 4.6Medium77.8100.0$0.092 × 81
Opus 5Medium77.2100.0$2.422 × 79
Grok 4.5Medium77.298.9$0.093 × 127
Grok 4.6High76.5100.0$0.092 × 81
SolMedium75.6100.0$0.432 × 82
AstraHigh74.7100.0$0.582 × 83
DeepSeek V4 Flash 0731High74.4100.0$0.052 × 78
LunaHigh73.5100.0$0.022 × 83
Grok 4.5Low72.198.3$0.091 × 43

Headline is exact ideal %. n is quality matches × scored decisions.

Look