Overall

Models play full games against classical engines. The bar is the share of scored decisions that match the engine oracle, pooled across Yatzi, Scotland Yard, and Splendor.

Quality vs cost

12 of 12 models
  • Anthropic
  • OpenAI
  • SpaceXAI
  • DeepSeek
  • Moonshot AI
80859095100$0.01$0.1$1$10Claude Fable 5 Medium · 91.4% · $1.66 · Pareto frontierClaude Opus 5 Medium · 81.1% · $1.66Claude Sonnet 5 Medium · 86.1% · $0.30 · Pareto frontierDeepSeek V4 Flash 0731 High · 80.7% · $0.02 · Pareto frontierGPT-5.6 Luna High · 77.7% · $0.03GPT-5.6 Sol Medium · 77.4% · $0.33GPT-6 Astra High · 83.5% · $0.59Grok 4.5 Medium · 80.9% · $0.06 · Pareto frontierGrok 4.5 Low · 77.3% · $0.07Grok 4.6 High · 81.8% · $0.07 · Pareto frontierGrok 4.6 Medium · 82.0% · $0.07 · Pareto frontierKimi K3 Medium · 82.0% · $0.14Claude Fable 5 MediumClaude Sonnet 5 MediumGPT-6 Astra HighGrok 4.6 MediumKimi K3 MediumGrok 4.6 HighClaude Opus 5 MediumGrok 4.5 MediumDeepSeek V4 Flash 0731 HighGPT-5.6 Luna HighGPT-5.6 Sol MediumGrok 4.5 Low
  • Pareto frontier: nothing plotted is both cheaper and better
  • Same model at different efforts
  • Anthropic
  • OpenAI
  • SpaceXAI
  • DeepSeek
  • Moonshot AI

Horizontal axis is logarithmic.

Look