Splendor

Two-player base Splendor against a strong engine. Gems, cards, nobles; first to the usual prestige target. The model’s buys, takes, and reserves are scored against the engine oracle. The headline is ideal-decision rate, not final prestige. See Methodology.

Quality is exact ideal %, 0–100. Protocol boardbench-0.1.

Quality vs cost

12 of 12 models
  • Anthropic
  • OpenAI
  • SpaceXAI
  • DeepSeek
  • Moonshot AI
80859095100$0.01$0.1$1$10Claude Fable 5 Medium · 96.0% · $1.33Claude Opus 5 Medium · 89.1% · $2.35Claude Sonnet 5 Medium · 91.5% · $0.40DeepSeek V4 Flash 0731 High · 96.4% · $0.02 · Pareto frontierGPT-5.6 Luna High · 84.4% · $0.05GPT-5.6 Sol Medium · 77.2% · $0.40GPT-6 Astra High · 91.8% · $0.78Grok 4.5 Medium · 92.9% · $0.09Grok 4.5 Low · 86.2% · $0.11Grok 4.6 High · 93.0% · $0.10Grok 4.6 Medium · 86.5% · $0.09Kimi K3 Medium · 85.0% · $0.25DeepSeek V4 Flash 0731 HighClaude Fable 5 MediumGrok 4.6 HighGrok 4.5 MediumGPT-6 Astra HighClaude Sonnet 5 MediumClaude Opus 5 MediumGrok 4.6 MediumGrok 4.5 LowKimi K3 MediumGPT-5.6 Luna HighGPT-5.6 Sol Medium
  • Pareto frontier: nothing plotted is both cheaper and better
  • Same model at different efforts
  • Anthropic
  • OpenAI
  • SpaceXAI
  • DeepSeek
  • Moonshot AI

Horizontal axis is logarithmic.

Table

ModelQualityLegalityCostn
DeepSeek V4 Flash 0731High96.496.6$0.021 × 28
Fable 5Medium96.0100.0$1.331 × 25
Grok 4.6High93.098.3$0.102 × 57
Grok 4.5Medium92.9100.0$0.091 × 28
AstraHigh91.8100.0$0.782 × 49
Sonnet 5Medium91.590.1$0.402 × 59
Opus 5Medium89.1100.0$2.352 × 55
Grok 4.6Medium86.5100.0$0.092 × 52
Grok 4.5Low86.290.6$0.111 × 29
Kimi K3Medium85.096.8$0.252 × 60
LunaHigh84.484.1$0.051 × 32
SolMedium77.2100.0$0.402 × 57

Headline is exact ideal %. n is quality matches × scored decisions.

Look