Overall Models play full games against classical engines. The bar is the share of scored decisions that match the engine oracle, pooled across Yatzi , Scotland Yard , and Splendor .
Quality vs cost 12 of 12 models Anthropic
OpenAI
SpaceXAI
DeepSeek
Moonshot AI
Clear Select all Reset (to default) Save Load 80 85 90 95 100 $0.01 $0.1 $1 $10 Claude Fable 5 Medium · 91.4% · $1.66 · Pareto frontier Claude Opus 5 Medium · 81.1% · $1.66 Claude Sonnet 5 Medium · 86.1% · $0.30 · Pareto frontier DeepSeek V4 Flash 0731 High · 80.7% · $0.02 · Pareto frontier GPT-5.6 Luna High · 77.7% · $0.03 GPT-5.6 Sol Medium · 77.4% · $0.33 GPT-6 Astra High · 83.5% · $0.59 Grok 4.5 Medium · 80.9% · $0.06 · Pareto frontier Grok 4.5 Low · 77.3% · $0.07 Grok 4.6 High · 81.8% · $0.07 · Pareto frontier Grok 4.6 Medium · 82.0% · $0.07 · Pareto frontier Kimi K3 Medium · 82.0% · $0.14 Claude Fable 5 Medium Claude Sonnet 5 Medium GPT-6 Astra High Grok 4.6 Medium Kimi K3 Medium Grok 4.6 High Claude Opus 5 Medium Grok 4.5 Medium DeepSeek V4 Flash 0731 High GPT-5.6 Luna High GPT-5.6 Sol Medium Grok 4.5 Low Pareto frontier: nothing plotted is both cheaper and better Same model at different efforts Anthropic OpenAI SpaceXAI DeepSeek Moonshot AI Horizontal axis is logarithmic.