Scotland Yard

The model is Mr. X. Three strong engine detectives hunt. Hidden movement, periodic reveals, ticket economy. Only Mr. X is quality-scored. The headline is how often Mr. X’s choice matches the oracle, not whether the game ended in escape. See Methodology.

Quality is exact ideal %, 0–100. Protocol boardbench-0.1.

Quality vs cost

13 of 13 models
  • Anthropic
  • OpenAI
  • SpaceXAI
  • DeepSeek
  • Moonshot AI
60708090100$0.001$0.01$0.1$1$10Claude Fable 5 Medium · 100.0% · $2.07 · Pareto frontierClaude Opus 5 Medium · 66.7% · $0.21Claude Sonnet 5 Medium · 93.8% · $0.08 · Pareto frontierDeepSeek V4 Flash 0731 High · 82.8% · <$0.01 · Pareto frontierGPT-5.6 Luna High · 86.7% · $0.01 · Pareto frontierGPT-5.6 Sol Medium · 85.0% · $0.18GPT-6 Astra High · 93.8% · $0.40Grok 4.5 High · 66.7% · $0.01Grok 4.5 Medium · 84.8% · $0.02Grok 4.5 Low · 66.7% · $0.01Grok 4.6 High · 60.0% · $0.01Grok 4.6 Medium · 85.3% · $0.04Kimi K3 Medium · 81.5% · $0.03Claude Fable 5 MediumClaude Sonnet 5 MediumGPT-6 Astra HighGPT-5.6 Luna HighGrok 4.6 MediumGPT-5.6 Sol MediumGrok 4.5 MediumDeepSeek V4 Flash 0731 HighKimi K3 MediumClaude Opus 5 MediumGrok 4.5 HighGrok 4.5 LowGrok 4.6 High
  • Pareto frontier: nothing plotted is both cheaper and better
  • Same model at different efforts
  • Anthropic
  • OpenAI
  • SpaceXAI
  • DeepSeek
  • Moonshot AI

Horizontal axis is logarithmic.

Table

ModelQualityLegalityCostn
Fable 5Medium100.0100.0$2.071 × 26
Sonnet 5Medium93.8100.0$0.083 × 32
AstraHigh93.8100.0$0.402 × 32
LunaHigh86.7100.0$0.012 × 15
Grok 4.6Medium85.3100.0$0.042 × 34
SolMedium85.0100.0$0.182 × 20
Grok 4.5Medium84.8100.0$0.023 × 33
DeepSeek V4 Flash 0731High82.8100.0<$0.014 × 29
Kimi K3Medium81.5100.0$0.034 × 27
Opus 5Medium66.7100.0$0.212 × 9
Grok 4.5High66.7100.0$0.011 × 3
Grok 4.5Low66.7100.0$0.011 × 3
Grok 4.6High60.0100.0$0.012 × 10

Headline is exact ideal %. n is quality matches × scored decisions.

Look