Scotland Yard
The model is Mr. X. Three strong engine detectives hunt. Hidden movement, periodic reveals, ticket economy. Only Mr. X is quality-scored. The headline is how often Mr. X’s choice matches the oracle, not whether the game ended in escape. See Methodology.
- Claude Fable 5Medium100.0$2.07
- Claude Sonnet 5Medium93.8$0.08
- GPT-6 AstraHigh93.8$0.40
- GPT-5.6 LunaHigh86.7$0.01
- Grok 4.6Medium85.3$0.04
- GPT-5.6 SolMedium85.0$0.18
- Grok 4.5Medium84.8$0.02
- DeepSeek V4 Flash 0731High82.8<$0.01
- Kimi K3Medium81.5$0.03
- Claude Opus 5Medium66.7$0.21
- Grok 4.5High66.7$0.01
- Grok 4.5Low66.7$0.01
- Grok 4.6High60.0$0.01
Quality is exact ideal %, 0–100. Protocol boardbench-0.1.
Quality vs cost
13 of 13 models
- Anthropic
- OpenAI
- SpaceXAI
- DeepSeek
- Moonshot AI
- Pareto frontier: nothing plotted is both cheaper and better
- Same model at different efforts
- Anthropic
- OpenAI
- SpaceXAI
- DeepSeek
- Moonshot AI
Horizontal axis is logarithmic.
Table
| Model | Quality | Legality | Cost | n |
|---|---|---|---|---|
| Fable 5Medium | 100.0 | 100.0 | $2.07 | 1 × 26 |
| Sonnet 5Medium | 93.8 | 100.0 | $0.08 | 3 × 32 |
| AstraHigh | 93.8 | 100.0 | $0.40 | 2 × 32 |
| LunaHigh | 86.7 | 100.0 | $0.01 | 2 × 15 |
| Grok 4.6Medium | 85.3 | 100.0 | $0.04 | 2 × 34 |
| SolMedium | 85.0 | 100.0 | $0.18 | 2 × 20 |
| Grok 4.5Medium | 84.8 | 100.0 | $0.02 | 3 × 33 |
| DeepSeek V4 Flash 0731High | 82.8 | 100.0 | <$0.01 | 4 × 29 |
| Kimi K3Medium | 81.5 | 100.0 | $0.03 | 4 × 27 |
| Opus 5Medium | 66.7 | 100.0 | $0.21 | 2 × 9 |
| Grok 4.5High | 66.7 | 100.0 | $0.01 | 1 × 3 |
| Grok 4.5Low | 66.7 | 100.0 | $0.01 | 1 × 3 |
| Grok 4.6High | 60.0 | 100.0 | $0.01 | 2 × 10 |
Headline is exact ideal %. n is quality matches × scored decisions.