Scrabble
A board, the mover’s rack and the unseen tiles, ENABLE lexicon. Simulation scores every legal play by equity, not just points; the answer is a square and a word. See Methodology.
Fewer puzzles than the other models, so results are noisy.
Quality is the mean oracle score of the moves played, × 100. The best move scores 100; an illegal move or no move scores 0. Protocol boardbench-0.2.
Quality vs cost
4 of 4 models
- Anthropic
- OpenAI
- Pareto frontier: nothing plotted is both cheaper and better
- Anthropic
- OpenAI
Horizontal axis is logarithmic.
Table
| Model | Quality | Legality | Cost | n |
|---|---|---|---|---|
| Haiku 5.5Medium | 32.5 | 100.0 | <$0.01 | 9 |
| Sol (fewer samples)Medium | 14.1 | 33.3 | $0.06 | 3 |
| Sonnet 5.5 (fewer samples)Medium | 13.1 | 33.3 | $0.01 | 3 |
| Luna highMedium | 9.7 | 22.2 | <$0.01 | 9 |
Headline is the mean move score × 100. n is answers: puzzles × samples.