Scrabble

A board, the mover’s rack and the unseen tiles, ENABLE lexicon. Simulation scores every legal play by equity, not just points; the answer is a square and a word. See Methodology.

Fewer puzzles than the other models, so results are noisy.

Quality is the mean oracle score of the moves played, × 100. The best move scores 100; an illegal move or no move scores 0. Protocol boardbench-0.2.

Quality vs cost

4 of 4 models
  • Anthropic
  • OpenAI
10152025303540$0.001$0.01$0.1Claude Haiku 5.5 Medium · 32.5% · <$0.01 · Pareto frontierClaude Sonnet 5.5 Medium · 13.1% · $0.01GPT-6 Luna high Medium · 9.7% · <$0.01GPT-6.1 Sol Medium · 14.1% · $0.06Claude Haiku 5.5 MediumGPT-6.1 Sol MediumClaude Sonnet 5.5 MediumGPT-6 Luna high Medium
  • Pareto frontier: nothing plotted is both cheaper and better
  • Anthropic
  • OpenAI

Horizontal axis is logarithmic.

Table

ModelQualityLegalityCostn
Haiku 5.5Medium32.5100.0<$0.019
Sol (fewer samples)Medium14.133.3$0.063
Sonnet 5.5 (fewer samples)Medium13.133.3$0.013
Luna highMedium9.722.2<$0.019

Headline is the mean move score × 100. n is answers: puzzles × samples.

Look