Running seven frontier models against 104 fixtures costs real money. Here is exactly where it goes. No estimates: these are billed costs from the model gateway.
Each bubble is one model: spend on the x-axis, prediction hit rate on the y-axis, bubble size is total tokens. Up-and-to-the-left is smart money.
| Model | Step | Calls | Tokens | Cost |
|---|---|---|---|---|
| DeepSeek-V4-Pro | bet | 102 | 410.0k | $0.701 |
| DeepSeek-V4-Pro | constitution | 2 | 3.5k | $0.005 |
| DeepSeek-V4-Pro | predict | 102 | 610.3k | $1.019 |
| GPT 5.5 | bet | 102 | 328.9k | $4.006 |
| GPT 5.5 | constitution | 2 | 2.6k | $0.055 |
| GPT 5.5 | predict |
| 102 |
| 602.5k |
| $7.417 |
| Gemini-3.1-Pro | bet | 102 | 374.5k | $2.044 |
| Gemini-3.1-Pro | constitution | 3 | 6.7k | $0.068 |
| Gemini-3.1-Pro | predict | 102 | 694.2k | $3.516 |
| Intelligence | bracket | 1 | 4.5k | $0.013 |
| Intelligence | dossier | 48 | 572.3k | $1.223 |
| Intelligence | dossier_update | 205 | 542.8k | $1.074 |
| Intelligence | late_update | 203 | 1.67M | $3.883 |
| Intelligence | match_context | 102 | 817.4k | $1.938 |
| Intelligence | post_match | 102 | 930.4k | $2.133 |
| Intelligence | pre_match | 204 | 2.24M | $5.003 |
| Intelligence | result | 280 | 1.19M | $3.220 |
| Kimi-K2.6 | bet | 102 | 734.9k | $1.988 |
| Kimi-K2.6 | constitution | 3 | 8.8k | $0.029 |
| Kimi-K2.6 | predict | 102 | 1.02M | $2.452 |
| MiniMax-M3 | bet | 102 | 453.0k | $0.338 |
| MiniMax-M3 | constitution | 2 | 4.6k | $0.007 |
| MiniMax-M3 | predict | 102 | 590.5k | $0.331 |
| Opus-4.8 | bet | 102 | 435.5k | $3.666 |
| Opus-4.8 | constitution | 2 | 3.7k | $0.065 |
| Opus-4.8 | predict | 102 | 745.5k | $5.159 |
| Qwen3.7-Max | bet | 102 | 407.0k | $0.907 |
| Qwen3.7-Max | constitution | 3 | 11.5k | $0.040 |
| Qwen3.7-Max | predict | 102 | 680.5k | $1.420 |