VALUE FRONTIER
Performance
per dollar.
Only models with a verified category score and current USD API price appear. Each category stands alone; there is no universal winner.
overall evidence vs standard API price
Green points are efficient within this eligible cohort: no other plotted model has both a lower price and an equal or higher score. Prices use the stated 75/25 input/output mix.
| Model | Category score | Blended price | Confidence | Frontier |
|---|---|---|---|---|
| Gemini 3.8 Flash | 51 | $1.50 | Moderate confidence · 3 independent domains | Efficient in this cohort |
| Grok 4.6 | 52 | $3.00 | Moderate confidence · 2 independent domains | Efficient in this cohort |
| Grok 4.5 | 44 | $3.00 | Moderate confidence · 2 independent domains | — |
| GPT-6 Sol | 48 | $4.00 | Moderate confidence · 3 independent domains | — |
| Claude Sonnet 5.5 | 88 | $4.00 | Moderate confidence · 2 independent domains | Efficient in this cohort |
| GPT-5.6 Sol | 53 | $8.00 | Moderate confidence · 3 independent domains | — |
| Claude Opus 5.5 | 86 | $8.00 | Moderate confidence · 3 independent domains | — |
| GPT-6 Astra | 60 | $20.00 | Moderate confidence · 4 independent domains | — |
Median standard API price
2026-09-292026-09-29
- 2026-09-29: 3 USD / 1M · Median of 24 priced models
One observation; no trend can be inferred.
Qualified performance per dollar
No dated qualified overall scores yet; no trend is shown.
ARC-AGI-3 Standard benchmark frontier
2026-09-292026-09-29
- 2026-09-29: 62.713 % · gpt-6-astra · ARC-AGI-3 Semi-Private Standard harness best effort 3 Semi-Private · source ↗
One observation; no trend can be inferred.