01 / EVIDENCE
Benchmark breakdown
| Benchmark | Version | Method | Configuration | Result | Test date | Verified | Source type | Source |
|---|---|---|---|---|---|---|---|---|
| CursorBench v3.2 | CursorBench v3.2 | Publisher-reported configuration; see source | — | 66.7 % | 2026-08-12 | 2026-09-29 | provider first party | View source ↗ |
| DeepSWE v1.1 | DeepSWE v1.1 | Publisher-reported configuration; see source | — | 54 % | 2026-08-12 | 2026-09-29 | provider first party | View source ↗ |
| APEX-Agents | APEX-Agents | Publisher-reported configuration; see source | — | 47.1 % | 2026-08-12 | 2026-09-29 | provider first party | View source ↗ |
| Terminal-Bench v3.0 | Terminal-Bench v3.0 | Publisher-reported configuration; see source | — | 15.7 % | 2026-08-12 | 2026-09-29 | provider first party | View source ↗ |
| FrontierCode v1.1 Extended | FrontierCode v1.1 Extended | Publisher-reported configuration; see source | — | 56.6 % | 2026-08-12 | 2026-09-29 | provider first party | View source ↗ |
| GDPVal-AA v2 | GDPVal-AA v2 | Publisher-reported configuration; see source | — | 1526 Elo | 2026-08-12 | 2026-09-29 | provider first party | View source ↗ |
| SkillsBench Vals OpenHands with skills | Vals 2026-09-27 | Vals OpenHands harness with skills | — | 66.03 % | 2026-09-27 | 2026-09-29 | independent lab | View source ↗ |
| LiveBench code_completion | 2026-06-25 | Public table; published model effort variant | grok-4.5 | 69.565 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench code_generation | 2026-06-25 | Public table; published model effort variant | grok-4.5 | 67.606 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench javascript | 2026-06-25 | Public table; published model effort variant | grok-4.5 | 72.727 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench AMPS_Hard | 2026-06-25 | Public table; published model effort variant | grok-4.5 | 99 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench olympiad | 2026-06-25 | Public table; published model effort variant | grok-4.5 | 89.239 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | grok-4.5 | 82.692 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | grok-4.5 | 94 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
02 / COST
Price & speed
Input / 1M tokens $2.00
Output / 1M tokens $6.00
Cached input / 1M —
Batch input / 1M —
Batch output / 1M —
Speed No comparable observation yet
Price verified 2026-09-29
Price scope Standard short-context rate; long-context and fast rates differ
Price source Official pricing ↗
02A / PERFORMANCE
Measured speed
No sourced speed measurement at a stated effort level yet.
03 / PROFILE
Capabilities
04 / CONTEXT
Strengths & limitations
Coverage: 73% · Moderate confidence · 14 observations · 3 source domains (2 independent).
05 / HISTORY
Version & price history
Version grok.4.5 · released 2026-07-16.
2026-09-29: $2.00 input / $6.00 output per 1M tokens · Price source ↗
06 / SOURCES
Source records
- Grok 4.5 official documentation ↗ · verified 2026-09-29
- xAI official pricing ↗ · verified 2026-09-29
DATED EVIDENCE
History
Standard API price history
- 2026-09-29: 3 USD / 1M · 75% input / 25% output USD per 1M · source ↗
One observation; no trend can be inferred.
CursorBench v3.2 / CursorBench v3.2 / %
- 2026-08-12: 66.7 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
DeepSWE v1.1 / DeepSWE v1.1 / %
- 2026-08-12: 54 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
APEX-Agents / APEX-Agents / %
- 2026-08-12: 47.1 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
Terminal-Bench v3.0 / Terminal-Bench v3.0 / %
- 2026-08-12: 15.7 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
FrontierCode v1.1 Extended / FrontierCode v1.1 Extended / %
- 2026-08-12: 56.6 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
GDPVal-AA v2 / GDPVal-AA v2 / Elo
- 2026-08-12: 1,526 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
SkillsBench Vals OpenHands with skills / Vals 2026-09-27 / %
- 2026-09-27: 66.03 result · Vals OpenHands harness with skills · source ↗
One observation; no trend can be inferred.
LiveBench code_completion / 2026-06-25 / %
- 2026-09-30: 69.565 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench code_generation / 2026-06-25 / %
- 2026-09-30: 67.606 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench javascript / 2026-06-25 / %
- 2026-09-30: 72.727 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench AMPS_Hard / 2026-06-25 / %
- 2026-09-30: 99 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench olympiad / 2026-06-25 / %
- 2026-09-30: 89.239 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench theory_of_mind / 2026-06-25 / %
- 2026-09-30: 82.692 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench zebra_puzzle / 2026-06-25 / %
- 2026-09-30: 94 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
agents category score history
- 2026-09-30: 4 score · Scoring 3.0.0
One observation; no trend can be inferred.
07 / EXPLORE