THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

xAI / Grok

Grok 4.6

Version grok.4.6 · Released 2026-08-12 · Available

Compare this model ↗
52QUANTUM SCORE

Explore xAI history →

Coding36
Reasoning65
Math66
Research—
Agents60
Multimodal—
Speed—
Value44

01 / EVIDENCE

Benchmark breakdown

BenchmarkVersionMethodConfigurationResultTest dateVerifiedSource typeSource
CursorBench v3.2CursorBench v3.2Publisher-reported configuration; see source—69.9 %2026-08-122026-09-29provider first partyView source ↗
DeepSWE v1.1DeepSWE v1.1Publisher-reported configuration; see source—65.9 %2026-08-122026-09-29provider first partyView source ↗
APEX-AgentsAPEX-AgentsPublisher-reported configuration; see source—57.5 %2026-08-122026-09-29provider first partyView source ↗
Terminal-Bench v3.0Terminal-Bench v3.0Publisher-reported configuration; see source—26 %2026-08-122026-09-29provider first partyView source ↗
FrontierCode v1.1 ExtendedFrontierCode v1.1 ExtendedPublisher-reported configuration; see source—61.3 %2026-08-122026-09-29provider first partyView source ↗
GDPVal-AA v2GDPVal-AA v2Publisher-reported configuration; see source—1753 Elo2026-08-122026-09-29provider first partyView source ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—2.11 %2026-09-292026-09-29benchmark maintainerView source ↗
LiveBench code_completion2026-06-25Public table; published model effort variantgrok-4.676.087 %not published2026-09-30benchmark maintainerView source ↗
LiveBench code_generation2026-06-25Public table; published model effort variantgrok-4.677.465 %not published2026-09-30benchmark maintainerView source ↗
LiveBench javascript2026-06-25Public table; published model effort variantgrok-4.672.727 %not published2026-09-30benchmark maintainerView source ↗
LiveBench AMPS_Hard2026-06-25Public table; published model effort variantgrok-4.697 %not published2026-09-30benchmark maintainerView source ↗
LiveBench olympiad2026-06-25Public table; published model effort variantgrok-4.691.213 %not published2026-09-30benchmark maintainerView source ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgrok-4.686.538 %not published2026-09-30benchmark maintainerView source ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgrok-4.695.5 %not published2026-09-30benchmark maintainerView source ↗

02 / COST

Price & speed

Input / 1M tokens $2.00

Output / 1M tokens $6.00

Cached input / 1M —

Batch input / 1M —

Batch output / 1M —

Speed No comparable observation yet

Price verified 2026-09-29

Price scope Standard short-context rate; long-context and fast rates differ

Price source Official pricing ↗

02A / PERFORMANCE

Measured speed

No sourced speed measurement at a stated effort level yet.

03 / PROFILE

Capabilities

visiontool use

04 / CONTEXT

Strengths & limitations

Coverage: 73% · Moderate confidence · 14 observations · 3 source domains (2 independent).

05 / HISTORY

Version & price history

Version grok.4.6 · released 2026-08-12.

2026-09-29: $2.00 input / $6.00 output per 1M tokens · Price source ↗

06 / SOURCES

Source records

DATED EVIDENCE

History

Standard API price history

2026-09-292026-09-29
  • 2026-09-29: 3 USD / 1M · 75% input / 25% output USD per 1M · source ↗

One observation; no trend can be inferred.

CursorBench v3.2 / CursorBench v3.2 / %

2026-08-122026-08-12
  • 2026-08-12: 69.9 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

DeepSWE v1.1 / DeepSWE v1.1 / %

2026-08-122026-08-12
  • 2026-08-12: 65.9 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

APEX-Agents / APEX-Agents / %

2026-08-122026-08-12
  • 2026-08-12: 57.5 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

Terminal-Bench v3.0 / Terminal-Bench v3.0 / %

2026-08-122026-08-12
  • 2026-08-12: 26 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

FrontierCode v1.1 Extended / FrontierCode v1.1 Extended / %

2026-08-122026-08-12
  • 2026-08-12: 61.3 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

GDPVal-AA v2 / GDPVal-AA v2 / Elo

2026-08-122026-08-12
  • 2026-08-12: 1,753 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

ARC-AGI-3 Semi-Private Standard harness best effort / 3 Semi-Private / %

2026-09-292026-09-29
  • 2026-09-29: 2.11 result · ARC Prize Standard harness; best published effort per model · source ↗

One observation; no trend can be inferred.

LiveBench code_completion / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 76.087 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench code_generation / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 77.465 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench javascript / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 72.727 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench AMPS_Hard / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 97 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench olympiad / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 91.213 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench theory_of_mind / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 86.538 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench zebra_puzzle / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 95.5 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

reasoning category score history

2026-09-302026-09-30
  • 2026-09-30: 3 score · Scoring 3.0.0

One observation; no trend can be inferred.

07 / EXPLORE

Similar models