THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

Anthropic / Claude

Claude Sonnet 5.5

Version claude-sonnet-5-5 · Released 2026-09-28 · Available

Compare this model ↗
88QUANTUM SCORE

Explore Anthropic history →

Coding100
Reasoning100
Math84
Research—
Agents80
Multimodal—
Speed—
Value65

01 / EVIDENCE

Benchmark breakdown

BenchmarkVersionMethodConfigurationResultTest dateVerifiedSource typeSource
Terminal-Bench 4.0 (reported best effort)Terminal-Bench 4.0 (reported best effort)Publisher-reported configuration; see source—70.6 %2026-09-282026-09-29provider first partyView source ↗
CursorBench 4.0 (reported best effort)CursorBench 4.0 (reported best effort)Publisher-reported configuration; see source—55.5 %2026-09-282026-09-29provider first partyView source ↗
FrontierCode 1.1 Main (reported best effort)FrontierCode 1.1 Main (reported best effort)Publisher-reported configuration; see source—46.2 %2026-09-282026-09-29provider first partyView source ↗
GDPval-AA v2.1 (reported best effort)GDPval-AA v2.1 (reported best effort)Publisher-reported configuration; see source—1844 Elo2026-09-282026-09-29provider first partyView source ↗
AA-Briefcase v1.1 (reported best effort)AA-Briefcase v1.1 (reported best effort)Publisher-reported configuration; see source—1811 Elo2026-09-282026-09-29provider first partyView source ↗
Humanity Last Exam (with tools; reported best effort)Humanity Last Exam (with tools; reported best effort)Publisher-reported configuration; see source—64.5 %2026-09-282026-09-29provider first partyView source ↗
OSWorld 2.1 (partial; reported best effort)OSWorld 2.1 (partial; reported best effort)Publisher-reported configuration; see source—80.1 %2026-09-282026-09-29provider first partyView source ↗
Chartography (no tools; reported best effort)Chartography (no tools; reported best effort)Publisher-reported configuration; see source—61.6 %2026-09-282026-09-29provider first partyView source ↗
Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed4.0Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures—62.63 %2026-09-292026-09-29independent labView source ↗
LiveBench code_completion2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effort86.957 %not published2026-09-30benchmark maintainerView source ↗
LiveBench code_generation2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effort95.775 %not published2026-09-30benchmark maintainerView source ↗
LiveBench javascript2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effort77.273 %not published2026-09-30benchmark maintainerView source ↗
LiveBench AMPS_Hard2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effort98 %not published2026-09-30benchmark maintainerView source ↗
LiveBench olympiad2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effort92.468 %not published2026-09-30benchmark maintainerView source ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effort86.538 %not published2026-09-30benchmark maintainerView source ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effort100 %not published2026-09-30benchmark maintainerView source ↗

02 / COST

Price & speed

Input / 1M tokens $2.00

Output / 1M tokens $10.00

Cached input / 1M —

Batch input / 1M —

Batch output / 1M —

Speed No comparable observation yet

Price verified 2026-09-29

Price scope Base input/output rates; caching and batch tiers differ

Price source Official pricing ↗

02A / PERFORMANCE

Measured speed

No sourced speed measurement at a stated effort level yet.

03 / PROFILE

Capabilities

tool usevision

04 / CONTEXT

Strengths & limitations

Coverage: 73% · Moderate confidence · 16 observations · 3 source domains (2 independent).

05 / HISTORY

Version & price history

Version claude-sonnet-5-5 · released 2026-09-28.

2026-09-29: $2.00 input / $10.00 output per 1M tokens · Price source ↗

06 / SOURCES

Source records

DATED EVIDENCE

History

Standard API price history

2026-09-292026-09-29
  • 2026-09-29: 4 USD / 1M · 75% input / 25% output USD per 1M · source ↗

One observation; no trend can be inferred.

Terminal-Bench 4.0 (reported best effort) / Terminal-Bench 4.0 (reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 70.6 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

CursorBench 4.0 (reported best effort) / CursorBench 4.0 (reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 55.5 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

FrontierCode 1.1 Main (reported best effort) / FrontierCode 1.1 Main (reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 46.2 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

GDPval-AA v2.1 (reported best effort) / GDPval-AA v2.1 (reported best effort) / Elo

2026-09-282026-09-28
  • 2026-09-28: 1,844 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

AA-Briefcase v1.1 (reported best effort) / AA-Briefcase v1.1 (reported best effort) / Elo

2026-09-282026-09-28
  • 2026-09-28: 1,811 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

Humanity Last Exam (with tools; reported best effort) / Humanity Last Exam (with tools; reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 64.5 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

OSWorld 2.1 (partial; reported best effort) / OSWorld 2.1 (partial; reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 80.1 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

Chartography (no tools; reported best effort) / Chartography (no tools; reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 61.6 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed / 4.0 / %

2026-09-292026-09-29
  • 2026-09-29: 62.63 result · Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures · source ↗

One observation; no trend can be inferred.

LiveBench code_completion / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 86.957 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench code_generation / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 95.775 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench javascript / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 77.273 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench AMPS_Hard / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 98 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench olympiad / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 92.468 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench theory_of_mind / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 86.538 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench zebra_puzzle / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 100 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

coding category score history

2026-09-302026-09-30
  • 2026-09-30: 100 score · Scoring 3.0.0

One observation; no trend can be inferred.

07 / EXPLORE

Similar models