THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

Anthropic / Claude

Claude Opus 5.5

Version claude-opus-5-5 · Released 2026-09-22 · Available

Compare this model ↗
86QUANTUM SCORE

Explore Anthropic history →

Coding87
Reasoning97
Math92
Research—
Agents100
Multimodal—
Speed—
Value52

01 / EVIDENCE

Benchmark breakdown

BenchmarkVersionMethodConfigurationResultTest dateVerifiedSource typeSource
Terminal-Bench 4.0 (reported best effort)Terminal-Bench 4.0 (reported best effort)Publisher-reported configuration; see source—66.4 %2026-09-282026-09-29provider first partyView source ↗
CursorBench 4.0 (reported best effort)CursorBench 4.0 (reported best effort)Publisher-reported configuration; see source—57.8 %2026-09-282026-09-29provider first partyView source ↗
FrontierCode 1.1 Main (reported best effort)FrontierCode 1.1 Main (reported best effort)Publisher-reported configuration; see source—54.4 %2026-09-282026-09-29provider first partyView source ↗
GDPval-AA v2.1 (reported best effort)GDPval-AA v2.1 (reported best effort)Publisher-reported configuration; see source—1846 Elo2026-09-282026-09-29provider first partyView source ↗
AA-Briefcase v1.1 (reported best effort)AA-Briefcase v1.1 (reported best effort)Publisher-reported configuration; see source—1822 Elo2026-09-282026-09-29provider first partyView source ↗
Humanity Last Exam (with tools; reported best effort)Humanity Last Exam (with tools; reported best effort)Publisher-reported configuration; see source—67.7 %2026-09-282026-09-29provider first partyView source ↗
OSWorld 2.1 (partial; reported best effort)OSWorld 2.1 (partial; reported best effort)Publisher-reported configuration; see source—81.8 %2026-09-282026-09-29provider first partyView source ↗
Chartography (no tools; reported best effort)Chartography (no tools; reported best effort)Publisher-reported configuration; see source—64.4 %2026-09-282026-09-29provider first partyView source ↗
Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed4.0Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures—58.08 %2026-09-292026-09-29independent labView source ↗
ProofBench v1.1 Vals 100 proof tasks1.1Vals 100 proof tasks; ceiling effect at top scores—100 %2026-09-292026-09-29independent labView source ↗
ProgramBench Vals 200 public tasks raw pass rateVals 2026-09-27Vals 200 public tasks; raw pass rate; mini-SWE-agent—87 %2026-09-272026-09-29independent labView source ↗
Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Opus Max with default fallback; Astra Max—61 %2026-09-292026-09-29independent labView source ↗
AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Opus Max with default fallback; Astra Max—70 %2026-09-292026-09-29independent labView source ↗
LiveBench code_completion2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effort86.957 %not published2026-09-30benchmark maintainerView source ↗
LiveBench code_generation2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effort91.549 %not published2026-09-30benchmark maintainerView source ↗
LiveBench javascript2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effort81.818 %not published2026-09-30benchmark maintainerView source ↗
LiveBench AMPS_Hard2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effort98 %not published2026-09-30benchmark maintainerView source ↗
LiveBench olympiad2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effort93.251 %not published2026-09-30benchmark maintainerView source ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effort84.615 %not published2026-09-30benchmark maintainerView source ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effort100 %not published2026-09-30benchmark maintainerView source ↗

02 / COST

Price & speed

Input / 1M tokens $4.00

Output / 1M tokens $20.00

Cached input / 1M —

Batch input / 1M —

Batch output / 1M —

Speed See measured observations below

Price verified 2026-09-29

Price scope Base input/output rates; caching and batch tiers differ

Price source Official pricing ↗

02A / PERFORMANCE

Measured speed

MetricResultEffortMethodSource
output tokens per second92 tokens/sMaxArtificial Analysis model comparison; output speedView source ↗
time to first token694 secondsMaxArtificial Analysis model comparison; includes reasoningView source ↗
end to end response time699.43 secondsMaxArtificial Analysis model comparison; includes reasoningView source ↗

03 / PROFILE

Capabilities

tool usevision

04 / CONTEXT

Strengths & limitations

Coverage: 73% · Moderate confidence · 20 observations · 4 source domains (3 independent).

05 / HISTORY

Version & price history

Version claude-opus-5-5 · released 2026-09-22.

2026-09-29: $4.00 input / $20.00 output per 1M tokens · Price source ↗

06 / SOURCES

Source records

DATED EVIDENCE

History

Standard API price history

2026-09-292026-09-29
  • 2026-09-29: 8 USD / 1M · 75% input / 25% output USD per 1M · source ↗

One observation; no trend can be inferred.

Terminal-Bench 4.0 (reported best effort) / Terminal-Bench 4.0 (reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 66.4 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

CursorBench 4.0 (reported best effort) / CursorBench 4.0 (reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 57.8 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

FrontierCode 1.1 Main (reported best effort) / FrontierCode 1.1 Main (reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 54.4 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

GDPval-AA v2.1 (reported best effort) / GDPval-AA v2.1 (reported best effort) / Elo

2026-09-282026-09-28
  • 2026-09-28: 1,846 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

AA-Briefcase v1.1 (reported best effort) / AA-Briefcase v1.1 (reported best effort) / Elo

2026-09-282026-09-28
  • 2026-09-28: 1,822 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

Humanity Last Exam (with tools; reported best effort) / Humanity Last Exam (with tools; reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 67.7 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

OSWorld 2.1 (partial; reported best effort) / OSWorld 2.1 (partial; reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 81.8 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

Chartography (no tools; reported best effort) / Chartography (no tools; reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 64.4 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed / 4.0 / %

2026-09-292026-09-29
  • 2026-09-29: 58.08 result · Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures · source ↗

One observation; no trend can be inferred.

ProofBench v1.1 Vals 100 proof tasks / 1.1 / %

2026-09-292026-09-29
  • 2026-09-29: 100 result · Vals 100 proof tasks; ceiling effect at top scores · source ↗

One observation; no trend can be inferred.

ProgramBench Vals 200 public tasks raw pass rate / Vals 2026-09-27 / %

2026-09-272026-09-27
  • 2026-09-27: 87 result · Vals 200 public tasks; raw pass rate; mini-SWE-agent · source ↗

One observation; no trend can be inferred.

Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) / AA Index 4.3.2 / %

2026-09-292026-09-29
  • 2026-09-29: 61 result · Opus Max with default fallback; Astra Max · source ↗

One observation; no trend can be inferred.

AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) / AA Index 4.3.2 / %

2026-09-292026-09-29
  • 2026-09-29: 70 result · Opus Max with default fallback; Astra Max · source ↗

One observation; no trend can be inferred.

LiveBench code_completion / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 86.957 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench code_generation / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 91.549 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench javascript / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 81.818 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench AMPS_Hard / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 98 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench olympiad / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 93.251 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench theory_of_mind / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 84.615 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench zebra_puzzle / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 100 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

output tokens per second / Max / tokens/s / Artificial Analysis model comparison; output speed

2026-09-292026-09-29
  • 2026-09-29: 92 measured · Artificial Analysis model comparison; output speed · source ↗

One observation; no trend can be inferred.

time to first token / Max / seconds / Artificial Analysis model comparison; includes reasoning

2026-09-292026-09-29
  • 2026-09-29: 694 measured · Artificial Analysis model comparison; includes reasoning · source ↗

One observation; no trend can be inferred.

end to end response time / Max / seconds / Artificial Analysis model comparison; includes reasoning

2026-09-292026-09-29
  • 2026-09-29: 699.43 measured · Artificial Analysis model comparison; includes reasoning · source ↗

One observation; no trend can be inferred.

coding category score history

2026-09-302026-09-30
  • 2026-09-30: 82 score · Scoring 3.0.0

One observation; no trend can be inferred.

reasoning category score history

2026-09-302026-09-30
  • 2026-09-30: 100 score · Scoring 3.0.0

One observation; no trend can be inferred.

math category score history

2026-09-302026-09-30
  • 2026-09-30: 100 score · Scoring 3.0.0

One observation; no trend can be inferred.

agents category score history

2026-09-302026-09-30
  • 2026-09-30: 100 score · Scoring 3.0.0

One observation; no trend can be inferred.

07 / EXPLORE

Similar models