Anthropic / Claude
Claude Opus 5.5
Version claude-opus-5-5 · Released 2026-09-22 · Available
Compare this model ↗01 / EVIDENCE
Benchmark breakdown
| Benchmark | Version | Method | Configuration | Result | Test date | Verified | Source type | Source |
|---|---|---|---|---|---|---|---|---|
| Terminal-Bench 4.0 (reported best effort) | Terminal-Bench 4.0 (reported best effort) | Publisher-reported configuration; see source | — | 66.4 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| CursorBench 4.0 (reported best effort) | CursorBench 4.0 (reported best effort) | Publisher-reported configuration; see source | — | 57.8 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| FrontierCode 1.1 Main (reported best effort) | FrontierCode 1.1 Main (reported best effort) | Publisher-reported configuration; see source | — | 54.4 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| GDPval-AA v2.1 (reported best effort) | GDPval-AA v2.1 (reported best effort) | Publisher-reported configuration; see source | — | 1846 Elo | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| AA-Briefcase v1.1 (reported best effort) | AA-Briefcase v1.1 (reported best effort) | Publisher-reported configuration; see source | — | 1822 Elo | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| Humanity Last Exam (with tools; reported best effort) | Humanity Last Exam (with tools; reported best effort) | Publisher-reported configuration; see source | — | 67.7 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| OSWorld 2.1 (partial; reported best effort) | OSWorld 2.1 (partial; reported best effort) | Publisher-reported configuration; see source | — | 81.8 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| Chartography (no tools; reported best effort) | Chartography (no tools; reported best effort) | Publisher-reported configuration; see source | — | 64.4 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed | 4.0 | Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures | — | 58.08 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| ProofBench v1.1 Vals 100 proof tasks | 1.1 | Vals 100 proof tasks; ceiling effect at top scores | — | 100 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| ProgramBench Vals 200 public tasks raw pass rate | Vals 2026-09-27 | Vals 200 public tasks; raw pass rate; mini-SWE-agent | — | 87 % | 2026-09-27 | 2026-09-29 | independent lab | View source ↗ |
| Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) | AA Index 4.3.2 | Opus Max with default fallback; Astra Max | — | 61 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) | AA Index 4.3.2 | Opus Max with default fallback; Astra Max | — | 70 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| LiveBench code_completion | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | 86.957 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench code_generation | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | 91.549 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench javascript | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | 81.818 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench AMPS_Hard | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | 98 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench olympiad | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | 93.251 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | 84.615 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | 100 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
02 / COST
Price & speed
Input / 1M tokens $4.00
Output / 1M tokens $20.00
Cached input / 1M —
Batch input / 1M —
Batch output / 1M —
Speed See measured observations below
Price verified 2026-09-29
Price scope Base input/output rates; caching and batch tiers differ
Price source Official pricing ↗
02A / PERFORMANCE
Measured speed
| Metric | Result | Effort | Method | Source |
|---|---|---|---|---|
| output tokens per second | 92 tokens/s | Max | Artificial Analysis model comparison; output speed | View source ↗ |
| time to first token | 694 seconds | Max | Artificial Analysis model comparison; includes reasoning | View source ↗ |
| end to end response time | 699.43 seconds | Max | Artificial Analysis model comparison; includes reasoning | View source ↗ |
03 / PROFILE
Capabilities
04 / CONTEXT
Strengths & limitations
Coverage: 73% · Moderate confidence · 20 observations · 4 source domains (3 independent).
05 / HISTORY
Version & price history
Version claude-opus-5-5 · released 2026-09-22.
2026-09-29: $4.00 input / $20.00 output per 1M tokens · Price source ↗
06 / SOURCES
Source records
- Claude Opus 5.5 official documentation ↗ · verified 2026-09-29
- Anthropic official pricing ↗ · verified 2026-09-29
DATED EVIDENCE
History
Standard API price history
- 2026-09-29: 8 USD / 1M · 75% input / 25% output USD per 1M · source ↗
One observation; no trend can be inferred.
Terminal-Bench 4.0 (reported best effort) / Terminal-Bench 4.0 (reported best effort) / %
- 2026-09-28: 66.4 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
CursorBench 4.0 (reported best effort) / CursorBench 4.0 (reported best effort) / %
- 2026-09-28: 57.8 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
FrontierCode 1.1 Main (reported best effort) / FrontierCode 1.1 Main (reported best effort) / %
- 2026-09-28: 54.4 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
GDPval-AA v2.1 (reported best effort) / GDPval-AA v2.1 (reported best effort) / Elo
- 2026-09-28: 1,846 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
AA-Briefcase v1.1 (reported best effort) / AA-Briefcase v1.1 (reported best effort) / Elo
- 2026-09-28: 1,822 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
Humanity Last Exam (with tools; reported best effort) / Humanity Last Exam (with tools; reported best effort) / %
- 2026-09-28: 67.7 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
OSWorld 2.1 (partial; reported best effort) / OSWorld 2.1 (partial; reported best effort) / %
- 2026-09-28: 81.8 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
Chartography (no tools; reported best effort) / Chartography (no tools; reported best effort) / %
- 2026-09-28: 64.4 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed / 4.0 / %
- 2026-09-29: 58.08 result · Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures · source ↗
One observation; no trend can be inferred.
ProofBench v1.1 Vals 100 proof tasks / 1.1 / %
- 2026-09-29: 100 result · Vals 100 proof tasks; ceiling effect at top scores · source ↗
One observation; no trend can be inferred.
ProgramBench Vals 200 public tasks raw pass rate / Vals 2026-09-27 / %
- 2026-09-27: 87 result · Vals 200 public tasks; raw pass rate; mini-SWE-agent · source ↗
One observation; no trend can be inferred.
Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) / AA Index 4.3.2 / %
- 2026-09-29: 61 result · Opus Max with default fallback; Astra Max · source ↗
One observation; no trend can be inferred.
AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) / AA Index 4.3.2 / %
- 2026-09-29: 70 result · Opus Max with default fallback; Astra Max · source ↗
One observation; no trend can be inferred.
LiveBench code_completion / 2026-06-25 / %
- 2026-09-30: 86.957 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench code_generation / 2026-06-25 / %
- 2026-09-30: 91.549 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench javascript / 2026-06-25 / %
- 2026-09-30: 81.818 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench AMPS_Hard / 2026-06-25 / %
- 2026-09-30: 98 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench olympiad / 2026-06-25 / %
- 2026-09-30: 93.251 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench theory_of_mind / 2026-06-25 / %
- 2026-09-30: 84.615 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench zebra_puzzle / 2026-06-25 / %
- 2026-09-30: 100 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
output tokens per second / Max / tokens/s / Artificial Analysis model comparison; output speed
- 2026-09-29: 92 measured · Artificial Analysis model comparison; output speed · source ↗
One observation; no trend can be inferred.
time to first token / Max / seconds / Artificial Analysis model comparison; includes reasoning
- 2026-09-29: 694 measured · Artificial Analysis model comparison; includes reasoning · source ↗
One observation; no trend can be inferred.
end to end response time / Max / seconds / Artificial Analysis model comparison; includes reasoning
- 2026-09-29: 699.43 measured · Artificial Analysis model comparison; includes reasoning · source ↗
One observation; no trend can be inferred.
coding category score history
- 2026-09-30: 82 score · Scoring 3.0.0
One observation; no trend can be inferred.
reasoning category score history
- 2026-09-30: 100 score · Scoring 3.0.0
One observation; no trend can be inferred.
math category score history
- 2026-09-30: 100 score · Scoring 3.0.0
One observation; no trend can be inferred.
agents category score history
- 2026-09-30: 100 score · Scoring 3.0.0
One observation; no trend can be inferred.
07 / EXPLORE