Anthropic / Claude
Claude Sonnet 5.5
Version claude-sonnet-5-5 · Released 2026-09-28 · Available
Compare this model ↗01 / EVIDENCE
Benchmark breakdown
| Benchmark | Version | Method | Configuration | Result | Test date | Verified | Source type | Source |
|---|---|---|---|---|---|---|---|---|
| Terminal-Bench 4.0 (reported best effort) | Terminal-Bench 4.0 (reported best effort) | Publisher-reported configuration; see source | — | 70.6 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| CursorBench 4.0 (reported best effort) | CursorBench 4.0 (reported best effort) | Publisher-reported configuration; see source | — | 55.5 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| FrontierCode 1.1 Main (reported best effort) | FrontierCode 1.1 Main (reported best effort) | Publisher-reported configuration; see source | — | 46.2 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| GDPval-AA v2.1 (reported best effort) | GDPval-AA v2.1 (reported best effort) | Publisher-reported configuration; see source | — | 1844 Elo | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| AA-Briefcase v1.1 (reported best effort) | AA-Briefcase v1.1 (reported best effort) | Publisher-reported configuration; see source | — | 1811 Elo | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| Humanity Last Exam (with tools; reported best effort) | Humanity Last Exam (with tools; reported best effort) | Publisher-reported configuration; see source | — | 64.5 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| OSWorld 2.1 (partial; reported best effort) | OSWorld 2.1 (partial; reported best effort) | Publisher-reported configuration; see source | — | 80.1 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| Chartography (no tools; reported best effort) | Chartography (no tools; reported best effort) | Publisher-reported configuration; see source | — | 61.6 % | 2026-09-28 | 2026-09-29 | provider first party | View source ↗ |
| Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed | 4.0 | Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures | — | 62.63 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| LiveBench code_completion | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | 86.957 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench code_generation | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | 95.775 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench javascript | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | 77.273 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench AMPS_Hard | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | 98 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench olympiad | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | 92.468 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | 86.538 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | 100 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
02 / COST
Price & speed
Input / 1M tokens $2.00
Output / 1M tokens $10.00
Cached input / 1M —
Batch input / 1M —
Batch output / 1M —
Speed No comparable observation yet
Price verified 2026-09-29
Price scope Base input/output rates; caching and batch tiers differ
Price source Official pricing ↗
02A / PERFORMANCE
Measured speed
No sourced speed measurement at a stated effort level yet.
03 / PROFILE
Capabilities
04 / CONTEXT
Strengths & limitations
Coverage: 73% · Moderate confidence · 16 observations · 3 source domains (2 independent).
05 / HISTORY
Version & price history
Version claude-sonnet-5-5 · released 2026-09-28.
2026-09-29: $2.00 input / $10.00 output per 1M tokens · Price source ↗
06 / SOURCES
Source records
- Claude Sonnet 5.5 official documentation ↗ · verified 2026-09-29
- Anthropic official pricing ↗ · verified 2026-09-29
DATED EVIDENCE
History
Standard API price history
- 2026-09-29: 4 USD / 1M · 75% input / 25% output USD per 1M · source ↗
One observation; no trend can be inferred.
Terminal-Bench 4.0 (reported best effort) / Terminal-Bench 4.0 (reported best effort) / %
- 2026-09-28: 70.6 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
CursorBench 4.0 (reported best effort) / CursorBench 4.0 (reported best effort) / %
- 2026-09-28: 55.5 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
FrontierCode 1.1 Main (reported best effort) / FrontierCode 1.1 Main (reported best effort) / %
- 2026-09-28: 46.2 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
GDPval-AA v2.1 (reported best effort) / GDPval-AA v2.1 (reported best effort) / Elo
- 2026-09-28: 1,844 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
AA-Briefcase v1.1 (reported best effort) / AA-Briefcase v1.1 (reported best effort) / Elo
- 2026-09-28: 1,811 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
Humanity Last Exam (with tools; reported best effort) / Humanity Last Exam (with tools; reported best effort) / %
- 2026-09-28: 64.5 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
OSWorld 2.1 (partial; reported best effort) / OSWorld 2.1 (partial; reported best effort) / %
- 2026-09-28: 80.1 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
Chartography (no tools; reported best effort) / Chartography (no tools; reported best effort) / %
- 2026-09-28: 61.6 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed / 4.0 / %
- 2026-09-29: 62.63 result · Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures · source ↗
One observation; no trend can be inferred.
LiveBench code_completion / 2026-06-25 / %
- 2026-09-30: 86.957 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench code_generation / 2026-06-25 / %
- 2026-09-30: 95.775 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench javascript / 2026-06-25 / %
- 2026-09-30: 77.273 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench AMPS_Hard / 2026-06-25 / %
- 2026-09-30: 98 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench olympiad / 2026-06-25 / %
- 2026-09-30: 92.468 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench theory_of_mind / 2026-06-25 / %
- 2026-09-30: 86.538 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench zebra_puzzle / 2026-06-25 / %
- 2026-09-30: 100 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
coding category score history
- 2026-09-30: 100 score · Scoring 3.0.0
One observation; no trend can be inferred.
07 / EXPLORE