Anthropic / Claude
Claude Fable 5.1
Version claude-fable-5-1 · Released 2026-09-01 · Available
Compare this model ↗01 / EVIDENCE
Benchmark breakdown
| Benchmark | Version | Method | Configuration | Result | Test date | Verified | Source type | Source |
|---|---|---|---|---|---|---|---|---|
| Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed | 4.0 | Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures | — | 50 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| ProofBench v1.1 Vals 100 proof tasks | 1.1 | Vals 100 proof tasks; ceiling effect at top scores | — | 100 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| ProgramBench Vals 200 public tasks raw pass rate | Vals 2026-09-27 | Vals 200 public tasks; raw pass rate; mini-SWE-agent | — | 82.7 % | 2026-09-27 | 2026-09-29 | independent lab | View source ↗ |
| LiveBench code_completion | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | 82.609 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench code_generation | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | 90.141 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench javascript | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | 68.182 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench AMPS_Hard | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | 99 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench olympiad | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | 92.968 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | 80.769 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | 100 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
02 / COST
Price & speed
Input / 1M tokens $10.00
Output / 1M tokens $50.00
Cached input / 1M —
Batch input / 1M —
Batch output / 1M —
Speed No comparable observation yet
Price verified 2026-09-29
Price scope Base input/output rates; caching and batch tiers differ
Price source Official pricing ↗
02A / PERFORMANCE
Measured speed
No sourced speed measurement at a stated effort level yet.
03 / PROFILE
Capabilities
04 / CONTEXT
Strengths & limitations
Coverage: 73% · Limited evidence · 10 observations · 2 source domains (2 independent).
Why no Overall AIQuantumScore?
Qualified category evidence: coding, reasoning, math, agents. An overall score is withheld because the independent observation threshold has not been met. Evidence still missing or lacking comparable overlap: research, multimodal, speed.
05 / HISTORY
Version & price history
Version claude-fable-5-1 · released 2026-09-01.
2026-09-29: $10.00 input / $50.00 output per 1M tokens · Price source ↗
06 / SOURCES
Source records
- Claude Fable 5.1 official documentation ↗ · verified 2026-09-29
- Anthropic official pricing ↗ · verified 2026-09-29
DATED EVIDENCE
History
Standard API price history
- 2026-09-29: 20 USD / 1M · 75% input / 25% output USD per 1M · source ↗
One observation; no trend can be inferred.
Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed / 4.0 / %
- 2026-09-29: 50 result · Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures · source ↗
One observation; no trend can be inferred.
ProofBench v1.1 Vals 100 proof tasks / 1.1 / %
- 2026-09-29: 100 result · Vals 100 proof tasks; ceiling effect at top scores · source ↗
One observation; no trend can be inferred.
ProgramBench Vals 200 public tasks raw pass rate / Vals 2026-09-27 / %
- 2026-09-27: 82.7 result · Vals 200 public tasks; raw pass rate; mini-SWE-agent · source ↗
One observation; no trend can be inferred.
LiveBench code_completion / 2026-06-25 / %
- 2026-09-30: 82.609 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench code_generation / 2026-06-25 / %
- 2026-09-30: 90.141 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench javascript / 2026-06-25 / %
- 2026-09-30: 68.182 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench AMPS_Hard / 2026-06-25 / %
- 2026-09-30: 99 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench olympiad / 2026-06-25 / %
- 2026-09-30: 92.968 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench theory_of_mind / 2026-06-25 / %
- 2026-09-30: 80.769 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench zebra_puzzle / 2026-06-25 / %
- 2026-09-30: 100 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
coding category score history
- 2026-09-30: 8 score · Scoring 3.0.0
One observation; no trend can be inferred.
math category score history
- 2026-09-30: 100 score · Scoring 3.0.0
One observation; no trend can be inferred.
07 / EXPLORE