THE EVIDENCE / REASONING
Reasoning
benchmarks.
Multi-step problem solving and inference. Results are grouped by benchmark and unit. Different tests are never merged into a raw leaderboard.
01 / RESULTS
Published results
| Benchmark | Version | Method | Configuration | Model | Result | Test date | Verified | Source type | Source |
|---|---|---|---|---|---|---|---|---|---|
| Humanity Last Exam (with tools; reported best effort) | Humanity Last Exam (with tools; reported best effort) | Publisher-reported configuration; see source | — | claude-sonnet-5-5 | 64.5 % | 2026-09-28 | 2026-09-29 | provider first party | Source ↗ |
| Humanity Last Exam (with tools; reported best effort) | Humanity Last Exam (with tools; reported best effort) | Publisher-reported configuration; see source | — | claude-opus-5-5 | 67.7 % | 2026-09-28 | 2026-09-29 | provider first party | Source ↗ |
| ARC-AGI-3 Provider Adapter (High) | ARC-AGI-3 Provider Adapter (High) | Publisher-reported configuration; see source | — | gpt-6-astra | 99.9 % | 2026-09-03 | 2026-09-29 | provider first party | Source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | gpt-5.6-sol | 7.78 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | Source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | gpt-5.6-terra | 0.8 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | Source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | gpt-5.6-luna | 0.18 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | Source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | gpt-6-astra | 62.7128 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | Source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | gpt-6-sol | 4.6236 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | Source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | gpt-6-luna | 0.1942 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | Source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | gemini-3-8-flash | 10.3663 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | Source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | grok-4-6 | 2.11 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | Source ↗ |
| Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) | AA Index 4.3.2 | Opus Max with default fallback; Astra Max | — | claude-opus-5-5 | 61 % | 2026-09-29 | 2026-09-29 | independent lab | Source ↗ |
| Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) | AA Index 4.3.2 | Opus Max with default fallback; Astra Max | — | gpt-6-astra | 55 % | 2026-09-29 | 2026-09-29 | independent lab | Source ↗ |
| Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) | AA Index 4.3.2 | Gemini High | — | gemini-3-8-flash | 48 % | 2026-09-29 | 2026-09-29 | independent lab | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gemini-3.1-pro-preview-high | gemini-3-1-pro-preview | 80.769 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gemini-3.1-pro-preview-high | gemini-3-1-pro-preview | 85.25 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gemini-3.5-flash-high | gemini-3-5-flash | 80.769 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gemini-3.5-flash-high | gemini-3-5-flash | 77.25 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | grok-4.5 | grok-4-5 | 82.692 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | grok-4.5 | grok-4-5 | 94 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gpt-5.6-sol-max | gpt-5.6-sol | 84.615 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gpt-5.6-sol-max | gpt-5.6-sol | 100 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gpt-5.6-terra-max | gpt-5.6-terra | 86.538 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gpt-5.6-terra-max | gpt-5.6-terra | 100 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gpt-5.6-luna-max | gpt-5.6-luna | 73.077 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gpt-5.6-luna-max | gpt-5.6-luna | 93.5 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gemini-3.5-flash-lite-high | gemini-3-5-flash-lite | 50 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gemini-3.5-flash-lite-high | gemini-3-5-flash-lite | 32.75 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gemini-3.6-flash-high | gemini-3-6-flash | 78.846 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gemini-3.6-flash-high | gemini-3-6-flash | 85.75 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | deepseek-v4-pro-0813 | deepseek-v4-pro | 84.615 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | deepseek-v4-pro-0813 | deepseek-v4-pro | 96.75 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | grok-4.6 | grok-4-6 | 86.538 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | grok-4.6 | grok-4-6 | 95.5 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gemini-3.7-flash-high | gemini-3-7-flash | 82.692 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gemini-3.7-flash-high | gemini-3-7-flash | 94.5 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | claude-fable-5-1 | 80.769 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | claude-fable-5-1-max-effort | claude-fable-5-1 | 100 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gemini-3.8-flash-high | gemini-3-8-flash | 76.923 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gemini-3.8-flash-high | gemini-3-8-flash | 98.25 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | gpt-6-astra | 84.615 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | gpt-6-astra | 100 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | deepseek-v4.1-flash-max | deepseek-v4-1-flash | 80.769 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | deepseek-v4.1-flash-max | deepseek-v4-1-flash | 100 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | claude-opus-5-5 | 84.615 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | claude-opus-5-5-max-effort | claude-opus-5-5 | 100 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gpt-6-sol-max | gpt-6-sol | 84.615 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gpt-6-sol-max | gpt-6-sol | 100 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | gpt-6-luna | 73.077 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | gpt-6-luna | 96 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | claude-sonnet-5-5 | 86.538 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | claude-sonnet-5-5-max-effort | claude-sonnet-5-5 | 100 % | not published | 2026-09-30 | benchmark maintainer | Source ↗ |
Limitations: Prompt wording and hidden test contamination can change outcomes. Publication methods, prompts, sampling and model versions may differ. Check each cited source before interpreting a result.