THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

THE EVIDENCE / REASONING

Reasoning
benchmarks.

Multi-step problem solving and inference. Results are grouped by benchmark and unit. Different tests are never merged into a raw leaderboard.

01 / RESULTS

Published results

BenchmarkVersionMethodConfigurationModelResultTest dateVerifiedSource typeSource
Humanity Last Exam (with tools; reported best effort)Humanity Last Exam (with tools; reported best effort)Publisher-reported configuration; see source—claude-sonnet-5-564.5 %2026-09-282026-09-29provider first partySource ↗
Humanity Last Exam (with tools; reported best effort)Humanity Last Exam (with tools; reported best effort)Publisher-reported configuration; see source—claude-opus-5-567.7 %2026-09-282026-09-29provider first partySource ↗
ARC-AGI-3 Provider Adapter (High)ARC-AGI-3 Provider Adapter (High)Publisher-reported configuration; see source—gpt-6-astra99.9 %2026-09-032026-09-29provider first partySource ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—gpt-5.6-sol7.78 %2026-09-292026-09-29benchmark maintainerSource ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—gpt-5.6-terra0.8 %2026-09-292026-09-29benchmark maintainerSource ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—gpt-5.6-luna0.18 %2026-09-292026-09-29benchmark maintainerSource ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—gpt-6-astra62.7128 %2026-09-292026-09-29benchmark maintainerSource ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—gpt-6-sol4.6236 %2026-09-292026-09-29benchmark maintainerSource ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—gpt-6-luna0.1942 %2026-09-292026-09-29benchmark maintainerSource ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—gemini-3-8-flash10.3663 %2026-09-292026-09-29benchmark maintainerSource ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—grok-4-62.11 %2026-09-292026-09-29benchmark maintainerSource ↗
Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Opus Max with default fallback; Astra Max—claude-opus-5-561 %2026-09-292026-09-29independent labSource ↗
Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Opus Max with default fallback; Astra Max—gpt-6-astra55 %2026-09-292026-09-29independent labSource ↗
Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Gemini High—gemini-3-8-flash48 %2026-09-292026-09-29independent labSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgemini-3.1-pro-preview-highgemini-3-1-pro-preview80.769 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgemini-3.1-pro-preview-highgemini-3-1-pro-preview85.25 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgemini-3.5-flash-highgemini-3-5-flash80.769 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgemini-3.5-flash-highgemini-3-5-flash77.25 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgrok-4.5grok-4-582.692 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgrok-4.5grok-4-594 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgpt-5.6-sol-maxgpt-5.6-sol84.615 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgpt-5.6-sol-maxgpt-5.6-sol100 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgpt-5.6-terra-maxgpt-5.6-terra86.538 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgpt-5.6-terra-maxgpt-5.6-terra100 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgpt-5.6-luna-maxgpt-5.6-luna73.077 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgpt-5.6-luna-maxgpt-5.6-luna93.5 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgemini-3.5-flash-lite-highgemini-3-5-flash-lite50 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgemini-3.5-flash-lite-highgemini-3-5-flash-lite32.75 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgemini-3.6-flash-highgemini-3-6-flash78.846 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgemini-3.6-flash-highgemini-3-6-flash85.75 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantdeepseek-v4-pro-0813deepseek-v4-pro84.615 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantdeepseek-v4-pro-0813deepseek-v4-pro96.75 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgrok-4.6grok-4-686.538 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgrok-4.6grok-4-695.5 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgemini-3.7-flash-highgemini-3-7-flash82.692 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgemini-3.7-flash-highgemini-3-7-flash94.5 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantclaude-fable-5-1-max-effortclaude-fable-5-180.769 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantclaude-fable-5-1-max-effortclaude-fable-5-1100 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgemini-3.8-flash-highgemini-3-8-flash76.923 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgemini-3.8-flash-highgemini-3-8-flash98.25 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgpt-6-astra-maxgpt-6-astra84.615 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgpt-6-astra-maxgpt-6-astra100 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-maxdeepseek-v4-1-flash80.769 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-maxdeepseek-v4-1-flash100 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effortclaude-opus-5-584.615 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effortclaude-opus-5-5100 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgpt-6-sol-maxgpt-6-sol84.615 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgpt-6-sol-maxgpt-6-sol100 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgpt-6-luna-maxgpt-6-luna73.077 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgpt-6-luna-maxgpt-6-luna96 %not published2026-09-30benchmark maintainerSource ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effortclaude-sonnet-5-586.538 %not published2026-09-30benchmark maintainerSource ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effortclaude-sonnet-5-5100 %not published2026-09-30benchmark maintainerSource ↗

Limitations: Prompt wording and hidden test contamination can change outcomes. Publication methods, prompts, sampling and model versions may differ. Check each cited source before interpreting a result.