THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

THE EVIDENCE / AGENTS

Agents
benchmarks.

Completion of multi-step tool-using tasks. Results are grouped by benchmark and unit. Different tests are never merged into a raw leaderboard.

01 / RESULTS

Published results

BenchmarkVersionMethodConfigurationModelResultTest dateVerifiedSource typeSource
APEX-AgentsAPEX-AgentsPublisher-reported configuration; see source—grok-4-657.5 %2026-08-122026-09-29provider first partySource ↗
APEX-AgentsAPEX-AgentsPublisher-reported configuration; see source—grok-4-547.1 %2026-08-122026-09-29provider first partySource ↗
Terminal-Bench v3.0Terminal-Bench v3.0Publisher-reported configuration; see source—grok-4-626 %2026-08-122026-09-29provider first partySource ↗
Terminal-Bench v3.0Terminal-Bench v3.0Publisher-reported configuration; see source—grok-4-515.7 %2026-08-122026-09-29provider first partySource ↗
GDPVal-AA v2GDPVal-AA v2Publisher-reported configuration; see source—grok-4-61753 Elo2026-08-122026-09-29provider first partySource ↗
GDPVal-AA v2GDPVal-AA v2Publisher-reported configuration; see source—grok-4-51526 Elo2026-08-122026-09-29provider first partySource ↗
GDPval-AA v2.1 (reported best effort)GDPval-AA v2.1 (reported best effort)Publisher-reported configuration; see source—claude-sonnet-5-51844 Elo2026-09-282026-09-29provider first partySource ↗
GDPval-AA v2.1 (reported best effort)GDPval-AA v2.1 (reported best effort)Publisher-reported configuration; see source—claude-opus-5-51846 Elo2026-09-282026-09-29provider first partySource ↗
GDPval-AA v2.1 (reported best effort)GDPval-AA v2.1 (reported best effort)Publisher-reported configuration; see source—gpt-6-sol1487 Elo2026-09-282026-09-29provider first partySource ↗
AA-Briefcase v1.1 (reported best effort)AA-Briefcase v1.1 (reported best effort)Publisher-reported configuration; see source—claude-sonnet-5-51811 Elo2026-09-282026-09-29provider first partySource ↗
AA-Briefcase v1.1 (reported best effort)AA-Briefcase v1.1 (reported best effort)Publisher-reported configuration; see source—claude-opus-5-51822 Elo2026-09-282026-09-29provider first partySource ↗
AA-Briefcase v1.1 (reported best effort)AA-Briefcase v1.1 (reported best effort)Publisher-reported configuration; see source—gpt-6-sol1483 Elo2026-09-282026-09-29provider first partySource ↗
OSWorld 2.1 (partial; reported best effort)OSWorld 2.1 (partial; reported best effort)Publisher-reported configuration; see source—claude-sonnet-5-580.1 %2026-09-282026-09-29provider first partySource ↗
OSWorld 2.1 (partial; reported best effort)OSWorld 2.1 (partial; reported best effort)Publisher-reported configuration; see source—claude-opus-5-581.8 %2026-09-282026-09-29provider first partySource ↗
SkillsBench Vals OpenHands with skillsVals 2026-09-27Vals OpenHands harness with skills—deepseek-v4-1-flash69.8 %2026-09-272026-09-29independent labSource ↗
SkillsBench Vals OpenHands with skillsVals 2026-09-27Vals OpenHands harness with skills—grok-4-566.03 %2026-09-272026-09-29independent labSource ↗
SkillsBench Vals OpenHands with skillsVals 2026-09-27Vals OpenHands harness with skills—gemini-3-7-flash65.89 %2026-09-272026-09-29independent labSource ↗
AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Opus Max with default fallback; Astra Max—claude-opus-5-570 %2026-09-292026-09-29independent labSource ↗
AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Opus Max with default fallback; Astra Max—gpt-6-astra68 %2026-09-292026-09-29independent labSource ↗
AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Gemini High—gemini-3-8-flash60 %2026-09-292026-09-29independent labSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgemini-3.1-pro-preview-highgemini-3-1-pro-preview59.091 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgemini-3.5-flash-highgemini-3-5-flash63.636 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgrok-4.5grok-4-572.727 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgpt-5.6-sol-maxgpt-5.6-sol63.636 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgpt-5.6-terra-maxgpt-5.6-terra68.182 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgpt-5.6-luna-maxgpt-5.6-luna63.636 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgemini-3.5-flash-lite-highgemini-3-5-flash-lite59.091 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgemini-3.6-flash-highgemini-3-6-flash63.636 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantdeepseek-v4-pro-0813deepseek-v4-pro68.182 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgrok-4.6grok-4-672.727 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgemini-3.7-flash-highgemini-3-7-flash68.182 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantclaude-fable-5-1-max-effortclaude-fable-5-168.182 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgemini-3.8-flash-highgemini-3-8-flash72.727 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgpt-6-astra-maxgpt-6-astra63.636 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-maxdeepseek-v4-1-flash81.818 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantclaude-opus-5-5-max-effortclaude-opus-5-581.818 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgpt-6-sol-maxgpt-6-sol63.636 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantgpt-6-luna-maxgpt-6-luna63.636 %not published2026-09-30benchmark maintainerSource ↗
LiveBench javascript2026-06-25Public table; published model effort variantclaude-sonnet-5-5-max-effortclaude-sonnet-5-577.273 %not published2026-09-30benchmark maintainerSource ↗

Limitations: Tool sets and time budgets differ across evaluations. Publication methods, prompts, sampling and model versions may differ. Check each cited source before interpreting a result.