THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

HEAD TO HEAD

Make the
right call.

Pick two to four models. Compare category evidence, price, context, speed and capabilities side by side.

01 / SCORECARD

Coding focus · side by side

MetricGPT-6 LunaGPT-6.1 Sol
Quantum Score46QUANTUM SCOREInsufficient verified data61% coverage
Confidence / coverageModerate confidence · 73% · 3 independent domainsLimited evidence · 61% · 3 independent domains
Context window1,050,000 tokens1,050,000 tokens
Modalities / toolstext, image · tool use, visiontext, image · tool use, vision
Coding3987 +48 vs first
Reasoning5284 +32 vs first
Math75No data
REVAL™
ⓘREVAL™ Research Evaluation Index — our proprietary 0–100 benchmark measuring real-world AI research performance. Methodology →
Not testedNot tested
Agents20No data
MultimodalNo data48
Speed122.4 tok/sMax · 2026-10-0164 tok/sMax · 2026-10-01
Value4652 +6 vs first
Input / 1M$0.10$2.00
Output / 1M$0.50$10.00
Measured output speed122.4 tok/sMax · 2026-10-0164 tok/sMax · 2026-10-01

THE DECISION

What this comparison means

Quality for your coding focus

1 current independent evaluations use matching versions, methods, and test settings across every selected model. These show performance on the named tests; they do not establish a universal best model.

GPT-6.1 Sol leads on Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed

4.0 · Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures

  • GPT-6 Luna: 13.636 % · max effort · Evidence ↗ · verified 2026-10-01
  • GPT-6.1 Sol: 55.051 % · max effort · Evidence ↗ · verified 2026-10-01

Different benchmarks test different tasks. Effort levels are shown because they can affect results and cost. Verification dates indicate when we checked a source; they are not necessarily evaluation dates.

Your estimated monthly API cost

Enter usage in millions of tokens. Estimates use the listed standard input and output prices in USD.

GPT-6 Luna

$2.00

Lowest estimated cost among these models

Pricing source ↗ · verified 2026-10-01

Price scope

Standard short-context token rate; long-context and service tiers differ

GPT-6.1 Sol

$40.00

Pricing source ↗ · verified 2026-10-01

Price scope

Standard short-context token rate; long-context and service tiers differ

This estimates API token charges. Subscription plans, caching, batch discounts, tools, taxes, and differences in tokens used per task can change the bill. Cost alone does not establish equivalent quality.

Tradeoffs and evidence gaps

GPT-6 Luna ↗

Listed capabilities: tool use, vision.

Independent benchmark records cover reasoning, coding, agents, math. Coverage does not guarantee comparable tests for every selected model.

Review capabilities, dates, and source records →

GPT-6.1 Sol ↗

Listed capabilities: tool use, vision.

Independent benchmark records cover multimodal, coding, reasoning. Coverage does not guarantee comparable tests for every selected model.

Review capabilities, dates, and source records →

How we assess scores and missing evidence →

Keep this comparison

Create a free account to save these models and get alerts when their verified evidence or prices change.

Sign in to save ↗

02 / EVIDENCE

Published benchmark evidence

BenchmarkGPT-6 LunaGPT-6.1 Sol
ARC-AGI-3 Semi-Private Standard harness best effort · 3 Semi-Private · ARC Prize Standard harness; best published effort per model0.1942 %
benchmark maintainer ↗
52.72904 %
benchmark maintainer ↗
Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed · 4.0 · Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures13.636 %
independent lab ↗
55.051 %
independent lab ↗

AS THE EVIDENCE CHANGED

Comparison history

Dated releases, price events, and benchmark publications. A price or test change can alter the factual comparison; this timeline does not pick a winner.

2026-09-29 · GPT-6.1 Sol released ↗ · source

2026-09-22 · GPT-6 Sol and Luna added to the API ↗ · source

2026-10-01 verified (test date not published) · gpt-6-luna entered LiveBench code_completion (2026-06-25) at 80.435 % · source ↗

2026-10-01 verified (test date not published) · gpt-6-luna entered LiveBench code_generation (2026-06-25) at 77.465 % · source ↗

2026-10-01 verified (test date not published) · gpt-6-luna entered LiveBench javascript (2026-06-25) at 63.636 % · source ↗

2026-10-01 verified (test date not published) · gpt-6-luna entered LiveBench AMPS_Hard (2026-06-25) at 98 % · source ↗

2026-10-01 verified (test date not published) · gpt-6-luna entered LiveBench olympiad (2026-06-25) at 90.376 % · source ↗

2026-10-01 verified (test date not published) · gpt-6-luna entered LiveBench theory_of_mind (2026-06-25) at 73.077 % · source ↗

2026-10-01 verified (test date not published) · gpt-6-luna entered LiveBench zebra_puzzle (2026-06-25) at 96 % · source ↗

2026-10-01 verified (test date not published) · gpt-6.1-sol entered Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed (4.0) at 55.051 % · source ↗

2026-10-01 verified (test date not published) · gpt-6-luna entered Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed (4.0) at 13.636 % · source ↗

2026-10-01 verified (test date not published) · gpt-6.1-sol entered ARC-AGI-3 Semi-Private Standard harness best effort (3 Semi-Private) at 52.72904 % · source ↗

2026-09-30 verified (test date not published) · gpt-6.1-sol entered MedXpertQA MM Mercor public leaderboard (MedXpertQA MM Mercor 200 public tasks) at 85.7 % · source ↗

2026-09-30 verified (test date not published) · gpt-6.1-sol entered MedXpertQA MM Mercor 51 held-out tasks (MedXpertQA MM Mercor Extended 51 held-out tasks) at 64.7 % · source ↗

An overall score appears only after enough current, comparable, independently sourced results exist. Pricing and individual benchmark results remain visible meanwhile.