THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

OpenAI / GPT-6

GPT-6 Sol

Version gpt-6-sol · Released 2026-09-22 · Available

Compare this model ↗
48QUANTUM SCORE

Explore OpenAI history →

Coding39
Reasoning67
Math79
Research—
Agents20
Multimodal—
Speed—
Value37

01 / EVIDENCE

Benchmark breakdown

BenchmarkVersionMethodConfigurationResultTest dateVerifiedSource typeSource
GDPval-AA v2.1 (reported best effort)GDPval-AA v2.1 (reported best effort)Publisher-reported configuration; see source—1487 Elo2026-09-282026-09-29provider first partyView source ↗
AA-Briefcase v1.1 (reported best effort)AA-Briefcase v1.1 (reported best effort)Publisher-reported configuration; see source—1483 Elo2026-09-282026-09-29provider first partyView source ↗
Chartography (no tools; reported best effort)Chartography (no tools; reported best effort)Publisher-reported configuration; see source—53.6 %2026-09-282026-09-29provider first partyView source ↗
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—4.6236 %2026-09-292026-09-29benchmark maintainerView source ↗
ProgramBench Vals 200 public tasks raw pass rateVals 2026-09-27Vals 200 public tasks; raw pass rate; mini-SWE-agent—81.9 %2026-09-272026-09-29independent labView source ↗
LiveBench code_completion2026-06-25Public table; published model effort variantgpt-6-sol-max80.435 %not published2026-09-30benchmark maintainerView source ↗
LiveBench code_generation2026-06-25Public table; published model effort variantgpt-6-sol-max83.099 %not published2026-09-30benchmark maintainerView source ↗
LiveBench javascript2026-06-25Public table; published model effort variantgpt-6-sol-max63.636 %not published2026-09-30benchmark maintainerView source ↗
LiveBench AMPS_Hard2026-06-25Public table; published model effort variantgpt-6-sol-max98 %not published2026-09-30benchmark maintainerView source ↗
LiveBench olympiad2026-06-25Public table; published model effort variantgpt-6-sol-max91.365 %not published2026-09-30benchmark maintainerView source ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgpt-6-sol-max84.615 %not published2026-09-30benchmark maintainerView source ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgpt-6-sol-max100 %not published2026-09-30benchmark maintainerView source ↗

02 / COST

Price & speed

Input / 1M tokens $2.00

Output / 1M tokens $10.00

Cached input / 1M —

Batch input / 1M —

Batch output / 1M —

Speed No comparable observation yet

Price verified 2026-09-29

Price scope Standard short-context token rate; long-context and service tiers differ

Price source Official pricing ↗

02A / PERFORMANCE

Measured speed

No sourced speed measurement at a stated effort level yet.

03 / PROFILE

Capabilities

tool usevision

04 / CONTEXT

Strengths & limitations

Coverage: 73% · Moderate confidence · 12 observations · 4 source domains (3 independent).

05 / HISTORY

Version & price history

Version gpt-6-sol · released 2026-09-22.

2026-09-29: $2.00 input / $10.00 output per 1M tokens · Price source ↗

06 / SOURCES

Source records

DATED EVIDENCE

History

Standard API price history

2026-09-292026-09-29
  • 2026-09-29: 4 USD / 1M · 75% input / 25% output USD per 1M · source ↗

One observation; no trend can be inferred.

GDPval-AA v2.1 (reported best effort) / GDPval-AA v2.1 (reported best effort) / Elo

2026-09-282026-09-28
  • 2026-09-28: 1,487 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

AA-Briefcase v1.1 (reported best effort) / AA-Briefcase v1.1 (reported best effort) / Elo

2026-09-282026-09-28
  • 2026-09-28: 1,483 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

Chartography (no tools; reported best effort) / Chartography (no tools; reported best effort) / %

2026-09-282026-09-28
  • 2026-09-28: 53.6 result · Publisher-reported configuration; see source · source ↗

One observation; no trend can be inferred.

ARC-AGI-3 Semi-Private Standard harness best effort / 3 Semi-Private / %

2026-09-292026-09-29
  • 2026-09-29: 4.624 result · ARC Prize Standard harness; best published effort per model · source ↗

One observation; no trend can be inferred.

ProgramBench Vals 200 public tasks raw pass rate / Vals 2026-09-27 / %

2026-09-272026-09-27
  • 2026-09-27: 81.9 result · Vals 200 public tasks; raw pass rate; mini-SWE-agent · source ↗

One observation; no trend can be inferred.

LiveBench code_completion / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 80.435 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench code_generation / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 83.099 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench javascript / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 63.636 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench AMPS_Hard / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 98 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench olympiad / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 91.365 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench theory_of_mind / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 84.615 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench zebra_puzzle / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 100 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

coding category score history

2026-09-302026-09-30
  • 2026-09-30: 0 score · Scoring 3.0.0

One observation; no trend can be inferred.

reasoning category score history

2026-09-302026-09-30
  • 2026-09-30: 7 score · Scoring 3.0.0

One observation; no trend can be inferred.

07 / EXPLORE

Similar models