01 / EVIDENCE
Benchmark breakdown
| Benchmark | Version | Method | Configuration | Result | Test date | Verified | Source type | Source |
|---|---|---|---|---|---|---|---|---|
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | 0.1942 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | View source ↗ |
| LiveBench code_completion | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | 80.435 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench code_generation | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | 77.465 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench javascript | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | 63.636 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench AMPS_Hard | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | 98 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench olympiad | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | 90.376 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | 73.077 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gpt-6-luna-max | 96 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
02 / COST
Price & speed
Input / 1M tokens $0.10
Output / 1M tokens $0.50
Cached input / 1M $0.01
Batch input / 1M $0.05
Batch output / 1M $0.25
Speed No comparable observation yet
Price verified 2026-09-29
Price scope Standard short-context token rate; long-context and service tiers differ
Price source Official pricing ↗
02A / PERFORMANCE
Measured speed
No sourced speed measurement at a stated effort level yet.
03 / PROFILE
Capabilities
04 / CONTEXT
Strengths & limitations
Coverage: 73% · Limited evidence · 8 observations · 2 source domains (2 independent).
Why no Overall AIQuantumScore?
Qualified category evidence: coding, reasoning, math, agents. An overall score is withheld because the independent observation threshold has not been met. Evidence still missing or lacking comparable overlap: research, multimodal, speed.
05 / HISTORY
Version & price history
Version gpt-6-luna · released 2026-09-22.
2026-09-29: $0.10 input / $0.50 output per 1M tokens · Price source ↗
06 / SOURCES
Source records
- GPT-6 Luna official documentation ↗ · verified 2026-09-29
- OpenAI official pricing ↗ · verified 2026-09-29
DATED EVIDENCE
History
Standard API price history
- 2026-09-29: 0.2 USD / 1M · 75% input / 25% output USD per 1M · source ↗
One observation; no trend can be inferred.
ARC-AGI-3 Semi-Private Standard harness best effort / 3 Semi-Private / %
- 2026-09-29: 0.194 result · ARC Prize Standard harness; best published effort per model · source ↗
One observation; no trend can be inferred.
LiveBench code_completion / 2026-06-25 / %
- 2026-09-30: 80.435 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench code_generation / 2026-06-25 / %
- 2026-09-30: 77.465 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench javascript / 2026-06-25 / %
- 2026-09-30: 63.636 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench AMPS_Hard / 2026-06-25 / %
- 2026-09-30: 98 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench olympiad / 2026-06-25 / %
- 2026-09-30: 90.376 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench theory_of_mind / 2026-06-25 / %
- 2026-09-30: 73.077 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench zebra_puzzle / 2026-06-25 / %
- 2026-09-30: 96 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
reasoning category score history
- 2026-09-30: 0 score · Scoring 3.0.0
One observation; no trend can be inferred.
07 / EXPLORE