01 / EVIDENCE
Benchmark breakdown
| Benchmark | Version | Method | Configuration | Result | Test date | Verified | Source type | Source |
|---|---|---|---|---|---|---|---|---|
| FrontierMath Tier 4 v2 | FrontierMath Tier 4 v2 | Publisher-reported configuration; see source | — | 98 % | 2026-09-03 | 2026-09-29 | provider first party | View source ↗ |
| ARC-AGI-3 Provider Adapter (High) | ARC-AGI-3 Provider Adapter (High) | Publisher-reported configuration; see source | — | 99.9 % | 2026-09-03 | 2026-09-29 | provider first party | View source ↗ |
| ARC-AGI-3 Semi-Private Standard harness best effort | 3 Semi-Private | ARC Prize Standard harness; best published effort per model | — | 62.7128 % | 2026-09-29 | 2026-09-29 | benchmark maintainer | View source ↗ |
| Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed | 4.0 | Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures | — | 59.6 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| ProofBench v1.1 Vals 100 proof tasks | 1.1 | Vals 100 proof tasks; ceiling effect at top scores | — | 99 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| ProgramBench Vals 200 public tasks raw pass rate | Vals 2026-09-27 | Vals 200 public tasks; raw pass rate; mini-SWE-agent | — | 85.4 % | 2026-09-27 | 2026-09-29 | independent lab | View source ↗ |
| Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) | AA Index 4.3.2 | Opus Max with default fallback; Astra Max | — | 55 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) | AA Index 4.3.2 | Opus Max with default fallback; Astra Max | — | 68 % | 2026-09-29 | 2026-09-29 | independent lab | View source ↗ |
| LiveBench code_completion | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | 80.435 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench code_generation | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | 80.282 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench javascript | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | 63.636 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench AMPS_Hard | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | 98 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench olympiad | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | 92.166 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench theory_of_mind | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | 84.615 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
| LiveBench zebra_puzzle | 2026-06-25 | Public table; published model effort variant | gpt-6-astra-max | 100 % | not published | 2026-09-30 | benchmark maintainer | View source ↗ |
02 / COST
Price & speed
Input / 1M tokens $10.00
Output / 1M tokens $50.00
Cached input / 1M $1.00
Batch input / 1M $5.00
Batch output / 1M $25.00
Speed See measured observations below
Price verified 2026-09-29
Price scope Standard short-context token rate; long-context and service tiers differ
Price source Official pricing ↗
02A / PERFORMANCE
Measured speed
| Metric | Result | Effort | Method | Source |
|---|---|---|---|---|
| output tokens per second | 55 tokens/s | Max | Artificial Analysis model comparison; output speed | View source ↗ |
| time to first token | 290.98 seconds | Max | Artificial Analysis model comparison; includes reasoning | View source ↗ |
| end to end response time | 300.06 seconds | Max | Artificial Analysis model comparison; includes reasoning | View source ↗ |
03 / PROFILE
Capabilities
04 / CONTEXT
Strengths & limitations
Coverage: 73% · Moderate confidence · 15 observations · 5 source domains (4 independent).
05 / HISTORY
Version & price history
Version gpt-6-astra · released 2026-09-03.
2026-09-29: $10.00 input / $50.00 output per 1M tokens · Price source ↗
06 / SOURCES
Source records
- GPT-6 Astra official documentation ↗ · verified 2026-09-29
- OpenAI official pricing ↗ · verified 2026-09-29
DATED EVIDENCE
History
Standard API price history
- 2026-09-29: 20 USD / 1M · 75% input / 25% output USD per 1M · source ↗
One observation; no trend can be inferred.
FrontierMath Tier 4 v2 / FrontierMath Tier 4 v2 / %
- 2026-09-03: 98 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
ARC-AGI-3 Provider Adapter (High) / ARC-AGI-3 Provider Adapter (High) / %
- 2026-09-03: 99.9 result · Publisher-reported configuration; see source · source ↗
One observation; no trend can be inferred.
ARC-AGI-3 Semi-Private Standard harness best effort / 3 Semi-Private / %
- 2026-09-29: 62.713 result · ARC Prize Standard harness; best published effort per model · source ↗
One observation; no trend can be inferred.
Terminal-Bench 4.0 Vals mini-SWE-agent avg@3 fallbacks failed / 4.0 / %
- 2026-09-29: 59.6 result · Vals mini-SWE-agent avg@3; Anthropic fallbacks treated as failures · source ↗
One observation; no trend can be inferred.
ProofBench v1.1 Vals 100 proof tasks / 1.1 / %
- 2026-09-29: 99 result · Vals 100 proof tasks; ceiling effect at top scores · source ↗
One observation; no trend can be inferred.
ProgramBench Vals 200 public tasks raw pass rate / Vals 2026-09-27 / %
- 2026-09-27: 85.4 result · Vals 200 public tasks; raw pass rate; mini-SWE-agent · source ↗
One observation; no trend can be inferred.
Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) / AA Index 4.3.2 / %
- 2026-09-29: 55 result · Opus Max with default fallback; Astra Max · source ↗
One observation; no trend can be inferred.
AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) / AA Index 4.3.2 / %
- 2026-09-29: 68 result · Opus Max with default fallback; Astra Max · source ↗
One observation; no trend can be inferred.
LiveBench code_completion / 2026-06-25 / %
- 2026-09-30: 80.435 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench code_generation / 2026-06-25 / %
- 2026-09-30: 80.282 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench javascript / 2026-06-25 / %
- 2026-09-30: 63.636 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench AMPS_Hard / 2026-06-25 / %
- 2026-09-30: 98 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench olympiad / 2026-06-25 / %
- 2026-09-30: 92.166 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench theory_of_mind / 2026-06-25 / %
- 2026-09-30: 84.615 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
LiveBench zebra_puzzle / 2026-06-25 / %
- 2026-09-30: 100 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗
One observation; no trend can be inferred.
output tokens per second / Max / tokens/s / Artificial Analysis model comparison; output speed
- 2026-09-29: 55 measured · Artificial Analysis model comparison; output speed · source ↗
One observation; no trend can be inferred.
time to first token / Max / seconds / Artificial Analysis model comparison; includes reasoning
- 2026-09-29: 290.98 measured · Artificial Analysis model comparison; includes reasoning · source ↗
One observation; no trend can be inferred.
end to end response time / Max / seconds / Artificial Analysis model comparison; includes reasoning
- 2026-09-29: 300.06 measured · Artificial Analysis model comparison; includes reasoning · source ↗
One observation; no trend can be inferred.
coding category score history
- 2026-09-30: 72 score · Scoring 3.0.0
One observation; no trend can be inferred.
reasoning category score history
- 2026-09-30: 77 score · Scoring 3.0.0
One observation; no trend can be inferred.
math category score history
- 2026-09-30: 94 score · Scoring 3.0.0
One observation; no trend can be inferred.
agents category score history
- 2026-09-30: 80 score · Scoring 3.0.0
One observation; no trend can be inferred.
07 / EXPLORE