THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

Google / Gemini

Gemini 3.8 Flash

Version gemini-3-8-flash · Released 2026-09-02 · Available

Compare this model ↗
51QUANTUM SCORE

Explore Google history →

Coding16
Reasoning62
Math95
Research—
Agents60
Multimodal—
Speed—
Value51

01 / EVIDENCE

Benchmark breakdown

BenchmarkVersionMethodConfigurationResultTest dateVerifiedSource typeSource
ARC-AGI-3 Semi-Private Standard harness best effort3 Semi-PrivateARC Prize Standard harness; best published effort per model—10.3663 %2026-09-292026-09-29benchmark maintainerView source ↗
Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Gemini High—48 %2026-09-292026-09-29independent labView source ↗
AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High)AA Index 4.3.2Gemini High—60 %2026-09-292026-09-29independent labView source ↗
LiveBench code_completion2026-06-25Public table; published model effort variantgemini-3.8-flash-high71.739 %not published2026-09-30benchmark maintainerView source ↗
LiveBench code_generation2026-06-25Public table; published model effort variantgemini-3.8-flash-high73.239 %not published2026-09-30benchmark maintainerView source ↗
LiveBench javascript2026-06-25Public table; published model effort variantgemini-3.8-flash-high72.727 %not published2026-09-30benchmark maintainerView source ↗
LiveBench AMPS_Hard2026-06-25Public table; published model effort variantgemini-3.8-flash-high99 %not published2026-09-30benchmark maintainerView source ↗
LiveBench olympiad2026-06-25Public table; published model effort variantgemini-3.8-flash-high92.159 %not published2026-09-30benchmark maintainerView source ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantgemini-3.8-flash-high76.923 %not published2026-09-30benchmark maintainerView source ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantgemini-3.8-flash-high98.25 %not published2026-09-30benchmark maintainerView source ↗

02 / COST

Price & speed

Input / 1M tokens $0.75

Output / 1M tokens $3.75

Cached input / 1M —

Batch input / 1M —

Batch output / 1M —

Speed No comparable observation yet

Price verified 2026-09-29

Price scope Introductory rate through 2026-12-31

Price source Official pricing ↗

02A / PERFORMANCE

Measured speed

No sourced speed measurement at a stated effort level yet.

03 / PROFILE

Capabilities

visiontool use

04 / CONTEXT

Strengths & limitations

Coverage: 73% · Moderate confidence · 10 observations · 3 source domains (3 independent).

05 / HISTORY

Version & price history

Version gemini-3-8-flash · released 2026-09-02.

2026-09-29: $0.75 input / $3.75 output per 1M tokens · Price source ↗

06 / SOURCES

Source records

DATED EVIDENCE

History

Standard API price history

2026-09-292026-09-29
  • 2026-09-29: 1.5 USD / 1M · 75% input / 25% output USD per 1M · source ↗

One observation; no trend can be inferred.

ARC-AGI-3 Semi-Private Standard harness best effort / 3 Semi-Private / %

2026-09-292026-09-29
  • 2026-09-29: 10.366 result · ARC Prize Standard harness; best published effort per model · source ↗

One observation; no trend can be inferred.

Humanity's Last Exam AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) / AA Index 4.3.2 / %

2026-09-292026-09-29
  • 2026-09-29: 48 result · Gemini High · source ↗

One observation; no trend can be inferred.

AutomationBench-AA Index v4.3.2 (Opus Max fallback, Astra Max, Gemini High) / AA Index 4.3.2 / %

2026-09-292026-09-29
  • 2026-09-29: 60 result · Gemini High · source ↗

One observation; no trend can be inferred.

LiveBench code_completion / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 71.739 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench code_generation / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 73.239 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench javascript / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 72.727 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench AMPS_Hard / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 99 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench olympiad / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 92.159 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench theory_of_mind / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 76.923 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench zebra_puzzle / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 98.25 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

reasoning category score history

2026-09-302026-09-30
  • 2026-09-30: 8 score · Scoring 3.0.0

One observation; no trend can be inferred.

agents category score history

2026-09-302026-09-30
  • 2026-09-30: 0 score · Scoring 3.0.0

One observation; no trend can be inferred.

07 / EXPLORE

Similar models