THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

DeepSeek / DeepSeek V4

DeepSeek V4.1 Flash

Version DeepSeek-V4.1-Flash · Released 2026-09-10 · Available

Compare this model ↗
Insufficient verified data73% coverage

Explore DeepSeek history →

Coding55
Reasoning92
Math69
Research—
Agents100
Multimodal—
Speed—
Value75

01 / EVIDENCE

Benchmark breakdown

BenchmarkVersionMethodConfigurationResultTest dateVerifiedSource typeSource
SkillsBench Vals OpenHands with skillsVals 2026-09-27Vals OpenHands harness with skills—69.8 %2026-09-272026-09-29independent labView source ↗
LiveBench code_completion2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-max82.609 %not published2026-09-30benchmark maintainerView source ↗
LiveBench code_generation2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-max77.465 %not published2026-09-30benchmark maintainerView source ↗
LiveBench javascript2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-max81.818 %not published2026-09-30benchmark maintainerView source ↗
LiveBench AMPS_Hard2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-max98 %not published2026-09-30benchmark maintainerView source ↗
LiveBench olympiad2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-max89.081 %not published2026-09-30benchmark maintainerView source ↗
LiveBench theory_of_mind2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-max80.769 %not published2026-09-30benchmark maintainerView source ↗
LiveBench zebra_puzzle2026-06-25Public table; published model effort variantdeepseek-v4.1-flash-max100 %not published2026-09-30benchmark maintainerView source ↗

02 / COST

Price & speed

Input / 1M tokens $0.30

Output / 1M tokens $1.20

Cached input / 1M —

Batch input / 1M —

Batch output / 1M —

Speed No comparable observation yet

Price verified 2026-09-29

Price scope API ID deepseek-flash. Peak cache-miss input rate; off-peak and cache-hit rates differ

Price source Official pricing ↗

02A / PERFORMANCE

Measured speed

No sourced speed measurement at a stated effort level yet.

03 / PROFILE

Capabilities

visiontool use

04 / CONTEXT

Strengths & limitations

Coverage: 73% · Limited evidence · 8 observations · 2 source domains (2 independent).

Why no Overall AIQuantumScore?

Qualified category evidence: coding, reasoning, math, agents. An overall score is withheld because the independent observation threshold has not been met. Evidence still missing or lacking comparable overlap: research, multimodal, speed.

See the benchmark coverage matrix →

05 / HISTORY

Version & price history

Version DeepSeek-V4.1-Flash · released 2026-09-10.

2026-09-29: $0.30 input / $1.20 output per 1M tokens · Price source ↗

06 / SOURCES

Source records

DATED EVIDENCE

History

Standard API price history

2026-09-292026-09-29
  • 2026-09-29: 0.525 USD / 1M · 75% input / 25% output USD per 1M · source ↗

One observation; no trend can be inferred.

SkillsBench Vals OpenHands with skills / Vals 2026-09-27 / %

2026-09-272026-09-27
  • 2026-09-27: 69.8 result · Vals OpenHands harness with skills · source ↗

One observation; no trend can be inferred.

LiveBench code_completion / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 82.609 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench code_generation / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 77.465 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench javascript / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 81.818 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench AMPS_Hard / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 98 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench olympiad / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 89.081 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench theory_of_mind / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 80.769 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

LiveBench zebra_puzzle / 2026-06-25 / %

2026-09-302026-09-30
  • 2026-09-30: 100 result · Public table; published model effort variant · evaluation date not published; verification date shown · source ↗

One observation; no trend can be inferred.

agents category score history

2026-09-302026-09-30
  • 2026-09-30: 100 score · Scoring 3.0.0

One observation; no trend can be inferred.

07 / EXPLORE

Similar models