Evidence checked 2026-10-01. This interpretation uses the cited records below; it is not a hands-on review.
What the evidence supports
There is not enough comparable independent evidence to recommend this version for a particular task yet.
0 current independent benchmark observations. Several tasks from one suite do not count as separate independent sources.
What remains uncertain
Comparable category scores are missing for coding, reasoning, agents, multimodal. Writing quality, factual accuracy, privacy and production reliability require separate evaluation.
An overall score is withheld because the evidence requirements are not met. Check the effort level and tool configuration before comparing results.
At 10M input + 2M output tokens, the base API estimate is $36.00/month. This excludes tools, caching, batch rates, subscriptions and taxes.
Output tokens per second measures generation after it starts. Reasoning time can make a fast-generating model take longer to finish. Use the measured speed and effort records below.
An overall score is withheld because only 0 benchmark source domains (minimum 3); only 0 independent observations (minimum 8); only 0% of the required weighted category coverage; only 0 qualified categories (minimum 4); only 0 independent source domains (minimum 2). Evidence still missing or lacking comparable overlap: coding, reasoning, math, research, agents, multimodal, speed.