HOW WE COUNT
The score
behind the score.
Version 3.0.0 · Updated 2026-09-29. Every published number comes from stored, attributed records.
Quantum Score
A model can receive a 0–100 score when current, comparable benchmark results meet the evidence thresholds. We normalize each result within its own benchmark, unit and direction, then average within a category. We never compare raw results from incompatible tests. A category is absent unless a compatible test has at least three distinct models and a benchmark maintainer or independent lab source. A sole publisher may report an exact benchmark cohort; independent domains are counted across the model before an overall score appears.
Coding22%
Reasoning15%
Math12%
Research5%
Agents12%
Multimodal12%
Speed10%
Value12%
Weights apply only to covered categories. Coverage is the sum of available category weights. A numeric overall score requires at least 60% weighted coverage, 4 represented benchmark categories, and 3 benchmark source domains, including 2 independent domains, plus 8 independent observations. Otherwise it is withheld as insufficient verified data.
Price and value
Input and output API prices are stored as USD per million tokens. When both are present, illustrative price = input × 0.75 + output × 0.25. Value = mean available performance score × 10 / (illustrative price + 10), capped at 100. Without at least two performance categories and both recently verified prices, value is unavailable. The 75/25 token mix is illustrative and does not represent every workload.
Speed evidence
Output throughput and time to first token are stored with effort level and method. They remain descriptive until a compatible cohort supports a speed category score. Max-effort response times include reasoning and are not general UI latency.
Freshness and corrections
Records older than 90 days since verification are excluded from score calculation. Source pages retain publication, test and verification dates. Corrections update the seed record, are reviewed through version control, and trigger a fresh import and ranking snapshot.
Scoring changelog
2026-09-29 — V3.0.0: independent-source classification, exact harness matching, performance ledger and stricter overall evidence threshold.