THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

DECISION GUIDE

Choose an AI for coding

Match the model to your development workflow, then compare independent evidence, effort and cost.

Start with the job you need done

Completing a function

Code completion and generation tests can help shortlist models. They do not establish that a model can navigate your repository or safely edit multiple files.

Fixing an existing codebase

Look for repository or terminal tasks with the same harness and tool access. Check whether retries, fallbacks and multiple attempts are included.

Running a coding agent

Agents need more than good code generation. Inspect terminal and tool-use evidence, then start with bounded tasks, review diffs and run your own tests.

Read the numbers without overreading them

A benchmark result

A reported pass rate belongs to a particular task set, harness and attempt policy. Two similar benchmark names can represent different tests.

A category score

Our coding score normalizes compatible results within the current model pool. A score of 100 means top relative performance on the covered evidence, not perfect coding.

Speed versus effort

More reasoning may improve difficult tasks while increasing delay and output tokens. Throughput alone does not measure how quickly a coding task finishes.

Inspect coding results and sources → · How scoring works → · Evidence gaps →

Models with comparable coding evidence

These are starting points from our current comparison pool. Category scores are relative to covered tests; they are not a universal recommendation.

Build my shortlist ↗ Compare evidence and estimate costs →

Run a small, useful trial

  1. Use the exact model version and reasoning effort you plan to deploy.
  2. Try a small set of representative bugs, features and unfamiliar files.
  3. Count successful tasks, review time, total token cost and time to finish.
  4. Keep your existing tests and human review before accepting changes.

Keep the prompts, model version and settings fixed across candidates. Include failures in your notes. Choose using your task results alongside the published evidence.

Keep your decision current

Save your comparison and follow the shortlisted models. Return when independent results, prices or capabilities change; a new announcement alone does not establish better performance.

What changed for my models? ↗ This week’s verified changes →

Explore another workflow: Choose an AI for research →Choose an AI for automation →