Completing a function
Code completion and generation tests can help shortlist models. They do not establish that a model can navigate your repository or safely edit multiple files.
DECISION GUIDE
Match the model to your development workflow, then compare independent evidence, effort and cost.
Code completion and generation tests can help shortlist models. They do not establish that a model can navigate your repository or safely edit multiple files.
Look for repository or terminal tasks with the same harness and tool access. Check whether retries, fallbacks and multiple attempts are included.
Agents need more than good code generation. Inspect terminal and tool-use evidence, then start with bounded tasks, review diffs and run your own tests.
A reported pass rate belongs to a particular task set, harness and attempt policy. Two similar benchmark names can represent different tests.
Our coding score normalizes compatible results within the current model pool. A score of 100 means top relative performance on the covered evidence, not perfect coding.
More reasoning may improve difficult tasks while increasing delay and output tokens. Throughput alone does not measure how quickly a coding task finishes.
Inspect coding results and sources → · How scoring works → · Evidence gaps →
These are starting points from our current comparison pool. Category scores are relative to covered tests; they are not a universal recommendation.
coding: 100/100 · Moderate confidence · 73% overall category coverage.
Read the practical guide → Follow updates →coding: 94/100 · Moderate confidence · 85% overall category coverage.
Read the practical guide → Follow updates →coding: 87/100 · Limited evidence · 61% overall category coverage.
Read the practical guide → Follow updates →coding: 75/100 · Moderate confidence · 85% overall category coverage.
Read the practical guide → Follow updates →Keep the prompts, model version and settings fixed across candidates. Include failures in your notes. Choose using your task results alongside the published evidence.
Save your comparison and follow the shortlisted models. Return when independent results, prices or capabilities change; a new announcement alone does not establish better performance.
What changed for my models? ↗ This week’s verified changes →Explore another workflow: Choose an AI for research →Choose an AI for automation →