THE SCOREBOARD FOR WHAT'S NEXTAI + QUANTUM, SCORED. COMPARED. EXPLAINED.
AIQUANTUMSCORE.

RESEARCH EVALUATION INDEX / V1.0.0

Research,
evaluated.

REVAL™ is our proprietary benchmark for measuring how effectively AI models conduct real-world research.

Methodology 1.0.0Latest evaluation: No completed evaluations yetView leaderboard →

Why REVAL exists

Research requires more than knowing an answer. A model must find evidence, verify claims, reconcile numbers, recognize conflicts and communicate a useful conclusion. REVAL evaluates that entire workflow against a controlled research assignment. It is published separately from third-party benchmark and provider data; it does not enter the existing Quantum Score.

Nine categories. One 0–100 scale.

Each criterion is graded against a frozen answer key or rubric, converted to a percentage and combined using the task’s preregistered criterion weights. Independent judge ratings are averaged within a criterion. Category scores average applicable completed tasks; the overall score is the weighted sum below. Missing evidence is never silently treated as a successful result. A score of 100 means full credit under this methodology, not perfect research in every setting.

CategoryWeightWhat we measure
Discovery18%Locate difficult, relevant, authoritative information through multi-step research.
Evidence Quality16%Use authoritative, diverse evidence with preference for primary sources.
Citation Integrity16%Attach citations that genuinely support the associated claims.
Factual Precision16%Accurate facts, dates, numbers, entities and relationships; unsupported claims lose credit.
Synthesis & Insight12%Combine evidence into a coherent, useful research result.
Contradiction Handling8%Identify conflicts, explain disagreements and represent uncertainty.
Completeness7%Address all material parts of the assignment.
Research Efficiency4%Research quality relative to measured tool calls, elapsed time, tokens and cost.
Freshness3%Retrieve and prioritize current evidence for time-sensitive assignments.

Private, rotating assignments

A published run must complete its entire preregistered bank: at least 24 distinct tasks, with at least two in each of 12 task families. The families cover difficult discovery, multi-hop research, primary-source retrieval, numerical reconciliation, academic research, current information, documents/PDFs, comparisons, conflicting evidence, long-form synthesis, citation verification and source-quality evaluation. Each scoring category needs at least four applicable tasks. Freshness applies to time-sensitive tasks; inapplicable criteria are declared before execution.

Questions, answer keys and hidden grading rubrics remain private. Each bank has an immutable version and content hash. Rotation preserves the category blueprint; scores from different bank versions are labelled separately and should not be treated as perfectly interchangeable.

Measured checks and independent judgment

Deterministic graders check structured answers against exact, alias-aware or tolerance-based numeric keys. Captured evidence records retain citation URLs, retrieval and publication dates, content hashes, claim-to-source links and verification outcomes. Citation support requires reviewing the actual source passage; a working URL alone earns no support credit.

Synthesis and contradiction handling require at least two identified judges with distinct provider/model versions per criterion. No single LLM controls the whole score. Judge outputs retain rubric versions and evidence references; a reviewer must confirm the artifacts and judge independence before publication. Source authority is assessed with a frozen rubric, not inferred from a model’s reputation.

Efficiency and unsupported claims

Tool calls and wall-clock execution are recorded. Tokens and API or estimated cost are retained when available, with provenance. Efficiency criteria compare research quality with preregistered resource budgets; unavailable token or cost telemetry cannot be invented. Factual and citation rubrics must explicitly include unsupported claims and contradictions as lost credit. All atomic numerator, denominator and criterion-weight measurements remain available internally for audit and supported future recalculation.

Confidence intervals

Where coverage permits, we compute a 95% stratified task-bootstrap interval using 2,000 resamples. Whole task results are resampled together within task family. The interval describes task-sampling uncertainty; it does not capture all judge bias, provider variability or future-world changes. A narrow interval is not proof of universal superiority. If resampling coverage is insufficient, no interval is shown.

Anti-gaming and retesting

Keep banks and answer keys outside public source and client bundles. Freeze tasks, evidence cutoffs, resource limits and grader versions before a run. Preserve failed attempts, raw responses and source snapshots. Judge identities and disagreements are recorded; reviewers investigate contamination or suspicious patterns. Test configurations, web/PDF tools and reasoning settings must be recorded so comparisons can be interpreted.

Retest exact model revisions after material provider updates, methodology changes or bank rotation. Never silently overwrite an earlier run. A newer reviewed result appears on the model page with its version and date; historical results remain retained. Atomic data may be regraded only where a new methodology supports those measurements, with a new recorded version. Major scoring changes require a new comparison cohort.

Current implementation status

The scoring, publication validation, private run schema and public result API are implemented. A generic external research-harness runner measures time and accepts tool/token/cost telemetry. Structured answer checks and aggregation are automated. An OpenAI web-search execution adapter and source snapshot capture are implemented but have not completed a live provider validation. Other provider/browser/PDF adapters, full source-passage verification, frozen authority/freshness/efficiency rubrics, independent judge integrations and the curated private task bank still require setup and validation before the first legitimate score can be published.

Public results API →

Speed on the leaderboard

Speed is raw measured output throughput in tokens per second, with a cited measurement method, effort setting and verification date. It is not a normalized 0–100 score. The latest verified measurement for the exact model version is shown; measurements older than 90 days are withheld. Output speed excludes the separate time-to-first-token measurement and does not represent full research completion time. Different effort, workload and serving conditions can affect comparability. Models without valid measurements display “Not measured”.