THE EVIDENCE / MULTIMODAL
Multimodal
benchmarks.
Understanding or producing information across modalities. Results are grouped by benchmark and unit. Different tests are never merged into a raw leaderboard.
01 / RESULTS
Published results
| Benchmark | Version | Method | Configuration | Model | Result | Test date | Verified | Source type | Source |
|---|---|---|---|---|---|---|---|---|---|
| Chartography (no tools; reported best effort) | Chartography (no tools; reported best effort) | Publisher-reported configuration; see source | — | claude-sonnet-5-5 | 61.6 % | 2026-09-28 | 2026-09-29 | provider first party | Source ↗ |
| Chartography (no tools; reported best effort) | Chartography (no tools; reported best effort) | Publisher-reported configuration; see source | — | claude-opus-5-5 | 64.4 % | 2026-09-28 | 2026-09-29 | provider first party | Source ↗ |
| Chartography (no tools; reported best effort) | Chartography (no tools; reported best effort) | Publisher-reported configuration; see source | — | gpt-6-sol | 53.6 % | 2026-09-28 | 2026-09-29 | provider first party | Source ↗ |
Limitations: Image, audio and video tests are distinct and are not pooled as raw results. Publication methods, prompts, sampling and model versions may differ. Check each cited source before interpreting a result.