How we score models
Found this helpful? Share it:
Found this helpful? Share it:
Every model page shows three kinds of numbers: a single capability score, individual benchmark results, and the real speed, reliability, and cost we measure. Here is where each one comes from, so you can trust what you read and pick the right model with confidence 📊
Capability is a single score from 0 to 100 that answers one question: how capable is this model overall. It comes from the Epoch Capabilities Index, published by Epoch AI, an independent AI research group. Epoch builds it from more than 50 public benchmarks, so no single test can skew it, and idapt rescales that number onto a friendly 0 to 100 range.
We show a capability score only when there is real data behind it. When a model is too new or too niche for a reliable number, we leave capability blank instead of guessing, and that model simply does not appear on the capability leaderboard. Some scores show a small range or a "provisional" note when the underlying data is still thin.
Capability is what Auto uses to pick a strong model for you, and what the capability leaderboard ranks by.
Below capability, each model page lists individual public benchmarks by their real name, grouped by skill:
Coding: SWE-bench Verified, Aider Polyglot, LiveCodeBench, and SciCode.
Math: MATH-500, AIME, and FrontierMath.
Reasoning and knowledge: MMLU-Pro, GPQA Diamond, HLE (Humanity's Last Exam), and SimpleQA.
Agentic: Terminal-Bench, ARC-AGI-2, and the METR autonomy time horizon.
These are the published results from each benchmark, shown as they are. We do not blend them into an invented index of our own: a coding leaderboard ranks by a real coding benchmark, never by a made-up composite. Each score links out to the benchmark it came from so you can check it yourself.
Speed and reliability are not benchmarks, they are measurements. idapt records how every model actually performs on our own traffic and reports:
Output speed in tokens per second, and time to first token.
Uptime and success rate.
How often a cached response saves you money.
The real cost per run.
These numbers reflect how idapt's users actually use each model, not a controlled lab test. Treat them as a real-world signal, not a fixed benchmark: a model that mostly gets short prompts looks different from one used for long documents. We compute them over recent windows (for example, the last 7 days) so they stay current.
Open any score's source popover and you see exactly where it came from and the date we checked it. Values we positioned by hand are labeled "Est.", so an estimate never hides among measured results. Cost always tracks live provider pricing, so the cost band you see matches what you would actually pay.
idapt only publishes data we are allowed to share. Capability is the open Epoch index, the benchmark scores are public results, and the speed and cost numbers are our own measurements. We do not use closed or proprietary scoring services whose terms forbid republishing their numbers, so nothing on a model page is a black box you cannot trace back to its source.
Related articles
Was this helpful?