Skip to main content
idapt
HomeCodeAI ModelsPricing
Sign inStart free trial
  • Home
  • Pricing
  • AI Models
  • Image models
  • Voice models
  • Video models
  • Rankings
  • New models
  • Model status
  • Multi-Model Chat
  • Agents
  • Computers
  • Drive
  • Automations
  • AI Gateway
  • All features →
  • LLM cost calculator
  • Token counter
  • All free tools →
  • Blog
  • Use cases
  • Comparisons
  • Best of
  • Benchmarks
  • Changelog
  • Help center
  • FAQ
  • Privacy
  • Compare all models
  • Support
  • idapt Code
  • Developers
  • Quickstarts
  • API reference
  • API pricing
  • CLI
  • MCP
  • Downloads
  • Desktop
  • Badges and embeds
© idapt[email protected]TermsPrivacy PolicyLegal noticeReport content
X (Twitter)
Help Center
📊

How we score models

Found this helpful? Share it:

Every model page shows three kinds of numbers: a single capability score, individual benchmark results, and the real speed, reliability, and cost we measure. Here is where each one comes from, so you can trust what you read and pick the right model with confidence 📊

The capability score

Capability is a single score from 0 to 100 that answers one question: how capable is this model overall. It comes from the Epoch Capabilities Index, published by Epoch AI, an independent AI research group. Epoch builds it from more than 50 public benchmarks, so no single test can skew it, and idapt rescales that number onto a friendly 0 to 100 range.

We show a capability score only when there is real data behind it. When a model is too new or too niche for a reliable number, we leave capability blank instead of guessing, and that model simply does not appear on the capability leaderboard. Some scores show a small range or a "provisional" note when the underlying data is still thin.

Capability is what Auto uses to pick a strong model for you, and what the capability leaderboard ranks by.

Named benchmark scores

Below capability, each model page lists individual public benchmarks by their real name, grouped by skill:

  • Coding: SWE-bench Verified, Aider Polyglot, LiveCodeBench, and SciCode.

  • Math: MATH-500, AIME, and FrontierMath.

  • Reasoning and knowledge: MMLU-Pro, GPQA Diamond, HLE (Humanity's Last Exam), and SimpleQA.

  • Agentic: Terminal-Bench, ARC-AGI-2, and the METR autonomy time horizon.

These are the published results from each benchmark, shown as they are. We do not blend them into an invented index of our own: a coding leaderboard ranks by a real coding benchmark, never by a made-up composite. Each score links out to the benchmark it came from so you can check it yourself.

Speed, reliability, and real cost

Speed and reliability are not benchmarks, they are measurements. idapt records how every model actually performs on our own traffic and reports:

  • Output speed in tokens per second, and time to first token.

  • Uptime and success rate.

  • How often a cached response saves you money.

  • The real cost per run.

These numbers reflect how idapt's users actually use each model, not a controlled lab test. Treat them as a real-world signal, not a fixed benchmark: a model that mostly gets short prompts looks different from one used for long documents. We compute them over recent windows (for example, the last 7 days) so they stay current.

Every number shows its source

Open any score's source popover and you see exactly where it came from and the date we checked it. Values we positioned by hand are labeled "Est.", so an estimate never hides among measured results. Cost always tracks live provider pricing, so the cost band you see matches what you would actually pay.

What we do not use

idapt only publishes data we are allowed to share. Capability is the open Epoch index, the benchmark scores are public results, and the speed and cost numbers are our own measurements. We do not use closed or proprietary scoring services whose terms forbid republishing their numbers, so nothing on a model page is a black box you cannot trace back to its source.

FAQs

Related articles

🧠

Choosing models

Browse 200+ AI models, read their cost bands, set the speed dial, and pick the right model for each conversation.

🔀

Comparing models side-by-side

Get multiple AI responses to the same message and compare them in tabs.

Up next

Goal mode

Set a goal and the agent keeps working toward it, turn after turn, until it is met or it needs you.

Was this helpful?