Skip to main content
idapt
HomeCodeAI ModelsPricing
Sign inStart free trial
  • Home
  • Pricing
  • AI Models
  • Image models
  • Voice models
  • Video models
  • Rankings
  • New models
  • Model status
  • Multi-Model Chat
  • Agents
  • Computers
  • Drive
  • Automations
  • AI Gateway
  • All features →
  • LLM cost calculator
  • Token counter
  • All free tools →
  • Blog
  • Use cases
  • Comparisons
  • Best of
  • Benchmarks
  • Demos
  • Changelog
  • Help center
  • FAQ
  • Privacy
  • Compare all models
  • Support
  • idapt Code
  • Developers
  • Quickstarts
  • API reference
  • API pricing
  • CLI
  • MCP
  • Downloads
  • Desktop
  • Badges and embeds
© idapt[email protected]TermsPrivacy PolicyLegal noticeReport content
X (Twitter)

Claude Sonnet 5 vs Grok Build 0.1 vs Kimi K3

Models that run tools well: the benchmark band leads with Terminal-Bench and the agent-loop demos show real multi-step work where captures exist.Last reviewed 2026-07-17.

Claude Sonnet 5 vs Grok Build 0.1 vs Kimi K3

At a glance

Anthropic's fifth-generation Sonnet — near-Opus intelligence at Sonnet speed and price.

801M ctx$2.00/$10.00 per M#26/92

xAI's agentic build model for autonomous software engineering.

256K ctx$1.00/$2.00 per M

MoonshotAI's 2.8T-parameter flagship for complex coding, knowledge work, and long-horizon agent runs.

1.0M ctx$3.00/$15.00 per M
Demos
All demos

No shared demos for these models yet.

AI Output
All AI outputs

No captured outputs for these models yet.

Benchmarks
See rankings
The paired significance test (McNemar, in "Measured on Idapt") appears only when exactly two models are pinned.
Reasoning
CI 72.7-84.7Provisional
CodingInsufficient data
AgenticInsufficient data
Sources:Epoch ECI·Epoch AI·OpenRouter· as of 2026-07-10
Capability
CapabilityECI80.1
Reasoning & Knowledge
Graduate Science90.5%Expert-Level—Factual Recall25.0%
Math
Competition Math94.7%
Coding
Real-World Coding—
Agentic
Terminal Tasks—
Reasoning
AgenticInsufficient data
ReasoningInsufficient data
MathInsufficient data
Sources:Provider·OpenRouter· as of 2026-06-15
Cheaper69%
Capability
Capability—
Reasoning & Knowledge
Graduate Science—Expert-Level—Factual Recall—
Math
Competition Math—
Coding
Real-World Coding70.8%
Agentic
Terminal Tasks—
Reasoning
CodingInsufficient data
MathInsufficient data
Sources:Provider·OpenRouter· as of 2026-07-16
Capability
Capability—
Reasoning & Knowledge
Graduate Science3%93.5%Expert-Level43.5%Factual Recall—
Math
Competition Math—
Coding
Real-World Coding—
Agentic
Terminal Tasks88.3%
Specs
1M
Context
128K
Max output
Jun 2026
Released
256K
Context
16K
Max output
May 2026
Released
1.0M
Context
16K
Max output
Jul 2026
Released
Capabilities
Vision
Audio
Reasoning
Vision
Audio
Reasoning
Vision
Audio
Reasoning
Pricing
Live pricing, shown before you run
Input$2.00/M tokens
Output$10.00/M tokens
Cache read$0.20/M tokens
Cache write$2.50/M tokens
Web search$0.01/search
Live pricing, shown before you run
Input50%$1.00/M tokens
Output80%$2.00/M tokens
Cache read$0.20/M tokens
Web search50%$0.005/search
Live pricing, shown before you run
Input$3.00/M tokens
Output$15.00/M tokens
Cache read$0.30/M tokens
Chat with Claude Sonnet 5Go to model
Chat with Grok Build 0.1Go to model
Chat with Kimi K3Go to model

Frequently asked

Which of Claude Sonnet 5, Grok Build 0.1, and Kimi K3 is the cheapest?

Grok Build 0.1 has the lowest blended list price. Per 1M tokens: Claude Sonnet 5 at $2.00 in / $10.00 out; Grok Build 0.1 at $1.00 in / $2.00 out; Kimi K3 at $3.00 in / $15.00 out.

Which of Claude Sonnet 5, Grok Build 0.1, and Kimi K3 has the largest context window?

Kimi K3 leads with 1M tokens. The other two: Claude Sonnet 5 at 1M, Grok Build 0.1 at 256K.

Do Claude Sonnet 5, Grok Build 0.1, and Kimi K3 support image input?

Yes, all three accept images.

Where can I run Claude Sonnet 5, Grok Build 0.1, and Kimi K3 side by side?

This idapt comparison runs all three in aligned columns: the same prompt, plus benchmarks, specs, and prices per model. Open any model in chat from its column to keep working with it.

More three-way comparisons

GPT 5.6 Sol vs Claude Opus 4.8 vs Gemini 3.1 ProClaude Fable 5 vs GPT 5.6 Sol vs Grok 4.5GPT 5.6 Luna vs Gemini 3.5 Flash vs Claude Haiku 4.5Claude Sonnet 5 vs Kimi K2.7 Code vs Qwen3 Coder NextDeepSeek V4 Pro vs Qwen3.7 Max vs Nemotron 3 UltraMiniMax M3 vs DeepSeek V4 Flash vs Gemini 3.1 Flash LiteMuse Spark 1.1 vs Kimi K3 vs Grok 4.5DeepSeek V4 Flash vs Gemini 3.5 Flash vs Qwen3.7 PlusGLM 5.2 vs MiniMax M3 vs DeepSeek V4 Pro