Google's high-speed thinking model for agentic workflows, chat, and coding.
MoonshotAI's coding-focused Kimi model for end-to-end software engineering.
Anthropic's fifth-generation Sonnet — near-Opus intelligence at Sonnet speed and price.
Fast and efficient with near-frontier intelligence.
OpenAI's new flagship for demanding end-to-end work: advanced analysis, software engineering, deep research, scientific work, and document creation.
Alibaba's updated Qwen3.8 Max flagship snapshot, post-trained for coding and agentic work.
Meta's latest multimodal reasoning model for long-running agentic, multi-agent, and coding workflows.
Google's most intelligent Flash model, with significant gains over 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Anthropic's upgraded Mythos-class flagship, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work.
Z.ai's efficient multimodal model for fast coding and long-horizon agent tasks.
Z.ai's large-scale reasoning flagship for complex software engineering and long-horizon agent tasks.
The GA release of DeepSeek's open 1.6T-parameter MoE flagship for advanced reasoning and long-horizon agents.
xAI's smartest model, with frontier performance on coding, knowledge work, and STEM.
Anthropic's flagship for demanding reasoning, coding, and long-horizon agentic work.
Google's high-efficiency model for focused subagent work and multi-agent workflows.
MoonshotAI's 2.8T-parameter flagship for complex coding, knowledge work, and long-horizon agent runs.
OpenAI's fastest, most affordable GPT-5.6 tier for high-volume, latency-sensitive work.
NVIDIA's open frontier reasoning model for orchestration, coding, and deep research.
Alibaba's cost-efficient Qwen3.7 model for long-context multimodal work.
MiniMax's multimodal foundation model for long-horizon agentic work.
xAI's agentic build model for autonomous software engineering.
Efficiency-optimized DeepSeek V4 at 284B total / 13B active for fast, high-throughput inference.
Google's fourth-generation open dense model with native multimodal understanding.
Efficient coding agent with 80B parameters, only 3B active per token.
Fast agentic model with 2M context window and tool calling.
OpenAI's high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads.
OpenAI's most advanced agentic coding model with 400K context.
Google's frontier model with a large leap in core reasoning.
xAI's newest flagship model with industry-leading speed and agentic tool calling.
OpenAI's most capable model with adaptive reasoning for complex tasks.
Anthropic's previous-generation Sonnet, with strong performance across coding, agents, and professional work.
Google's flagship model for high-precision multimodal reasoning.
Alibaba's flagship Qwen3.7 model, built for agent-centric workloads.
Z.ai's frontier reasoning model for long-context agents and software engineering.
MoonshotAI's next-generation multimodal model built for long-horizon coding and multi-agent orchestration.
High-speed thinking model for agentic workflows and coding.
Zhipu AI's most capable model with a major leap in coding and long-horizon task performance.
OpenAI's powerful general-purpose frontier model.
Open-source 1.6T-parameter MoE with 49B active, built for advanced reasoning and long-horizon agents.
Qwen's latest hybrid model combining linear attention with sparse MoE routing.
xAI's reasoning model built for agentic workflows and high-accuracy instruction-following.
GPT-5.4 distilled into a fast, cost-efficient model for high-throughput workloads.
MoonshotAI's native multimodal model with state-of-the-art visual coding.
MiniMax's most capable model with advanced reasoning and long context.
OpenAI's most powerful reasoning model — top across math, coding, and science.
Ultra-lightweight GPT-5.4 variant built for speed-critical, high-volume tasks.
Zhipu AI's flagship fifth-generation model.
o4-mini with reasoning effort locked to high for maximum accuracy.
Compact GPT-5 variant balancing speed and intelligence.
MiniMax's most capable model, built for autonomous real-world productivity with subagent collaboration.
Experimental DeepSeek V3.2 variant with DeepSeek Sparse Attention.
Google's high-efficiency model for high-volume use cases at half the cost of Gemini 3 Flash.
Flagship reasoning model for complex multi-step tasks.
Zhipu AI's advanced model with strong reasoning and multimodal support.
Updated DeepSeek R1 with performance on par with OpenAI o1.
Ultra-fast and lightweight GPT-5 variant for low-latency tasks.
OpenAI's budget-friendly reasoning model — fast and surprisingly capable.
Fast workhorse model with built-in thinking and audio support.
Qwen's flagship 235B MoE model with 22B active parameters.
OpenAI's flagship GPT-4.1 with advanced instruction following and long context.
Distilled reasoning model based on Llama 3.3 70B using DeepSeek R1 outputs.
Mid-sized model competitive with GPT-4o at lower latency and cost.
Qwen's top commercial model for the most demanding tasks.
Microsoft's compact research model with strong STEM reasoning.
Fastest and cheapest model in the GPT-4.1 series for low-latency applications.
Qwen's mid-tier commercial model for balanced performance.
Large 72B vision-language model from Qwen 2.5 generation.
Meta's efficient 70B model matching Llama 3.1 405B on key benchmarks.
Fast and cost-efficient Qwen model for high-throughput applications.
Inception's latest diffusion LLM and the fastest reasoning model available, producing and refining tokens in parallel instead of one at a time.
OpenAI's GPT-6 Astra served in pro reasoning mode for the hardest problems.
IBM's compact dense reasoning model for mathematics, code generation, and multilingual dialogue.
Tencent's Hunyuan 4 preview, built for coding agents and complex tool-use workflows.
Alibaba's fast multimodal Qwen3.8 model for high-volume agentic work.
Experimental vision-enabled DeepSeek V4 Flash with image understanding added on top of the 0731 snapshot.
Tencent's most compact translation model, covering 33 language pairs plus five Chinese dialect and minority-language pairs.
Tencent's flagship translation model, specialized in 33 language pairs plus five Chinese dialect and minority-language pairs.
Compact translation model from Tencent covering 33 language pairs plus five Chinese dialect and minority-language pairs.
Open-weight dense 27B vision-language model from Qwen for long-running agent tasks.
Google's multimodal model for fast agentic workflows, coding, and complex multi-step reasoning.
ByteDance's multimodal model for coding and long-horizon agent workflows.
The open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total.
ByteDance's agentic coding model for end-to-end software work.
NVIDIA's open 30B MoE built for high-throughput agentic workloads.
Meta's open-weight 30B multimodal model, distilled from Muse Spark for autonomous agents on consumer hardware.
The GA release of DeepSeek's efficiency-optimized V4 Flash, re-post-trained for coding, reasoning, and agent workflows.
Thinking Machines Lab's smaller, more efficient open-weight multimodal model, with 12B active parameters out of 276B total.
Alibaba's ultra-affordable vision-language model for multimodal agents and visual coding.
Thinking Machines Lab's open-weight multimodal model with 41B active parameters out of 975B total.
OpenAI's balanced GPT-5.6 tier — GPT-5.5-class quality for everyday work at half the price.
OpenAI's GPT-5.6 flagship — built for complex reasoning, agentic coding, and long-running professional work.
StepFun's efficient multimodal MoE model for high-throughput analysis and agents.
Xiaomi's flagship model for general agentic capabilities, complex software engineering, and long-horizon tasks.
Xiaomi's native omnimodal model delivering Pro-level agentic performance at roughly half the inference cost.
Efficient Mixture-of-Experts Gemma 4 with only 4B active parameters per token.
Zhipu AI's first native multimodal agent foundation model.
Mistral's next-gen small model, unifying Magistral's reasoning, Pixtral's vision, and Devstral's coding into one system.
Zhipu AI's fast-inference GLM 5 variant optimized for agent workflows.
ByteDance's cost-efficient enterprise model for production-scale workloads.
Compact 9B vision-language model from the Qwen3.5 family.
Refined for natural conversations — smoother, more useful, and more directly helpful.
ByteDance's compact model for latency-sensitive, high-concurrency workloads.
Efficient vision-language MoE with linear attention and 3B active parameters.
Dense 27B vision-language model with linear attention for fast response times.
Qwen's second-strongest model — text capabilities exceeding Qwen3 235B, vision surpassing Qwen3 VL 235B.
Ultra-fast Qwen3.5 variant with 1M-token context at rock-bottom pricing.
Qwen3.5 Plus variant from February 2025 — optimized for balanced performance.
Qwen's largest MoE model with 397B total parameters, 17B active.
StepFun's fast and efficient language model.
MiniMax M2 variant optimized for helpful, engaging responses.
Fast variant of GLM 4.7 for low-latency applications.
Optimized for software engineering and coding workflows.
Fast variant of Seed 1.6 for low-latency applications.
ByteDance's flagship language model with top reasoning capabilities.
Updated MiniMax M2 model with enhanced capabilities.
NVIDIA's efficient 30B MoE Nemotron model with 3B active parameters.
Specialized coding and agentic reasoning model from Mistral and All Hands AI.
Zhipu AI's GLM 4.6 with native vision support.
Amazon's fast, cost-effective reasoning model for everyday workloads.
Mistral's most capable model with sparse MoE architecture.
Arcee AI's compact model for fast, efficient inference.
Open-source model with GPT-5 class reasoning performance.
Kimi K2 with always-on extended reasoning for complex problem solving.
Amazon's most capable multimodal model for complex reasoning tasks.
32B dense vision-language model for high-quality multimodal tasks.
Compact vision-language model with always-on reasoning for visual tasks.
NVIDIA's mid-size Nemotron model based on Llama 3.3 at 49B scale.
Efficient 30B MoE vision-language model with 3B active parameters.
Qwen's largest vision-language model with 235B MoE architecture.
Enhanced coding model with superior performance on complex engineering tasks.
GPT-5 optimized for code generation and software engineering tasks.
Fast variant of Qwen3 Coder for speed-prioritized coding tasks.
Next-generation 80B MoE model with always-on reasoning mode.
Qwen Plus with always-on extended thinking for July 2025.
September 2025 checkpoint of Kimi K2 with improved capabilities.
July 2025 Qwen3 30B MoE model with always-on extended thinking.
NousResearch's efficient 70B reasoning model based on Llama 4.
NousResearch's flagship large-scale reasoning model based on Llama 4.
Updated enterprise-grade model delivering frontier performance at lower cost.
GLM 4.5 with vision capabilities for multimodal understanding.
OpenAI's compact open-weights model for efficient inference.
Mistral's dedicated code model — August 2025 update.
Efficient 30B MoE coding model with only 3B active parameters.
Zhipu AI's capable mid-tier model with strong bilingual performance.
Lightweight GLM 4.5 variant for fast, cost-efficient inference.
Qwen's primary code generation model for software development workflows.
Ultra-fast and cost-efficient for high-throughput tasks.
July 2025 update to Qwen3 235B A22B with improved capabilities.
Baidu's large vision-language MoE model — 424B total, 47B active.
Latest Mistral Small — efficient 24B model for fast, cost-effective tasks.
Arcee AI's flagship model for complex enterprise tasks.
Efficient 30B MoE language model with 3B active parameters per token.
Compact Qwen3 model for fast, efficient inference.
Mid-size dense Qwen3 model balancing capability and efficiency.
Dense 32.8B parameter model optimized for reasoning and dialogue.
Compact reasoning model optimized for fast, cost-efficient STEM performance.
Meta's high-capacity multimodal MoE model with 128 experts.
Meta's efficient multimodal MoE model with 16 experts.
Cohere's most capable model for enterprise and agentic use cases.
AionLabs' flagship reasoning model for complex tasks.
AionLabs' compact reasoning model at reduced cost.
Amazon's ultra-fast text-only model for low-latency tasks.
Compact 11B vision model from Meta for multimodal tasks.
xAI's subagent variant of Grok 4.20 for collaborative deep research.
Speedy reasoning model built for agentic coding.
xAI's flagship reasoning model with parallel tool calling.
Distilled reasoning model based on Qwen 2.5 32B using DeepSeek R1 outputs.
Qwen's dedicated reasoning model at 32B scale.
32B vision-language model from Qwen 2.5 with strong visual understanding.
High-performance code generation model from Mistral and All Hands AI.
Efficient coding agent at compact scale from Mistral and All Hands AI.
Mistral Small optimized for creative writing and content generation.
Baidu's flagship ERNIE 4.5 MoE model — 300B total, 47B active.
Compact ERNIE 4.5 model with always-on thinking mode.
Compact vision-language MoE model from Baidu — 28B total, 3B active.
Xiaomi's flagship 1T-parameter model deeply optimized for agentic scenarios.
Xiaomi's frontier omni-modal model processing image, video, and audio natively.
NVIDIA's largest Nemotron model based on Llama 3.1 at 253B scale.
NVIDIA's compact 12B vision-language Nemotron model.
Arcee AI's dedicated reasoning model for complex analytical tasks.
Arcee AI's compact vision-language model for fast image understanding.
Inception AI's flagship general-purpose model.
Inception AI's code-specialized Mercury model.
TNG Tech's hybrid reasoning model combining DeepSeek R1 and T2 architectures.
Compact model from the self-improving Ornith family for agentic coding.
Flagship model from the self-improving Ornith family for agentic coding.
Mistral's small, efficient open reasoning model.
Microsoft's 14B open reasoning model that rivals much larger models.
NVIDIA's open 30B MoE model with strong reasoning and agentic ability.
Cohere's first developer model for agentic software engineering.