Cloud inference provider surfaced as a serving endpoint for hosted model routes.
35 text
Z.ai's frontier reasoning model for long-context agents and software engineering.
MoonshotAI's coding-focused Kimi model for end-to-end software engineering.
MiniMax's multimodal foundation model for long-horizon agentic work.
Open-source 1.6T-parameter MoE with 49B active, built for advanced reasoning and long-horizon agents.
Efficiency-optimized DeepSeek V4 at 284B total / 13B active for fast, high-throughput inference.
MoonshotAI's next-generation multimodal model built for long-horizon coding and multi-agent orchestration.
Zhipu AI's most capable model with a major leap in coding and long-horizon task performance.
Zhipu AI's fast-inference GLM 5 variant optimized for agent workflows.
Efficient vision-language MoE with linear attention and 3B active parameters.
Dense 27B vision-language model with linear attention for fast response times.
Qwen's second-strongest model — text capabilities exceeding Qwen3 235B, vision surpassing Qwen3 VL 235B.
Qwen's largest MoE model with 397B total parameters, 17B active.
MiniMax's most capable model with advanced reasoning and long context.
Zhipu AI's flagship fifth-generation model.
Efficient coding agent with 80B parameters, only 3B active per token.
MoonshotAI's native multimodal model with state-of-the-art visual coding.
Updated MiniMax M2 model with enhanced capabilities.
Zhipu AI's advanced model with strong reasoning and multimodal support.
Open-source model with GPT-5 class reasoning performance.
Kimi K2 with always-on extended reasoning for complex problem solving.
MiniMax's flagship MoE model with 1T parameters.
Efficient 30B MoE vision-language model with 3B active parameters.
Zhipu AI's GLM 4.6 generation model.
Experimental DeepSeek V3.2 variant with DeepSeek Sparse Attention.
Qwen's largest vision-language model with 235B MoE architecture.
Next-generation 80B MoE model with always-on reasoning mode.
September 2025 checkpoint of Kimi K2 with improved capabilities.
July 2025 Qwen3 30B MoE model with always-on extended thinking.
Large hybrid reasoning model supporting thinking and non-thinking modes.
Qwen's primary code generation model for software development workflows.
July 2025 update to Qwen3 235B A22B with improved capabilities.
Updated DeepSeek R1 with performance on par with OpenAI o1.
Compact Qwen3 model for fast, efficient inference.
Dense 32.8B parameter model optimized for reasoning and dialogue.
March 2024 iteration of the DeepSeek V3 685B MoE model.