AI inference provider for text and multimodal models, commonly exposed through OpenRouter endpoints and direct API integrations.
57 text
Z.ai's frontier reasoning model for long-context agents and software engineering.
MoonshotAI's coding-focused Kimi model for end-to-end software engineering.
MiniMax's multimodal foundation model for long-horizon agentic work.
StepFun's efficient multimodal MoE model for high-throughput analysis and agents.
Open-source 1.6T-parameter MoE with 49B active, built for advanced reasoning and long-horizon agents.
Efficiency-optimized DeepSeek V4 at 284B total / 13B active for fast, high-throughput inference.
Xiaomi's flagship model for general agentic capabilities, complex software engineering, and long-horizon tasks.
MoonshotAI's next-generation multimodal model built for long-horizon coding and multi-agent orchestration.
Zhipu AI's most capable model with a major leap in coding and long-horizon task performance.
Efficient Mixture-of-Experts Gemma 4 with only 4B active parameters per token.
Google's fourth-generation open dense model with native multimodal understanding.
MiniMax's most capable model, built for autonomous real-world productivity with subagent collaboration.
Dense 27B vision-language model with linear attention for fast response times.
Qwen's second-strongest model — text capabilities exceeding Qwen3 235B, vision surpassing Qwen3 VL 235B.
Qwen's largest MoE model with 397B total parameters, 17B active.
MiniMax's most capable model with advanced reasoning and long context.
Zhipu AI's flagship fifth-generation model.
Efficient coding agent with 80B parameters, only 3B active per token.
MoonshotAI's native multimodal model with state-of-the-art visual coding.
Fast variant of GLM 4.7 for low-latency applications.
Updated MiniMax M2 model with enhanced capabilities.
Zhipu AI's advanced model with strong reasoning and multimodal support.
NVIDIA's efficient 30B MoE Nemotron model with 3B active parameters.
Zhipu AI's GLM 4.6 with native vision support.
Open-source model with GPT-5 class reasoning performance.
Kimi K2 with always-on extended reasoning for complex problem solving.
MiniMax's flagship MoE model with 1T parameters.
Efficient 30B MoE vision-language model with 3B active parameters.
Zhipu AI's GLM 4.6 generation model.
Experimental DeepSeek V3.2 variant with DeepSeek Sparse Attention.
Qwen's largest vision-language model with 235B MoE architecture.
Next-generation 80B MoE model with always-on reasoning mode.
September 2025 checkpoint of Kimi K2 with improved capabilities.
Large hybrid reasoning model supporting thinking and non-thinking modes.
GLM 4.5 with vision capabilities for multimodal understanding.
OpenAI's budget-friendly reasoning model — fast and surprisingly capable.
OpenAI's compact open-weights model for efficient inference.
Efficient 30B MoE coding model with only 3B active parameters.
Lightweight GLM 4.5 variant for fast, cost-efficient inference.
Qwen's primary code generation model for software development workflows.
July 2025 update to Qwen3 235B A22B with improved capabilities.
MoonshotAI's large MoE model — 1T total parameters, 32B active.
Baidu's large vision-language MoE model — 424B total, 47B active.
MiniMax's previous flagship model — strong long context and multimodal.
Updated DeepSeek R1 with performance on par with OpenAI o1.
Meta's high-capacity multimodal MoE model with 128 experts.
Meta's efficient multimodal MoE model with 16 experts.
March 2024 iteration of the DeepSeek V3 685B MoE model.
Google's open multimodal model with 128K context and 140+ language support.
Distilled reasoning model based on Llama 3.3 70B using DeepSeek R1 outputs.
Open-source reasoning model matching OpenAI o1 performance.
Original DeepSeek V3 — the model that put DeepSeek on the map.
Meta's efficient 70B model matching Llama 3.1 405B on key benchmarks.
Large 72B model from the Qwen 2.5 generation.
Compact 8B model from Meta's Llama 3.1 series.
12B model from Mistral and NVIDIA — efficient and multilingual.
Microsoft's MoE model built on Mixtral 8x22B architecture.