Inference provider serving open models with cost-optimized endpoint configurations across multiple quantization levels.
27 text
Z.ai's frontier reasoning model for long-context agents and software engineering.
MoonshotAI's coding-focused Kimi model for end-to-end software engineering.
Open-source 1.6T-parameter MoE with 49B active, built for advanced reasoning and long-horizon agents.
Efficiency-optimized DeepSeek V4 at 284B total / 13B active for fast, high-throughput inference.
MoonshotAI's next-generation multimodal model built for long-horizon coding and multi-agent orchestration.
Zhipu AI's most capable model with a major leap in coding and long-horizon task performance.
Efficient Mixture-of-Experts Gemma 4 with only 4B active parameters per token.
Google's fourth-generation open dense model with native multimodal understanding.
Compact 9B vision-language model from the Qwen3.5 family.
Efficient vision-language MoE with linear attention and 3B active parameters.
Dense 27B vision-language model with linear attention for fast response times.
Qwen's second-strongest model — text capabilities exceeding Qwen3 235B, vision surpassing Qwen3 VL 235B.
MiniMax's most capable model with advanced reasoning and long context.
Zhipu AI's flagship fifth-generation model.
StepFun's fast and efficient language model.
MoonshotAI's native multimodal model with state-of-the-art visual coding.
Open-source model with GPT-5 class reasoning performance.
Efficient 30B MoE vision-language model with 3B active parameters.
Experimental DeepSeek V3.2 variant with DeepSeek Sparse Attention.
Large hybrid reasoning model supporting thinking and non-thinking modes.
OpenAI's budget-friendly reasoning model — fast and surprisingly capable.
OpenAI's compact open-weights model for efficient inference.
Efficient 30B MoE coding model with only 3B active parameters.
Lightweight GLM 4.5 variant for fast, cost-efficient inference.
Updated DeepSeek R1 with performance on par with OpenAI o1.
Dense 32.8B parameter model optimized for reasoning and dialogue.
March 2024 iteration of the DeepSeek V3 685B MoE model.