Inference infrastructure provider serving open and commercial models with a focus on high-throughput GPU-backed endpoints.
55 text
Z.ai's frontier reasoning model for long-context agents and software engineering.
MoonshotAI's coding-focused Kimi model for end-to-end software engineering.
NVIDIA's open frontier reasoning model for orchestration, coding, and deep research.
StepFun's efficient multimodal MoE model for high-throughput analysis and agents.
Open-source 1.6T-parameter MoE with 49B active, built for advanced reasoning and long-horizon agents.
Efficiency-optimized DeepSeek V4 at 284B total / 13B active for fast, high-throughput inference.
Xiaomi's flagship model for general agentic capabilities, complex software engineering, and long-horizon tasks.
MoonshotAI's next-generation multimodal model built for long-horizon coding and multi-agent orchestration.
Zhipu AI's most capable model with a major leap in coding and long-horizon task performance.
Efficient Mixture-of-Experts Gemma 4 with only 4B active parameters per token.
Google's fourth-generation open dense model with native multimodal understanding.
MiniMax's most capable model, built for autonomous real-world productivity with subagent collaboration.
Compact 9B vision-language model from the Qwen3.5 family.
Efficient vision-language MoE with linear attention and 3B active parameters.
Dense 27B vision-language model with linear attention for fast response times.
Qwen's largest MoE model with 397B total parameters, 17B active.
MiniMax's most capable model with advanced reasoning and long context.
Zhipu AI's flagship fifth-generation model.
StepFun's fast and efficient language model.
MoonshotAI's native multimodal model with state-of-the-art visual coding.
Fast variant of GLM 4.7 for low-latency applications.
Zhipu AI's advanced model with strong reasoning and multimodal support.
NVIDIA's efficient 30B MoE Nemotron model with 3B active parameters.
Open-source model with GPT-5 class reasoning performance.
NVIDIA's mid-size Nemotron model based on Llama 3.3 at 49B scale.
Efficient 30B MoE vision-language model with 3B active parameters.
Zhipu AI's GLM 4.6 generation model.
Qwen's largest vision-language model with 235B MoE architecture.
Large hybrid reasoning model supporting thinking and non-thinking modes.
OpenAI's budget-friendly reasoning model — fast and surprisingly capable.
OpenAI's compact open-weights model for efficient inference.
Qwen's primary code generation model for software development workflows.
July 2025 update to Qwen3 235B A22B with improved capabilities.
Latest Mistral Small — efficient 24B model for fast, cost-effective tasks.
Updated DeepSeek R1 with performance on par with OpenAI o1.
Efficient 30B MoE language model with 3B active parameters per token.
Mid-size dense Qwen3 model balancing capability and efficiency.
Dense 32.8B parameter model optimized for reasoning and dialogue.
Meta's high-capacity multimodal MoE model with 128 experts.
Meta's efficient multimodal MoE model with 16 experts.
March 2024 iteration of the DeepSeek V3 685B MoE model.
Compact open model for efficient on-device or edge deployment.
Mid-size open multimodal model balancing capability and efficiency.
Google's open multimodal model with 128K context and 140+ language support.
January 2025 Mistral Small 24B instruction-tuned model.
Microsoft's compact research model with strong STEM reasoning.
Original DeepSeek V3 — the model that put DeepSeek on the map.
Meta's efficient 70B model matching Llama 3.1 405B on key benchmarks.
Compact 11B vision model from Meta for multimodal tasks.
Large 72B model from the Qwen 2.5 generation.
NousResearch's 70B model based on Llama 3.1.
NousResearch's large model based on Llama 3.1 405B.
Meta's 70B model from the Llama 3.1 generation.
Compact 8B model from Meta's Llama 3.1 series.
12B model from Mistral and NVIDIA — efficient and multilingual.