Inference provider built on custom LPU hardware, known for very high token throughput on open-weight models.
7 text
September 2025 checkpoint of Kimi K2 with improved capabilities.
OpenAI's budget-friendly reasoning model — fast and surprisingly capable.
OpenAI's compact open-weights model for efficient inference.
Dense 32.8B parameter model optimized for reasoning and dialogue.
Meta's efficient multimodal MoE model with 16 experts.
Meta's efficient 70B model matching Llama 3.1 405B on key benchmarks.
Compact 8B model from Meta's Llama 3.1 series.