Qwen's largest vision-language model with 235B MoE architecture.
Add a second model to compare.
This model hasn't been benchmarked yet.