32B vision-language model from Qwen 2.5 with strong visual understanding.
Add a second model to compare.
This model hasn't been benchmarked yet.