Earlier Qwen vision-language model with multimodal capabilities.
Add a second model to compare.
This model hasn't been benchmarked yet.