Compact 7B vision-language model from Qwen 2.5.
Add a second model to compare.
This model hasn't been benchmarked yet.