32B dense vision-language model for high-quality multimodal tasks.
Add a second model to compare.
This model hasn't been benchmarked yet.