Compact 11B vision model from Meta for multimodal tasks.
Add a second model to compare.
This model hasn't been benchmarked yet.