Efficient 30B MoE vision-language model with 3B active parameters.
Add a second model to compare.
This model hasn't been benchmarked yet.