InclusionAI's hybrid vision-language model, layering native visual perception on the Ling 3.0 Flash MoE (124B total, 5.5B active).
Add a second model to compare.