Frontier open-source vision-language model
Qwen3 VL 235B is Alibaba's most capable vision-language model, combining the 235B MoE architecture with advanced multimodal understanding. It sets new records for open-weight models on visual reasoning, document analysis, and image comprehension benchmarks.
Try Qwen3 VL 235BSourced from official model cards and academic papers. Higher is better.