Leading open-source vision-language model
Qwen2.5 VL 72B is one of the most capable open-source vision-language models, able to analyse images, charts, documents, and screenshots with strong accuracy. It outperforms GPT-4V on several vision benchmarks.
Try Qwen2.5 VL 72BSourced from official model cards and academic papers. Higher is better.