Compact multimodal model for visual tasks
Llama 3.2 11B Vision adds image understanding to a compact, efficient model. It is ideal for document analysis, image captioning, and visual question answering at a very low cost.
Try Llama 3.2 11B VisionSourced from official model cards and academic papers. Higher is better.