Multimodal flagship — text, vision, and audio
GPT-4o ('o' for omni) is OpenAI's multimodal model that processes text, images, and audio natively. It matches GPT-4 Turbo on text tasks while being twice as fast and half the price.
Try GPT-4oSourced from official model cards and academic papers. Higher is better.