GPT-4o (“o” stands for “omni”) is a new-generation multimodal model designed for more natural human–computer interaction. It can accept any combination of text, audio, image, and video as input, and generate multimodal outputs including text, audio, and images. With audio response latency as low as 232 milliseconds on average around 320 milliseconds, it approaches real human conversational speed. The model delivers strong performance in English text and code, significantly improved multilingual understanding, and outstanding capabilities in visual and audio perception, while offering faster API performance and substantially reduced cost for real-time and complex multimodal applications.
← Models
GPT 4o
GPT 4o Compare
Compare pricing, specifications, performance, and benchmarks for up to four models.
+ Add model · 1/4
GPT 4o
OpenAI · text, image → text
Input$2.50 /M
Output$10.00 /M
Pick a second model to start comparing.
Popular comparisons
Related model match-ups readers also look at.
