GPT-4o (“o” stands for “omni”) is a new-generation multimodal model designed for more natural human–computer interaction. It can accept any combination of text, audio, image, and video as input, and generate multimodal outputs including text, audio, and images. With audio response latency as low as 232 milliseconds on average around 320 milliseconds, it approaches real human conversational speed. The model delivers strong performance in English text and code, significantly improved multilingual understanding, and outstanding capabilities in visual and audio perception, while offering faster API performance and substantially reduced cost for real-time and complex multimodal applications.
Pricing
Input Modalities
- Text
- Vision
Output Modalities
- Text
Capabilities
- Tools
- Tool calling
- Structured outputs
Providers
Azure gpt-4o
Pricing$2.500$10.000
Cache Read$1.25/M tokens
Web Search$0.025/request
Context128K
Max output16K
Latency1.6S
Throughput34.1TPS
Uptime
99.97% uptime 3 days ago
99.93% uptime 2 days ago
99.95% uptime yesterday
OpenAI gpt-4o
Pricing$2.500$10.000
Cache Read$1.25/M tokens
Web Search$0.025/request
Context128K
Max output16K
Latency0.6S
Throughput52.2TPS
Uptime
0.00% uptime 3 days ago
0.00% uptime 2 days ago
0.00% uptime yesterday
Performance for gpt-4o
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
