GPT-5.4 supports configuring reasoning strength only through the /responses endpoint. To make lower-overhead reasoning available directly in the /chat endpoint, the GPT-5.4-Low model is provided. This model is based on GPT-5.4 with reasoning_effort preset to low. This model is designed for use cases that are sensitive to response latency and cost. By adopting a lighter reasoning strategy, it delivers stable responses with lower latency and higher throughput. It is well suited for high-concurrency conversations, real-time interactions, basic Q&A, and scenarios where deep reasoning is not required.
Pricing
Input Modalities
- Text
- Vision
Output Modalities
- Text
Capabilities
- Thinking
- Web
- Tools
- Tool calling
- Structured outputs
Try this model
Python
