GPT-6.1 Sol has five reasoning effort levels: low, medium (the default), high, xhigh, and max. The big change from GPT-6 Sol is that none and minimal are gone, so low is now the floor.
This one setting drives both latency and cost. You never see reasoning tokens, but you pay for them at the output rate ($10 per million for 6.1 Sol on AIHubMix), and they take up room in the context window. Pick the wrong level and you can easily pay several times more than you need to.
What each level is for
This table combines OpenAI's descriptions from the reasoning guide with our own recommendations:
| Level | OpenAI's description | Good fit | Poor fit |
|---|---|---|---|
| low | Efficient reasoning with a modest latency increase | Intermediate steps in tool loops, search, light planning, classification, extraction, support replies | Hard debugging, multi-file refactors |
| medium (default) | The balanced default for most workloads | Everyday coding, code review, writing, typical agent tasks | Latency-critical interactive use |
| high | Hard reasoning, complex debugging, deep planning | Tricky bugs, architecture work, SWE-style tasks | High-QPS online traffic |
| xhigh | Deep research, async work, long agent runs | Background jobs, research reports | Anything your evals haven't justified |
| max | Maximum reasoning for the hardest tasks | Computer use, research-grade problems, offline heavy lifting | Almost all day-to-day requests |
OpenAI calls effort "a tuning knob, not the primary way to recover quality." When results are bad, look at the prompt, the tool definitions, and the context first. Turn up the effort only after that.
What the benchmarks say about each level
These are OpenAI's own numbers, compiled by Vellum and DataCamp. Treat them as a guide, not gospel.
For coding, high is usually enough. On DeepSWE v1.1, 6.1 Sol at high scores 75.2, matching Astra at high (74.8), for about $1.50 per task. Its cost curve reaches 72% to 75% for between $0.50 and $1.50 per task. GPT-6 Sol maxed out at 68.8 for about $2.60.
Going above medium doesn't always help. On AutomationBench, 6.1 Sol scores 35.4% at medium and only about 36.0% at higher effort. For business automation, the extra reasoning tokens bought almost nothing.
Save max for computer use and research. On OSWorld 2.0, max gets 71.4 at around $1.30 per task. Terminal-Bench Science at max runs $5.47 per task: far cheaper than Astra's $23.80, but an order of magnitude above most other workloads.
Low got better too. The factual error rate at low dropped from 11.4% on 6 Sol to 7.7%. If you bumped tasks up to medium on 6 Sol because low wasn't accurate enough, try low again.
Starting points by use case
| Use case | Start at | Notes |
|---|---|---|
| Calls that used none on 6 Sol | low | OpenAI's official mapping. If latency is critical, consider GPT-6 Luna, which still supports none |
| Calls that used minimal | low | Start at low and compare results |
| Chatbots, customer support | low | Ask the model for a one-line preamble to get the first token out sooner |
| RAG Q&A | low → medium | Retrieval quality matters more than effort |
| IDE coding assistant | medium | The default works |
| Automated bug fixing, SWE agents | high | DeepSWE shows high already matches Astra |
| Computer use, browser agents | high → max | OSWorld peaks at max |
| Background research, long runs | xhigh | Only when evals show a gain. Batch halves the cost if you don't need results right away |
| The hardest research problems | max, or just use Astra | OpenAI recommends Astra here too |
The none and minimal mappings come from OpenAI's GPT-6 migration guide.
Changing effort mid-conversation
A common pattern is to plan at high, execute at low, and switch back to high when something breaks.
The catch: editing reasoning.effort on the request breaks your prompt cache. Effort is part of the cached prefix (see the prompt caching guide). On 6.1 Sol a cache read costs $0.10 per million tokens, while rewriting the cache costs $2.50. That's a 25× difference.
The GPT-6 family fixes this with a configuration_update input item. Leave the request-level reasoning.effort alone and insert an update before the next user message. Here's what that looks like through the AIHubMix Responses API:
from openai import OpenAI
import os
client = OpenAI(
api_key=os.environ["AIHUBMIX_API_KEY"],
base_url="https://aihubmix.com/v1",
)
# Turn 1: plan at high effort
r1 = client.responses.create(
model="gpt-6.1-sol",
reasoning={"effort": "high"},
input="Figure out why this repo's tests are failing and propose a fix plan.",
)
# Turn 2: drop to low for execution, without touching request-level effort
r2 = client.responses.create(
model="gpt-6.1-sol",
reasoning={"effort": "high"}, # unchanged, so the cache survives
previous_response_id=r1.id,
input=[
{"type": "configuration_update", "reasoning": {"effort": "low"}},
{"role": "user", "content": "Carry out step 1 of the plan."},
],
)
Theconfiguration_updateshape above is illustrative. Check OpenAI's reasoning guide for the exact schema. AIHubMix's Responses docs currently list four levels (minimal through high), so send a small test request to confirm thatxhigh,max, andconfiguration_updatepass through.
Some limits to know:
- It only works in standard single-agent mode, and it only changes effort.
- Two updates can't sit next to each other.
- It doesn't mix with automatic compaction or truncation. Explicit compaction is fine, but add a fresh update afterwards.
- The response's
reasoning.effortfield still shows the request-level value, so don't read too much into it.
Settings that go with effort
Leave room in max_output_tokens. The cap includes reasoning tokens. Set it too low and the model can stop before producing any visible text. You get status: incomplete and still pay for input and reasoning. OpenAI suggests reserving at least 25,000 tokens to start, and more for xhigh or max.
Track reasoning tokens. They're in usage.output_tokens_details.reasoning_tokens. Looking at the distribution per effort level beats tuning by feel.
Turn on summaries if you need visibility. reasoning.summary: "auto" returns a summary of the model's reasoning (you may need to verify your organization first). The raw reasoning is never exposed.
Pro mode is a separate switch. It makes the model do more work, billed at standard rates but with more tokens overall. Use it for hard problems where extra latency is acceptable.
A tuning process you can actually run
- Pull 20 to 50 real tasks into an eval set.
- Run a baseline at
medium. Record success rate, average reasoning tokens, and P95 latency. - Try
low. If success barely drops, switch. - Try
high. If success clearly improves, use high for that task type only. - Only reach for
xhighormaxwhen high falls short, and compare against GPT-6 Astra at high while you're at it. Sometimes a bigger model beats more effort. On AIHubMix that's just a change to themodelparameter, and the model list shows what's available. - Route by task type instead of using one global setting.
The short version: start at medium, save money with low, use high for coding, and make xhigh and max prove themselves in your evals.
Next up: what the $2 / $10 price tag really means once caching, long context, and billing multipliers come into play.
FAQ
What's the default reasoning effort for GPT-6.1 Sol? medium. If you don't set it, that's what you get.
Why does none return an error? 6.1 Sol doesn't support none or minimal. OpenAI recommends switching to low. If you really need none, stay on GPT-6 Sol or GPT-6 Luna.
How are reasoning tokens billed? At the output rate ($10 per million for 6.1 Sol), and they count against the context window. Check usage.output_tokens_details.reasoning_tokens for the actual numbers.
Is the parameter name the same in Chat Completions and Responses? No. Chat Completions uses a top-level reasoning_effort. Responses uses a nested reasoning: {"effort": ...}. Mixing them up returns Unsupported parameter.
Does higher effort always give better results? No. On AutomationBench, going above medium only moved the score from 35.4% to about 36.0%, while costing noticeably more. Test on your own tasks.
Can I change effort partway through a conversation? Yes. Use a configuration_update item and leave request-level reasoning.effort unchanged so the prompt cache stays valid. It doesn't work with automatic compaction or truncation.
Why did I get status: incomplete with no output? Most likely max_output_tokens was too low and reasoning used it all up. OpenAI recommends reserving at least 25,000 tokens.
Keep reading: the GPT-6.1 Sol series
- How much of your bill is reasoning? Effort is only one lever. Caching, the 272K threshold, and Batch discounts all shape the final number. See What GPT-6.1 Sol Really Costs: Beyond the $2 / $10 Price Tag.
- Code that used
none? Switching tolowis just the start. Tool calls, sampling parameters, and cache settings need changes too: Migrating to GPT-6.1 Sol: 9 Things That Can Go Wrong. - How much better is 6.1 Sol than 6 Sol and Astra? Specs and benchmarks side by side in GPT-6.1 Sol vs GPT-6 Sol: A One-Week Upgrade That Nearly Catches Astra.



