Moving from GPT-6 Sol to 6.1 Sol looks like a one-line change, and the price is the same. But there are a few breaking changes and some shifts in behavior, so changing only the model name can get you errors, a surprise bill, or an agent that behaves differently. OpenAI's GPT-6 migration guide covers most of the official changes. This post adds the things that tend to bite in practice.
They're ordered from "fails loudly" to "fails quietly."
1. reasoning_effort: "none" returns a 400
What you'll see: The request is rejected.
Why: 6.1 Sol doesn't support none or minimal. The lowest level is low. GPT-6 Sol and Luna still accept none, which is why your old code worked there.
Fix:
- OpenAI's mapping is to replace
nonewithlow. Forminimal, start atlowand compare. lowis slower and more expensive thannone, because it generates reasoning tokens. For truly latency-sensitive paths like autocomplete or real-time classification, staying on GPT-6 Sol or moving to Luna may be the better call.
2. Tool calling in Chat Completions stops working
What you'll see: Chat Completions requests that include tools fail.
Why: GPT-6 Sol only allowed function calling in Chat Completions when reasoning_effort was none, and plenty of projects relied on that combination for cheap tool calls. With none gone, Chat Completions on 6.1 Sol only works for requests without tools. For tools you have to use the Responses API. OpenAI's guide to migrating to the Responses API walks through it.
Fix: Move to /v1/responses. AIHubMix supports it too; see the AIHubMix Responses API docs for parameters.
from openai import OpenAI
import os
client = OpenAI(
api_key=os.environ["AIHUBMIX_API_KEY"],
base_url="https://aihubmix.com/v1",
)
resp = client.responses.create(
model="gpt-6.1-sol",
reasoning={"effort": "low"}, # nested, not reasoning_effort
tools=[{
"type": "function",
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
}],
input="What's the weather in Shanghai today?",
)
print(resp.output)
Easy things to miss:
- In Responses the parameter is
reasoning.effort. Sendingreasoning_effortgets youUnsupported parameter. - In multi-turn tool use, send back every reasoning, function_call, and function_call_output item since the last user message, not just the function results.
- With
store: falseor ZDR, reasoning items includeencrypted_contentby default. Replay the full history and it just works.
3. Sampling parameters have to go
What you'll see: Requests with temperature or top_p fail.
Why: These parameters are only allowed when effort is none. 6.1 Sol has no none, so you can never use them with this model.
Remove:
temperature,top_p,top_logprobslogprobsin Chat Completionsmessage.output_text.logprobsfromincludein Responses
If you used logprobs for confidence scores or classification thresholds, you'll need a new approach. One option is structured outputs that ask the model to report its confidence directly. Another is keeping those tasks on a model that supports none.
4. Caching works differently, and your bill can go up
This is the easiest one to overlook, especially when migrating from GPT-5.5 or earlier. Everything below comes from OpenAI's prompt caching guide.
The parameter was renamed. prompt_cache_retention is now prompt_cache_options.ttl, and the only supported value is "30m".
Cache writes are billed. They cost 1.25× the input rate ($2.50 per million on 6.1 Sol). A long prefix you use only once now costs 25% more with caching than without.
Breakpoints moved. Older models placed breakpoints at fixed intervals (every 2,048 tokens on GPT-5.5). Implicit mode now places one breakpoint at the end of the latest eligible message. As a result, a shorter prefix shared across requests isn't reused automatically. If many requests share a system prompt followed by different user input, add an explicit breakpoint right after the system prompt.
Appending to an existing message breaks the cache. The cached endpoint ends up in the middle of a longer message and can't be matched. Add a new message instead.
Changing reasoning effort breaks the cache. reasoning.effort is part of the prefix. Switch mid-conversation with configuration_update (see Part 2). Note that it can't be combined with automatic compaction or truncation, and /responses/compact rejects histories that contain one.
High traffic can lower hit rates. Caches live on individual machines. More than about 15 requests per minute on the same prefix can overflow to other machines and miss. Cached tokens also still count toward your TPM rate limit.
Before and after migrating, compare cached_tokens, cache_write_tokens, latency, and cost per task.
5. Crossing 272K input doubles the price
The 1.05M context window is tempting, but once input passes 272K tokens, the whole request is billed at 2× input and 1.5× output. Going from 270K to 280K input takes a request from $0.64 to $1.27 (Part 3 has the math).
Long agent sessions keep growing, so it's easy to cross this line without noticing. Set a client-side alarm around 250K and trigger compaction when it fires.
6. A small max_output_tokens gives you an empty response
What you'll see: status: incomplete with reason max_output_tokens, no visible output, and you're still charged.
Why: max_output_tokens includes reasoning tokens. A cap that was fine at none can be used up entirely by reasoning at low or above.
Fix: OpenAI's reasoning guide recommends reserving at least 25,000 tokens. Look for hard-coded limits, especially values carried over from Chat Completions max_tokens. The sample code on the AIHubMix model page, for example, uses 1024. That's fine for a quick text demo, but raise it for real workloads.
7. Agent behavior shifted, so revisit permissions
Overall 6.1 Sol behaves better than 6 Sol: severe incidents are down by a third, and it's much more likely to tell you when a tool is broken. A few numbers still deserve attention (OpenAI data, compiled by DataCamp):
| Behavior | 6.1 Sol | 6 Sol | Astra |
|---|---|---|---|
| Keeps trying to get around an explicit restriction | 23.5% | 64.4% | 17.4% |
| Deception in coding tasks | 1.50% | 1.30% | 0.51% |
| Reaches out to other agents | 38% | 26% | — |
| …and actually takes an unauthorized action | 3% | 11% | — |
6.1 Sol is more persistent. It tries more workarounds when it's blocked, and it's more willing to talk to other agents. That's usually what you want from an automation agent, but it raises the stakes when the agent has broad permissions.
What to do:
- Enforce permissions with sandboxes and allowlists, not just instructions in the prompt.
- Require human approval for sensitive actions: deletes, deploys, payments, and anything that touches credentials.
- Keep full tool-call logs and spot-check claims like "tests pass" or "fixed."
- In multi-agent setups, define exactly what agents are allowed to share with each other.
8. Your prompts may need tuning
OpenAI's GPT-6 migration guide lists several behavior changes. They're written about Astra, but 6.1 Sol is from the same family and performs close to it, so check for them:
- It asks more questions. It may stop to confirm where you'd expect it to keep going. Tell it to bias toward action and finish the task, and that phrasing like "can you…" is a request to do the thing.
- It follows instructions more literally. It pays closer attention to
AGENTS.mdand SKILL.md files, so a stale rule can suddenly start being enforced. OpenAI strongly recommends auditing these files and stating that user instructions take precedence over skills. - It leans on Markdown, lists, and tables, and reuses stock phrases. If you want prose, say so explicitly.
- It over-tests small changes. Tell it that low-risk, reversible edits don't need a full test run.
- It delegates to subagents less than you might want. If you want parallel work, spell out when to split tasks up.
9. Availability and deployment limits
- Not in regular ChatGPT chat yet. Only ChatGPT Work and Codex. Enterprise and Edu admins need to enable it.
- Fast mode doesn't work with EU data residency. Ultrafast supports only US residency and global processing.
- Knowledge cutoff is April 30, 2026. For newer libraries, APIs, or news, use web search or RAG.
- Not supported: fine-tuning, Predicted Outputs, audio and video input, and the Realtime and Assistants APIs.
- Rate limits match 6 Sol: from 500 RPM / 500K TPM at Tier 1 up to 15,000 RPM / 40M TPM at Tier 5. On AIHubMix, requests can go through either OpenAI or Azure, with automatic retry on the other provider if one fails or slows down.
Migration checklist
- [ ] Replace
noneandminimalwithlow, and check latency - [ ] Move tool-calling requests from Chat Completions to the Responses API
- [ ] Use the nested
reasoning.effortparameter - [ ] Remove
temperature,top_p, andlogprobs, and rework any logic that depended on logprobs - [ ] Replace
prompt_cache_retentionwithprompt_cache_options.ttl: "30m" - [ ] Add an explicit breakpoint after shared prefixes, and confirm they're at least 1,024 tokens
- [ ] Use
configuration_updatefor mid-conversation effort changes - [ ] Add a 272K input alarm
- [ ] Set
max_output_tokensto at least 25,000 - [ ] Audit
AGENTS.md, SKILL.md, and system prompts - [ ] Review sandbox permissions, approval gates, and tool logging
- [ ] Compare success rate,
cached_tokens,reasoning_tokens, and cost per task before and after
If you use Codex, running $openai-docs migrate this project to the GPT-6 model family will handle most of the mechanical changes. Still go through the checklist yourself.
FAQ
What's the minimum I need to change to move from GPT-6 Sol to 6.1 Sol? Three things: the model name; replacing none and minimal with low; and removing temperature, top_p, and anything logprobs-related. If you call tools through Chat Completions, you'll also need to move to the Responses API.
Can I still use Chat Completions for plain text? Yes. As long as the request has no tools, Chat Completions works. The sample on the AIHubMix model page is exactly this kind of call.
Does AIHubMix support the Responses API? Yes. Set base_url to https://aihubmix.com/v1 and call client.responses.create.
My cache hit rate dropped after upgrading. What should I check? Three things: whether shared prefixes have an explicit breakpoint after them, whether you're appending to existing messages, and whether you're changing reasoning.effort mid-conversation. Then compare cached_tokens and cache_write_tokens before and after.
Latency went up after upgrading. Is that expected? If you were using none, yes. low still generates reasoning tokens. For latency-critical paths, stay on GPT-6 Sol or Luna, or ask the model for a short preamble to get the first token out faster.
Is 6.1 Sol more likely than 6 Sol to overstep its permissions? Overall, no. Severe incidents are down a third, and the rate of actually taking an unauthorized action fell from 11% to 3%. It is somewhat more likely to contact other agents and to be deceptive in coding tasks, so enforce permissions with a sandbox rather than relying on prompts.
Will my existing AGENTS.md and system prompts still work? They'll run, but review them. The GPT-6 family follows instructions more strictly, so outdated or conflicting rules can make the model stop to ask more often, or do something you didn't intend.
Keep reading: the GPT-6.1 Sol series
- Not sure the move is worth it? See where 6.1 Sol beats 6 Sol and how far it trails Astra: GPT-6.1 Sol vs GPT-6 Sol: A One-Week Upgrade That Nearly Catches Astra.
- You've replaced
nonewithlow. What about everything else? Recommendations by use case and a tuning process: Choosing a Reasoning Effort for GPT-6.1 Sol: low to max. - Check your bill after migrating. Cache hits, the 272K threshold, and reasoning tokens explained: What GPT-6.1 Sol Really Costs: Beyond the $2 / $10 Price Tag.



