Codex Compaction Fails With "response protection is unavailable": It's the Web Search, Not Your Context

AIHubMix8 min read
Codex Compaction Fails With "response protection is unavailable": It's the Web Search, Not Your Context

Since October 8, 2026, long Codex sessions have been dying the moment they try to compact. The error reads stream disconnected before completion: response protection is unavailable, retries don't help, and the conversation can't continue. Most reports involve gpt-6.1-sol, but gpt-6-astra and gpt-5.6-sol fail the same way. It shows up behind self-hosted proxies, and it also shows up in the official Codex app with a direct ChatGPT login.

The short answer: the context isn't too big, and your network isn't the problem. ChatGPT's Codex backend now rejects any request that replays an earlier web search in its history without also declaring the web search tool. Codex's compaction request declares no tools at all. So a single web search early in a session is enough to make every later compaction fail.

What to do right now:

  • New sessions: set web_search = "disabled" until the backend or Codex changes.
  • Stuck sessions: move the work into a fresh session. The old one won't compact.
  • If you maintain your own gateway in front of Codex: add a web search declaration when the history contains a search. The code is below.

The rest of this post explains how the trigger was found, which community fixes don't work, and how to rescue a stuck session.

What the failure looks like

You'll see one of these in Codex:

stream disconnected before completion: response protection is unavailable
Error running remote compact task: stream disconnected before completion: stream closed before response.completed

The second message is the generic version. Codex doesn't always show the upstream error, so a compaction that fails with "stream closed before response.completed" can be the same problem. Codex's local log often has Failed to run pre-sampling compact next to it.

Underneath, the upstream sends one of two things. Sometimes it's an HTTP 502 with this body, which a ChatGPT Plus user on gpt-6.1-sol also posted to OpenAI's developer forum:

{"message": "response protection is unavailable", "type": "internal_error"}

Other times it's an HTTP 200 whose stream ends in a response.failed event with code: "upstream_error" and no response.completed. Anything that only checks the HTTP status misses the second form and reports it as a stream that ended early.

The pattern is the same everywhere:

  • Normal turns keep working. Only compaction fails.
  • Codex retries about five times, then gives up. In one gateway's logs, each failed attempt took 2 to 13 seconds and recorded zero tokens.
  • Resuming the session runs into the same compaction, so the session stays stuck.
  • A fresh session works until it hits the same conditions.

The trigger: a search in history, no search tool in the request

Codex has a built-in web search tool. When the model uses it, a web_search_call item goes into the conversation history. Every later request sends that history back.

Normal turns declare the tool in tools, so the upstream accepts the replayed search. A compaction request is different. Codex sends the whole history with tools: [], because a summary doesn't need tools. That applies to local compaction, which Codex runs for custom providers, and to remote compaction v2.

Around October 6, the ChatGPT Codex backend started rejecting requests that replay a web_search_call but don't declare web_search. Developers who captured the raw requests narrowed it down. Changing or removing the search item's ID makes no difference. Declaring only a function tool still fails. Declaring web_search makes the same request succeed.

The same result shows up with the official, unmodified codex-cli signed in with a ChatGPT account. A history with three search items failed to compact. The same history with only those items removed compacted fine. A second reporter in that thread captured a compaction request with one web_search_call and tools: []. Adding a web_search declaration fixed it, and so did removing the search item.

That explains why it looks like a context-size problem. Compaction only runs on long sessions, and long sessions are the ones most likely to have used search at some point. Web search is also on by default in Codex (the default mode is "cached"), so a session can contain a search the user never asked for.

The evidence

At least four groups ran A/B tests independently and got the same result. The rows below combine tests posted on the openai/codex tracker and in the issue trackers of several open-source gateway projects. OpenAI hasn't confirmed the rule or responded in any of those threads.

Request history Tools declared Result
Messages and reasoning only None Completes
Includes function-call output None Completes
Includes a web_search_call None Fails
Includes a web_search_call Function tool only Fails
Includes a web_search_call web_search Completes
Same, search items removed None Completes

The tests rule out size. One tester set the auto-compaction threshold to 2,000 tokens with -c model_auto_compact_token_limit=2000. A session with no tool use compacted fine. A session that had searched once failed six times in a row. Another tester found that removing reasoning items didn't help, while removing the single search item did.

In one test, the version with web_search declared completed in 2.63 seconds and made zero new search calls. Declaring the tool doesn't make the model search again.

What the community tried, and what actually works

The LINUX DO thread that prompted this post went through most of the usual guesses:

Suggestion Helps? Why
Shrink context, compact earlier No Size isn't the trigger
Connect by IP, change nginx No The error comes from upstream
Switch to WebSocket Unreliable Same request body either way
Log in to ChatGPT directly No Official login fails too
New session, hand over context Stopgap Breaks again after a search
Disable web search Yes, for new sessions No search, no trigger
Declare web_search in the request Yes Satisfies the check

A few of these need more explanation.

Nginx and IP. Timeouts and idle connections can cause other "stream disconnected" errors, but they can't produce this message. Maintainers of the proxies involved confirmed the message isn't generated by the proxy; it comes from the upstream provider. If the message is there, the network got the request through and back.

WebSocket. The gateway patches that worked had to cover the WebSocket path as well as HTTP, because WebSocket requests carry the same body. A session that "recovered" after switching transport most likely had no search in its history.

Official login. Reports on the openai/codex tracker include the desktop app and codex-cli signed in directly with a ChatGPT account, with no proxy involved. As of October 10, neither Codex 0.162.1 nor the 0.163.0 alphas mention a fix, and the issues have no maintainer response.

What you can do now

Disable web search for new sessions

Turn search off in ~/.codex/config.toml:

web_search = "disabled"

The Codex config reference lists four values: disabled, cached (the default), indexed and live. Sessions started with --yolo or another full-access sandbox default to live, so set it explicitly.

Apply this to new sessions. An old session already has web_search_call items in its history. With search disabled, normal turns stop declaring the tool too, so those turns may start failing as well. That follows from the rule above, but nobody has reported testing it yet.

The cost is that Codex can't search the web. For a session that needs search, open a separate short one, or have the model write findings to a file that the long session reads.

Rescue a stuck session

Nothing on your side will make a stuck session compact while the backend behaves this way. To keep the work:

  1. Leave the stuck session as it is. Its history is still on disk under ~/.codex/sessions/.
  2. Start a new session with web search disabled.
  3. Point it at the old session file, or paste a short handover: the goal, the files changed, decisions made, and what's left.
  4. Stop sending prompts to the old session. Every attempt retries compaction and fails again.

If you maintain your own gateway

The fix that works follows one rule. If input contains a web_search_call and no web_search* tool is declared, add one. If the caller declared no tools, set tool_choice to "none" so the tool can't run. Leave input alone.

def declare_replayed_web_search(body: dict) -> dict:
    """Let the upstream accept a replayed web_search_call in a tool-less request."""
    items = body.get("input")
    if not isinstance(items, list):
        return body
    if not any(isinstance(i, dict) and i.get("type") == "web_search_call" for i in items):
        return body

    tools = body.get("tools") or []
    if any(isinstance(t, dict) and str(t.get("type", "")).startswith("web_search") for t in tools):
        return body

    caller_had_tools = bool(tools)
    # Cached index only: the declaration exists to satisfy the check, not to search.
    body["tools"] = tools + [{"type": "web_search", "external_web_access": False}]
    if not caller_had_tools:
        body["tool_choice"] = "none"
    return body

Responses Lite requests are an exception. They return a 400 if web_search sits in top-level tools. The working patches put it in the first additional_tools input item instead, and keep any trailing compaction_trigger last. The alternative, stripping search items out of compaction requests, also works, but the summary then loses whatever the search found.

If you'd rather not depend on the subscription backend

Every report we found goes through ChatGPT's subscription backend, the endpoint Codex uses with a ChatGPT login. Its rules aren't documented, and this one changed without notice.

Codex can also use the public Responses API with an API key. We haven't seen a report of this error on that path, but we haven't run the reproduction there ourselves either, so treat that as an observation, not a guarantee. To set Codex up with an AIHubMix key, follow the Codex CLI tutorial:

model = "gpt-6.1-sol"
model_provider = "aihubmix"

[model_providers.aihubmix]
name = "AIHubMix"
base_url = "https://aihubmix.com/v1"
wire_api = "responses"
env_key = "AIHUBMIX_API_KEY"

API pricing applies per token, not per subscription. Rates are on the gpt-6.1-sol model page.

Checklist

  • The error text contains "response protection is unavailable", or compaction fails with "stream closed before response.completed".
  • Normal turns still work and only compaction fails.
  • The session used web search at some point, possibly without being asked.
  • New sessions have web_search set to disabled.
  • Stuck work has been moved to a new session with a handover note.
  • A gateway you maintain adds a web_search declaration when the history contains a search.
  • Nginx and proxy network settings have been left alone. They aren't the cause.

FAQ

What does "response protection is unavailable" mean in Codex?
It's an error from ChatGPT's Codex backend, not from Codex itself or your proxy. Since early October 2026 it appears when a request replays an earlier web search but doesn't declare the web search tool, which is exactly what a compaction request does.

Is my context window too large?
No. Testers reproduced it with a 2,000-token compaction threshold, and sessions of the same size without a search compact normally. It only looks size-related because compaction only runs in long sessions.

Does it affect the official Codex app with a ChatGPT login?
Yes. Users of the desktop app and codex-cli report it on direct ChatGPT logins with no proxy. As of October 10, 2026, no Codex release mentions a fix.

How can I tell whether a session will hit it?
If the session used web search at any point, its next compaction will most likely fail. Codex records each search as a web_search_call item in the session files under the .codex/sessions folder.

Will disabling web search fix a session that's already stuck?
Probably not. The old history still contains the search, and once search is disabled, normal turns may stop declaring the tool and fail too. Use the setting for new sessions and move the work over.

Does switching to WebSocket or changing nginx help?
No. The request body is the same over either transport, and the error comes from upstream, so network settings can't remove it. Other stream errors can come from timeouts, but not this one.

Does it happen when Codex uses an API key instead of a ChatGPT subscription?
All the reports so far involve the subscription backend. No one has reported it on the public Responses API, though that hasn't been tested directly.

Keep reading: the GPT-6.1 Sol series

Sources