What GPT-6.1 Sol Really Costs: Beyond the $2 / $10 Price Tag

AIHubMix6 min read
What GPT-6.1 Sol Really Costs: Beyond the $2 / $10 Price Tag

GPT-6.1 Sol has the same list price as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. Your actual bill depends on four other things: how often you hit the cache, whether you cross 272K input tokens, which service tier you use, and how many reasoning tokens the model burns. This post goes through each one.


The full price list

Standard rates (per 1M tokens)

GPT-6.1 SolGPT-6 SolGPT-6 AstraGPT-6 Luna
Input$2.00$2.00$10.00$0.10
Cached input$0.10$0.20$1.00—
Cache writes$2.50$2.50——
Output (incl. reasoning)$10.00$10.00$50.00$0.50
Astra's cached rate comes from third-party reporting. Its cache-write rate should be $12.50 under OpenAI's 1.25× rule. For Luna's cache pricing, see OpenAI's pricing page.

Multipliers

These rules come from the GPT-6.1 Sol model page:

ConditionWhat happens
More than 272K input tokensThe whole request bills at 2× input and cache, 1.5× output ($4 / $15, cache reads $0.20, cache writes $5)
Batch / FlexHalf the standard rate
Fast modeDouble the standard rate (not available with EU data residency)
Regional processing+10%
Tool calls (search, computer use, etc.)Charged per call. AIHubMix lists web search at $0.01 per request
Ultrafast (coming to Codex)Not officially priced yet. One report says 6× standard, unconfirmed

On AIHubMix, the OpenAI and Azure routes cost the same as OpenAI direct, and the over-272K tier is already listed.


The real change: cache reads at 5% of input

The list price didn't move. Cache reads did, from 10% of the input rate ($0.20) down to 5% ($0.10). OpenAI's prompt caching docs put it plainly: cache reads are 0.1× input on most models and "0.05× on GPT-6.1 Sol."

This matters because agents resend the same material on every step: system prompt, tool definitions, repo context, and conversation history. Most of their input is repeated content. For these workloads, the cached rate is effectively your real input price.

Worked example: one coding-agent task

Assumptions:

  • A 200K-token fixed prefix (system prompt, tools, repo context)
  • 50 steps, each adding 2K new input tokens and producing 3K output tokens (including reasoning)
  • To keep it simple, the prefix stays the same size, is written to cache on step 1, and hits on all 49 remaining steps
6.1 Sol, no cache6 Sol + cache6.1 Sol + cacheAstra + cache
Initial 200K cache write—$0.50$0.50$2.50
49 cache reads—$1.96$0.98$9.80
Prefix billed as regular input$20.00———
100K new input (at write rate)$0.20$0.25$0.25$1.25
150K output$1.50$1.50$1.50$7.50
Total$21.70$4.21$3.23$21.05

Three takeaways:

  1. Caching is the biggest lever. Same model, same work, and the uncached run costs nearly 7× more.
  2. 6.1 Sol is about 23% cheaper than 6 Sol here, while also scoring higher.
  3. For cache-heavy work, Astra costs more than the 5× list-price gap suggests. In this example it's 6.5×, because Astra discounts cached input by 90% and 6.1 Sol discounts it by 95%.

That lines up with OpenAI's reported per-task costs (compiled by Vellum and The Next Web): about $1.50 vs $7.70 on DeepSWE, $1.30 vs $9.30 on OSWorld, and $5.47 vs $23.80 on Terminal-Bench Science.

One caveat: OpenAI says Astra usually needs fewer output tokens to finish the same task. Compare models on cost per task, not cost per token.

Cache writes aren't free

Cache writes on 6.1 Sol cost 1.25× the input rate. If a prefix is only used once, caching makes it 25% more expensive. Using the formula from OpenAI's docs:

  • One write plus one read: 1.35× regular input cost (1.3× on 6.1 Sol), versus 2× without caching
  • One write plus nine reads: about 2.15× (about 1.7× on 6.1 Sol), versus 10×

The minimum cacheable prefix is 1,024 tokens. OpenAI's break-even math says a prefix of 102 tokens or fewer never pays off. If you expect 10 or more reuses, padding anything over 221 tokens up to 1,024 comes out cheaper.

For one-off calls such as single-turn classification or bulk extraction, use explicit mode (prompt_cache_options.mode: "explicit") with no breakpoints. You won't be charged for any writes.


The 272K cliff

6.1 Sol accepts 1.05M tokens of context, but once input goes past 272K, the entire request moves to the higher rate, not just the tokens over the line.

RequestInputOutputCost
A270K10K0.27 × $2 + 0.01 × $10 = $0.64
B280K10K0.28 × $4 + 0.01 × $15 = $1.27

Ten thousand more input tokens nearly doubles the bill.

How to stay under it:

  • Count tokens client-side and compact or trim before you reach 272K
  • Retrieve the relevant parts of long documents instead of sending the whole thing
  • When you genuinely need long context, put the fixed part in a cached prefix

How reasoning effort shows up on the bill

Reasoning tokens bill at the output rate, $10 per million. Moving the same task from low to max can multiply reasoning tokens by ten or more.

OpenAI's reported per-task costs give a sense of scale (these are different benchmarks, so only compare the magnitudes):

  • AutomationBench (medium): ~$0.30
  • GDP.pdf: ~$0.38
  • DeepSWE (high): ~$1.50
  • Terminal-Bench Science (max): ~$5.47

Part 2 covers how to pick a level. One reminder here: changing reasoning.effort mid-conversation invalidates the cache. Use configuration_update instead.


How it compares at this price point

ModelInput / outputCached input
GPT-6.1 Sol$2 / $10$0.10
Claude Sonnet 5.5$2 / $10$0.20
GPT-6 Astra$10 / $50$1.00

6.1 Sol and Claude Sonnet 5.5 share a list price, and 6.1 Sol's cached rate is half of Sonnet's. They're strong at different things, though. On AutomationBench, for example, Sonnet 5.5 scores well ahead. When two models cost the same per token, the one with the lower cost per task on your workload is the cheaper one.

The AIHubMix model list shows current prices, and one API key works for all of them, so side-by-side tests are easy.

Competitor prices come from third-party sources and may change. Check the official pricing pages before relying on them.

Cost checklist

  1. Put stable content first. System prompt, tool definitions, and reference material go at the top. Timestamps and per-user data go at the end.
  2. Append, don't rewrite. Summarizing, truncating, and compacting all restart the cache.
  3. Keep the tool list stable. Don't add or remove tools per request. Use tool_choice or allowed_tools instead.
  4. Switch effort with configuration_update, not the request parameter.
  5. Stay under 272K input.
  6. Send non-urgent work through Batch or Flex for half price.
  7. Route by task. Luna for simple high-volume work, 6.1 Sol for most things, Astra for the hardest problems.
  8. Watch three numbers: cached_tokens, cache_write_tokens, and reasoning_tokens.

The bottom line: the list price didn't change, but cache reads got 50% cheaper. For agent workloads, that makes 6.1 Sol the best-value model in the GPT-6 family, as long as you use caching well and stay under 272K.

Next up: the problems you're likely to hit when you switch.


FAQ

How much does GPT-6.1 Sol cost? $2 per million input tokens, $10 per million output tokens, $0.10 for cached input, and $2.50 for cache writes. Requests over 272K input tokens bill at $4 input and $15 output for the whole request.

Is it cheaper through AIHubMix or OpenAI directly? Same price. AIHubMix's OpenAI and Azure routes both match OpenAI's official rates.

Why is my bill higher than I estimated? Four common reasons: reasoning tokens bill at the output rate; you crossed 272K input; you missed the cache or paid write charges on one-off prefixes; or Fast mode or regional processing is on.

Do cache writes cost extra? Written tokens bill at 1.25× the input rate. That replaces the normal input charge rather than adding to it. A prefix pays for itself after two uses. For one-off prefixes, explicit mode lets you skip writes entirely.

How much cheaper is 6.1 Sol than Astra? Five times cheaper on list price. For cache-heavy agent work, the gap can reach 6 to 7× because 6.1 Sol discounts cached input more steeply. Astra often uses fewer output tokens per task, though, so compare cost per task.

Can Batch be combined with caching? Batch and Flex are both half price. For how they stack with caching, check OpenAI's pricing page.


Keep reading: the GPT-6.1 Sol series


Sources