The price change sits in the cache

OpenAI released GPT-6.1 Sol on September 29 for coding and other complex work. The company describes its performance as close to GPT-6 Astra at a lower cost. That is OpenAI's positioning, not a result Model Current has reproduced.

The standard API rates for ordinary input and output tokens are $2 and $10 per million. Those are the same headline rates as GPT-6 Sol, released a week earlier. The measurable price change is for cached input, which is previously supplied text the service can reuse. Its rate falls from $0.20 to $0.10 per million tokens.

That matters most to applications that send the same long instructions or documents repeatedly and actually get cache hits. It tells us little about a one-off request. OpenAI also lists separate charges for cache writes and some tools, and higher rates for prompts over 272,000 input tokens. A team's total cost still depends on its requests, outputs, tools and retries.

Check the call before changing the model name

The new model identifier is gpt-6.1-sol. OpenAI's model page says tool calling should use the Responses API. Chat Completions remains available without tool calling. An application relying on tools through a different call path cannot assume the model name is a drop-in replacement.

Reasoning settings need attention too. GPT-6 Sol accepted 'none'; GPT-6.1 Sol supports low through max but neither 'none' nor 'minimal'. The right effort level depends on the task, so a migration should be tested with representative work rather than only a successful sample prompt.

OpenAI lists a 1.05 million-token context window and 128,000 maximum output tokens. Those limits describe what the API accepts, not whether a very long answer is useful or economical. The model accepts text and image input, not audio or video input.

The useful comparison is the finished job

OpenAI says GPT-6.1 Sol is near Astra for difficult work. The published model pages show a much lower standard input and output price than Astra, but they do not establish that every real task will finish at the same quality or require the same number of attempts.

For a fair migration test, keep a small set of actual jobs, define what counts as a correct result, and compare completed-task cost, latency and review effort. Watch cache-hit rates separately. A lower cached-input rate is real; an overall saving is something each user still has to measure.

Sources

  1. OpenAI API changelog: September 29 releasePrimary dated record for GPT-6.1 Sol release, standard prices, multi-agent beta and Responses API tool-calling requirement.
  2. OpenAI API: GPT-6.1 Sol model pagePrimary current specifications, effort levels, modalities, context limits, prices and longer-prompt qualifications.
  3. OpenAI API: GPT-6 Sol model pageDirect comparison for unchanged uncached input/output rates, older $0.20 cached-input rate and different reasoning support.