Reviewed by Jonathan West · Updated Sep 30, 2026

GPT-6.1 Sol API Pricing: Token Rates and Cost Models

OpenAI published token rates of $2.00 per million input, $0.10 for cached context, and $10.00 for output.

Reviewed by Jonathan West · Updated Sep 30, 2026

On September 29, 2026, OpenAI introduced GPT-6.1 Sol at DevDay 2026 as an upgraded frontier model built for agentic coding, computer use, and professional document workflows. The model carries the official Application Programming Interface (API) identifier gpt-6.1-sol and is available to developers alongside ChatGPT Work and Codex users.

Unlike the standard API default or the earlier GPT-6 Sol released a week prior, GPT-6.1 Sol approaches the performance of OpenAI's flagship GPT-6 Astra at roughly one-fifth the token price. On DeepSWE v1.1 for real-world software engineering, GPT-6.1 Sol matches GPT-6 Astra while reducing per-task cost, and on GDP.pdf it exceeds Claude Opus 5.5 by Anthropic on multi-page document reasoning at less than half the operational cost.

For technical leaders and operations teams in regulated sectors like legal, healthcare, and financial services, this release changes the unit economics of autonomous agent pipelines. Workloads that previously required rationing top-tier models can now execute long-context analysis and multi-step tool calls with low input cache pricing.


Standard Token Rates for GPT-6.1 Sol

OpenAI charges $2.00 per million standard input tokens and $10.00 per million output tokens for GPT-6.1 Sol. Prompt caching drops the input rate to $0.10 per million tokens for cached context, which represents a 95 percent discount compared to standard input rates and a 50 percent cut from GPT-6 Sol's cache pricing.

The model is accessed via the model identifier gpt-6.1-sol in the chat completions and assistants endpoints. OpenAI announced that a specialized variant named GPT-6.1 Sol Ultrafast will launch in the coming days, providing up to 8x faster token generation inside Codex.

OpenAI has not published official context window limits or maximum output token ceilings on the launch announcement page. Teams evaluating migration from GPT-6 Sol should verify whether the earlier 1.05 million token window applies to 6.1 before deploying massive prompts.

  • Standard input tokens: $2.00 per million tokens.
  • Cached input tokens: $0.10 per million tokens.
  • Standard output tokens: $10.00 per million tokens.
  • Ultrafast speed: up to 8x faster token generation coming soon in Codex.

How Context Reuse Shapes Effective Token Costs

Prompt caching cuts the base input cost of GPT-6.1 Sol from $2.00 down to $0.10 per million tokens. For systems with large system instructions, broad schemas, or indexed document reference blocks, cached calls remove the majority of recurring input billing.

Workflows that run iterative agent loops send the entire conversational history on every turn. When 80 percent of a request's input tokens hit the cache, effective input pricing drops to $0.48 per million tokens instead of the flat $2.00 standard rate.

In our implementations for clients managing automated document pipelines, cache hit consistency is the single largest variable determining whether a deployment stays within budget. Designing prompts with static prefixes placed before variable user data ensures requests reliably trigger the $0.10 cached rate.


GPT-6.1 Sol API Pricing Compared to Rival Models

OpenAI prices GPT-6.1 Sol at roughly one-fifth the standard token rates of its flagship sibling GPT-6 Astra. While Astra targets the most demanding scientific research, GPT-6.1 Sol delivers competitive results across enterprise tasks at a fraction of the expenditure.

Benchmark data shows significant operational savings against Claude Opus 5.5 by Anthropic and earlier OpenAI models. On the GDP.pdf evaluation covering complex financial, legal, and medical documents, GPT-6.1 Sol scored higher than Claude Opus 5.5 with fallbacks while operating at less than half the task cost.

On the AutomationBench 1.0.6 test measuring multi-step business actions across 47 enterprise tools, GPT-6.1 Sol scored 2.2 percentage points higher than Claude Opus 5.5 at medium reasoning effort while running at roughly one-third the cost. On Terminal-Bench Science 0.1, the average task cost was $5.47 for GPT-6.1 Sol compared to $23.21 for Claude Opus 5.5 and $23.80 for GPT-6 Astra.

  • AutomationBench 1.0.6: beats Claude Opus 5.5 by 2.2 points at roughly one-third the per-task cost.
  • GDP.pdf document parsing: outperforms Claude Opus 5.5 at less than half the cost across tested reasoning settings.
  • Terminal-Bench Science 0.1: averages $5.47 per task versus $23.21 for Claude Opus 5.5 and $23.80 for Astra.
  • OSWorld 2.0 computer use: scores within 2.1 points of Astra while costing roughly one-seventh as much per task.

Worked Workload Cost Models for GPT-6.1 Sol

Real application pricing depends heavily on the ratio of cached input tokens to newly generated output tokens. Below are three realistic enterprise scenarios modeled against OpenAI's official $2.00 input, $0.10 cached input, and $10.00 output pricing.

A tier-one customer support assistant handles 100,000 inquiries monthly, with 4,000 prompt tokens per ticket where 3,000 are cached enterprise policies and 400 tokens are generated in response. Standard input costs $0.20, cached input costs $0.03, and output costs $0.40, totaling $0.63 per 1,000 tickets or $63.00 per month.

A legal document review pipeline processing 5,000 contracts monthly averages 40,000 tokens of contract text with 35,000 tokens cached across multiple review passes and generates 2,000 tokens of structured analysis. Input costs $0.01 for new text, $0.0035 for cached text, and output costs $0.02, leading to $0.0335 per contract or $167.50 monthly.

  • Support Assistant: 100,000 conversations with 75% cache hit rate costs approximately $63.00 monthly.
  • Contract Pipeline: 5,000 complex agreements reviewed with 87% cache reuse costs roughly $167.50 monthly.
  • Coding Agent: 1,000 multi-turn bug-fixing sessions averaging 50 tool iterations costs approximately $275.00 monthly.
  • Factuality: 7.7% factual error rate on hard evals reduces expensive human remediation cycles by 32% versus GPT-6 Sol.

Audience Fit, Operating Limits, and Inversion Conditions

GPT-6.1 Sol is poorly suited for ultra-low-latency classification tasks where small models like GPT-4o mini or Haiku deliver adequate results at a fraction of a cent. High-volume, single-turn routing that requires neither reasoning nor multi-step tool integration does not justify a $2.00 per million input baseline.

Our recommendation of GPT-6.1 Sol would change if competing frontier labs drop flagship output pricing below $5.00 per million tokens without requiring tiered commitments, or if specialized scientific simulation remains strictly necessary. For extreme theorem proving and frontline scientific discovery, OpenAI explicitly advises using GPT-6 Astra despite its 4x higher per-task cost.

Check your organization's current OpenAI account spend tier on the OpenAI platform dashboard to verify your rate limits before moving production agent traffic to the gpt-6.1-sol API identifier.

Frequently Asked Questions

  • OpenAI set GPT-6.1 Sol pricing at $2.00 per million standard input tokens, $0.10 per million cached input tokens, and $10.00 per million output tokens. Cached inputs receive a 95 percent discount compared to standard input calls.
  • GPT-6.1 Sol runs at roughly one-fifth the standard token price of GPT-6 Astra. On benchmarked business and software-engineering tasks, it approaches Astra's accuracy while reducing operational costs by 75 to 85 percent.
  • The official model identifier to pass in API requests is gpt-6.1-sol. Developers can use this identifier in OpenAI SDK calls across completions and agent workflows.
  • OpenAI made GPT-6.1 Sol available to Plus, Pro, Business, Enterprise, and Education accounts inside ChatGPT Work and Codex on September 29, 2026. The vendor noted it is not yet available in the standard ChatGPT Chat interface.
  • GPT-6.1 Sol Ultrafast is an upcoming configuration announced by OpenAI that will deliver up to 8x faster token generation speeds inside Codex compared to standard generation rates.
  • Prompt caching reduces input token expenses from $2.00 per million tokens to $0.10 per million tokens. This cuts the cost of cached context by 95 percent relative to standard inputs and 50 percent compared to the prior GPT-6 Sol cache rate.

Optimize Your AI Architecture and Compliance

Layer3Labs helps regulated small and mid-sized businesses design, integrate, and monitor production AI workflows. Book a 30-minute consultation to review your pipeline economics and compliance posture.

Book a Consultation