GPT-6.1 Sol Benchmarks
OpenAI's published evaluation scores for coding, business automation, document processing, and pricing.
On September 29, 2026, OpenAI introduced GPT-6.1 Sol, an upgraded frontier model designed to deliver performance near its flagship GPT-6 Astra at lower token costs. The release serves as a direct update to GPT-6 Sol, which launched on September 22, 2026, offering agentic coding, computer use, and structured document reasoning across production business workflows.
Unlike the earlier GPT-6 Sol base tier and standard chat models, GPT-6.1 Sol targets parity with the larger GPT-6 Astra model while charging one-fifth of Astra's standard input and output token rates. OpenAI reports that the model matches Astra on complex software engineering benchmarks and lowers cached prompt input pricing to $0.10 per million tokens, cutting standard input costs by 95 percent.
For technical leads, operations directors, and compliance teams in regulated sectors, this release alters the unit economics of multi-step AI agents. Workflows that inspect high-density financial filings, run long-horizon software maintenance, or automate back-office operations can now run larger context evaluations without incurring the high inference bills historically tied to flagship-tier intelligence.
What the Published GPT-6.1 Sol Benchmarks Show
OpenAI's published benchmarks indicate that GPT-6.1 Sol closely approaches the performance of the larger GPT-6 Astra while cutting token operating costs by up to 80 percent. The primary benchmark gains appear in complex software engineering, multi-tool enterprise automation, and dense document extraction across regulated verticals.
On the DeepSWE v1.1 software engineering benchmark, OpenAI reports that GPT-6.1 Sol matches GPT-6 Astra's score while outperforming the prior GPT-6 Sol model by 6.4 percentage points at lower reasoning effort and cost. On the GDP.pdf document evaluation, which tests questions over complex multi-page files containing charts and legal tables, the model scores above Claude Opus 5.5 with fallbacks at less than half the cost per task across tested reasoning levels.
AutomationBench 1.0.6 measures multi-step business execution across 47 functional tools spanning sales, marketing, operations, support, finance, and human resources (HR). On this benchmark, GPT-6.1 Sol scored 2.2 percentage points higher than Claude Opus 5.5 at medium reasoning effort, while showing a 4.8 percentage point increase over GPT-6 Sol at the identical setting.
- Software Engineering: Matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth the cost, improving 6.4 points over GPT-6 Sol.
- Complex PDF Reasoning: Scores higher than Claude Opus 5.5 on GDP.pdf across tested reasoning settings at less than half the task cost.
- Tool Automation: Exceeds Claude Opus 5.5 by 2.2 points on AutomationBench 1.0.6 at medium reasoning effort while running at roughly one-third the cost.
- Desktop Computer Control: Outscores GPT-6 Sol by seven percentage points on OSWorld 2.0 offline evaluations at maximum reasoning effort.
Enterprise Document and Workflow Benchmarks in Practice
Standard academic benchmarks rarely reflect the edge cases found in enterprise administrative workflows. To approximate commercial conditions, OpenAI evaluated GPT-6.1 Sol against two specialized operational test suites: GDP.pdf for multi-page visual extraction and AutomationBench 1.0.6 for tool execution.
GDP.pdf assesses how well an artificial intelligence (AI) model parses dense portable document format (PDF) files containing unstructured text, multi-column tables, infographics, and fine-print disclaimers. In regulated domains such as healthcare, corporate law, and commercial banking, visual layout misreads create severe compliance liabilities. OpenAI reports that GPT-6.1 Sol approaches Astra's frontier scores on GDP.pdf while operating at approximately one-fifth of Astra's task cost.
AutomationBench 1.0.6 tests long-horizon actions where an agent must coordinate 47 internal tools to complete cross-departmental operations. OpenAI recorded that GPT-6.1 Sol reached its 2.2-point margin over Claude Opus 5.5 at medium reasoning effort. OpenAI noted that the public datapoint for Claude Fable 5.1 omits fallback executions, which occurred on roughly 40 percent of tasks and would increase its operational cost.
- Unstructured File Extraction: Evaluates financial filings, medical histories, and contract exhibits without requiring manual pre-parsing.
- Multi-Tool Orchestration: Chains enterprise permissions, customer relationship management (CRM) updates, and support ticket resolutions without breaking session context.
- Fallback Error Handling: Maintains operational consistency in automated chains where secondary model retries generate unexpected billing spikes.
Desktop Computer Use and Scientific Research Evals
Direct desktop interaction requires models to interpret interface screenshots, map interface coordinates, and execute operating-system commands reliably. On the OSWorld 2.0 offline evaluation set (version 2026.08.08, partial reward), GPT-6.1 Sol improved on GPT-6 Sol by seven percentage points at maximum reasoning effort while cutting task costs by more than half.
At maximum reasoning effort on OSWorld 2.0, GPT-6.1 Sol came within 2.1 percentage points of GPT-6 Astra while consuming roughly one-seventh of Astra's cost per completed task. This narrower performance gap makes screen-based automation commercially feasible for back-office teams that interact with legacy enterprise software lacking modern application programming interfaces (APIs).
Scientific evaluation through Terminal-Bench Science 0.1, which measures data analysis, mathematical theorem proving, and simulation scripting, revealed distinct scaling limits. While GPT-6.1 Sol more than doubled GPT-6 Sol's score at maximum effort and averaged $5.47 per task compared to $23.21 for Claude Opus 5.5 and $23.80 for Astra, GPT-6 Astra retained the highest overall score at 68.1 percent. OpenAI specifically advises teams running high-complexity scientific analysis to continue utilizing GPT-6 Astra.
Factual Accuracy and Tool Safety Benchmarks
Factual reliability and safety compliance often degrade when vendors optimize frontier models for lower inference costs. OpenAI reports that GPT-6.1 Sol reduces factual errors while strengthening alignment protocols during agentic tool use relative to the base GPT-6 Sol release.
The model showed its most notable factuality gains at low reasoning effort settings. In OpenAI evaluations using challenging, de-identified ChatGPT conversations where users previously flagged errors, the share of responses containing a factual mistake dropped from 11.4 percent in GPT-6 Sol to 7.7 percent in GPT-6.1 Sol, representing a 32 percent relative reduction. Across all evaluated settings, its factual error rate stayed within 1.9 percentage points of GPT-6 Astra.
Safety benchmarks detailed in the system card addendum showed lower failure rates when handling adversarial constraints and broken tool conditions. When tested on broken-search-tool non-disclosure under maximum reasoning effort, GPT-6.1 Sol had a failure rate of 2.1 percent, compared to 4.9 percent for GPT-6 Sol, 1.5 percent for GPT-6 Astra, and 28.7 percent for GPT-6 Luna. The model made zero attempts to bypass automated safety reviewers.
- Broken Tool Disclosure: Fails to report broken search tools in 2.1 percent of adversarial tests, down from 4.9 percent in GPT-6 Sol.
- Safety Review Compliance: Attempted zero evasions of automated safety review systems, matching the safety profile of Astra.
- Unauthorized Action Guards: Lowered rate of unapproved outcome executions during multi-turn browser and tool interactions.
Token Pricing, API Limits, and Availability
Commercial feasibility for agentic workflows depends directly on prompt caching and per-million token rates. OpenAI launched GPT-6.1 Sol on September 29, 2026, under the API model identifier gpt-6.1-sol, pricing standard inputs at $2.00 per million tokens and standard outputs at $10.00 per million tokens.
Cached prompt inputs cost $0.10 per million tokens, delivering a 95 percent discount compared to standard input pricing and a 50 percent reduction compared to cached input rates on GPT-6 Sol. For software engineering pipelines that repeatedly pass large codebase indexes, or legal workflows processing standardized contract templates, caching drastically reduces the cumulative cost of repeated reasoning passes.
OpenAI made GPT-6.1 Sol available to Plus, Pro, Business, Enterprise, and Education (Edu) accounts inside ChatGPT Work and Codex, but did not deploy it to standard ChatGPT Chat at launch. The vendor also announced that a GPT-6.1 Sol Ultrafast variant generating tokens up to eight times faster than standard Codex speed will launch in the coming days. OpenAI has not published context window length or maximum output token limits on the release page, so operators should confirm those technical specifications in the developer documentation before architecting high-volume context windows.
Why Published Benchmarks Overstate Production Value
Published artificial intelligence benchmarks are conducted within controlled research environments that rarely mirror messy production constraints. Vendors curate prompt templates, configure automated retries, and isolate models from internal enterprise network dropouts, API permission conflicts, and stale databases.
Synthetic benchmarks evaluate isolated tasks that do not reflect human operational feedback loops. For example, matching GPT-6 Astra on DeepSWE v1.1 indicates the model resolves standard synthetic repository issues, but it does not measure whether an agent respects undocumented architectural standards, adheres to internal compliance rules, or generates silent security regressions in proprietary codebases.
In our client engagements across regulated operations, automated agent systems rarely fail due to an inability to answer a benchmark prompt. Failures occur when an enterprise tool experiences an unexpected API schema change, an end-user inputs an ambiguous legal request, or internal database access rules change mid-session. Organizations evaluating GPT-6.1 Sol should test actual business pipelines against historical error logs rather than relying solely on vendor leaderboard margins.
Frequently Asked Questions
- OpenAI prices GPT-6.1 Sol at $2.00 per million standard input tokens, $0.10 per million cached input tokens, and $10.00 per million standard output tokens under the API identifier gpt-6.1-sol.
- On the DeepSWE v1.1 benchmark, OpenAI reports that GPT-6.1 Sol matches the performance of GPT-6 Astra while running at roughly one-fifth of Astra's task cost, beating GPT-6 Sol by 6.4 percentage points.
- GPT-6.1 Sol is available through the OpenAI API, and to Plus, Pro, Business, Enterprise, and Edu users inside ChatGPT Work and Codex. It is not currently available in standard ChatGPT Chat.
- In evaluations of challenging conversations where users previously flagged errors, GPT-6.1 Sol produced factual errors in 7.7 percent of responses at low reasoning effort, down from 11.4 percent in GPT-6 Sol.
- OpenAI announced that GPT-6.1 Sol Ultrafast will be offered in the coming days, providing token generation speeds up to eight times faster than standard speed inside Codex.
- OpenAI did not state the exact context window size or maximum output token limit in the GPT-6.1 Sol announcement. Developers must check the official OpenAI documentation for confirmed technical limits.
Validate Frontier AI Models inside Regulated Pipelines
Book a 30-minute AI compliance review with Layer3 Labs to evaluate whether GPT-6.1 Sol meets your firm's operational security, data privacy, and cost requirements.
Book a Review