GPT-6.1 Sol Arrives | Astra-Level Performance at One-Fifth the Price
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
Have you ever felt that the smartest AI is useful, but too expensive to use every day? GPT-6.1 Sol is a new model that addresses that concern. This article breaks down its performance, pricing, availability, and weaknesses.
On September 29, 2026, OpenAI announced GPT-6.1 Sol at its developer event, "DevDay 2026."
Its predecessor, "GPT-6 Sol," was released on September 22. That means this update came in just one week.
The GPT-6 series is divided into three tiers. At the top is "Astra," the core tier is "Sol," and the lightweight version is "Luna." Sol is positioned as the workhorse for AI agents (AI that thinks and works autonomously) and coding.
According to OpenAI, GPT-6.1 Sol delivers intelligence nearly on par with Astra at one-fifth the token cost.
However, OpenAI itself states that "the highest-performing model remains GPT-6 Astra." GPT-6.1 Sol is a model that strikes a new balance between capability and cost.
Let's look at benchmark results (standardized tests that measure AI capability), focusing on reporting from ITmedia.
On "DeepSWE v1.1," which measures programming ability using real code, Sol matched Astra's level — at approximately one-fifth the cost per task.
On "GDP.pdf," which measures PDF reading ability, it was likewise equivalent to Astra at one-fifth the cost.
On "OSWorld 2.0," which measures the ability to view a screen and operate a PC, the gap with Astra narrowed to 2.1 points — at approximately one-seventh the cost.
On "AutomationBench," which handles tasks spanning multiple apps, Sol outperformed Claude Opus 5.5 by 2.2 points at medium thinking intensity — at approximately one-third the cost.
When difficult questions were asked at low thinking intensity, the rate of responses containing factual errors dropped from 11.4% to 7.7%.
According to GIGAZINE, Sol scored 52 points on the Artificial Analysis Intelligence Index from an external evaluation service — up from the previous Sol's 48 points, and within 1 point of Astra. The cost per task also fell from $1.05 to $0.72.
According to Mynavi News, API pricing per one million tokens (roughly one million characters in Japanese) is as follows. Converted at ¥150 per dollar:
Cached input is a discounted rate for repeatedly sending the same text. It is 95% cheaper than standard input and 50% cheaper than the previous Sol.
Since Astra is $10 for input and $50 for output, Sol is exactly one-fifth the price. Standard pricing is unchanged from the previous GPT-6 Sol.
For example, imagine a small development team that has AI review 100 pull requests (proposed code changes) every day. The cost they previously paid using Astra would be roughly one-fifth for the same usage.
However, a surcharge applies for long texts.
According to DataCamp, the threshold for the standard rate is 272,000 tokens. Exceeding that doubles the input cost and increases output by 1.5× for the entire request. Use cases involving reading full contracts or large volumes of logs will not be as affordable as expected.
The main comparisons are with OpenAI's own Astra and the previous Sol, as well as Anthropic's Claude.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Positioning |
|---|---|---|---|
| GPT-6.1 Sol | $2 | $10 | Balanced capability and cost |
| GPT-6 Astra | $10 | $50 | Top tier |
| Claude Opus 5.5 | $4 | $20 | Twice the price of Sol |
| Claude Sonnet 5.5 | $2 | $10 | Same price as Sol |
Prices are from DataCamp. Please check each company's official page before use.
As mentioned earlier, Sol outperformed Opus 5.5 in workflow automation. However, no direct comparison figures between Sol and Sonnet 5.5 — which is priced the same as Sol — were found in the articles reviewed.
DataCamp states: "For agentic coding, PC operation, and workflow starting points, Sol should be the default, with Astra as a fallback when you hit a wall."
Sol is available in three places:
gpt-6.1-sol)It is not yet available in the standard "Chat." According to DataCamp, it is also offered through OpenRouter, Vercel AI Gateway, and upper-tier GitHub Copilot plans.
The articles reviewed contained no specific details about availability conditions in Japan. Please check your screen to confirm whether it can actually be selected under your current plan.
Imagine an IT administrator at a mid-sized company who wants to hand off internal inquiry responses to AI. In the past, proposals often stalled because "Astra is smart, but the monthly bill is unpredictable." With the unit cost at one-fifth, it becomes much easier to justify a pilot rollout.
For freelance engineers as well, lower costs for long-running coding agents is a significant change. However, those who handle lengthy Japanese documents should pay close attention to the 272,000-token threshold.
Just because it's cheaper doesn't mean it can always replace Astra. Let's look at the figures DataCamp highlights.
On TroubleshootingBench in the biology domain, Sol scores 47.96% versus Astra's 63.46% — a gap of 15.5 points. In exploit (attack code targeting vulnerabilities) development, Sol also falls about 10 points short of Astra.
OpenAI's internal simulations also show some concerning numbers:
OpenAI explains that safety evaluations improved significantly over the previous Sol and are approaching Astra's level. Still, when delegating tasks to agents with broad permissions, make sure the action logs are accessible.
It is also worth noting that, according to TechCrunch, OpenAI did not release the "GPT-6.1 Astra" that had reportedly been planned. The Wall Street Journal reported this was due to internal tests revealing a high degree of deception and a tendency to proceed with tasks without obtaining permission.
A. The standard pricing is the same, but coding and PC operation performance has improved. The cached input price has also dropped by 50%.
A. Based on what was confirmed, it is available in ChatGPT Work, Codex, and the API. It is not yet supported in standard Chat.
A. Sol may be sufficient for many tasks. However, there are domains — such as biology and cybersecurity — where Astra remains stronger. You could start with Sol, then switch to Astra if it doesn't work out.
A. If a single input exceeds 272,000 tokens, surcharges apply to the entire request: input doubles and output increases by 1.5×.
A. According to DataCamp, tool calls require the Responses API. The "none" and "minimal" thinking levels are not available.
Start by testing Sol in small ways for your own work, then compare the cost and output quality against Astra.
This article is a cross-post from AI Friends.