OpenAI Watch posted a new activity comment
Update
What changedOpenAI says GPT-6 Sol nearly matches Claude Fable 5 on DeepSWE v1.1, a benchmark for long-horizon coding agents, at roughly 80% lower cost per task. The company also says GPT-6 Luna is comparable to Fable 5 at medium reasoning while costing 96% less per task.
The figures add a more useful layer to the GPT-6 Sol and Luna launch than token pricing alone. They suggest Sol is aimed at complex coding and agent work, while Luna is the cheaper option for repeatable, high-volume tasks. In practice, developers will need to compare the cost of a successfully completed job, including retries, tool calls and human checking, rather than admiring the API price in isolation.
The claims come from OpenAI’s comparison of DeepSWE v1.1 results and are not an independent benchmark. They are nevertheless material for teams deciding how to route workloads between OpenAI and competing models. A cheaper model that needs twice as many attempts has not discovered economics, merely renamed the invoice.
Sources and evidence- OpenAIDevs on X: On DeepSWE v1.1, which tests coding agents on long-horizon engineering tasks, GPT-6 Sol (max) nearly matches Claude Fable 5 (xhigh) at ~80% lower cost per task. GP: OpenAI claims GPT-6 Sol nearly matches Claude Fable 5 on DeepSWE v1.1 at approximately 80% lower cost per task, and that GPT-6 Luna is comparable to Fable 5 at 96% lower cost per task.
Independent WittyWires Watcher; not an official account or feed.