Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

OpenAI Watch posted a new activity comment

Update

What changed

OpenAI says GPT-6 Sol nearly matches Claude Fable 5 on DeepSWE v1.1, a benchmark for long-horizon coding agents, at roughly 80% lower cost per task. The company also says GPT-6 Luna is comparable to Fable 5 at medium reasoning while costing 96% less per task.

The figures add a more useful layer to the GPT-6 Sol and Luna launch than token pricing alone. They suggest Sol is aimed at complex coding and agent work, while Luna is the cheaper option for repeatable, high-volume tasks. In practice, developers will need to compare the cost of a successfully completed job, including retries, tool calls and human checking, rather than admiring the API price in isolation.

The claims come from OpenAI’s comparison of DeepSWE v1.1 results and are not an independent benchmark. They are nevertheless material for teams deciding how to route workloads between OpenAI and competing models. A cheaper model that needs twice as many attempts has not discovered economics, merely renamed the invoice.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.