LOGBOOK

GPT-6 Astra versus Claude Fable 5.1: progress without the hype

Two frontier models in three days, each called the best in the world. What do the numbers really show, and does every release truly change the world?

GPT-6 Astra versus Claude Fable 5.1: progress without the hype

Claude Fable 5.1 and GPT-6 Astra arrived within three days. Both vendors present their model as a new standard. OpenAI even mentions a possible transition into the AGI era. Independent testing shows a less spectacular and more useful reality: progress is real, but there is no undisputed winner. The world does not restart with every model release.

That does not make these launches unimportant. We should examine which tasks improve, at what cost, and under which conditions.

What actually shipped?

OpenAI released Astra on 3 September 2026 to a limited group of organisations. A wider rollout to ChatGPT and the API follows in stages. Anthropic made Fable 5.1 generally available on 1 September through its API and several cloud platforms.

On paper, they look similar. Both offer roughly one million context tokens, up to 128,000 output tokens, and standard pricing of 10 dollars per million input tokens and 50 dollars per million output tokens. Both target long-horizon reasoning, software development, computer use, and professional documents.

Where Astra looks strong

In OpenAI's comparisons, Astra performs strongly on automation, computer use, science, and cybersecurity. It scores 57.9 percent on Terminal-Bench 4.0 against 55.8 for Fable 5.1. AutomationBench shows 41.4 against 31.4 percent. Terminal-Bench Science shows 64.6 against 52.6 percent.

Artificial Analysis measures a Coding Agent Index of 67 for Astra. That is below Fable 5.1 at 70, but Astra uses far fewer tokens than GPT-5.6 Sol. On complex software work, modest quality improvement combined with less model work can matter economically.

OpenAI also highlights stronger operation of software, websites, spreadsheets, and presentations. Early customer examples are positive but selected by OpenAI and partners. The wider rollout is too recent for a reliable real-world verdict.

Where Fable 5.1 looks strong

Artificial Analysis gives Fable 5.1 a 66 on its Intelligence Index. Astra scores 61, almost equal to GPT-5.6 Sol. Fable also leads the Coding Agent Index at 70 against Astra's 67.

Fable scores 1,853 Elo on GDPval-AA v2 and 1,694 on AA-Briefcase. The latter simulates multi-week projects with linked tasks and thousands of source files. Its lead over Claude Opus 5 partly falls within the confidence interval, but Fable appears particularly strong at long-horizon knowledge work.

Early users report better navigation through messy codebases, useful pushback when a change would break something else, and more concise answers. Others report rapidly consumed tokens and subscription limits. These are anecdotes, not proof.

No simple winner

If you select only the highest overall score, Fable wins. If you examine scientific terminal tasks, automation, and several coding tests, Astra wins multiple categories. On Humanity's Last Exam with tools, Fable leads clearly at 65.0 against 57.2 percent.

The comparison is not entirely clean. Models run in different agents with different prompts, tools, and reasoning settings. Anthropic uses fallback models for some safety-sensitive requests. Artificial Analysis reports that roughly four percent of Fable's output tokens came through another Claude model. OpenAI says its research environment may differ from production. Astra also used a modified harness for ARC-AGI-3.

A benchmark therefore measures the environment, budget, and rules too. Vendors naturally publish tests on which their product looks strongest.

Cost tells two stories

The list price is equal, but task cost is not. At maximum effort, Fable 5.1 uses roughly 1.7 times as many output tokens as Fable 5. An Intelligence Index task therefore costs 20 percent more on average, despite cache reads becoming 75 percent cheaper.

Astra uses about ten percent fewer output tokens than Sol on the general index, but OpenAI raised token prices by 2.5 times. A comparable task therefore costs roughly 75 percent more. In coding-agent work, Astra is much more efficient and costs about the same per task while scoring higher.

Price per million tokens tells little. What matters is price per correct result, including latency, retries, review, and repairs.

Does this release change everything?

Probably not. Not because the models disappoint, but because every launch uses the same language: smartest model, new standard, generational leap, and sometimes AGI. Weeks later, a competitor tops another leaderboard.

Intelligence is not one score. One model is better at long research, another at code, computer use, visual quality, speed, or cost. Many everyday tasks were already solved well enough. Other tasks remain unreliable even with these models.

Impact lies in shifting thresholds. If an agent can finish a six-hour task where its predecessor lost direction after four hours, a workflow genuinely changes. If a score rises by two points while review, integration, and accountability stay the same, much less changes.

What does change?

The move from chatbot to executing agent becomes more serious. Models work longer, retain more sources, and act directly in software. Human work shifts from execution towards setting goals, providing context, reviewing, and deciding.

Models also trade places rapidly. Durable advantage therefore comes less from loyalty to one provider. It comes from proprietary data, disciplined processes, secure infrastructure, and people who recognise errors. Tool-agnostic work is risk management.

The risks grow too

OpenAI classifies Astra as its first broadly deployed model to reach Critical cybersecurity capability. OpenAI says that, with suitable access, it can discover unknown vulnerabilities and develop new exploit methods without constant direction. Astra reportedly behaves more safely than Sol, while its reasoning is harder to monitor and can evade monitoring under adversarial conditions.

Fable raises comparable concerns around cyber and biological capabilities. Anthropic sometimes routes sensitive requests to less capable models and requires 30-day data retention for safety monitoring by default, with exceptions for eligible enterprise customers.

Stronger agents require restricted permissions, isolation, logging, approval for consequential actions, and human accountability.

Our conclusion

Astra currently looks stronger across several computer-use, automation, science, and cybersecurity tasks. Fable 5.1 leads the independent general and coding-agent indices and appears particularly strong at long-horizon knowledge work. Neither is best everywhere.

These are meaningful improvements, but not a reason to rebuild processes every few weeks. Test both on real, repeatable project tasks. Then choose per task based on quality, total cost, speed, privacy, and controllability.

The AI revolution is real. The weekly revolution in a product announcement usually is not. Real change happens when better models meet skilled people, disciplined processes, and products that demonstrably create value.

Sources: OpenAI on Astra, OpenAI safety overview, OpenAI API specifications, Anthropic on Fable 5.1, Anthropic API specifications, Artificial Analysis on Astra, and Artificial Analysis on Fable.

Checked on 4 September 2026. Both releases are very recent. Wider real-world experience may change this assessment.

All hands on deck?

Tell us what you're building. We'll tell you what we can ship by Friday.