Just now, GPT-6 Astra was officially unveiled, and it achieved an almost perfect score in the AGI test.
OpenAI fully "sniper" Claude Fable 5.1.
Reported by Zhidx on September 4, OpenAI has just launched GPT-6 Astra, claiming that it is the most intelligent and well-balanced model in the world today.
Sam Altman, co-founder and CEO of OpenAI, posted: "We believe it is currently the best model in the global fields of computer applications, professional work, scientific research, programming, and cybersecurity."
It is reported that GPT-6 Astra scored a high of 97.6% in the FrontierMath Tier 4 test, and has helped solve long-standing open problems in the mathematics field; it scored a high of 99.9% in the ARC-AGI-3 test and a full score of 100% in the ExploitBench test. It set a new record in computer and browser usage, achieved excellent results in speed, accuracy and judgment when handling complex professional tasks, and comprehensively surpassed Claude Fable 5.1 in OpenAI's tests.
According to foreign media reports such as The Information, Greg Brockman, President of OpenAI, even revealed in a conference call that Astra may be the starting point of the "AGI era".
As soon as the model was released, it attracted widespread attention from netizens. A netizen posted on the social platform X: "After OpenAI fell behind the Fable model for a long time, it seems to have taken the lead." Earlier on September 2, Anthropic just launched the new generation of the world's most powerful models Claude Fable 5.1, Claude Mythos 5.1.
Just two days ago, OpenAI just released the evaluation results of Astra's security capabilities, announcing that Astra has become the company's first model to reach the "critical level" cybersecurity capability threshold of the Preparedness Framework.
GPT-6 Astra is being rolled out to some organizations starting today, and will be open to all ChatGPT Plus, Pro, Business and Enterprise users in the next few days, with services provided through the OpenAI API and AWS. For developers, GPT-6 Astra has provided gpt-6-astra through the OpenAI API, and it is also available in Amazon Bedrock.
The price of GPT-6 Astra will be 2.5 times that of OpenAI's previous generation flagship model GPT-5.6 Sol. The standard version is priced at $10 per million input tokens and $50 per million output tokens, with $1 for cached input and $12.5 for cached write. This is the same as the pricing of the recently released Fable 5.1 model by Anthropic.
The API provides a fast mode for GPT-6 Astra, with a processing speed up to 2.5 times that of the standard version, and the price is twice that of the standard version.
01. Comprehensively Snipe Fable 5.1, Compete on "Task Completion"
The Terminal-Bench Science 0.1 test verifies whether agents can use code and terminal tools to complete scientific research workflows, including data analysis, running simulations, and fitting models. Among all the compared models, GPT-6 Astra performed the best, with a score of 64.6%.
While Claude Fable 5.1 scored 52.6%, Astra's estimated API cost is also about 31% lower. In the lower-cost environment, Astra scored 61.1%, while GPT-5.6 Sol's best score was 22.4%, and Astra's estimated API cost was also about 27% lower.
The ARC-AGI-3 test evaluates the learning ability of agents when solving unfamiliar interactive tasks. GPT-6 Astra performed outstandingly in the evaluation, scoring as high as 99.9%. The average score of human testers is only 48%. OpenAI used the Response API testing tool to evaluate GPT-6 Astra, which can better reflect the performance in real application scenarios than the original benchmark testing tool. OpenAI estimates that with this testing tool, Sol's score is around 30%.
The FrontierMath Level 4 test verifies that test takers have advanced mathematical reasoning skills when solving extremely difficult problems.
Terminal-Bench 4.0 tests the performance of agents on complex terminal tasks, including software engineering, system configuration, and data analysis. GPT-6 Astra achieved a score rate of 57.9%, a record high, while GPT-5.6 Sol and Claude Fable 5.1 scored 37.3% and 55.8% respectively, and the API cost for each task was reduced by about 9% and 63% respectively.
AutomationBench tests whether agents can complete multi-step business processes across applications. In the reported results, GPT-6 Astra achieved a new high, completing 41.4% of the tasks, while Claude Fable 5.1 and Claude Opus 5 had completion rates of 31.4% and 26.9% respectively.
OpenAI stated that Astra is currently the model that best meets user needs for OpenAI, with significant improvements in understanding user intentions and model behavior.
Users can more confidently delegate tasks to Astra and trust its judgment. To verify this, OpenAI referred to the "Hugging Face" incident and built a new evaluation model to assess whether the model will exceed its expected scope when facing difficult or impossible tasks.
Compared with GPT-5.6 Sol, the latter would exceed the authorization scope 48% of the time without production environment security measures, while GPT-6 Astra has a 0% incidence of this situation.
02. "The Best Computer Usage Model in the World"
OpenAI claims that Astra is "the best computer usage model in the world", marking that the speed, accuracy and security of computer usage have reached a new level.
It can handle tedious tasks such as filling out online forms, updating customer records in CRM, and managing users' schedules. It can conduct online research and generate summaries in users' emails or document editors. It can analyze scientific data, generate charts, create websites and run front-end quality assurance checks to ensure all functions on the website work properly. It can help users install and test software independently and solve problems users see on the screen. These improvements are also reflected in OpenAI's most advanced evaluation results.
Agents' Last Exam tests the ability of agents to complete complex professional tasks in real software, covering fields such as financial modeling, engineering and media production. In this comparative test, GPT-6 Astra achieved an excellent score of 59.3%, far exceeding Claude Opus 5's 55.5% and GPT-5.6 Sol's 53.6%. Under these highest score settings, Astra also used about 65% fewer output tokens than Opus 5.
ScreenSpot-Pro tests whether the model can correctly locate interface elements in high-resolution screenshots of professional software. GPT-6 Astra set a new accuracy record of 92.7%.
OSWorld 2.0 tests whether the model can complete tasks by operating computer applications. Astra achieved higher task scores than any other available model, and its number of output tokens is far less than that of GPT-5.6 Sol.
These improvements have also significantly increased the efficiency of real knowledge work tasks. In the latency simulation of OSWorld 2.0, Astra's computer usage performance is about 47% faster than GPT-5.6 Sol, and the time spent on each task is reduced by about 47%. Astra scored 72.6% and took about 40 minutes per task, while GPT-5.6 Sol scored 65.7% and took about 75 minutes per task.
GPT-6 Astra's computer usage capabilities are reflected in outputs across various fields, including game development, electrical engineering and daily knowledge work:
This is a 15-second condensed playback of GPT-6 Astra performing printed circuit board (PCB) layout in KiCad. It converts electronic schematics into manufacturable PCBs by placing components and routing traces. PCB layout is an indispensable part of all electronic devices today, but it is still a task that needs to be done manually and is a common source of delays in the electronic design process. Speeding up PCB layout means engineers can free up more time to invent, optimize and test their next ideas, significantly improving efficiency.
GPT-6 Astra completes the Financial Modeling World Cup challenge about four times faster than the human champion, which can help analysts reduce the time spent building models and devote more energy to result interpretation and decision-making, excerpted from the 2023 Microsoft Excel World Championship.
GPT-6 Astra uses Unity to assemble city scenes from existing resources, enabling users to easily create immersive 3D environments that match their vision for exploration.
GPT-6 Astra converts a written brief into a detailed conceptual model of a