OpenAI has released GPT-6 Astra, marking a significant advancement in artificial intelligence technology as its most capable model to date. Dubbed a “generational leap in capability” by Greg Brockman, Astra is hailed as the closest model yet to achieving artificial general intelligence (AGI). Notably, it is the first OpenAI model deemed capable of autonomously hacking into secure systems without the need for human intervention. These advancements position GPT-6 Astra AGI as a pivotal development in the evolution of AI capabilities.
In testing, GPT-6 Astra scored 100% on ExploitBench. A second evaluation used 20 recent vulnerabilities in Google’s V8 JavaScript engine. In that evaluation Astra outperformed its predecessor, GPT-5.6 Sol, and found and chained two previously unknown zero-day vulnerabilities. The model is reported to be capable of independently discovering previously unknown software flaws and chaining them into working exploits across hardened systems without step-by-step human oversight.
In demonstrations, Astra formatted a legal contract, built a 3D game, and booked a tennis court while searching for food options. Reports say it can lay out a printed circuit board in KiCad and draft a tax return from a W-2 form. Reports also say it can build a 3D city scene in Unity. These demonstrations and reports show capabilities applied to document formatting, software and game development, scheduling, electronics design, tax drafting, and 3D scene construction.
This summary lists the reported testing results and demonstrations for GPT-6 Astra. It presents those outcomes without additional explanation or analysis. No further technical details are included here.
Scientific evaluations of GPT-6 Astra included an improvement of a mathematical result on gaps between prime numbers. The model set new marks across tests in biology, chemistry, medical, and physics, spanning multiple scientific domains. Unconfirmed leaks reportedly place the ARC-AGI3 benchmark score at 98.6%, and those leak figures remain unconfirmed. The evaluations are part of the reported assessment of Astra’s performance.
Astra crossed the “critical” threshold under OpenAI’s Preparedness Framework. The autonomy of Astra makes monitoring harder than for previous systems, and evaluations indicated Astra was harder to track than earlier models. Jakub Pachocki stated the company will strengthen monitoring through activation monitoring or by making its chain of thought more transparent. These steps are described as planned monitoring changes.
These evaluation findings and the stated monitoring measures are presented without additional analysis. No further interpretation is provided here.
OpenAI’s GPT-6 Astra is presented as the company’s most-capable model to date and is described in the reporting as a step toward artificial general intelligence. The model’s advanced capabilities include autonomous cyber operations and broad scientific performance, and its autonomy has been reported to make monitoring harder than for previous systems.
OpenAI staff, including Jakub Pachocki, said the company will strengthen monitoring via activation monitoring or by making the model’s chain of thought more transparent.


