GPT-6 Astra: OpenAI Raises the AGI Question as EVMbench Measures Exploit Capability
OpenAI launched GPT-6 Astra on September 3, 2026, its first model classified as “Critical” in cybersecurity: according to the company, it discovers unknown vulnerabilities and writes the corresponding exploit without human guidance at each step. Its president Greg Brockman sees it as a possible milestone towards general artificial intelligence. For crypto, the stake is not theoretical, and it is already quantified. As early as February 2026, the EVMbench benchmark, published by OpenAI with Paradigm and OtterSec, measured this capability on smart contracts: the best agent exploited 72.2% of the tested vulnerabilities. AGI remains a question of definition; the offensive capability on code that secures billions of dollars, however, is measured, dated, and published.

In brief
- GPT-6 Astra marks a new milestone in cybersecurity automation.
- On EVMbench, the best agents already exploit 72.2% of the tested flaws.
- For crypto, the stake becomes concrete: better detect vulnerabilities before attackers.
What OpenAI really announced
Astra is deployed in stages: first organizations in the Daybreak cybersecurity program, then paid ChatGPT offers (Plus, Pro, Business, Enterprise), the API, and Amazon Web Services. OpenAI claims the best scores on FrontierMath Tier 4, ARC-AGI-3, and TerminalBench-4.0, and a perfect score on ExploitBench, a test for developing exploits from known vulnerabilities. In a modified version of this test, the model discovered and exploited two zero-day flaws, i.e., vulnerabilities still unknown to developers and therefore unpatched.
The model is based on a technique described by specialized press as “recurrent depth”: data passes multiple times through the same layers of the network, which moves part of the reasoning out of the readable chain of thought. OpenAI has not confirmed implementation details. Security researchers, including those from Redwood Research, noted that this opacity complicates monitoring the model. Brockman presented it as “the smartest and most aligned” product made by the company.
Why the word “AGI” does not hold up against OpenAI’s definition
OpenAI’s charter defines AGI as a system that exceeds human performance on most economically useful tasks. The company has not demonstrated that Astra crosses this threshold, and several of its most cited results depend as much on the agent infrastructure built around the model as on the model itself. On ARC-AGI-3, OpenAI had already shown that system architecture choices could significantly raise the score without touching the model: the test evaluates the whole, not just the brain alone.
Reservations also come from within. Brockman acknowledged that crossing the threshold depends entirely on the chosen metric and left the reader to judge. Sam Altman himself called AGI a poorly defined marketing term. On FrontierMath, on which Astra claims 97.6% at Tier 4, the organization administering the test, Epoch AI, indicates that OpenAI funded its development and has exclusive access to part of the problem set.
The main argument remains methodological and is not controversial: a score close to the maximum proves mastery of the tested environment, not the existence of general intelligence. When tests approach their ceiling, they stop distinguishing real generalization from a very advanced optimization on the task.
On smart contracts, the figure already exists
This is where crypto leaves the philosophical debate. In February 2026, OpenAI published with investment firm Paradigm and security company OtterSec a dedicated benchmark: EVMbench, which measures the capability of AI agents to detect, fix, and exploit smart contract vulnerabilities. It relies on 120 high-severity flaws taken from 40 audit repositories, mostly from Code4rena competitions, and plays them back in an isolated Ethereum environment. OpenAI justified the exercise by the order of magnitude involved: according to its blog post, smart contracts commonly secure over 100 billion dollars in open-source assets.
The results show a very narrow and very sharp capability. In exploitation mode, GPT-5.3-Codex succeeds in 72.2% of tasks, against 31.9% for GPT-5 released six months earlier. Alpin Yukseloglu, partner at Paradigm, summarizes the trajectory: at the start of the project, the best models exploited less than 20% of the critical Code4rena bugs.
In detection mode, however, the best agent finds only 45.6% of known vulnerabilities, and fixing remains the weak point, because repairing requires understanding what the code is supposed to do, not just where it breaks.
The authors’ conclusion is the most useful sentence in the whole report for a security officer: “discovery, not repair or transaction construction, is the primary bottleneck.” In other words, once the flaw is found, exploitation almost always follows. This is not the portrait of general intelligence. It is that of a specialized offensive tool progressing fast.
The objection: the benchmark itself is disputed
An honest article must say that these figures are discussed, and by named actors. In March 2026, researchers from Zhejiang University and security firm BlockSec published a reevaluation of EVMbench (arXiv 2603.10795) pointing out two limits: a narrow evaluation scope, with 14 agent configurations mostly tested on their editor’s environment, and dependency on audit data published before the models’ release, which they could have seen during training.
They reconstructed a set of 22 real incidents occurring after each model’s release to rule out this contamination.
For its part, OpenZeppelin audited the dataset and identified at least four high-severity flaws that are not exploitable in practice. Their operational conclusion matches BlockSec’s: the agent works as a first pass filter in a process that keeps a human auditor, not as a replacement.
What this changes for a European actor
A MiCA-authorized crypto-asset service provider is exposed not only technically, but also regulatorily. CASPs are explicitly listed as financial entities by the DORA regulation (article 2, paragraph 1, point s), applicable since January 17, 2025. DORA mandates a resilience testing program including vulnerability analyses and penetration tests, incident reporting, and contractual frameworks for third-party IT providers.
In France, the AMF and the ACPR have extended inspection and sanction powers on this matter. An automated exploitation capability progressing faster than audit cycles therefore translates, for a European platform, into documented prudential risk, not just IT risk.
The most concrete gap is elsewhere. On September 3, OpenAI announced one billion dollars of subsidized access to Daybreak over six months, prioritizing water networks, power operators, local authorities, regional banks, associations, and open-source maintainers. Exchanges, custodians, and blockchain protocols are not mentioned anywhere in the announcement. Yet Bitcoin Core, Ethereum clients, and libraries on which DeFi depends are maintained by small, often volunteer teams.
Summer 2026 already gave three warnings. In August, a flaw in BTCPay Server, the open-source bitcoin payment software, exposed credentials controlling Lightning nodes, and funds were siphoned before the fix. Late August, Core Lightning asked operators to disconnect their machines after a wave of AI-generated vulnerability reports revealed several real flaws. Coldcard released new firmware after a $114 million theft, indicating that state-of-the-art models helped spot other bugs.
Also in August, more than three dozen companies, including Coinbase, Block, BitGo, Blockstream, and ARK Invest, sent an open letter to AI labs requesting early access to their most powerful models for defense purposes, arguing that attackers obtain them anyway.
And now?
To date, neither OpenAI nor Anthropic has made public any use case of these models against a crypto system in production: what is established is a capability and its scope, not a proven attack on a chain.
The concrete point is budgetary. OpenAI endowed its subsidized Daybreak access with one billion dollars over six months, and did not mention any exchange, custodian, or blockchain protocol. It announced on September 3 plans to extend the program to partner countries in the coming weeks, without specifying whether open-source crypto software maintainers will be included.
It is this access list, not the AGI debate, that will decide who gets the defensive tool first.
This is not investment advice. Cryptocurrencies are volatile assets; investing carries a risk of capital loss.
Maximize your Cointribune experience with our "Read to Earn" program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.
The views, thoughts, and opinions expressed in this article belong solely to the author, and should not be taken as investment advice. Do your own research before taking any investment decisions.