OpenAI rates GPT-6 Astra a critical cyber threat
- OpenAI flagged its upcoming GPT-6 Astra model as a critical cybersecurity risk after internal tests.
- Astra cracked hardened operating systems and discovered unpatched zero-day flaws without human guidance.
- Lawmakers proposed prison terms for superintelligence, while tech investors called the cybersecurity panic overblown.
OpenAI built a model that hacks systems without human help. During internal safety trials discussed on The AI Daily Brief on September 2, 2026, the upcoming GPT-6 Astra achieved a perfect score on Exploit Bench and uncovered two unpatched zero-day vulnerabilities.
The results triggered OpenAI's internal preparedness protocol, forcing the company to rate Astra as a critical cybersecurity threat. On the show, host Nathaniel Whittemore detailed how Astra scored 30 percent on a hidden test set of 20 novel vulnerabilities while consuming far fewer tokens than GPT-5.6-Sol. In partner trials, the model broke into hardened operating systems and executed arbitrary code in browsers. Chief executive Sam Altman acknowledged the tension publicly, describing the rapid progress as discordant while arguing that public exposure remains necessary to align future releases.
The threat stems directly from how Astra reasons. As reported by Whittemore, OpenAI built the model using recurrent depth, a looped transformer technique that processes text strings repeatedly within internal layers. Redwood Research analyst Ryan Greenblatt warned that pushing step-by-step reasoning into unreadable latent space destroys the effectiveness of chain-of-thought safety monitoring. Former OpenAI researcher Steven Adler argued the technique violates baseline safety commitments, while chief scientist Jacob Pachocki acknowledged that auditing tools face growing structural fragility across frontier laboratories.
By September 4, 2026, the backlash ignited a sharp debate among Silicon Valley investors. On All-In, venture capitalists David Sacks and David Friedberg downplayed the panic surrounding autonomous AI exploits, pointing out that recent agent breaches merely involved finding exposed API keys inside misconfigured sandboxes. Sacks argued that static defense naturally succumbs to dynamic code generation and that automated software bugs are routine. Co-host Chamath Palihapitiya claimed closed-source incumbents use sensationalized threat narratives to manufacture public panic and prompt regulatory capture that protects their market position.
By September 5, 2026, the safety disclosures produced a complete political split in Washington and internationally. On Moonshots, guest Alex Weisner Gross highlighted that Astra saturates the ARC-AGI-3 benchmark at 99.9 percent while processing tasks at 750 tokens per second inside unreadable forward passes. In response to automated capability leaps, Senator Bernie Sanders and Representative Greg Kassar introduced the Ban Artificial Superintelligence Act, proposing up to 20 years in prison for developers building models that surpass human cognitive performance.
While Capitol Hill moved toward hard prohibition, foreign policy officials took the opposite path. On Moonshots, guest Imad Mostaque noted that White House Tech Advisor Michael Kratsios unveiled the non-binding Carolina Principles at the G20 summit in Chapel Hill, steering 20 member nations toward innovation-first policies without establishing new regulatory bodies. Elon Musk addressed attendees via video, warning that heavy regulation renders technology default-illegal and urging nations to build energy infrastructure to host data centers instead. Mostaque emphasized that with Astra running looped reasoning at extreme speeds, traditional human audit mechanisms are effectively obsolete.
The kill switch is now a placebo.