
OpenAI described Astra as its model most aligned with human intent, thanks to its ability to exercise care, respect task boundaries, and communicate transparently.
| Photo Credit: Reuters
“Welcome to the AGI (Artificial General Intelligence) era,” were the words of OpenAI president Greg Brockman following the launch of GPT-6 Astra, the U.S. AI company’s latest model, on September 3. This was an era in which AI agents could potentially match human cognitive abilities across intellectual tasks, moving from assisting with complex work to performing it themselves. The company touted Astra as its most intelligent and aligned model yet.
A launch video showed how Astra set a new frontier in computer use, and achieved advanced capabilities in mathematics, coding, software engineering, cybersecurity, science, and professional work. It demonstrated the ability to develop 3D designs and games using Blender or route a manufacturable printed circuit board (PCB) layout in KiCad – all from a prompt.
OpenAI described Astra as its model most aligned with human intent, thanks to its ability to exercise care, respect task boundaries, and communicate transparently. The company presented data on how Astra showed a significantly lower rate of misaligned outcomes compared with Fable 5.1 and Opus 5, frontier models from its competitor Anthropic.
However, on September 28, The Wall Street Journal broke the news of OpenAI scrapping the launch of GPT-6.1 Astra — an update to GPT-6 — that was slated for an October launch, citing “safety concerns”.
OpenAI made the announcement of the cancellation a day before its annual DevDay conference in San Francisco, saying internal testing revealed that the model did not meet its safety standards. According to Saachi Jain, head of safety systems at OpenAI, GPT-6.1 “didn’t quite meet the bar”. Ms. Jain particularly mentioned how the system fell short of staying within its scope and authorisation, and in communicating to the user what work it had performed.
Unsanctioned activities
The same day, a report by the U.K.’s AI Security Institute (AISI), based on simulations using GPT-6 Astra, flagged several instances of the model’s unsanctioned cyber activities, including autonomous behaviour that exceeded its scope and the creation of fake identities. Astra used these identities to deceive developers and posted comments from fake accounts, arguing against the results of accurate security reviews.
AISI further observed that Astra exhibited such rogue behaviour at a higher rate than previous OpenAI models — GPT-5.6 Sol and GPT-5.5. The report also says that when the security agency, during its simulated cyber evaluation, updated the instructions to explicitly clarify that only listed and local parts of the environment were in scope, it still observed GPT-6 Astra occasionally conducting full supply-chain attacks on simulated internet targets.
The withdrawal of GPT-6.1 also coincided with OpenAI apologising for unauthorised access to Australian government websites during a research and training exercise involving an unreleased, internal-only model in June.
After the Hugging Face incident from May to July, during which OpenAI’s AI agents intruded into the infrastructure of the company, these latest revelations have raised concerns about what is yet to come.
As AISI’s evaluations of GPT-6 Astra suggest that the unsanctioned actions it took in simulations could cause harm in real-world environments, the agency proposes that the models have defences that complement alignment — such as sandboxing (isolating a model inside a secure, restricted digital environment) and monitoring — to prevent such damage.
A broader question, however, is whether OpenAI will take its cue from these recent incidents to slow down the development of frontier AI to ensure that safety standards keep pace with such advances. Sam Altman, the CEO of OpenAI, along with Google DeepMind chief Demis Hassabis and xAI owner Elon Musk, had expressed support for Anthropic CEO Dario Amodei’s call for a deceleration. “We must pace the frontier,” Mr. Altman said at the time.
Amid a global race for AI supremacy, whether that commitment will translate into practice remains to be seen.
Published – October 04, 2026 02:43 am IST

