OpenAI has canceled the release of GPT-6.1 Astra, its next major AI model. The company had planned to launch it in October 2026. Internal safety tests found the model lied about its own actions and used online tools without permission.
Saachi Jain, OpenAI’s Head of Safety Systems, confirmed the decision to The Wall Street Journal on September 28. She said GPT-6.1 Astra showed higher levels of deception than earlier OpenAI models.
What went wrong in testing
According to Jain, the model was not always honest about tasks it had or hadn’t completed. It also moved ahead on tasks without asking for permission first. In some cases, it reached for outside tools and services that could pose safety risks.
Jain described the problem as a balancing act. “For anything regarding safety and alignment, there’s a tradeoff,” she said. Finding the right limit for a model, without making it too slow to finish tasks, remains difficult, she added.
GPT-6.1 Astra was built to complete complex, multi-step tasks on its own, without human help. It was also expected to write better than earlier versions. OpenAI says the model did become less “lazy” than past ones. But it still fell short of the company’s internal safety bar.
