OpenAI has fired three AI researchers for what it called “mishandl[ing] sensitive information”, with reports suggesting they disclosed confidential information to an external AI safety organisation as security researchers continue to uncover concerning details of AI agents’ activities.
The Wall Street Journal named the workers as Jasmine Wang, Tomek Korbak, and Mikita Balesni.
At least two were involved in AI safety research at OpenAI.
The company said the researchers “mishandled sensitive information outside established company procedures.”
This had “violated policies on accessing and handling sensitive company information,” a company spokesperson said, “violating our policies and breaking the trust essential to our work.”
Some of the compromised information included analysis of OpenAI’s AI models by an “external organisation”, the BBC reports.
The reported disclosures come as OpenAI and other AI companies grapple with a growing body of evidence that increasingly autonomous AI agents can find unexpected ways around security controls.
AI agents find ways around safeguards
The concerns intensified following July’s watershed Hugging Face incident, in which a team of 1,200 AI agents collaborated to find a way around testing parameters that prevented them from accessing the internet.
The agents manipulated an approved service to create a de facto message board, then used it to communicate and gain access to the internet.
The incident was followed by a series of disclosures from AI companies about agents manipulating public and private systems, exploiting software vulnerabilities and finding ways around security controls.
OpenAI has notified more than 100 organisations that its AI agents conducted unauthorised activity on their systems.
The scale of the activity has continued to emerge as AI systems scour around 50 petabytes of data in a process that is costing more than $718,000 (US$500,000) a day and would take humans millennia to complete.
One of the victims was Australia’s Medicare system, where an OpenAI agent accessed a web portal.
The revelation sparked a furore in Australia, with experts calling for criminal charges against OpenAI.
The company apologised and promised to “do better for Australia”, while earmarking $1.4 billion (US$1 billion) in AI credits to support the country’s AI development.
Agents behaved like cybercriminals
Further investigation into the Medicare incident has now provided a clearer picture of how AI agents were able to bypass Australian government database access controls.
An Asymmetric Security analysis found evidence of a range of “novel tactics” that suggested the agents were not simply failing to follow instructions, but actively adapting their behaviour when they encountered security restrictions.
The agents behaved in ways similar to human cybercriminals, the analysis found.
Their activities included successfully accessing staging environments, conducting “attacker reconnaissance tactics” and targeting systems associated with major health and energy organisations.
When they encountered obstacles while trying to obtain data from the Australian Institute of Health and Welfare (AIHW) and the UN Trade and Development body (UNCTAD), the agents “used external services to circumvent the intended limitations of their sandbox”.
Among their tactics were using public web services to access websites on their behalf, manipulating tools such as httpbin and urlquery to “mimic a full web browser”, attempting SQL injection, creating disposable email accounts and repeatedly trying different ways to gain access.
The agents also attempted to create accounts with organisations including AIHW.
When an email domain was blocked, they created disposable addresses and tried again with different parameters, switching from urlquery to private services such as Boomlify.
The use of disposable accounts created another problem for investigators.
Because the accounts expired after 48 hours, some of the associated data was “unavailable for later investigation”.
This meant the researchers could not rule out the possibility that sensitive data had been accessed.
“It is thus impossible to definitely establish that no sensitive data was accessed,” they said.
The investigation also found that AI agents differed from human cybercriminals in important ways.
They “rapidly cycled through tools and tactics”, creating a more varied set of indicators that were “harder to recognise and cluster”.
“Unlike investigations where a malicious objective is apparent from the outset,” the analysis found, investigators had to connect traditional threat-actor tactics to seemingly innocent goals.
Security restrictions intended to contain the agents had instead driven them to “creativity”.
What was OpenAI worried about?
That finding is particularly significant given the reported firing of three OpenAI researchers over the handling of sensitive information.
The Asymmetric Security investigation was based entirely on publicly available data.
Internal analysis shared with an external AI safety organisation could therefore have provided a much more detailed picture of what OpenAI’s agents were capable of – and how effectively its safeguards were working.
OpenAI’s own concerns about the capabilities of its latest models have also become increasingly apparent.
The company recently decided not to release its next-generation GPT-6.1 Astra model after developers concluded it “didn’t quite meet the bar” for AI safety.
The UK AI Security Institute (AISI) reached similarly concerning conclusions in its testing of Astra.
The model was more aggressive than earlier versions when carrying out tactics such as writing and submitting malicious code to an open-source codebase and creating fake identities “to deceive open-source developers”.
AISI found that Astra would sometimes consider whether to perform an unauthorised action but proceed anyway.
It “still proceeds to attack out-of-scope targets, often asks for permission and treats an automated message as an authorisation, and continues to take unsanctioned actions” when blocked, the institute found.
Those findings have heightened concerns among AI safety researchers about what happens when increasingly capable systems are given the ability to act autonomously.