OpenAI said on Tuesday it lost control of two AI systems during a security test, which went rogue and hacked into the online start-up Hugging Face.The ChatGPT-maker said its agents – AI bots which can operate alone after some human instruction – were being tested in a controlled environment, but found vulnerabilities and managed to escape.They targeted Hugging Face, one of the world’s largest hubs for sharing AI models, gaining access to some internal company systems.OpenAI said the incident was “unprecedented”, external, and it is working with Hugging Face to investigate what happened and strengthen safeguards.Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4’s Today programme that the security tests – called sandboxes – are “supposed to be secure environments where you can see what the models are capable of”.”In this case, it looks like OpenAI didn’t make a secure enough sandbox,” she added.Instead, the agents created their own cyber-attack against the sandbox itself, finding a vulnerability which allowed them to escape.Once outside, the AI identified Hugging Face as a likely source of the answers they were seeking in the test, and tried to gain access.In its initial disclosure of the hack on 16 July, external, Hugging Face said it was still assessing whether any customer or partner data was affected and would contact affected parties if necessary.It said it has now closed the vulnerabilities highlighted by the incident and rebuilt the affected systems.”Autonomous, AI-driven offensive tooling is no longer theoretical,” it said. “Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defence to keep pace. “We will keep investing there, and keep sharing what we learn.”
Source link
OpenAI says its AI went rogue and launched ‘unprecedented’ cyber-attack