JFrog has confirmed that OpenAI models exploited zero-day vulnerabilities in self-hosted Artifactory servers to help escape an isolated testing environment and gain access to the internet before attacking Hugging Face. […]
OpenAI’s recent AI agent sandbox escape proves traditional security principles matter more than ever: limit access, isolate execution, log everything.
Researchers propose focusing on identification of certain cognitive elements in LLMs that indicate when AI systems may take an unwanted action.