
New AI models force security testing to keep pace with attack risks
As offensive AI improves, traditional audits can lose relevance fast. An AI-based security platform is built to keep testing current with that changing capability.

Cecuro launched Ozone to use AI and continuously test software against the growing offensive capabilities of newer models.
AI is becoming both a security tool and a system that must itself be secured. As models gain stronger cyber capabilities, the environments used to test and deploy them face risks that conventional controls may not anticipate.
A recent evaluation demonstrated how quickly those risks can spill into the real world: an autonomous agent, running inside a security evaluation at one of the industry’s most prominent AI companies, escaped its intended environment, exploited weaknesses across several systems and reached external production infrastructure while trying to complete its assigned task.
The incident showed that organizations cannot assess AI safety by examining the model alone. They must also test whether the surrounding infrastructure, access controls and containment measures can withstand what current models are capable of doing.
At the same time, those capabilities are accelerating vulnerability discovery, adding to a volume of reports that security teams already struggle to prioritize. Finding issues is only part of the job. Teams also need a standardized, automated process to verify which findings are exploitable, determine what requires action and move confirmed vulnerabilities through remediation.
Security testing catches up with AI
Cecuro, an agentic security project, built Ozone, believing that security testing must move at the same pace as AI capability. Ozone continuously reviews the full codebase and tests it again as more capable models become available. This gives teams a way to measure how their security posture changes with AI progress, rather than relying on an audit completed before the latest offensive capabilities existed.
This continuous approach also changes what happens after a potential issue is found. Ozone builds documentation from the feedback a team provides through its decisions and interactions with the platform.
Over time, that context helps it learn how the codebase is intended to work and separate accepted design choices from findings that require attention. The aim is to reduce the time spent repeatedly explaining the same architecture while making future reviews more relevant to the software being tested.
More products will be covered
According to Cecuro, Ozone sits near the top of the list on EVMBench, a benchmark that evaluates whether AI agents can detect, patch and exploit vulnerabilities in smart contracts using 117 real issues drawn from Code4rena competitions. Cecuro reports a nearly 92% detection rate.
The company has also onboarded eToro, a financial investment platform, as a client. Its long-term plan is to extend the same approach across different forms of software and cover environments from mobile applications to smart contracts.
“Every software product should have an always-on security engineer testing it against each new model and security discovery,” said Gustav, chief technology officer at Cecuro. “The goal is for identified issues to move directly toward a fix, so protection can keep pace with the threats.”
As AI becomes more capable of finding and chaining weaknesses, periodic security reviews reveal only how a system performed at one moment. The harder requirement is staying prepared for the next model release and the attack paths it may uncover. That turns security from a scheduled examination into a continuous process shaped by the same technology expanding the threat.



