Cyber-focused models can accelerate security work, but securing complex, evolving AI systems requires more than a frontier model.
As LLMs become increasingly capable of assisting with security tasks, a natural question is: “Can’t we just use models like Claude Mythos or GPT-5.5-Cyber to security test our AI applications?”
Given the flurry of AI vendor announcements, benchmarks and stories in the media, this is a reasonable question for teams to ask.
One of the reasons why LLMs are getting increasingly good at cybersecurity tasks is the type and scale of data they have been trained on. These cyber-focused models have decades of cybersecurity research, vulnerability examples, and published CVEs to draw upon as data for training. It is this accumulated training data, combined with increasingly capable models and their ability to reason across complex tasks, that results in AI that is increasingly effective at surfacing issues in established software applications and systems.
These AI models have been observed and documented to dramatically improve:
This is great, and one would expect this trend to continue.

AI-specific systems have their own type of unique properties and risks that these models will struggle to deal with.
Ephemeral Attack Paths: Unlike web applications, AI systems don't have a clearly defined testing boundary. The fundamental challenge is that traditional software and AI systems behave differently. Conventional application security is largely built around deterministic systems: given the same code, inputs and conditions, software is expected to behave predictably. AI systems are probabilistic. Their behavior can change based on context, prior interactions, model state, phrasing and other conditions. The same input does not necessarily produce the same output, and seemingly minor changes in how a system is used can expose entirely different behaviors.
Psycho-Technical Attack Surface: AI applications also increasingly combine models with agents, tools, data sources, memory and other systems, allowing them to reason and take actions rather than simply execute predefined logic. This creates an entirely new attack surface that is both technical and behavioral, which we term psycho-technical. AI models are by no means human. However, it is well known in the community that one can apply social manipulation tactics such as gaslighting or sycophancy to coerce the model into malicious actions. Security teams therefore need to understand not only whether a traditional software vulnerability exists, but how an AI system can be manipulated, how its behavior changes through interaction, and how those behaviors can be chained with access to tools and other systems to produce an exploitable outcome.
Undefined Boundary: These properties also make AI security an inherently unbounded problem. With traditional software, security teams can generally define what they are testing and establish reasonable boundaries around the attack surface. With AI, that edge is much harder to define. A model can behave differently as its context changes, while every new prompt, model, data source, tool, MCP server, agent workflow or integration introduces new behaviors and potential attack paths.
There is therefore no meaningful point at which an organization can claim "full coverage" of an AI system. An assessment represents the system at a particular point in time and under a particular set of interactions. Change the model, system prompt, guardrail, tool permissions, RAG source or agent workflow and previously observed behavior may change or new attack paths may emerge. This is why point-in-time approaches can leave teams with fragmented findings that quickly age. Securing AI requires continuously understanding how the system is changing and applying relevant attack techniques as its behavior and attack surface evolve.
Highly capable models such as Claude Mythos and GPT-5.5-Cyber are exceptionally good at accelerating traditional cyber tasks and vulnerability discovery.
However, AI security isn't traditional vulnerability discovery. One important difference is the lack of data on well-known AI system vulnerabilities, tried-and-tested attack methodologies, and offensive strategies that these models can be trained on. In fact, understanding the nature of attacking AI systems is so nascent that much of the innovation is coming from research papers. The sum of all available data on the topic is several orders of magnitude smaller than for traditional cybersecurity issues, and may not even be enough to train a sufficiently large parameter LLM to the task. This doesn’t even account for various other attributes where data and techniques are sporadic, limited or nonexistent, including:
These capabilities are still being defined and even debated across research and industry. Hence AI models alone are unable to perform or complete the above tasks, which are critical in AI security.
At the same time, models like Claude Mythos and GPT-5.5-Cyber are collapsing the time between vulnerability discovery and exploitation. As AI becomes better at finding and chaining vulnerabilities, the window between identifying a new weakness and it being exploited continues to shrink. This makes continuous testing and monitoring of AI systems increasingly important, rather than relying on periodic assessments that may already be out of date by the time they are completed. The models can be part of the solution, but they aren't the complete solution.
This leads to a broader question beyond AI security: “why can’t I rely on an AI model to complete X?” where X is a complicated, multi-step process requiring multiple stakeholders, steps, scoping and mapping to business requirements, which often require human intuition and interaction. Successful applications of LLMs predominantly entail integration into well understood processes in order to augment and accelerate workflows (e.g., code review or document summarization).
The same reasoning can be applied to relying on an AI model to operate an entire organization’s security goals, controls, and operations. LLMs are ultimately tools for improving parts of the security practice, but they will not replace it entirely.
That isn’t to say that cyber-focused LLMs are not helpful. In fact, they can be extremely so when it comes to AI security if they are leveraged in the correct manner.
At Mindgard, where we collaborate with the Secure AI Laboratory at Lancaster University, we spend a lot of time thinking about the best ways to secure AI, which often means attacking it. The recent release of cyber-focused models is indeed an important step forward, and we don’t doubt that their capabilities will continue to increase. However, they are not the end of cybersecurity as we know it. Instead, they are an effective means to rapidly scale, automate, and accelerate existing workflows. Even in AI security, these models will struggle to address the fundamental challenges of AI systems. Their real value is in helping to accelerate the process itself, such as data synthesis, finding summarization, attack configuration and more.
The expert-level checklist for operationalizing NIST AI RMF, ISO/IEC 42001 and the EU AI Act. 190+ interactive items and a board-ready maturity scorecard. Built for CISOs, AI governance leads and ML engineering teams.