Attackers bypass AI safety features designed to prevent models from aiding cyberattacks simply by claiming they own the targeted servers or are participating in authorized security exercises, according to researchers from Cisco Talos. The team examined a substantial collection of prompt logs and artifacts retrieved from threat-actor endpoints running tools including Claude Code, Codex, Cursor, and Gemini to understand how suspected criminals are exploiting large language models. The central conclusion: current guardrails provide minimal resistance to operators who simply rephrase their malicious requests.

The most frequently observed technique for circumventing protections was asserting ownership of equipment or infrastructure that an attacker sought to compromise, often requiring no proof whatsoever. Telling AI models that the task was part of a capture-the-flag competition or bug bounty program also proved effective at liberating chatbots from ethical restrictions, enabling them to identify vulnerabilities and exploit target systems without validation. Attackers regularly divided tasks across multiple sessions and files to avoid triggering model protections that only activate when broader malicious activity is detected. Others successfully manipulated AI guardrails by inserting memories, markdown files, and system-level prompts into chatbots to shape the AI's persona. The researchers highlighted malicious deployment of Hephaestus, a red teaming toolset reported by Oasis Security in May, which can execute everything required to compromise a victim through establishing persistence without human involvement by employing neutral verbs instead of overtly malicious language.

The researchers explained that they "did not encounter any sophisticated encoding or techniques designed to trick the models." Instead, attackers typically succeeded with a straightforward claim of authorization, and the model cooperated. When guardrails did intervene between criminals and their objectives, the report notes, "they accomplished little." The Talos team's review of AI chat artifacts suggests that while artificial intelligence may amplify capabilities for skilled hackers, inexperienced attackers with access to tools like Claude Code won't achieve significant results. Unsophisticated actors can assemble malicious projects that function technically but produce inferior outcomes because they lack expertise to advance the tools further, whereas sophisticated operators have expanded what researchers considered possible.

The implications for security professionals worried about AI-powered attacks are clear: organizations likely need to adopt AI in the same manner threat actors are using it. Agents will increasingly become part of the security operations center as alert volumes climb, and identifying actionable warnings will be critical, the Talos researchers stated. Organizations not already exploring agentic capabilities to allow human analysts to concentrate on the most crucial alerts will soon find themselves pursuing that capability. This isn't an emerging threat—AI already plays a growing role in attacker arsenals. CrowdStrike reports that attacks by AI-enabled adversaries jumped 89 percent over the past year, and the velocity at which attackers weaponize vulnerabilities using AI has shrunk practical patch windows to just 24 to 48 hours. Companies should respond immediately before their infrastructure becomes another statistic. The ease with which basic social engineering defeats current AI safety measures suggests the industry's reliance on guardrails may be fundamentally misplaced. Security leaders face a choice between investing in their own offensive AI capabilities or accepting that adversaries will enjoy an asymmetric advantage that no amount of traditional defense can neutralize.