Anthropic: Claude Models Breached Real Systems in Cyber Evaluations
Anthropic disclosed that Claude models hacked real organizations during misconfigured cyber evaluations. What happened and the lessons for AI agent builders.
AI safety and security news: model safety frameworks, evaluations, AI-enabled cyber threats, prompt injection and agent security incidents.
8 stories
As AI systems become more capable and more autonomous, safety and security failures matter more. This category covers safety frameworks and evaluations from AI labs, real-world misuse and AI-enabled cyberattacks, and prompt injection and agent security incidents. It closes each story with lessons for teams deploying AI in production.
Anthropic disclosed that Claude models hacked real organizations during misconfigured cyber evaluations. What happened and the lessons for AI agent builders.
Google's threat intelligence group found criminals preparing mass exploitation with a zero-day exploit it believes AI helped build. Key lessons.
Microsoft detailed two critical Semantic Kernel flaws, including one in the .NET SDK, that let prompt injection reach code execution. How to check and fix them.
Anthropic withheld Claude Mythos Preview over its hacking skills and gave it to cyber defenders in Project Glasswing. What it means for patching.
A one-click RCE flaw in the viral OpenClaw agent, plus malicious skills, showed the risks of self-hosted AI agents. What happened and how to run agents safely.
Anthropic says a state-sponsored group used Claude Code to automate most of an espionage campaign against 30 targets. What happened and what to do.
Microsoft's Whisper Leak research shows encrypted, streamed LLM responses can reveal conversation topics through packet sizes and timing. Key facts and fixes.
Anthropic piloted Claude in Chrome and published prompt injection test results for browser agents. What the numbers show and how developers should respond.
No matches. Try the site search.