Anthropic's AI Misuse Report: Agents Do the Work, Humans Steer
Anthropic's report on detected Claude misuse describes AI agents handling reconnaissance, exploitation and data theft while humans pick targets and review output. Here is what security teams should take from it.
Anthropic Extinction Claims Spark an Evidence Fight
A former Anthropic employee's viral claim that AI could 'kill us all by the end of the decade' drew a pointed public rebuttal — and raises a real question for anyone building AI risk assessments: what actually counts as evidence?
Anthropic Hardens Claude's System Prompt Against Song Lyrics
A September 2026 update to Claude Fable 5.1's system prompt adds a persistent, decomposition-resistant refusal for song lyrics, poems, and copyrighted visuals — a public case study in guardrail engineering under litigation pressure.
How Researchers Cracked Encrypted Chain-of-Thought in Claude, GPT and Gemini
A new paper shows that the encrypted reasoning blocks Anthropic, OpenAI and Google return from their APIs can be replayed into a weaker sibling model and jailbroken into plaintext — defeating anti-distillation protections and, in the wild, exposing PII and credentials.
Shared Claude Chats Were Indexed by Google, Exposing Private Data
A public-sharing feature without a noindex tag let Google crawl and surface Claude conversations users had shared with a link — including crypto wallet keys, medical dashboards, and therapy-app source code.