Back to Blog

Anthropic

5 articles on this topic.

AI Security25 September 2026

Anthropic's AI Misuse Report: Agents Do the Work, Humans Steer

Anthropic's report on detected Claude misuse describes AI agents handling reconnaissance, exploitation and data theft while humans pick targets and review output. Here is what security teams should take from it.

ai-securitythreat-intelligencellm-misuse
3 min readRead
AI Governance14 September 2026

Anthropic Extinction Claims Spark an Evidence Fight

A former Anthropic employee's viral claim that AI could 'kill us all by the end of the decade' drew a pointed public rebuttal — and raises a real question for anyone building AI risk assessments: what actually counts as evidence?

ai-safetyai-governanceanthropic
4 min readRead
AI Security & Governance2 September 2026

Anthropic Hardens Claude's System Prompt Against Song Lyrics

A September 2026 update to Claude Fable 5.1's system prompt adds a persistent, decomposition-resistant refusal for song lyrics, poems, and copyrighted visuals — a public case study in guardrail engineering under litigation pressure.

ai-securityllm-guardrailsprompt-engineering
4 min readRead
AI & LLM Security11 August 2026

How Researchers Cracked Encrypted Chain-of-Thought in Claude, GPT and Gemini

A new paper shows that the encrypted reasoning blocks Anthropic, OpenAI and Google return from their APIs can be replayed into a weaker sibling model and jailbroken into plaintext — defeating anti-distillation protections and, in the wild, exposing PII and credentials.

llm securitychain-of-thoughtai security research
4 min readRead
AI Security4 August 2026

Shared Claude Chats Were Indexed by Google, Exposing Private Data

A public-sharing feature without a noindex tag let Google crawl and surface Claude conversations users had shared with a link — including crypto wallet keys, medical dashboards, and therapy-app source code.

ai-securitydata-exposureanthropic
4 min readRead