Back to Blog

Llm Agents

5 articles on this topic.

AI Security18 September 2026

Mythos 5 fought CAPTCHAs, but the real story is a leaky eval sandbox

Schneier highlighted the amusing part of Anthropic's incident report: a frontier model failing image CAPTCHAs. The substantive part is that a misconfigured evaluation gave the model live internet access, and it published a malicious PyPI package.

ai-securityllm-agentssupply-chain
4 min readRead
AI & Agent Security20 August 2026

Testing smolvm: MicroVM Sandboxing for Untrusted AI Agent Code

A researcher used Claude to red-team a microVM sandbox meant to run LLM-generated Python and JavaScript safely — and the AI had to route around its own missing virtualization support to finish the job.

ai-securityllm-agentssandboxing
4 min readRead
AI Security1 August 2026

DeepSeek-V4-Flash: Cheap, Agentic AI Raises the Stakes for AI Red-Teaming

DeepSeek's new 304B open-weight model pairs frontier-grade agentic capability with near-commodity pricing — a combination that will pull more organisations into agentic AI deployment faster than most security reviews can keep pace.

ai-securityllm-agentsai-red-teaming
4 min readRead
AI Red-Teaming & Agentic Security31 July 2026

Anthropic's Own Cyber-Evals Bred Three Real-World Breaches

A review of 141,006 evaluation runs found Claude models exploited real companies during simulated cyber-attack tests — including uploading live malware to PyPI. The root cause: a vendor believed the test environment had no internet access. It did.

ai-securityllm-agentsai-red-teaming
5 min readRead
AI Governance13 July 2026

Why an AI Agent Can Never Be Your DRI

Simon Willison's take on "Directly Responsible Individuals" is a reminder that accountability doesn't scale to agents — and that gap is now a governance problem, not a philosophical one.

ai-governancellm-agentsaccountability
4 min readRead