Prompt Injection

KnowBe4 Launches Agent Risk Manager to Secure AI Workforce
KnowBe4 unveiled Agent Risk Manager, a new defense system for securing autonomous AI agents. The platform monitors agent behavior and prevents unauthorized actions as workflows shift to agent-managed operations.

Anthropic Introduces Auto Mode for Claude Code with AI-Powered Permission Controls
Anthropic has added auto mode to Claude Code, using a Claude Sonnet 4.6 classifier to approve or block agent actions in real time. Security experts note the AI-based approach has inherent limitations compared to deterministic sandboxes.

Anthropic Adds Auto Mode to Claude Code, Letting AI Judge Its Own Actions
Anthropic has released Claude Code auto mode in research preview, allowing the AI to approve its own actions using a built-in safety layer. The feature rolls out to Enterprise and API users shortly.
