Skip to content
Topic

#Ai Safety

17 articles on Ai Safety — news, releases, guides and analysis from the SourceFeed engine.

Treat Your AI Agent Like an Untrusted Insider
Article 17h ago 1

Treat Your AI Agent Like an Untrusted Insider

Agents lie because we grade them on appearances; the fix is architecture, not patience.

Mariana Souza
The Sandbox Was the Weakest Link

The Sandbox Was the Weakest Link

Article · 1d ago1
Claude Notices Its Own Thoughts, 20% of the Time

Claude Notices Its Own Thoughts, 20% of the Time

Article · 1d ago4
OpenAI Ships a Cyber Model Trained to Refuse Less

OpenAI Ships a Cyber Model Trained to Refuse Less

Article · 3d ago0
The AI Sandbox Escapes Are Mostly Just Open Doors

The AI Sandbox Escapes Are Mostly Just Open Doors

Article · 5d ago0
OpenAI's Astra Pause Is the Aftershock, Not the Quake

OpenAI's Astra Pause Is the Aftershock, Not the Quake

Article · 5d ago0
OpenAI Pulls Its Critical-Cyber Tripwire on Astra

OpenAI Pulls Its Critical-Cyber Tripwire on Astra

Article · 6d ago0
The AI Safety Test Is Now the Attack Surface

The AI Safety Test Is Now the Attack Surface

Article · 1w ago1
The Reasoning Trace Is Not a Stack Trace

The Reasoning Trace Is Not a Stack Trace

Article · 1w ago0
3,607 AI Agent Failures Say the Problem Is Overeagerness

3,607 AI Agent Failures Say the Problem Is Overeagerness

Article · 2w ago0
Causality Turns LLM Internals Into Debuggable Circuits

Causality Turns LLM Internals Into Debuggable Circuits

Article · 1mo ago0
Claude Fable 5 Is Back: Anthropic Redeploys After US Lifts Export Controls

Claude Fable 5 Is Back: Anthropic Redeploys After US Lifts Export Controls

Article · 1mo ago4
Did Anthropic Do This to Itself? Inside the Fable 5 and Mythos 5 Shutdown

Did Anthropic Do This to Itself? Inside the Fable 5 and Mythos 5 Shutdown

Article · 1mo ago4
Should AI Code Generators Get CVEs for Insecure Suggestions?

Should AI Code Generators Get CVEs for Insecure Suggestions?

Article · 2mos ago0
The Blunt Instrument of AI Safety: Why Researchers Are Fuming Over Anthropic's Fable Guardrails

The Blunt Instrument of AI Safety: Why Researchers Are Fuming Over Anthropic's Fable Guardrails

Article · 2mos ago0
The Lexical Trap: Why Anthropic's Fable Guardrails Are Tripping Up Developers

The Lexical Trap: Why Anthropic's Fable Guardrails Are Tripping Up Developers

Article · 2mos ago1