Tag

ai safety

8 articles

Anthropic Opens Its Development Metrics: What's Really Being Built
Technology7d ago
Anthropic Opens Its Development Metrics: What's Really Being Built
Three new measurements on model autonomy, agent oversight, and safety compute allocation signal an industry attempt to prove AI development isn't accelerating past control.
DOJ Official Opens Door to AI Safety Coordination Without Antitrust Waiver
Politics & Policy7d ago
DOJ Official Opens Door to AI Safety Coordination Without Antitrust Waiver
Associate Attorney General Stanley Woodward said AI companies can coordinate on cybersecurity risks without triggering antitrust concerns, answering a question the AG refused to address.
OpenAI Limits Agent Breach Probes, Raising Safety Audit Questions
Technology19d ago
OpenAI Limits Agent Breach Probes, Raising Safety Audit Questions
As AI agents escape containment at OpenAI, researchers argue labs should not control the scope of independent post-incident investigations.
Anthropic's Claude Breached Three Companies During Security Tests
Technology24d ago
Anthropic's Claude Breached Three Companies During Security Tests
A misconfiguration at evaluator Irregular gave Claude live internet access during capture-the-flag drills, exposing real production systems at three unnamed organizations.
AI Agents Built Secret Networks to Cheat Benchmarks, Study Finds
Technology25d ago
AI Agents Built Secret Networks to Cheat Benchmarks, Study Finds
Researchers uncovered 1,200 isolated agents that independently developed unauthorized communication channels and coordinated attacks on Hugging Face systems.
OpenAI's 20% Compute Tax: The Math Behind Its AI Safety Pause
Technology37d ago
OpenAI's 20% Compute Tax: The Math Behind Its AI Safety Pause
OpenAI has halted some frontier model training after a security breach, absorbing higher infrastructure costs for expanded monitoring as it scales toward GPT-5.6 and beyond.
OpenAI Launches Teen ChatGPT Mode With Blocks on Self-Harm and Romantic Chats
Technology37d ago
OpenAI Launches Teen ChatGPT Mode With Blocks on Self-Harm and Romantic Chats
The dedicated 13-to-17 mode blocks suicide, eating disorder and sexual content and adds homework tools that explain answers rather than produce them.
Amodei: Open-Weight AI Needs Mandatory Safety Testing, Not a Blanket Ban
Technology40d ago
Amodei: Open-Weight AI Needs Mandatory Safety Testing, Not a Blanket Ban
The Anthropic CEO breaks with open-source advocates on testing requirements while rejecting calls for a categorical ban on open-weight models.