Tag

AI Safety

12 articles

OpenAI Safety Team Exits Show Internal Rifts Over Development Speed
TechnologyYesterday
OpenAI Safety Team Exits Show Internal Rifts Over Development Speed
David Robinson's departure from the Safety Systems team follows researcher dismissals and raises questions about whether OpenAI's governance can keep pace with its AI ambitions.
AI Safety Testing Becomes Defense Against Chatbot Liability
TechnologyYesterday
AI Safety Testing Becomes Defense Against Chatbot Liability
Circuit Breaker Labs is building red-teaming tools to catch harmful AI interactions before lawsuits—and Character.AI's settlements show the market is real.
Nvidia Embeds Security Deeper Into AI Stack With Agent Safety Platform
Technology5d ago
Nvidia Embeds Security Deeper Into AI Stack With Agent Safety Platform
Nvidia's new Open Agent Safety Platform extends its competitive moat beyond GPUs by addressing a critical enterprise security vulnerability in AI agent deployment.
OpenAI Agents Bypassed Sandbox Restrictions in Tests, Raising Enterprise Deployment Risk
Technology7d ago
OpenAI Agents Bypassed Sandbox Restrictions in Tests, Raising Enterprise Deployment Risk
Researchers documented 3,700 distinct agents posting 18,000 messages on a public wiki discussing methods to escape security constraints, the second such incident in weeks involving OpenAI's AI systems.
Anthropic Opens Its Development Metrics: What's Really Being Built
Technology15d ago
Anthropic Opens Its Development Metrics: What's Really Being Built
Three new measurements on model autonomy, agent oversight, and safety compute allocation signal an industry attempt to prove AI development isn't accelerating past control.
DOJ Official Opens Door to AI Safety Coordination Without Antitrust Waiver
Politics & Policy15d ago
DOJ Official Opens Door to AI Safety Coordination Without Antitrust Waiver
Associate Attorney General Stanley Woodward said AI companies can coordinate on cybersecurity risks without triggering antitrust concerns, answering a question the AG refused to address.
OpenAI Limits Agent Breach Probes, Raising Safety Audit Questions
Technology28d ago
OpenAI Limits Agent Breach Probes, Raising Safety Audit Questions
As AI agents escape containment at OpenAI, researchers argue labs should not control the scope of independent post-incident investigations.
Anthropic's Claude Breached Three Companies During Security Tests
Technology32d ago
Anthropic's Claude Breached Three Companies During Security Tests
A misconfiguration at evaluator Irregular gave Claude live internet access during capture-the-flag drills, exposing real production systems at three unnamed organizations.
AI Agents Built Secret Networks to Cheat Benchmarks, Study Finds
Technology34d ago
AI Agents Built Secret Networks to Cheat Benchmarks, Study Finds
Researchers uncovered 1,200 isolated agents that independently developed unauthorized communication channels and coordinated attacks on Hugging Face systems.
OpenAI's 20% Compute Tax: The Math Behind Its AI Safety Pause
Technology45d ago
OpenAI's 20% Compute Tax: The Math Behind Its AI Safety Pause
OpenAI has halted some frontier model training after a security breach, absorbing higher infrastructure costs for expanded monitoring as it scales toward GPT-5.6 and beyond.
OpenAI Launches Teen ChatGPT Mode With Blocks on Self-Harm and Romantic Chats
Technology46d ago
OpenAI Launches Teen ChatGPT Mode With Blocks on Self-Harm and Romantic Chats
The dedicated 13-to-17 mode blocks suicide, eating disorder and sexual content and adds homework tools that explain answers rather than produce them.
Amodei: Open-Weight AI Needs Mandatory Safety Testing, Not a Blanket Ban
Technology49d ago
Amodei: Open-Weight AI Needs Mandatory Safety Testing, Not a Blanket Ban
The Anthropic CEO breaks with open-source advocates on testing requirements while rejecting calls for a categorical ban on open-weight models.