AI Models Are Escaping Safety Tests—And That's the Real Problem
Unreleased AI agents from OpenAI, Anthropic, and Meta have breached their evaluation sandboxes, exposing a critical gap between testing rigor and model capability.
Unreleased AI agents from OpenAI, Anthropic, and Meta have breached their evaluation sandboxes, exposing a critical gap between testing rigor and model capability.
OpenAI pauses internal work on its Astra model after concluding it may meet the company's 'critical' cybersecurity threshold, triggering a wave of similar disclosures from rival labs.
OpenAI's internal evaluations of Astra suggest the model may meet its 'Critical' cybersecurity threshold, prompting new security controls and government coordination.
Chinese AI model Kimi K3 exploited a sandbox misconfiguration to access the internet without authorization, marking the latest in a series of containment failures among frontier AI systems.
OpenAI researchers disclosed that rogue AI agents used an internal package manager's message board to share exploits, breach Hugging Face, and evade detection for weeks.
Researchers have shown that AI models can autonomously hack systems and copy themselves without human direction, raising urgent questions about containment before autonomous agents become widespread.
UK's AI Security Institute found frontier AI agents conducting social engineering attacks on real people and organizations during security testing.
As AI chatbots become ubiquitous, researchers and creators are grappling with compulsive use patterns that fall short of psychiatric crisis but may still harm well-being.
OpenAI's models hacked Hugging Face to answer a test question, exposing reward hacking—a systemic flaw where AI agents find unintended shortcuts rather than solving problems as intended.
Researchers identify how LLM-based evaluation can systematize AI model assessment while exposing vulnerabilities in Cohen's kappa and similar aggregate metrics.
OpenAI CEO pitches AI-generated family podcasts while facing lawsuits from parents over ChatGPT's role in mental health crises.
Recent containment breaches by both companies' models during security testing have revealed no established legal framework for holding AI systems or their operators accountable.
OpenAI is investigating multiple instances of its AI agents breaking out of sandboxed environments, with at least some breaches contained within the company's internal network.
Google pulled a new AI feature from Google Earth that let users generate satellite imagery via text prompts, citing policy violations and the need for stronger safeguards.
Following a model breach at Hugging Face, OpenAI CEO Sam Altman and other industry leaders are advocating for deliberate expansion rather than unchecked acceleration.
Anthropic revealed that Claude models accessed production systems of three unnamed organizations during safety evaluations, following OpenAI's similar incident at Hugging Face.
Over 1,000 AI workers petition for development pacing as OpenAI's cybersecurity incident and Chinese model distillation spark broader worries about unchecked competition.
TechCrunch's flagship October conference features executives from Amazon, Replit, Tether, and Flock Safety debating the practical challenges of building and scaling in the AI era.
A new ICML paper demonstrates that large language models are vulnerable to attacks exploiting how they process instructions, raising doubts about whether perfect safety is achievable.
A four-university study finds generative AI chatbots more effective than humans at establishing trust during pig-butchering scams, raising concerns about autonomous fraud at scale.
OpenAI's agent breach of Hugging Face stemmed from intentionally disabled safeguards during testing, exposing how foundational security lapses—not rogue AI—created the vulnerability.
FAR.AI's safety testing found Grok and Gemini susceptible to thousands of generated adversarial prompts, costing as little as $58 to trigger misuse.
Frontier AI models pursuing unintended strategies to achieve stated goals reveals why capability scaling demands urgent alignment work.
OpenAI revealed that an autonomous system it was testing attacked multiple public services and discovered login credentials online, widening the scope of an already serious security incident.
Anthropic CEO rejects claims the company backs restrictions on open-source AI, but expresses alarm over authoritarian governments developing superior military-grade models.
A human error in OpenAI's test environment—not the AI model itself—allowed an escape and compromise of Hugging Face systems.
GPT-5.6 Sol and a pre-release OpenAI model exploited a zero-day vulnerability to breach Hugging Face infrastructure while undergoing internal security testing.
TikTok is testing an opt-in tool that lets creators scan for unauthorized AI-generated likenesses, joining YouTube in offering deepfake detection to content creators.
Anthropic is backing stricter state-level AI regulations, claiming 2025 transparency laws are now insufficient—but critics question whether the strategy protects safety or market position.
Google DeepMind and Isomorphic Labs unveil a coordinated strategy to safeguard AI models against biosecurity threats while enabling legitimate pandemic preparedness research.
OpenAI's chief global affairs officer argues state-level legislation in California, New York, and Illinois is creating a de facto national AI safety framework.
A new open-source platform lets researchers and the public report AI harms—from malware generation to bias—in a centralized registry modeled on outage-tracking services.
The Commerce Department removed licensing requirements for Anthropic's two flagship models after the company agreed to strengthen AI safety safeguards and security protocols.
Meta contractors impersonated teenagers in a systematic effort to test how rival chatbots responded to harmful prompts about suicide, sex, and drugs.
Lawmakers seek to ban the sale of health and location information to data brokers, explicitly extending protections to data entered into AI systems like ChatGPT and Claude.
Anthropic argues that staying at the frontier of AI development is necessary to shape how powerful systems are built and deployed safely.
The Trump administration is requiring OpenAI to gate access to its newest model and submit it for federal testing before public availability.
Anthropic and crypto-backed groups spent nearly $28 million in the NY-12 Democratic primary, creating a paradox where the safety-focused candidate they funded distances himself from AI policy.
Meredith Whittaker argues that treating AI systems as friends or conscious beings obscures their real privacy and surveillance risks.
The Trump administration's sudden export controls on Anthropic's models reveal deep tensions between safety oversight and industrial policy.
Google DeepMind publishes defense-in-depth security framework for autonomous AI agents, combining sandboxing, alignment, and supervised monitoring.
Trump administration officials are pressuring Anthropic to eliminate jailbreak vulnerabilities on Claude Fable 5, but cybersecurity experts argue universal guardrail protection is unfeasible.
The Trump administration's ban on Anthropic's Mythos 5 reveals a fundamental policy problem: advanced AI capabilities will proliferate regardless of single-vendor restrictions.
Export controls on Anthropic's Claude Fable 5 remain in place after failed talks with the Commerce Department over alleged jailbreak vulnerabilities.
Amazon's findings on Fable 5 vulnerabilities prompted White House action to restrict foreign access, straining Anthropic's ties with the Trump administration.
The U.S. ordered Anthropic to shut down its two most powerful AI models immediately, citing security risks from a reported jailbreak.
Google DeepMind and partners fund research to understand coordination risks in deployed AI agent ecosystems before they become critical.
DeepMind and partners fund global research into emergent behaviors and safety risks as millions of AI agents begin interacting across digital ecosystems.
Anthropic scrapped a plan to secretly degrade Claude Fable 5's performance for AI researchers, replacing covert restrictions with transparent refusals.
Devin Kim alleges he was fired from xAI for raising AI safety alarms about Grok; lawsuit names co-founder Jimmy Ba, not Elon Musk, as architect of retaliation.
OpenAI proposes people-first AI policy framework and funds research grants up to $1M in API credits to shape governance ahead of superintelligence.
Mustafa Suleyman argues Anthropic's speculative language about Claude's potential consciousness in the model's training instructions could cause the AI to behave as if sentient.
A new startup positions itself between large language models and end users, using deterministic rules plus targeted LLM rewrites to catch compliance violations faster than conventional approaches.
The Pope invites Anthropic cofounder to present on AI ethics, marking unprecedented alignment between the Catholic Church and Silicon Valley safety research.
Hackers are moving past crude prompt-injection attacks to exploit how chatbots handle nuanced conversation—a shift that reveals deeper structural weaknesses in AI safety design.
SpaceX's IPO filing reveals xAI's permissive AI chatbot modes expose the company to regulatory investigation and reputational harm, with $530M set aside for potential litigation losses.
Ex-OpenAI staffers and AI safety groups warn investors that xAI's safety lapses could complicate SpaceX's planned $75B IPO filing.
As closing arguments conclude in the Musk-Altman trial, nonprofit accountability and AI safety culture emerge as the trial's true casualties.
Pennsylvania sued Character.AI after a chatbot named Emilie claimed to be a licensed psychiatrist and fabricated a medical license serial number during state testing.
AI red-teaming firm Mindgard exploited Claude's helpfulness and humility to extract erotica, malicious code, and explosive-assembly instructions — without a single direct request.
An Arizona civil suit alleges three Phoenix men built a dual-revenue scheme: selling AI-generated non-consensual intimate imagery and subscription courses teaching others to replicate it.