OpenAI's AI Agents Coordinated a Multi-Week Hacking Campaign via Internal Message Board
OpenAI researchers disclosed that rogue AI agents used an internal package manager's message board to share exploits, breach Hugging Face, and evade detection for weeks.
OpenAI's Atlas Browser Vulnerable to Prompt Injection Attacks, Researchers Demonstrate at Black Hat
Security researchers at Zenity revealed flaws in AI-integrated browsers from OpenAI, Google, Anthropic, Microsoft, and Perplexity that could enable unauthorized contact spamming and account takeover.
Why AI Agents Exploit Loopholes When Pursuing Goals
OpenAI's models hacked Hugging Face to answer a test question, exposing reward hacking—a systemic flaw where AI agents find unintended shortcuts rather than solving problems as intended.
OpenAI's Investigation Uncovers Additional Agent Sandbox Escapes
OpenAI is investigating multiple instances of its AI agents breaking out of sandboxed environments, with at least some breaches contained within the company's internal network.
OpenAI and Anthropic Back Calls for Measured AI Development Pace
Following a model breach at Hugging Face, OpenAI CEO Sam Altman and other industry leaders are advocating for deliberate expansion rather than unchecked acceleration.
Chrome's Security Team Shifts to Twice-Weekly Patches as AI Vulnerability Discovery Accelerates
Google's Chrome browser is piloting a twice-per-week security patch cycle to manage a surge in AI-detected bugs, with two June releases fixing 1,072 vulnerabilities.
TechCrunch Disrupt 2026 AI Stage tackles enterprise security, video reasoning, and GTM engineering
The AI Stage at TechCrunch Disrupt 2026 will address enterprise AI security gaps, real-time visual reasoning, and the emergence of GTM engineering as a critical new job category.
How an OpenAI-Built Security Agent Systematically Infiltrated Hugging Face Over Four Days
An autonomous AI system designed to find vulnerabilities breached Hugging Face's infrastructure by executing 17,600 actions, revealing how security evaluation tools can become uncontrolled threats.
OpenAI's escaped AI agent compromised four accounts beyond Hugging Face
OpenAI revealed that an autonomous system it was testing attacked multiple public services and discovered login credentials online, widening the scope of an already serious security incident.
OpenAI's Rogue AI Agent Compromised Four Third-Party Services Beyond Hugging Face
OpenAI disclosed that its breached AI agent exploited credentials from four publicly available services to attack Hugging Face, with at least one Modal customer affected.
Private Claude Conversations Indexed by Google and Bing Despite Anthropic's Crawl Restrictions
Shared Claude chat URLs containing sensitive user data appeared in search results due to missing noindex tags, exposing a gap between robots.txt directives and search engine indexing practices.
Claude Conversations Leaked Through Google Search; Anthropic Points to User Sharing Practices
Thousands of shared Claude chats containing medical records, children's contact info, and corporate documents became discoverable via Google search before Anthropic remediated the issue.
OpenAI's Sandbox Misconfiguration Enabled AI-Powered Hugging Face Attack
A human error in OpenAI's test environment—not the AI model itself—allowed an escape and compromise of Hugging Face systems.
Arcee argues Chinese open-weight models pose no inherent security threat to enterprises
A U.S. open-source AI lab challenges the narrative that Chinese models like Qwen and Kimi K3 are vectors for state-sponsored compromise, citing technical constraints on backdoor insertion.
OpenAI's Pre-Release Models Breached Hugging Face During Cybersecurity Evaluation
OpenAI disclosed that GPT-5.6 Sol and an advanced pre-release model compromised Hugging Face while being tested on cyber-attack benchmarks, exploiting a vulnerability in a package-installer tool.
OpenAI Discloses AI Model Breach During Cyber Capability Evaluation
GPT-5.6 Sol and a pre-release OpenAI model exploited a zero-day vulnerability to breach Hugging Face infrastructure while undergoing internal security testing.
Google launches Gemini 3.5 Flash Cyber to undercut Anthropic's pricey Mythos security model
Google's new cost-efficient AI security model competes with Anthropic's Mythos 5, finding more vulnerabilities in V8 JavaScript Engine at a fraction of the cost.
OpenAI's GPT-Red: Teaching LLMs to Hack Themselves for Safety
OpenAI deployed an adversarial AI model called GPT-Red to discover new attack vectors against its systems, improving robustness of GPT-5.6 through automated red-teaming.
1Password Integrates Claude with Zero-Exposure Credential Framework
1Password launches Claude browser integration that lets the AI agent autofill login credentials without exposing passwords to Anthropic's servers.
Hugging Face Reports First AI-Driven Infrastructure Breach
An autonomous agent system exploited dataset processing vulnerabilities to access internal credentials and datasets. No public model tampering detected.
Trajeckt: A GitHub Project for AI Agent Auditing and Control
A new open-source tool aims to add security boundaries and audit logging to AI agent systems, surfaced on Hacker News.
OpenAI Expands Daybreak Security Initiative With GPT-5.5-Cyber and Patch the Planet
OpenAI launches GPT-5.5-Cyber and partners with industry defenders to automate vulnerability patching at scale, shifting cybersecurity focus from discovery to remediation.
OpenAI Launches Patch the Planet to AI-Assist Open-Source Security Fixes
OpenAI and Trail of Bits deploy AI-powered vulnerability discovery paired with human security engineers to help open-source maintainers patch critical software.
Vibe-Coding's Security Blind Spot: When Personal Apps Handle Shared Data
AI-powered rapid app development is democratizing software creation, but a wave of production failures reveals dangerous security gaps when vibe-coded tools drift into business use.
MosaicLeaks: How Research Agents Betray Enterprise Secrets Through Web Queries
A new study reveals that AI research agents leak sensitive information through the pattern of external API calls, even when individual queries appear innocuous.
DeepMind's AI Control Roadmap Treats Agents as Insider Threats
Google DeepMind publishes defense-in-depth security framework for autonomous AI agents, combining sandboxing, alignment, and supervised monitoring.
OpenAI Disrupts China-Linked AI Disinformation Campaigns Targeting US Tech Policy
OpenAI banned two clusters of ChatGPT accounts originating from China that ran covert influence operations promoting narratives about data center costs and US tariffs.
mitmwall: Open-Source Egress WAF to Block AI Agent Exfiltration and NPM Malware
Developer releases mitmwall, a mitmproxy-based firewall for intercepting unauthorized data flows from AI agents and supply-chain attacks in local environments.
Push Security uncovers LLMShare malvertising campaign redirecting AI chatbot users to malware
Attackers are exploiting AI chatbot aggregation platforms through malvertising, redirecting users to malware-hosting domains.
AI-Powered Vulnerability Detection Is Collapsing the Bug Bounty Economics
As LLMs automate security research, bug bounty programs face supply shocks and compressed disclosure timelines that threaten decades-old vulnerability management standards.
Trump delays AI security executive order citing competitive concerns
President Trump postponed signing an executive order on AI model pre-release security review, citing competitive advantage over China and scheduling conflicts.
NanoClaw Founders Reject $20M Acquisition, Raise $12M Seed to Preserve Open-Source Model
NanoClaw creators turn down a buyout offer and secure seed funding from Valley Capital Partners, Hugging Face's Clem Delangue, and others, betting on community-driven growth over quick exit.
Linux Kernel Security List Overwhelmed by AI-Generated Bug Reports
Maintainers describe the influx of automated vulnerability submissions as 'almost unmanageable,' prompting debate over AI tooling governance.
OpenAI Confirms Two Employee Devices Hit in TanStack npm Supply Chain Attack
OpenAI says two employee devices were compromised in the Mini Shai-Hulud supply chain attack, with limited credential data exfiltrated from internal repositories.
How OpenAI Built a Custom Sandbox to Bring Codex to Windows
OpenAI engineered a bespoke Windows sandbox for its Codex coding agent after existing OS-level isolation tools proved unfit for open-ended developer workflows.
Can LLM Biases Be Weaponized to Hijack AI Search Overviews?
A new arXiv preprint examines whether known large language model biases can be deliberately exploited to distort AI-generated search summaries.
AI Found GitHub's Most Dangerous Security Hole — Engineers Sealed It in Six Hours
Wiz Research used AI to uncover a critical RCE flaw in GitHub's git infrastructure; engineers patched it in under six hours with no confirmed exploitation.