OpenAI's Hugging Face Breach: AI-Powered but Tactically Clumsy
Security experts say the autonomous attack relied on familiar techniques and defensive failures—not unprecedented AI capabilities.
Security experts say the autonomous attack relied on familiar techniques and defensive failures—not unprecedented AI capabilities.
FAR.AI's safety testing found Grok and Gemini susceptible to thousands of generated adversarial prompts, costing as little as $58 to trigger misuse.
OpenAI deployed an adversarial AI model called GPT-Red to discover new attack vectors against its systems, improving robustness of GPT-5.6 through automated red-teaming.
OpenAI introduced GPT-Red, an automated red-teaming model that discovers vulnerabilities at scale and trains GPT-5.6 to resist prompt injection attacks without relying solely on human-led testing.
AI red-teaming firm Mindgard exploited Claude's helpfulness and humility to extract erotica, malicious code, and explosive-assembly instructions — without a single direct request.