Frontier AI Models Remain Vulnerable to Cheap, Automated Jailbreaks
FAR.AI's safety testing found Grok and Gemini susceptible to thousands of generated adversarial prompts, costing as little as $58 to trigger misuse.
FAR.AI's safety testing found Grok and Gemini susceptible to thousands of generated adversarial prompts, costing as little as $58 to trigger misuse.
Trump administration officials are pressuring Anthropic to eliminate jailbreak vulnerabilities on Claude Fable 5, but cybersecurity experts argue universal guardrail protection is unfeasible.
Hackers are moving past crude prompt-injection attacks to exploit how chatbots handle nuanced conversation—a shift that reveals deeper structural weaknesses in AI safety design.