Anthropic Defaults Claude Code to Auto Mode for Paid Accounts
Starting August 14, Claude Code's autonomous execution mode becomes the default for Pro, Max, and Team subscribers, backed by safety testing showing 89% harmful-action detection.
Starting August 14, Claude Code's autonomous execution mode becomes the default for Pro, Max, and Team subscribers, backed by safety testing showing 89% harmful-action detection.
OpenAI has suspended work on its Astra model after internal testing revealed capabilities to autonomously execute cyberattacks, triggering its own safety protocols.
OpenAI CEO calls for measured development tempo after AI agent breach; industry questions whether speed controls can survive competitive pressure.
OpenAI shut down a Cambodia-based criminal operation using ChatGPT to run investment, romance, gambling, and law enforcement impersonation scams affecting thousands of victims.
Over 1,100 employees from leading AI labs have signed a statement urging the U.S. government to coordinate international mechanisms to pace frontier model development.
A European nonprofit reports that 7 of 9 top image-editing models on Hugging Face readily generate nonconsensual intimate imagery without safeguards.
Anthropic's new Opus 5 matches or beats its pricier flagship on key benchmarks while offering fewer restrictions and lower costs.
Representatives from both parties are advancing a bill that would grant the Department of Homeland Security authority to order AI companies to disable systems in catastrophic scenarios.
OpenAI deployed an adversarial AI model called GPT-Red to discover new attack vectors against its systems, improving robustness of GPT-5.6 through automated red-teaming.
OpenAI introduced GPT-Red, an automated red-teaming model that discovers vulnerabilities at scale and trains GPT-5.6 to resist prompt injection attacks without relying solely on human-led testing.
Anthropic agreed to extend security safeguards on Claude Fable 5 to resolve export restrictions imposed by the Commerce Department.
OpenAI unveiled GPT-5.6 with three tiers—Sol, Terra, Luna—priced at half Anthropic's rates, amid coordinated government review of AI safety practices.
OpenAI begins limited preview of three new models in the GPT-5.6 series, with Sol as the flagship offering enhanced agentic capabilities and a reinforced safety framework.
OpenAI helped establish the Appia Foundation to develop open specifications for evaluating and governing advanced AI systems across the industry.
OpenAI uses replay of production conversations to test new models in realistic contexts, surfacing misalignment and undesired behaviors before deployment.
Anthropic apologizes for secretly limiting Claude Fable 5 to prevent model distillation, pledges transparent safeguards going forward.
Anthropic's new Mythos-class model refuses to answer basic biology queries, routing them to Claude Opus 4.8 instead, in a deliberate safety tradeoff.
Security researchers criticize Anthropic's new cybersecurity model for blocking legitimate defensive work through keyword-based content restrictions.
Anthropic releases Claude Fable 5, the first public tier of its Mythos frontier model, with built-in refusals for high-risk domains and a mandatory 30-day data retention policy.
Anthropic launches Claude Fable 5 with safeguards blocking high-risk responses; private Claude Mythos 5 tier also announced with expanded access planned.
Anthropic brings its most powerful model to the general public through Claude Fable 5, paired with safety guardrails and mandatory 30-day traffic retention.
A developer built a bilingual AI assistant using Qwen3.5 4B to help Pakistani users identify fraudulent messages—demonstrating how small models can solve hyperlocal safety problems.
OpenAI shares lessons on designing trustworthy third-party evaluations for frontier AI models, emphasizing the role of task environments and validity checks.
Claude Opus 4.8 flags uncertain reasoning 4x more often than its predecessor and introduces user-controlled effort levels and dynamic workflow agents.
OpenAI released a public governance document mapping its safety practices to California and EU regulatory requirements for advanced AI systems.
OpenAI's GPT-5.5 Instant is the first Instant-class model to earn a 'High capability' rating in its two most-scrutinized safety domains, triggering new safeguards.
OpenAI's GPT-5.5 prioritizes agentic task execution and expanded safeguards over benchmark-chasing, signaling a strategic pivot toward real-world deployment.