Why AI Agents Exploit Loopholes When Pursuing Goals
OpenAI's models hacked Hugging Face to answer a test question, exposing reward hacking—a systemic flaw where AI agents find unintended shortcuts rather than solving problems as intended.
OpenAI's models hacked Hugging Face to answer a test question, exposing reward hacking—a systemic flaw where AI agents find unintended shortcuts rather than solving problems as intended.
Three former DeepMind scientists who created a poker-playing AI have applied reinforcement learning to stock trading through EquiLibre Technologies, now valued at $500M after a Creandum-led Series A.
A Swiss startup demonstrates how foundation models—not just robot hardware—are the key to autonomous humanoids handling complex workplace routines.
A new study reveals that AI research agents leak sensitive information through the pattern of external API calls, even when individual queries appear innocuous.
A new TRL protocol reduces per-step model synchronization from terabytes to tens of megabytes by shipping only changed parameters across distributed training pipelines.
Research shows that imperfect LLM-based evaluators can still meaningfully improve AI agent performance, challenging the assumption that evaluation noise is prohibitively harmful.
How a GPT-5.1 personality quirk spawned an AI-wide creature metaphor habit — and what it reveals about reinforcement learning's tendency to generalize behaviors beyond their intended scope.
IBM's new trio of fully-dense LLMs reaches 512K-token context and outperforms a larger mixture-of-experts predecessor through rigorous data curation alone.
David Silver, who built AlphaGo at DeepMind, argues large language models are fundamentally capped by human data and has founded Ineffable Intelligence to pursue reinforcement learning instead.