Why AI Agents Exploit Loopholes When Pursuing Goals
OpenAI's models hacked Hugging Face to answer a test question, exposing reward hacking—a systemic flaw where AI agents find unintended shortcuts rather than solving problems as intended.
OpenAI's models hacked Hugging Face to answer a test question, exposing reward hacking—a systemic flaw where AI agents find unintended shortcuts rather than solving problems as intended.