LLM Judges Gave an Agent 0.85—Without Noticing It Never Opened the File
Tenure AI exposes how LLM-based evaluation can miss silent failures in agent behavior, scoring high on plausibility while missing ground-truth task completion.
Tenure AI exposes how LLM-based evaluation can miss silent failures in agent behavior, scoring high on plausibility while missing ground-truth task completion.