MUD evaluation frameworks reveal LLM-judge blind spots that aggregate metrics miss
Researchers identify how LLM-based evaluation can systematize AI model assessment while exposing vulnerabilities in Cohen's kappa and similar aggregate metrics.
Researchers identify how LLM-based evaluation can systematize AI model assessment while exposing vulnerabilities in Cohen's kappa and similar aggregate metrics.
OpenAI reveals that two API settings—preserved chain-of-thought and output compaction—nearly tripled benchmark performance, exposing how harness design shapes model evaluation.
Intel's benchmarking study reveals that enterprise AI agents succeed or fail based on infrastructure, not model performance alone.
A new 1M-human-rating benchmark reveals that conversational speech systems underperform on emotional nuance, speaker consistency, and accent handling despite saturation on latency metrics.
EEE and Community Evals now interoperate, unifying scattered evaluation results across 229K+ benchmark runs into a single standardized schema.
The first open far-field ASR benchmark measures speech recognition performance across simulated rooms with realistic acoustics, revealing significant performance gaps in noisy conditions.
Tenure AI exposes how LLM-based evaluation can miss silent failures in agent behavior, scoring high on plausibility while missing ground-truth task completion.
Miami-based startup Subquadratic released third-party benchmarks for SubQ, its new model claiming 12x context scaling and lower energy consumption than existing LLMs.
Hugging Face introduces tool-specific benchmarking methodology that measures not just correctness but token efficiency for coding agents interacting with library APIs.
Allen AI's olmo-eval extends the OLMES benchmark standard with flexible, composable evaluation infrastructure for model development iterations.
ServiceNow and Hugging Face benchmark ASR models on bilingual customer interactions, revealing significant performance gaps when speakers mix languages mid-sentence.
A new analysis shows that large language models excel at language tasks but struggle with seemingly simple visual reasoning—like reading analog clocks.
Hugging Face introduces private ASR evaluation datasets from Appen Inc. and DataoceanAI to block benchmaxxing, with scores visible via an opt-in toggle.
GitHub user erogol's BlaGPT offers an open-source research sandbox for evaluating LM architectures and components on compact datasets.