Allen AI releases olmo-eval, a development-loop evaluation workbench for large language models
Allen AI's olmo-eval extends the OLMES benchmark standard with flexible, composable evaluation infrastructure for model development iterations.
Allen AI's olmo-eval extends the OLMES benchmark standard with flexible, composable evaluation infrastructure for model development iterations.