BenchMIRT: What are LLM benchmarks actually measuring?
Hugging Face
Read full postBenchMIRT is a new method developed by AllenAI to analyze large language model (LLM) benchmarks at the level of individual prompts. It uses multidimensional item response theory to identify which underlying capabilities influence performance on specific benchmark questions, revealing that benchmarks often measure multiple abilities beyond their stated goals.


