evaluation benchmarks
Evaluation benchmarks are reference datasets and metrics used to objectively compare the performance of different models or approaches within a specific field, facilitating advancements by providing a standardized way to assess progress.
- FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation