evaluation suite
An evaluation suite is a collection of benchmarks, metrics, and datasets used to assess the performance of AI models consistently across different tasks and standards.
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
- MergeBench: A Benchmark for Merging Domain-Specialized LLMs
- MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
- Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
- VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs