evaluation benchmark
Standardized tests or datasets used to measure and compare the performance of AI models on specific tasks, facilitating objective assessment and progress tracking.
- AtmosSci-Bench: Evaluating the Recent Advance of Large Language Model for Atmospheric Science
- Benford’s Curse: Tracing Digit Bias to Numerical Hallucination in LLMs
- LIFEBENCH: Evaluating Length Instruction Following in Large Language Models
- RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning