evaluation frameworks
Structured approaches to assess the performance and effectiveness of AI models, often incorporating various metrics and benchmarks.
- AGI-Elo: How Far Are We From Mastering A Task?
- Fostering the Ecosystem of AI for Social Impact Requires Expanding and Strengthening Evaluation Standards
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- Scaling Physical Reasoning with the PHYSICS Dataset
- SynTSBench: Rethinking Temporal Pattern Learning in Deep Learning Models for Time Series