evaluation framework
An evaluation framework is a structured approach designed to assess the performance of AI models against defined criteria or benchmarks. It includes methodologies and metrics that guide the comparison of model predictions to actual outcomes, ensuring that the evaluation is consistent, fair, and reproducible.
- AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
- CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
- CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
- EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
- Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
- Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
- GC4NC: A Benchmark Framework for Graph Condensation on Node Classification with New Insights
- Generating Multi-Table Time Series EHR from Latent Space with Minimal Preprocessing
- IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
- Is Artificial Intelligence Generated Image Detection a Solved Problem?
- Latency NMS Attacks: Is It Real Life or Is It Just Fantasy?
- Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- Meta-World+: An Improved, Standardized, RL Benchmark
- Probing Hidden Knowledge Holes in Unlearned LLMs
- RADAR: Benchmarking Language Models on Imperfect Tabular Data
- Scalable Evaluation and Neural Models for Compositional Generalization
- TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
- TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language Models
- The Leaderboard Illusion
- ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
- TransferBench: Benchmarking Ensemble-based Black-box Transfer Attacks
- Unveiling the Learning Mind of Language Models: A Cognitive Framework and Empirical Study
- When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration