accuracy evaluation
The assessment of an AI model's predictive performance, typically measured by metrics such as precision, recall, F1 score, or accuracy rate, in order to determine how effectively the model can solve the tasks for which it was designed.
- AnomalyCoT: A Multi-Scenario Chain-of-Thought Dataset for Multimodal Large Language Models
- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
- EgoBlind: Towards Egocentric Visual Assistance for the Blind
- QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multimodal Outputs
- SeePhys: Does Seeing Help Thinking? – Benchmarking Vision-Based Physics Reasoning
- What Makes a Reward Model a Good Teacher? An Optimization Perspective
- miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward