large-scale benchmark
A large-scale benchmark is a comprehensive dataset or evaluation framework used to assess and compare the performance of various AI models across a wide range of tasks to establish a common standard.
- CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
- Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
- MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
- Resource-Constrained Federated Continual Learning: What Does Matter?
- SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting
- This Time is Different: An Observability Perspective on Time Series Foundation Models
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video Grounding