benchmarking
The process of evaluating and comparing the performance of AI models against standardized datasets and metrics to gauge improvements and establish best practices within the field.
- A Unified Framework for Provably Efficient Algorithms to Estimate Shapley Values
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
- ChemPile: A 250 GB Diverse and Curated Dataset for Chemical Foundation Models
- Don’t call it privacy-preserving or human-centric pose estimation if you don’t measure privacy
- EconGym: A Scalable AI Testbed with Diverse Economic Tasks
- FAIR Universe HiggsML Uncertainty Dataset and Competition
- Holistic Order Prediction in Natural Scenes
- IRRISIGHT: A Large-Scale Multimodal Dataset and Scalable Pipeline to Address Irrigation and Water Management in Agriculture
- Information Retrieval Induced Safety Degradation in AI Agents
- MIRA: Medical Time Series Foundation Model for Real-World Health Data
- Making Classic GNNs Strong Baselines Across Varying Homophily: A Smoothness–Generalization Perspective
- Merlin L48 Spectrogram Dataset
- NOVA: A Benchmark for Rare Anomaly Localization and Clinical Reasoning in Brain MRI
- NOVA: A Benchmark for Rare Anomaly Localization and Clinical Reasoning in Brain MRI
- OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata
- OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata
- PerturBench: Benchmarking Machine Learning Models for Cellular Perturbation Analysis
- STSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex
- Seeking and Updating with Live Visual Knowledge
- Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
- Synthesizing Performance Constraints for Evaluating and Improving Code Efficiency
- THUNDER: Tile-level Histopathology image UNDERstanding benchmark
- The Catechol Benchmark: Time-series Solvent Selection Data for Few-shot Machine Learning
- The Leaderboard Illusion
- The Rashomon Set Has It All: Analyzing Trustworthiness of Trees under Multiplicity
- Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
- Torch-Uncertainty: Deep Learning Uncertainty Quantification
- Towards precision protein-ligand affinity prediction benchmark: A Complete and Modification-Aware DAVIS Dataset
- Unified Transferability Metrics for Time Series Foundation Models
- 🎧MOSPA: Human Motion Generation Driven by Spatial Audio