benchmark evaluation
The process of systematically assessing an AI model against benchmark datasets to quantify its performance and compare it with other models.
- AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
- AugGen: Synthetic Augmentation using Diffusion Models Can Improve Recognition
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- BMW: Bidirectionally Memory bank reWriting for Unsupervised Person Re-Identification
- BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
- C-SEO Bench: Does Conversational SEO Work?
- CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
- CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
- CLEAR: Command Level Annotated Dataset for Ransomware Detection
- Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- CausalDynamics: A large‐scale benchmark for structural discovery of dynamical causal models
- ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- DISC: Dynamic Decomposition Improves LLM Inference Scaling
- DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
- DeepDiver: Adaptive Web-Search Intensity Scaling via Reinforcement Learning
- Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
- Dynamic Shadow Unveils Invisible Semantics for Video Outpainting
- EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
- EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis
- Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
- FastJAM: a Fast Joint Alignment Model for Images
- From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
- Generalizing Verifiable Instruction Following
- Generating and Checking DNN Verification Proofs
- InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
- Introducing FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark
- LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
- Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab
- MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
- Meta-Learning Objectives for Preference Optimization
- MigGPT: Harnessing Large Language Models for Automated Migration of Out-of-Tree Linux Kernel Patches Across Versions
- NavBench: Probing Multimodal Large Language Models for Embodied Navigation
- NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
- OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
- ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
- PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
- PhysDrive: A Multimodal Remote Physiological Measurement Dataset for In-vehicle Driver Monitoring
- PoLAR: Polar-Decomposed Low-Rank Adapter Representation
- PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
- Preserving Task-Relevant Information Under Linear Concept Removal
- ROOT: Rethinking Offline Optimization as Distributional Translation via Probabilistic Bridge
- RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
- Repo2Run: Automated Building Executable Environment for Code Repository at Scale
- Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset
- SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
- Template-Guided 3D Molecular Pose Generation via Flow Matching and Differentiable Optimization
- Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
- Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
- Two Causally Related Needles in a Video Haystack
- Weak-shot Keypoint Estimation via Keyness and Correspondence Transfer
- What’s in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
- WorldModelBench: Judging Video Generation Models As World Models
- macOSWorld: A Multilingual Interactive Benchmark for GUI Agents