evaluation metrics
Evaluation metrics are quantitative measures used to assess the performance of machine learning models. These can include accuracy, precision, recall, F1 score, and area under the curve (AUC), among others, depending on the specific task at hand.
- A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
- A Reliable Cryptographic Framework for Empirical Machine Unlearning Evaluation
- A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers
- ADPretrain: Advancing Industrial Anomaly Detection via Anomaly Representation Pretraining
- AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios
- Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences
- AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy
- Automatic Auxiliary Task Selection and Adaptive Weighting Boost Molecular Property Prediction
- BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model
- Correlation Dimension of Autoregressive Large Language Models
- CrossAD: Time Series Anomaly Detection with Cross-scale Associations and Cross-window Modeling
- Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation
- Diffusion Guided Adversarial State Perturbations in Reinforcement Learning
- Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation
- Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
- Don’t call it privacy-preserving or human-centric pose estimation if you don’t measure privacy
- Emergent Temporal Correspondences from Video Diffusion Transformers
- Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
- HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion
- ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
- IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
- Improving the Generation and Evaluation of Synthetic Data for Downstream Medical Causal Inference
- Information-Driven Design of Imaging Systems
- Information-Theoretic Reward Decomposition for Generalizable RLHF
- LawShift: Benchmarking Legal Judgment Prediction Under Statute Shifts
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- MUniverse: A Simulation and Benchmarking Suite for Motor Unit Decomposition
- Make Information Diffusion Explainable: LLM-based Causal Framework for Diffusion Prediction
- Metritocracy: Representative Metrics for Lite Benchmarks
- OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
- OpenLex3D: A Tiered Benchmark for Open-Vocabulary 3D Scene Representations
- OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
- PSBench: a large-scale benchmark for estimating the accuracy of protein complex structural models
- Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
- ProteinConformers: Benchmark Dataset for Simulating Protein Conformational Landscape Diversity and Plausibility
- Quantifying Generalisation in Imitation Learning
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
- Results of the Big ANN: NeurIPS’23 competition
- SMMILE: An expert-driven benchmark for multimodal medical in-context learning
- Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information Foraging
- SimWorld: An Open-ended Simulator for Agents in Physical and Social Worlds
- SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
- Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following
- Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders
- Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
- Tree of Preferences for Diversified Recommendation
- UMU-Bench: Closing the Modality Gap in Multimodal Unlearning Evaluation
- Uncertain Knowledge Graph Completion via Semi-Supervised Confidence Distribution Learning
- Understanding and Enhancing Message Passing on Heterophilic Graphs via Compatibility Matrix
- UniSite: The First Cross-Structure Dataset and Learning Framework for End-to-End Ligand Binding Site Detection
- VeriThoughts: Enabling Automated Verilog Code Generation using Reasoning and Formal Verification
- When Are Concepts Erased From Diffusion Models?
- When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions