performance evaluation
The assessment of an AI system's effectiveness based on predetermined metrics, essential for understanding its capabilities and limitations.
- $\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
- A Near-Optimal Algorithm for Decentralized Convex-Concave Finite-Sum Minimax Optimization
- AC-LoRA: (Almost) Training-Free Access Control Aware Multi-Modal LLMs
- AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios
- AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation
- ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data
- Accelerating Optimization via Differentiable Stopping Time
- Addressing Mark Imbalance in Integration-free Marked Temporal Point Processes
- AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?
- Alias-Free ViT: Fractional Shift Invariance via Linear Attention
- Angular Constraint Embedding via SpherePair Loss for Constrained Clustering
- Anomaly Detection by an Ensemble of Random Pairs of Hyperspheres
- BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
- BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision Models
- Boosting the Uniqueness of Neural Networks Fingerprints with Informative Triggers
- COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
- Can LLMs Outshine Conventional Recommenders? A Comparative Evaluation
- CellVerse: Do Large Language Models Really Understand Cell Biology?
- ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World
- CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
- Collapsing Taylor Mode Automatic Differentiation
- Compositional Neural Network Verification via Assume-Guarantee Reasoning
- ConnectomeBench: Can LLMs proofread the connectome?
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
- DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding
- Delving into Large Language Models for Effective Time-Series Anomaly Detection
- Detecting High-Stakes Interactions with Activation Probes
- Efficient Multimodal Dataset Distillation via Generative Models
- Enhancing Consistency of Flow-Based Image Editing through Kalman Control
- EuroSpeech: A Multilingual Speech Corpus
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- From Indicators to Insights: Diversity-Optimized for Medical Series-Text Decoding via LLMs
- IR-OptSet: An Optimization-Sensitive Dataset for Advancing LLM-Based IR Optimizer
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
- JAFAR: Jack up Any Feature at Any Resolution
- KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
- LIFEBENCH: Evaluating Length Instruction Following in Large Language Models
- Last-Iterate Convergence of Smooth Regret Matching$^+$ Variants in Learning Nash Equilibria
- Learning to Insert for Constructive Neural Vehicle Routing Solver
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- Linguini: A benchmark for language-agnostic linguistic reasoning
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
- LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
- MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
- MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
- MS-Bench: Evaluating LMMs in Ancient Manuscript Study through a Dunhuang Case Study
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
- NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
- NerfBaselines: Consistent and Reproducible Evaluation of Novel View Synthesis Methods
- OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
- Object-Centric Concept-Bottlenecks
- OmniTry: Virtual Try-On Anything without Masks
- On the Integration of Spatial-Temporal Knowledge: A Lightweight Approach to Atmospheric Time Series Forecasting
- PanTS: The Pancreatic Tumor Segmentation Dataset
- Procurement Auctions with Predictions: Improved Frugality for Facility Location
- PyraMotion: Attentional Pyramid-Structured Motion Integration for Co-Speech 3D Gesture Synthesis
- Quantifying Cross-Modality Memorization in Vision-Language Models
- Relieving the Over-Aggregating Effect in Graph Transformers
- Resource-Constrained Federated Continual Learning: What Does Matter?
- Rethinking Neural Combinatorial Optimization for Vehicle Routing Problems with Different Constraint Tightness Degrees
- Rethinking Verification for LLM Code Generation: From Generation to Testing
- Rotary Masked Autoencoders are Versatile Learners
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models
- Semi-supervised Graph Anomaly Detection via Robust Homophily Learning
- SolidGeo: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
- SpEx: A Spectral Approach to Explainable Clustering
- SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMs
- Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
- The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
- Towards A Generalist Code Embedding Model Based On Massive Data Synthesis
- Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
- Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
- UMA: A Family of Universal Models for Atoms
- Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
- VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
- VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
- What's Producible May Not Be Reachable: Measuring the Steerability of Generative Models
- When No Paths Lead to Rome: Benchmarking Systematic Neural Relational Reasoning
- WritingBench: A Comprehensive Benchmark for Generative Writing
- ZEUS: Zero-shot Embeddings for Unsupervised Separation of Tabular Data
- Zero-shot World Models via Search in Memory
- nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning