performance metrics
Performance metrics are quantitative measures used to evaluate the effectiveness of an AI model. Common metrics include accuracy, precision, recall, F1 score, and area under the ROC curve (AUC), which provide insight into how well the model is making predictions.
- 3D Visual Illusion Depth Estimation
- Benchmarking End-To-End Performance of AI-Based Chip Placement Algorithms
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
- Causal LLM Routing: End-to-End Regret Minimization from Observational Data
- Compiler-R1: Towards Agentic Compiler Auto-tuning with Reinforcement Learning
- ConStellaration: A dataset of QI-like stellarator plasma boundaries and optimization benchmarks
- Estimating Model Performance Under Covariate Shift Without Labels
- Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
- Learning (Approximately) Equivariant Networks via Constrained Optimization
- Learning (Approximately) Equivariant Networks via Constrained Optimization
- Memory Mosaics at scale
- Memory Mosaics at scale
- Meta-Learning Objectives for Preference Optimization
- Monitoring Risks in Test-Time Adaptation
- Multivariate Time Series Anomaly Detection with Idempotent Reconstruction
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation
- Near-Optimal Regret-Queue Length Tradeoff in Online Learning for Two-Sided Markets
- On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
- PaZO: Preconditioned Accelerated Zeroth-Order Optimization for Fine-Tuning LLMs
- PerturBench: Benchmarking Machine Learning Models for Cellular Perturbation Analysis
- Resounding Acoustic Fields with Reciprocity
- SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
- The Rashomon Set Has It All: Analyzing Trustworthiness of Trees under Multiplicity
- URB - Urban Routing Benchmark for RL-equipped Connected Autonomous Vehicles
- Weak-to-Strong Generalization under Distribution Shifts
- ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual Decoding