empirical evaluation
The process of assessing a model's performance through experimental validation, utilizing metrics and benchmarks to provide evidence for its effectiveness on specific tasks.
- $\textit{HiMaCon:}$ Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data
- A Circular Argument: Does RoPE need to be Equivariant for Vision?
- AccuQuant: Simulating Multiple Denoising Steps for Quantizing Diffusion Models
- Are Language Models Efficient Reasoners? A Perspective from Logic Programming
- Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection
- Bandit Guided Submodular Curriculum for Adaptive Subset Selection
- Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining
- CCL: Causal-aware In-context Learning for Out-of-Distribution Generalization
- Causal Mixture Models: Characterization and Discovery
- Chain of Execution Supervision Promotes General Reasoning in Large Language Models
- CoLT: The conditional localization test for assessing the accuracy of neural posterior estimates
- Competitive Advantage Attacks to Decentralized Federated Learning
- Conformal Prediction for Ensembles: Improving Efficiency via Score-Based Aggregation
- Constructing an Optimal Behavior Basis for the Option Keyboard
- Continual Optimization with Symmetry Teleportation for Multi-Task Learning
- Corporate Needs You to Find the Difference: Revisiting Submodular and Supermodular Ratio Optimization Problems
- Cost-aware LLM-based Online Dataset Annotation
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
- Discretization-free Multicalibration through Loss Minimization over Tree Ensembles
- ElliCE: Efficient and Provably Robust Algorithmic Recourse via the Rashomon Sets
- Enhancing Deep Batch Active Learning for Regression with Imperfect Data Guided Selection
- EquiTabPFN: A Target-Permutation Equivariant Prior Fitted Network
- Estimating Interventional Distributions with Uncertain Causal Graphs through Meta-Learning
- Exploring the limits of strong membership inference attacks on large language models
- Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial Inference
- Graph Few-Shot Learning via Adaptive Spectrum Experts and Cross-Set Distribution Calibration
- Improved Algorithms for Fair Matroid Submodular Maximization
- Improving Model-Based Reinforcement Learning by Converging to Flatter Minima
- Improving Regret Approximation for Unsupervised Dynamic Environment Generation
- Improving the Straight-Through Estimator with Zeroth-Order Information
- IneqSearch: Hybrid Reasoning for Olympiad Inequality Proofs
- Inexact Column Generation for Bayesian Network Structure Learning via Difference-of-Submodular Optimization
- Learning from Interval Targets
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
- Monoculture or Multiplicity: Which Is It?
- Nearly-Linear Time and Massively Parallel Algorithms for $k$-anonymity
- New Parallel and Streaming Algorithms for Directed Densest Subgraph
- Novel Exploration via Orthogonality
- PaZO: Preconditioned Accelerated Zeroth-Order Optimization for Fine-Tuning LLMs
- Permissioned LLMs: Enforcing Access Control in Large Language Models
- Private Zeroth-Order Optimization with Public Data
- ProtInvTree: Deliberate Protein Inverse Folding with Reward-guided Tree Search
- Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
- Quasi-Self-Concordant Optimization with $\ell_{\infty}$ Lewis Weights
- RUAGO: Effective and Practical Retain-Free Unlearning via Adversarial Attack and OOD Generator
- SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
- Scaling Speculative Decoding with Lookahead Reasoning
- Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
- Simple and Optimal Sublinear Algorithms for Mean Estimation
- Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss Landscapes
- Tensor Product Attention Is All You Need
- Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds
- Towards a Pairwise Ranking Model with Orderliness and Monotonicity for Label Enhancement
- Tracing Back the Malicious Clients in Poisoning Attacks to Federated Learning
- Transformers for Mixed-type Event Sequences
- Unveiling Concept Attribution in Diffusion Models
- Valid Selection among Conformal Sets
- Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
- \(\varepsilon\)-Optimally Solving Two-Player Zero-Sum POSGs