empirical results
Data gathered from practical experiments or observations, which help validate models and theories in AI and provide insights into their effectiveness.
- $\epsilon$-Seg: Sparsely Supervised Semantic Segmentation of Microscopy Data
- A Cautionary Tale on Integrating Studies with Disparate Outcome Measures for Causal Inference
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio–Language Models
- Accurate and Efficient Low-Rank Model Merging in Core Space
- AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
- Adversarial Diffusion for Robust Reinforcement Learning
- Algorithm- and Data-Dependent Generalization Bounds for Diffusion Models
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- BayeSQP: Bayesian Optimization through Sequential Quadratic Programming
- Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
- CDFlow: Building Invertible Layers with Circulant and Diagonal Matrices
- Characterizing the Expressivity of Fixed-Precision Transformer Language Models
- Conditioning Matters: Training Diffusion Policies is Faster Than You Think
- Conformal Online Learning of Deep Koopman Linear Embeddings
- Continual Multimodal Contrastive Learning
- Convex Approximation of Two-Layer ReLU Networks for Hidden State Differential Privacy
- Copresheaf Topological Neural Networks: A Generalized Deep Learning Framework
- Coupled Data and Measurement Space Dynamics for Enhanced Diffusion Posterior Sampling
- DeCaFlow: A deconfounding causal generative model
- Defending Multimodal Backdoored Models by Repulsive Visual Prompt Tuning
- Differentiable Structure Learning and Causal Discovery for General Binary Data
- Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models
- Diffusion Models Meet Contextual Bandits
- Diffusion Models and the Manifold Hypothesis: Log-Domain Smoothing is Geometry Adaptive
- Distance-informed Neural Processes
- Don’t Trade Off Safety: Diffusion Regularization for Constrained Offline RL
- E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor Products
- Edit Flows: Variable Length Discrete Flow Matching with Sequence-Level Edit Operations
- Efficient Randomized Experiments Using Foundation Models
- Enabling Differentially Private Federated Learning for Speech Recognition: Benchmarks, Adaptive Optimizers, and Gradient Clipping
- EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
- Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
- Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
- Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient Adjustment
- Fast Local Search Algorithms for Clustering with Adaptive Sampling and Bandit Strategies
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- Geometry-Aware Edge Pooling for Graph Neural Networks
- How Well Can Differential Privacy Be Audited in One Run?
- Implicit Generative Property Enhancer
- Improving Decision Trees through the Lens of Parameterized Local Search
- Integration Matters for Learning PDEs with Backwards SDEs
- KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
- Learning Expandable and Adaptable Representations for Continual Learning
- Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data
- Learning Shared Representations from Unpaired Data
- Learning from Demonstrations via Capability-Aware Goal Sampling
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
- Local Curvature Descent: Squeezing More Curvature out of Standard and Polyak Gradient Descent
- Mitigating Spurious Features in Contrastive Learning with Spectral Regularization
- MoFo: Empowering Long-term Time Series Forecasting with Periodic Pattern Modeling
- NaDRO: Leveraging Dual-Reward Strategies for LLMs Training on Noisy Data
- Non-Convex Tensor Recovery from Tube-Wise Sensing
- On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
- On the Existence and Complexity of Core-Stable Data Exchanges
- One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement Learning
- Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-Making
- Online Optimization for Offline Safe Reinforcement Learning
- OrdShap: Feature Position Importance for Sequential Black-Box Models
- Pre-Trained Policy Discriminators are General Reward Models
- Principled Model Routing for Unknown Mixtures of Source Domains
- Probabilistic Token Alignment for Large Language Model Fusion
- Protein Design with Dynamic Protein Vocabulary
- Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models
- ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
- Reliably detecting model failures in deployment without labels
- Residual Stream Analysis of Overfitting And Structural Disruptions
- Rethinking Fair Federated Learning from Parameter and Client View
- Rethinking Tokenized Graph Transformers for Node Classification
- Robust and Computation-Aware Gaussian Processes
- S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
- SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL
- SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly
- Sampling by averaging: A multiscale approach to score estimation
- Shortcuts and Identifiability in Concept-based Models from a Neuro-Symbolic Lens
- Simulation-Based Inference for Adaptive Experiments
- Spectral Graph Coarsening Using Inner Product Preservation and the Grassmann Manifold
- Spectral Perturbation Bounds for Low-Rank Approximation with Applications to Privacy
- Spectral Perturbation Bounds for Low-Rank Approximation with Applications to Privacy
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold
- Stepsize anything: A unified learning rate schedule for budgeted-iteration training
- Structured Initialization for Vision Transformers
- Table as a Modality for Large Language Models
- The Omni-Expert: A Computationally Efficient Approach to Achieve a Mixture of Experts in a Single Expert Model
- Thompson Sampling for Multi-Objective Linear Contextual Bandit
- Tight Asymptotics of Extreme Order Statistics
- Time-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series
- Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
- Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
- Understanding and Improving Adversarial Robustness of Neural Probabilistic Circuits
- Universally Invariant Learning in Equivariant GNNs
- Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search