performance gains
Improvements in an AI system's efficiency, accuracy, or effectiveness resulting from changes to models, data, or training processes.
- A Learning-Augmented Dynamic Programming Approach for Orienteering Problem with Time Windows
- AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
- ARIA: Training Language Agents with Intention-driven Reward Aggregation
- Automatic Auxiliary Task Selection and Adaptive Weighting Boost Molecular Property Prediction
- Balancing Multimodal Training Through Game-Theoretic Regularization
- Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
- Eliciting Reasoning in Language Models with Cognitive Tools
- End-to-End Vision Tokenizer Tuning
- Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
- FlashBias: Fast Computation of Attention with Bias
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
- Generative Data Augmentation via Diffusion Distillation, Adversarial Alignment, and Importance Reweighting
- Group-Level Data Selection for Efficient Pretraining
- Group-in-Group Policy Optimization for LLM Agent Training
- Improving Deep Learning for Accelerated MRI With Data Filtering
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame Projections
- MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
- Multi-Modal Interactive Agent Layer for Few-Shot Universal Cross-Domain Retrieval and Beyond
- Multi-Scale Finetuning for Encoder-based Time Series Foundation Models
- On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
- OpenCUA: Open Foundations for Computer-Use Agents
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks
- RLVR-World: Training World Models with Reinforcement Learning
- SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
- Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
- Thompson Sampling in Function Spaces via Neural Operators
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
- Towards Generalizable 3D Human Pose Estimation via Ensembles on Flat Loss Landscapes
- Towards Graph Foundation Models: Training on Knowledge Graphs Enables Transferability to General Graphs
- Understanding Differential Transformer Unchains Pretrained Self-Attentions
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
- Zero-shot protein stability prediction by inverse folding models: a free energy interpretation