model performance
A measure of how well an AI model accomplishes its intended tasks, typically evaluated through various metrics such as accuracy, precision, or recall.
- ARM: Adaptive Reasoning Model
- Abstain Mask Retain Core: Time Series Prediction by Adaptive Masking Loss with Representation Consistency
- Accelerating Block Coordinate Descent for LLM Finetuning via Landscape Expansion
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Adaptive Latent-Space Constraints in Personalized Federated Learning
- Antidistillation Sampling
- Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
- Augmenting Biological Fitness Prediction Benchmarks with Landscapes Features from GraphFLA
- Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
- Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents
- Benchmarking Large Language Models with Integer Sequence Generation Tasks
- Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
- Bigram Subnetworks: Mapping to Next Tokens in Transformer Language Models
- Bridging Theory and Practice in Link Representation with Graph Neural Networks
- Channel Matters: Estimating Channel Influence for Multivariate Time Series
- Compact Memory for Continual Logistic Regression
- Complexity Scaling Laws for Neural Models using Combinatorial Optimization
- Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
- Datasets, Documents, and Repetitions: The Practicalities of Unequal Data Quality
- DeltaFormer: Unlock the state space of Transformer
- Demystifying Network Foundation Models
- Dimensional Collapse in VQVAEs: Evidence and Remedies
- Distributionally Robust Feature Selection
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
- Enhanced Expert Merging for Mixture-of-Experts in Graph Foundation Models
- Enhancing Sample Selection Against Label Noise by Cutting Mislabeled Easy Examples
- Entropy-Calibrated Label Distribution Learning
- EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
- Factor Decorrelation Enhanced Data Removal from Deep Predictive Models
- Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
- Fairshare Data Pricing via Data Valuation for Large Language Models
- Fantastic Bugs and Where to Find Them in AI Benchmarks
- FedGPS: Statistical Rectification Against Data Heterogeneity in Federated Learning
- FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
- FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models
- Gated Integration of Low-Rank Adaptation for Continual Learning of Large Language Models
- How Different from the Past? Spatio-Temporal Time Series Forecasting with Self-Supervised Deviation Learning
- How Many Domains Suffice for Domain Generalization? A Tight Characterization via the Domain Shattering Dimension
- Hyper-Modality Enhancement for Multimodal Sentiment Analysis with Missing Modalities
- Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
- Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
- Is Limited Participant Diversity Impeding EEG-based Machine Learning?
- KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
- LLM at Network Edge: A Layer-wise Efficient Federated Fine-tuning Approach
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- Memorization in Graph Neural Networks
- Modeling the Economic Impacts of AI Openness Regulation
- NaDRO: Leveraging Dual-Reward Strategies for LLMs Training on Noisy Data
- OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
- On Transferring Transferability: Towards a Theory for Size Generalization
- PROFIT: A Specialized Optimizer for Deep Fine Tuning
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Prediction-Powered Semi-Supervised Learning with Online Power Tuning
- Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
- Privacy Reasoning in Ambiguous Contexts
- ProteinConformers: Benchmark Dataset for Simulating Protein Conformational Landscape Diversity and Plausibility
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks
- RCCDA: Adaptive Model Updates in the Presence of Concept Drift under a Constrained Resource Budget
- ROSE: Remove Objects with Side Effects in Videos
- RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
- Reinforced Context Order Recovery for Adaptive Reasoning and Planning
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- Rethinking Tokenized Graph Transformers for Node Classification
- Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
- Revisiting Semi-Supervised Learning in the Era of Foundation Models
- Revolutionizing Graph Aggregation: From Suppression to Amplification via BoostGCN
- Revolutionizing Training-Free NAS: Towards Efficient Automatic Proxy Discovery via Large Language Models
- SPFL: Sequential updates with Parallel aggregation for Enhanced Federated Learning under Category and Domain Shifts
- Scaling Laws for Optimal Data Mixtures
- SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
- Shapley-Based Data Valuation for Weighted $k$-Nearest Neighbors
- Spend Wisely: Maximizing Post-Training Gains in Iterative Synthetic Data Bootstrapping
- SubTrack++ : Gradient Subspace Tracking for Scalable LLM Training
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- The Best Instruction-Tuning Data are Those That Fit
- Two Causally Related Needles in a Video Haystack
- UMoE: Unifying Attention and FFN with Shared Experts
- What Makes a Reward Model a Good Teacher? An Optimization Perspective