performance comparison
The evaluation of different AI models or algorithms based on defined metrics, which helps determine their effectiveness and suitability for specific tasks or datasets. This often involves benchmarking against standard datasets.
- An Adaptive Algorithm for Bilevel Optimization on Riemannian Manifolds
- Approximately Aligned Decoding
- Conditional Distribution Compression via the Kernel Conditional Mean Embedding
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
- DUO: No Compromise to Accuracy Degradation
- Diffusion Beats Autoregressive in Data-Constrained Settings
- Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
- Exploiting Dynamic Sparsity in Einsum
- Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial Inference
- LLMs Encode Harmfulness and Refusal Separately
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
- QFFT, Question-Free Fine-Tuning for Adaptive Reasoning
- ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
- RidgeLoRA: Matrix Ridge Enhanced Low-Rank Adaptation of Large Language Models
- Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
- Sherlock: Self-Correcting Reasoning in Vision-Language Models
- Sound Logical Explanations for Mean Aggregation Graph Neural Networks
- Unveiling the Uncertainty in Embodied and Operational Carbon of Large AI Models through a Probabilistic Carbon Accounting Model
- rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset