performance benchmarks
Standardized tests or metrics used to evaluate the capabilities and reliability of AI models, facilitating comparison within the field.
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- Boosting Knowledge Utilization in Multimodal Large Language Models via Adaptive Logits Fusion and Attention Reallocation
- Boosting Knowledge Utilization in Multimodal Large Language Models via Adaptive Logits Fusion and Attention Reallocation
- Continuous Diffusion Model for Language Modeling
- DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing
- DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
- Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
- Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
- GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning
- Generative RLHF-V: Learning Principles from Multi-modal Human Preference
- Learning Neural Exposure Fields for View Synthesis
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
- Preference Distillation via Value based Reinforcement Learning
- Probing Hidden Knowledge Holes in Unlearned LLMs
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
- SteerConf: Steering LLMs for Confidence Elicitation
- The Underappreciated Power of Vision Models for Graph Structural Understanding