benchmark performance
The measured effectiveness of an AI model against standardized tasks, allowing for comparison with other models or techniques in the field.
- Activation-Informed Merging of Large Language Models
- Erasing Conceptual Knowledge from Language Models
- Flow Field Reconstruction with Sensor Placement Policy Learning
- Measuring AI Ability to Complete Long Software Tasks
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
- RESPIN-S1.0: A read speech corpus of 10000+ hours in dialects of nine Indian Languages
- SkyLadder: Better and Faster Pretraining via Context Window Scheduling
- Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
- T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning
- VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
- When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration