model evaluation
Model evaluation encompasses the processes and methodologies used to assess how well a machine learning model performs against standard metrics. It involves comparing the model to benchmarks and understanding its strengths and weaknesses.
- All that structure matches does not glitter
- Behavior Injection: Preparing Language Models for Reinforcement Learning
- Continuous Concepts Removal in Text-to-image Diffusion Models
- Corrector Sampling in Language Models
- Efficient Part-level 3D Object Generation via Dual Volume Packing
- Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
- From Black-box to Causal-box: Towards Building More Interpretable Models
- General-Reasoner: Advancing LLM Reasoning Across All Domains
- HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
- Introducing FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark
- Learning from positive and unlabeled examples -Finite size sample bounds
- Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab
- OmniBench: Towards The Future of Universal Omni-Language Models
- OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
- Panoptic Captioning: An Equivalence Bridge for Image and Text
- Probing Hidden Knowledge Holes in Unlearned LLMs
- Reverse-Annealed Sequential Monte Carlo for Efficient Bayesian Optimal Experiment Design
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- Stratify or Die: Rethinking Data Splits in Image Segmentation
- The Rashomon Set Has It All: Analyzing Trustworthiness of Trees under Multiplicity
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding