evaluation protocols
Evaluation protocols consist of standardized methods for assessing the performance and robustness of machine learning models, including metrics, benchmarks, and datasets used to ensure reliability and comparability of results across studies.
- DGCBench: A Deep Graph Clustering Benchmark
- DiffBreak: Is Diffusion-Based Purification Robust?
- NerfBaselines: Consistent and Reproducible Evaluation of Novel View Synthesis Methods
- PSI: A Benchmark for Human Interpretation and Response in Traffic Interactions
- PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
- Rethinking Evaluation of Infrared Small Target Detection
- Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models
- Toward Real-world Text Image Forgery Localization: Structured and Interpretable Data Synthesis