performance estimation
The assessment of how well an AI model or algorithm is likely to perform on unseen data, leveraging methodologies such as cross-validation, holdout validation, and benchmarking against established performance metrics.
- Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
- Establishing Best Practices in Building Rigorous Agentic Benchmarks
- Evaluating multiple models using labeled and unlabeled data
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness