model capabilities
Model capabilities refer to the range and scope of tasks that an AI model can successfully perform. This encompasses the model's robustness, generalization, interpretability, and effectiveness across different datasets and applications.
- BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
- ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
- Training Language Models to Reason Efficiently
- We Should Chart an Atlas of All the World's Models