reasoning benchmarks

Standard tests or datasets used to evaluate the reasoning capabilities of AI models, assessing their performance on tasks requiring logical deduction.

27 papers