Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving of Inequalities
algebraic rewritingam/gm inequalitycompositional reasoningcompositional settingdeepseek-prover-v2-7bformal proof assistantsgeneralization behaviorllmmathematical discoverymathematical inequalitiesmathematical intuitionmulti-step compositionperformance dropsyntactic correctnessvariable duplication
LLM-based formal proof assistants (e.g., in Lean) hold great promise for automating mathematical discovery. But beyond syntactic correctness, do these systems truly understand mathematical structure as humans do? We investigate this question in context of mathematical inequalities