complex reasoning tasks
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
- Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
- EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction Evolution
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO