multi-turn interaction
- DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning