Qipeng Guo
- Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
- Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
- Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
- Pre-Trained Policy Discriminators are General Reward Models
- World-aware Planning Narratives Enhance Large Vision-Language Model Planner