yelong shen
- Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
- SAS: Simulated Attention Score
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning