Le Yu
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- MobileODE: An Extra Lightweight Network