Yiming Dong
- AdaMSS: Adaptive Multi-Subspace Approach for Parameter-Efficient Fine-Tuning
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
- Stepsize anything: A unified learning rate schedule for budgeted-iteration training