Tianyu Zhang
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
- Majority of the Bests: Improving Best-of-N via Bootstrapping
- MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
- PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models