Jing Shao
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis
- Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- Systematic Reward Gap Optimization for Mitigating VLM Hallucinations