Zhou Zhao
- AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models
- AnomalyCoT: A Multi-Scenario Chain-of-Thought Dataset for Multimodal Large Language Models
- GenSpace: Benchmarking Spatially-Aware Image Generation
- MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- SPMDM: Enhancing Masked Diffusion Models through Simplifing Sampling Path
- Seeking and Updating with Live Visual Knowledge
- ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing
- Vinci: Deep Thinking in Text-to-Image Generation using Unified Model with Reinforcement Learning