Lei Zhang
- BurstDeflicker: A Benchmark Dataset for Flicker Removal in Dynamic Scenes
- DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
- DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
- Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence Patterns
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- InstructRestore: Region-Customized Image Restoration with Human Instructions
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
- MobileODE: An Extra Lightweight Network
- One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis
- PASS: Path-selective State Space Model for Event-based Recognition
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
- Polyline Path Masked Attention for Vision Transformer
- Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection
- Rethinking Out-of-Distribution Detection and Generalization with Collective Behavior Dynamics
- The Underappreciated Power of Vision Models for Graph Structural Understanding
- VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank