Kai Zhang
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
- A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking
- ARM: Adaptive Reasoning Model
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
- OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions
- Personalized Visual Content Generation in Conversational Systems
- Results of the Big ANN: NeurIPS’23 competition
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval