Mohit Bansal
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding