temporal understanding
- DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding
- IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
- NavBench: Probing Multimodal Large Language Models for Embodied Navigation
- TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language Models
- VideoLucy: Deep Memory Backtracking for Long Video Understanding