action recognition
- EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
- MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
- OSKAR: Omnimodal Self-supervised Knowledge Abstraction and Representation
- Prompt-guided Disentangled Representation for Action Recognition
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding