video question answering
Video question answering is a task in AI that involves providing answers to questions based on the content of a video. It requires the integration of computer vision and natural language processing to accurately interpret both visual and textual information.
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- OSKAR: Omnimodal Self-supervised Knowledge Abstraction and Representation
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
- Seeing the Arrow of Time in Large Multimodal Models
- Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
- Towards Understanding Camera Motions in Any Video