Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question Answering

Jianwen Jiang; Ziqiang Chen; Haojie Lin; Xibin Zhao; Yue Gao

doi:10.1609/aaai.v34i07.6766

Authors

Jianwen Jiang Tsinghua University
Ziqiang Chen Tsinghua University
Haojie Lin Tsinghua University
Xibin Zhao Tsinghua University
Yue Gao Tsinghua University

DOI:

https://doi.org/10.1609/aaai.v34i07.6766

Abstract

Understanding questions and finding clues for answers are the key for video question answering. Compared with image question answering, video question answering (Video QA) requires to find the clues accurately on both spatial and temporal dimension simultaneously, and thus is more challenging. However, the relationship between spatio-temporal information and question still has not been well utilized in most existing methods for Video QA. To tackle this problem, we propose a Question-Guided Spatio-Temporal Contextual Attention Network (QueST) method. In QueST, we divide the semantic features generated from question into two separate parts: the spatial part and the temporal part, respectively guiding the process of constructing the contextual attention on spatial and temporal dimension. Under the guidance of the corresponding contextual attention, visual features can be better exploited on both spatial and temporal dimensions. To evaluate the effectiveness of the proposed method, experiments are conducted on TGIF-QA dataset, MSRVTT-QA dataset and MSVD-QA dataset. Experimental results and comparisons with the state-of-the-art methods have shown that our method can achieve superior performance.

Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question Answering

Authors

DOI:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information