A Proposal-Based Approach for Activity Image-to-Video Retrieval


  • Ruicong Xu Shanghai Jiao Tong University
  • Li Niu Shanghai Jiao Tong University
  • Jianfu Zhang Shanghai Jiao Tong University
  • Liqing Zhang Shanghai Jiao Tong University




Activity image-to-video retrieval task aims to retrieve videos containing the similar activity as the query image, which is a challenging task because videos generally have many background segments irrelevant to the activity. In this paper, we utilize R-C3D model to represent a video by a bag of activity proposals, which can filter out background segments to some extent. However, there are still noisy proposals in each bag. Thus, we propose an Activity Proposal-based Image-to-Video Retrieval (APIVR) approach, which incorporates multi-instance learning into cross-modal retrieval framework to address the proposal noise issue. Specifically, we propose a Graph Multi-Instance Learning (GMIL) module with graph convolutional layer, and integrate this module with classification loss, adversarial loss, and triplet loss in our cross-modal retrieval framework. Moreover, we propose geometry-aware triplet loss based on point-to-subspace distance to preserve the structural information of activity proposals. Extensive experiments on three widely-used datasets verify the effectiveness of our approach.




How to Cite

Xu, R., Niu, L., Zhang, J., & Zhang, L. (2020). A Proposal-Based Approach for Activity Image-to-Video Retrieval. Proceedings of the AAAI Conference on Artificial Intelligence, 34(07), 12524-12531. https://doi.org/10.1609/aaai.v34i07.6941



AAAI Technical Track: Vision