Learning a Key-Value Memory Co-Attention Matching Network for Person Re-Identification


  • Yaqing Zhang Zhejiang University
  • Xi Li Zhejiang University
  • Zhongfei Zhang Zhejiang University




Person re-identification (Re-ID) is typically cast as the problem of semantic representation and alignment, which requires precisely discovering and modeling the inherent spatial structure information on person images. Motivated by this observation, we propose a Key-Value Memory Matching Network (KVM-MN) model that consists of key-value memory representation and key-value co-attention matching. The proposed KVM-MN model is capable of building an effective local-position-aware person representation that encodes the spatial feature information in the form of multi-head key-value memory. Furthermore, the proposed KVM-MN model makes use of multi-head co-attention to automatically learn a number of cross-person-matching patterns, resulting in more robust and interpretable matching results. Finally, we build a setwise learning mechanism that implements a more generalized query-to-gallery-image-set learning procedure. Experimental results demonstrate the effectiveness of the proposed model against the state-of-the-art.




How to Cite

Zhang, Y., Li, X., & Zhang, Z. (2019). Learning a Key-Value Memory Co-Attention Matching Network for Person Re-Identification. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01), 9235-9242. https://doi.org/10.1609/aaai.v33i01.33019235



AAAI Technical Track: Vision