TransZero: Attribute-Guided Transformer for Zero-Shot Learning

Shiming Chen; Ziming Hong; Yang Liu; Guo-Sen Xie; Baigui Sun; Hao Li; Qinmu Peng; Ke Lu; Xinge You

doi:10.1609/aaai.v36i1.19909

Authors

Shiming Chen Huazhong University of Science and Technology
Ziming Hong Huazhong University of Science and Technology
Yang Liu Alibaba Group
Guo-Sen Xie Inception Institute of Artificial Intelligence
Baigui Sun Alibaba Group
Hao Li Alibaba Group
Qinmu Peng Huazhong University of Science and Technology
Ke Lu University of Chinese Academy of Sciences
Xinge You Huazhong University of Science and Technology

DOI:

https://doi.org/10.1609/aaai.v36i1.19909

Keywords:

Computer Vision (CV), Machine Learning (ML)

Abstract

Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is learned from attribute descriptions shared between different classes, which are strong prior for localization of object attribute for representing discriminative region features enabling significant visual-semantic interaction. Although few attention-based models have attempted to learn such region features in a single image, the transferability and discriminative attribute localization of visual features are typically neglected. In this paper, we propose an attribute-guided Transformer network to learn the attribute localization for discriminative visual-semantic embedding representations in ZSL, termed TransZero. Specifically, TransZero takes a feature augmentation encoder to alleviate the cross-dataset bias between ImageNet and ZSL benchmarks and improve the transferability of visual features by reducing the entangled relative geometry relationships among region features. To learn locality-augmented visual features, TransZero employs a visual-semantic decoder to localize the most relevant image regions to each attributes from a given image under the guidance of attribute semantic information. Then, the locality-augmented visual features and semantic vectors are used for conducting effective visual-semantic interaction in a visual-semantic embedding network. Extensive experiments show that TransZero achieves a new state-of-the-art on three ZSL benchmarks. The codes are available at: https://github.com/shiming-chen/TransZero.

TransZero: Attribute-Guided Transformer for Zero-Shot Learning

Authors

DOI:

Keywords:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information

Subscription