Hierarchical Semantic Alignment for Image Clustering

Authors

  • Xingyu Zhu University of Science and Technology of China Nanyang Technological University
  • Beier Zhu Nanyang Technological University
  • Yunfan Li Sichuan University
  • Junfeng Fang National University of Singapore
  • Shuo Wang University of Science and Technology of China
  • Kesen Zhao Nanyang Technological University
  • Hanwang Zhang Nanyang Technological University

DOI:

https://doi.org/10.1609/aaai.v40i34.40156

Abstract

Image clustering is a classic problem in computer vision, which categorizes images into different groups. Recent studies utilize nouns as external semantic knowledge to improve clustering performance. However, these methods often overlook the inherent ambiguity of nouns, which can distort semantic representations and degrade clustering quality. To address this issue, we propose a hierarChical semAntic alignmEnt method for image clustering, dubbed CAE, which improves clustering performance in a training-free manner. In our approach, we incorporate two complementary types of textual semantics: caption-level descriptions, which convey fine-grained attributes of image content, and noun-level concepts, which represent high-level object categories. We first select relevant nouns from WordNet and descriptions from caption datasets to construct a semantic space aligned with image features. Then, we design a residual attention mechanism to further enhance the discriminability of this space. Finally, we combine the enhanced semantic and image features to perform clustering. Extensive experiments across 8 datasets demonstrate the effectiveness of our method, notably surpassing the state-of-the-art training-free approach with a 4.2% improvement in accuracy and a 2.9% improvement in adjusted rand index (ARI) on the ImageNet-1K dataset.

Downloads

Published

2026-03-14

How to Cite

Zhu, X., Zhu, B., Li, Y., Fang, J., Wang, S., Zhao, K., & Zhang, H. (2026). Hierarchical Semantic Alignment for Image Clustering. Proceedings of the AAAI Conference on Artificial Intelligence, 40(34), 29177–29185. https://doi.org/10.1609/aaai.v40i34.40156

Issue

Section

AAAI Technical Track on Machine Learning XI