Temporal Calibrating and Distilling for Scene-Text Aware Text-Video Retrieval

Zhiqian Zhao; Liang Li; Lei Shen; Xichun Sheng; Yaoqi Sun; Fang Kang; Chenggang Yan

doi:10.1609/aaai.v40i16.38335

Authors

Zhiqian Zhao Hangzhou Dianzi University, Hangzhou, China Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
Liang Li Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
Lei Shen Hangzhou Dianzi University, Hangzhou, China
Xichun Sheng Macao Polytechnic University, Macao, China
Yaoqi Sun Lishui University, Lishui, China
Fang Kang Center for Machine Vision and Signal Analysis, University of Oulu, Oulu, Finland
Chenggang Yan Hangzhou Dianzi University, Hangzhou, China

DOI:

https://doi.org/10.1609/aaai.v40i16.38335

Abstract

Existing text-video retrieval methods mainly focus on singlemodal video content (i.e., visual entities), often overlooking heterogeneous scene text that is ubiquitous in human environments. Although scene text in videos provides finegrained semantics for cross-modal retrieval, effectively utilizing it presents two key challenges: (1) Temporally dense scene text disrupts sync with sparse video frames, obstructing video understanding;(2) Redundant scene text and irrelevant video frames hinder the learning of discriminative temporal clues for retrieval. To address them, we propose a temporal scene-text calibrating and distilling (TCD) network for textvideo retrieval. Specifically, we first design a window-OCR captioner that aggregates dense scene text into OCR captions to facilitate feature interaction. Next, we devise a heterogeneous semantics calibration module that leverages scene text as a self-supervised signal to temporally align window-level OCR captions and frame-level video features. Further, we introduce a context-guided temporal clue distillation module to learn the complementary and relevant details between scene text and video modalities, thereby obtaining discriminative temporal clues for retrieval. Extensive experiments show that our TCD achieves state-of-the-art performance on three scene-text related benchmarks.

Temporal Calibrating and Distilling for Scene-Text Aware Text-Video Retrieval

Authors

DOI:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information