Contrastive Predictive Autoencoders for Dynamic Point Cloud Self-Supervised Learning

Xiaoxiao Sheng; Zhiqiang Shen; Gang Xiao

doi:10.1609/aaai.v37i8.26170

Authors

Xiaoxiao Sheng Shanghai Jiao Tong University
Zhiqiang Shen Shanghai Jiao Tong University
Gang Xiao Shanghai Jiao Tong University

DOI:

https://doi.org/10.1609/aaai.v37i8.26170

Keywords:

ML: Unsupervised & Self-Supervised Learning, CV: 3D Computer Vision, CV: Video Understanding & Activity Analysis, CV: Biometrics, Face, Gesture & Pose

Abstract

We present a new self-supervised paradigm on point cloud sequence understanding. Inspired by the discriminative and generative self-supervised methods, we design two tasks, namely point cloud sequence based Contrastive Prediction and Reconstruction (CPR), to collaboratively learn more comprehensive spatiotemporal representations. Specifically, dense point cloud segments are first input into an encoder to extract embeddings. All but the last ones are then aggregated by a context-aware autoregressor to make predictions for the last target segment. Towards the goal of modeling multi-granularity structures, local and global contrastive learning are performed between predictions and targets. To further improve the generalization of representations, the predictions are also utilized to reconstruct raw point cloud sequences by a decoder, where point cloud colorization is employed to discriminate against different frames. By combining classic contrast and reconstruction paradigms, it makes the learned representations with both global discrimination and local perception. We conduct experiments on four point cloud sequence benchmarks, and report the results on action recognition and gesture recognition under multiple experimental settings. The performances are comparable with supervised methods and show powerful transferability.

Contrastive Predictive Autoencoders for Dynamic Point Cloud Self-Supervised Learning

Authors

DOI:

Keywords:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information

Subscription