PAPER DIGEST
Most Influential IJCAI 2013 Paper · 2026-03 edition

Learning Discriminative Representations From RGB-D Video Data

Li Liu; Ling Shao

Venue
International Joint Conference on Artificial Intelligence (IJCAI) 2013
Recognition
Most Influential IJCAI 2013 Paper (Rank No. 11)
Edition
2026-03
Impact factor
6
Certificate ID
51dfbbfd88247993

Abstract

Recently, the low-cost Microsoft Kinect sensor, which can capture real-time high-resolution RGB and depth visual information, has attracted increasing attentions for a wide range of applications in computer vision. Existing techniques extract hand-tuned features from the RGB and the depth data separately and heuristically fuse them, which would not fully exploit the complementarity of both data sources. In this paper, we introduce an adaptive learning methodology to automatically extract (holistic) spatio-temporal features, simultaneously fusing the RGB and depth information, from RGBD video data for visual recognition tasks. We address this as an optimization problem using our proposed restricted graph-based genetic programming (RGGP) approach, in which a group of primitive 3D operators are first randomly assembled as graph-based combinations and then evolved generation by generation by evaluating on a set of RGBD video samples. Finally the best-performed combination is selected as the (near-)optimal representation for a pre-defined task. The proposed method is systematically evaluated on a new hand gesture dataset, SKIG, that we collected ourselves and the public MSRDailyActivity3D dataset, respectively. Extensive experimental results show that our approach leads to significant advantages compared with state-of-the-art handcrafted and machine-learned features.

Download PDF certificate