Source-linked AI summary

RGB-D-based Human Motion Recognition with Deep Learning: A Survey

Pichao Wang, Wanqing Li, Philip Ogunbona, Jun Wan, Sergio Escalera

arXiv:1711.08362v2cs.CV

TL;DR

RGB-D human motion recognition remains difficult under clutter, occlusion, viewpoint and lighting changes, execution-rate differences, and biometric variation. This survey organizes deep-learning approaches by modality and network architecture, analyzes their spatial-temporal-structural encoding and comparative performance, and identifies challenges and future opportunities. Its scope excludes RGB-D group activity recognition because relevant datasets are lacking.

  • Problem

    Human motion recognition remains challenging under background clutter, partial occlusion, viewpoint and lighting changes, execution-rate differences, and biometric variation.

  • Method

    The survey categorizes RGB-D motion-recognition methods by modality and discusses CNN, RNN, and other networks through spatial-temporal-structural encoding.

  • Results

    The survey analyzes relative performance across several commonly used RGB-D datasets and highlights associated challenges.

  • Takeaways & Limitations

    The survey presents a comprehensive overview of deep-learning-based RGB-D motion recognition and indicates numerous opportunities despite advances to date.

  • Takeaways & Limitations

    RGB-D group activity recognition is not covered because the field lacks datasets focused on that task.

Abstract

from arXiv · show

Human motion recognition is one of the most important branches of human-centered research activities. In recent years, motion recognition based on RGB-D data has attracted much attention. Along with the development in artificial intelligence, deep learning techniques have gained remarkable success in computer vision. In particular, convolutional neural networks (CNN) have achieved great success for image-based tasks, and recurrent neural networks (RNN) are renowned for sequence-based problems. Specifically, deep learning methods based on the CNN and RNN architectures have been adopted for motion recognition using RGB-D data. In this paper, a detailed overview of recent advances in RGB-D-based motion recognition is presented. The reviewed methods are broadly categorized into four groups, depending on the modality adopted for recognition: RGB-based, depth-based, skeleton-based and RGB+D-based. As a survey focused on the application of deep learning to RGB-D-based motion recognition, we explicitly discuss the advantages and limitations of existing techniques. Particularly, we highlighted the methods of encoding spatial-temporal-structural information inherent in video sequence, and discuss potential directions for future research.

1. Introduction

The survey reviews deep learning for RGB-D human motion recognition, motivated by persistent recognition challenges and the distinct properties of RGB, depth, and skeleton modalities. It organizes methods by modality and neural-network architecture while analyzing spatial-temporal-structural encoding, limitations, datasets, and future directions.

  • Human motion recognition remains challenging because of background clutter, partial occlusion, viewpoint and lighting changes, execution-rate differences, and biometric variation.
  • RGB-D sensing attracts attention because depth is less sensitive to illumination, provides 3D scene structure, and supports body-joint position estimation.
  • The paper provides comprehensive coverage of recent deep-learning methods, benchmark datasets, method advantages and limitations, recognition challenges, and potential research directions.
  • RGB, depth, and skeleton modalities offer different cues: appearance and optical flow, illumination-robust geometry and structure, and high-level joint positions.
  • The survey categorizes methods into RGB-based, depth-based, skeleton-based, and RGB+D-based groups, with further subdivisions for segmented and continuous recognition.
  • It reviews CNN-based, RNN-based, and other structured networks from the viewpoint of encoding spatial-temporal-structural information.

2. Benchmark Datasets

The survey reviews publicly available RGB-D benchmark datasets and organizes them by modality and dataset structure. It describes 15 large-scale datasets commonly used to evaluate deep learning methods.

  • RGB-D benchmark datasets are sourced mainly from motion-capture systems, structured-light cameras, and time-of-flight cameras.
  • The datasets cover RGB, depth, skeleton, and combined modalities.
  • Deep methods can estimate skeletons directly from single images or video sequences, including DeepPose, Deepercut, and Adversarial PoseNet.
  • 15 large-scale datasets commonly adopted for evaluating deep learning-based methods are described.
  • The survey divides datasets into segmented and continuous/online groups.

2.1. Segmented Datasets

Segmented datasets contain complete begin-to-end action or gesture samples, typically with one segment per action, and are mainly used for classification. The section surveys commonly used benchmarks and their recording conditions.

  • Segmented datasets: Segmented datasets contain whole begin-end action or gesture samples, with one segment corresponding to one action.
  • Segmented datasets: These datasets are mainly used for classification purposes.
  • CMU Mocap: CMU Mocap covers interactions, locomotion, uneven terrain, sports, and other human actions.
  • HDM05: HDM05 contains 2337 sequences for 130 actions performed by 5 non-professional actors, with 31 joints per frame.
  • MSR-Action3D: MSR-Action3D contains 20 actions performed three times by 10 subjects, with videos recorded from a fixed viewpoint.

2.1.4. MSRC-12

MSRC-12 provides gesture instruction data across text, images, video, and combinations, while MSRDailyActivity3D focuses on daily activities recorded near a sofa. The latter’s tracker produces noisy joint positions in these settings.

  • MSRC-12: MSRC-12 uses descriptive text, ordered static images, video, and combinations as gesture instruction modalities.
  • MSRC-12: MSRC-12 includes 30 participants distributed across text, images, video, video-with-text, and images-with-text conditions.
  • MSRC-12: Only skeleton data are made available for MSRC-12, which was captured using one Kinect sensor.
  • MSRDailyActivity3D: MSRDailyActivity3D focuses on daily activities performed by 10 actors while sitting on or standing near a sofa.
  • MSRDailyActivity3D: Joint positions extracted by the tracker are very noisy because actors are sitting on or standing close to the sofa.

2.1.6. UTKinect

UTKinect contains multi-view action recordings with substantial actor and occlusion variability. The section also describes G3D and SBU Kinect as additional interaction-oriented datasets with specialized annotations.

  • UTKinect: UTKinect contains 10 human actions performed twice by 10 subjects from a variety of views.
  • UTKinect: UTKinect is challenging because of high actor-dependent variability, human-object occlusions, and body parts leaving the field of view.
  • UTKinect: UTKinect provides action labels and sequence segmentation as ground truth.
  • G3D: G3D contains 20 gaming actions performed three times by each of 10 subjects for real-time recognition scenarios.
  • SBU Kinect Interaction Dataset: SBU Kinect contains eight interaction types performed by seven participants, with labels for segmented actions and active or inactive actors.

2.1.9. Berkeley MHAD

Berkeley MHAD is a multimodal human-action dataset collected in 2013, combining five sensing modalities and varied full-body, upper-extremity, and lower-extremity actions.

  • Berkeley MHAD was collected by the University of California at Berkeley and Johns Hopkins University in 2013 using five modalities.The modalities include an optical motion-capture system, multi-view stereo cameras, Kinect v1 cameras, wireless accelerometers, and microphones.
  • The dataset contains 12 subjects performing 11 actions five times each.
  • Its actions cover full-body movement, highly dynamic upper-extremity movements, and highly dynamic lower-extremity movements.Examples include jumping, waving, clapping, sitting down, and standing up.
  • Northwestern-UCLA Multiview Action 3D contains actions performed by 10 actors and captured simultaneously from three Kinect v1 viewpoints.The dataset was collected by Northwestern University and the University of California at Los Angeles and provides varied viewpoints.

2.1.11. ChaLearn LAP IsoGD

ChaLearn LAP IsoGD is a large Kinect v1 RGB-D dataset for segmented gesture recognition, with 47,933 sequences spanning 249 gestures and subject-independent splits.

  • ChaLearn LAP IsoGD is a large RGB-D dataset for segmented gesture recognition collected with a Kinect v1 camera.
  • 47,933 RGB-D depth sequences represent individual gesture instances in the dataset.
  • The dataset includes 249 gestures performed by 21 different individuals.
  • Training, validation, and test sets use different subjects so validation and test gestures do not come from subjects seen during training.

2.2. Continuous/Online Datasets

Continuous/online datasets contain one or more actions or gestures without known class boundaries, supporting detection, localization, and online prediction across varied modalities and views.

  • Continuous/online datasets: Continuous/online datasets may contain one or more actions or gestures whose boundaries between motion classes are unknown.
  • Continuous/online datasets: These datasets are mainly used for action detection, localization, and online prediction, but few datasets support this setting.
  • ChaLearn2014 Multimodal Gesture Recognition: ChaLearn2014 Multimodal Gesture Recognition combines RGB, depth, skeleton, and audio from Kinect v1 for continuous Italian gesture performances.
  • ChaLearn2014 Multimodal Gesture Recognition: ChaLearn2014 provides nearly 14K labeled gesture performances across 20 Italian gesture categories and 1,720,800 labeled frames.
  • ChaLearn LAP ConGD: ChaLearn LAP ConGD contains 47,933 gesture instances in 22,535 videos, with each video potentially containing one or more gestures.
  • PKU-MMD: PKU-MMD contains 1,076 long sequences across 51 action categories, nearly 20,000 action instances, and 5.4 million frames.
  • Benchmark coverage: Public benchmarks cover gestures, simple actions, daily activities, and human-object or human-human interactions across segmented and continuous settings.

3. RGB-based Motion Recognition with Deep Learning

RGB-based recognition methods use CNNs and RNNs to encode spatial, temporal, and structural information from video. The survey organizes these methods around temporal fusion, 3D convolution, dynamic-image encoding, and multiple streams.

  • RGB data provides shape, color, and texture cues that support direct use of image-based CNNs.
  • CNN-based Approach: CNN-based methods encode temporal information through frame-level fusion, 3D convolutions, dynamic images, or multiple streams.
  • CNN-based Approach: Long-term modeling methods address sequences lasting several seconds, while reducing spatial resolution or factoring 3D filters to control complexity.
  • CNN-based Approach: Two-stream networks combine raw frames with optical flow, while temporal segment networks sample snippets across long videos to capture long-range structure.

4. Depth-based Motion Recognition with Deep Learning

Depth-based recognition benefits from illumination and color invariance but is constrained by weaker appearance information and relatively small datasets. Methods therefore encode depth sequences as dynamic, structured, or pseudo-color images for CNN recognition.

  • Depth is insensitive to illumination variation, invariant to color and texture changes, and useful for estimating silhouettes, skeletons, and 3D structure.
  • Depth-based deep-learning results remain limited because depth maps lack color and texture, while available depth datasets are relatively small.
  • Dynamic Depth Images, Dynamic Depth Normal Images, and Dynamic Depth Motion Normal Images capture posture dynamics and 3D structure at different levels.
  • S2DDI aggregates global-to-fine-grained motion and structure across body parts, joints, and temporal scales with low construction cost and memory.
  • Structured dynamic images enable fine-tuning image-trained ConvNets for depth-sequence classification without training the models afresh.
  • Other approaches improve viewpoint tolerance or continuous recognition through synthetic views, weighted depth-motion maps, segmentation, and rank pooling.

5. Skeleton-based Motion Recognition with Deep Learning

Skeleton-based methods exploit joint positions as high-level motion features through CNN, RNN, and other architectures. The survey emphasizes representations and recurrent designs that model spatial joint dependencies alongside temporal evolution.

  • Skeleton data provides joint positions as relatively high-level features, while estimated skeletons remain vulnerable to viewpoint and occlusion errors.
  • CNN-based Approach: CNN-based methods convert skeleton sequences into images encoding joint distances, trajectories, coordinates, or motion for visual feature learning.
  • CNN-based Approach: Joint Distance Maps are reported as less sensitive to view variations than the compared skeleton representations.
  • RNN-based Approach: RNN-based methods model long-term temporal context and increasingly incorporate body-part structure, spatial joint dependencies, attention, or geometric features.
  • RNN-based Approach: Spatio-temporal LSTM explicitly models joint dependencies and recurrently analyzes spatial and temporal domains concurrently, with a trust gate for noisy inputs.
  • RNN-based Approach: Temporal sliding LSTMs use short-, medium-, and long-term subnetworks whose representations are hierarchically fused into higher-level representations.

6. RGB+D-based Motion Recognition with Deep Learning

RGB+D methods combine complementary modalities using CNN, RNN, and other architectures. Fusion may occur early, through shared representations, or later through score and feature integration, with some methods also using skeleton information or scene flow.

  • RGB, depth, and skeleton modalities have distinct properties, motivating methods that combine their strengths for motion recognition.
  • Fusion strategies include pyramidal 3D convolution, two-stream networks, scene-flow representations, and direct depth-channel integration.
  • c-ConvNet jointly optimizes ranking and softmax losses to strengthen discriminative features and reduce modality discrepancy.
  • Depth-skeleton models fuse features, body-part and object interactions, and temporal structure in end-to-end frameworks designed to improve viewpoint robustness.
  • RNN-based Approach: Privileged-information RNNs use skeleton sequences to improve depth-based parameter estimation through classification, regression, and refinement steps.
  • Multiple-channel fusion mechanisms are reported to outperform individual modules in the reviewed RGB+D methods.

7. Discussion

The survey compares deep-learning methods for RGB-D motion recognition across modalities, datasets, and recognition settings. It identifies modality- and task-dependent performance patterns, persistent temporal-encoding and data limitations, and future directions including multimodal, hybrid, and larger-scale learning.

  • Taxonomy: The survey organizes methods into segmented and continuous/online recognition, with four modality-based categories in each group.The categories are based on the adopted modalities and include RGB, depth, skeleton, and RGB+D approaches.
  • Performance Analysis: Performance is evaluated with accuracy for segmented recognition and Jaccard Index for continuous recognition.The Jaccard Index measures average relative overlap between true and predicted frame sequences for a gesture/action.
  • Performance Analysis: No single approach achieves the best performance across all datasets, while multimodal methods can outperform single-modality counterparts.The survey attributes this potential advantage to complementary properties of the modalities.
  • Performance Analysis: CNN-based methods tend to outperform RNN-based methods on some datasets, whereas RNN-based methods tend to perform well for continuous motion recognition.Combining CNN and RNN architectures also produces promising results, including C3D+ConvLSTM on ChaLearn LAP IsoGD.
  • Challenges: Temporal encoding remains unresolved because existing approaches neglect temporal order, impose rigid or short frame structures, incur high computational cost, lose information, or approximate sequence matching.The survey states that no perfect temporal-encoding method exists and that modeling temporal information remains a major challenge.
  • Challenges: Small labeled datasets, viewpoint variation, and occlusion limit practical recognition, especially on large complex and continuous datasets.Available datasets often use visible, restricted views, while occlusion is inevitable in practical interaction scenarios.
  • Future Research Directions: Future research directions include multimodal fusion, large-scale fine-grained and occlusion-based datasets, zero/one-shot learning, hybrid networks, GAN-based techniques, and joint spatial-temporal-structural modeling.These directions are motivated by the limitations identified in current deep-learning approaches.

8. Conclusion

The paper presents a comprehensive survey of deep-learning approaches for RGB-D human motion recognition. It groups methods by modality, analyzes spatial, temporal, and structural encoding, and uses the resulting insights to identify future research opportunities.

  • Scope: The survey covers RGB-D motion recognition methods and provides an overview of commonly used datasets.It also distinguishes this work from surveys focused mainly on datasets.
  • Taxonomy: Available methods are grouped into RGB-based, depth-based, skeleton-based, and RGB+D-based categories.The modalities have distinct features that lead to different deep-learning method choices.
  • Encoding Analysis: The paper defines spatial, temporal, and structural information in video sequences and analyzes methods through their spatio-temporal-structural encoding.The analysis considers the advantages and disadvantages of available methods.
  • Future Research: The survey uses its analysis of existing methods to describe potential future research directions.It characterizes the field as containing opportunities despite advances achieved to date.
Loading 1711.08362v2…