Source-linked AI summary

Vision-based Human Fall Detection Systems using Deep Learning: A Review

Ekram Alam, Abu Sufian, Paramartha Dutta, Marco Leo

arXiv:2207.10952v1cs.CVcs.AI

TL;DR

The paper addresses the need for non-intrusive fall detection as the elderly population grows and rapid assistance becomes increasingly important. It reviews deep-learning-based vision methods, benchmark datasets, and evaluation metrics, concluding that these methods generally achieve good results while facing data, computation, privacy, and security constraints.

  • Problem

    Growing elderly populations and the risks associated with delayed assistance motivate effective, non-intrusive human fall detection for assistive living.

  • Method

    The paper surveys deep-learning-based vision fall-detection techniques, benchmark human-fall datasets, and performance metrics, classifying methods as CNN, auto-encoder, LSTM, MLP, or hybrid.

  • Results

    The reviewed deep-learning techniques generally give good results, with reported examples including 93.3% true positive rate and 98.05% average accuracy.

  • Takeaways & Limitations

    The review identifies vision-based deep learning as a substantial body of work for fall detection and outlines future attention to privacy, security, and deployment constraints.

  • Takeaways & Limitations

    Deep-learning systems are data-hungry and computationally demanding, complicating deployment on embedded IoT devices and potentially requiring cloud data transfer.

Abstract

from arXiv · show

Human fall is one of the very critical health issues, especially for elders and disabled people living alone. The number of elder populations is increasing steadily worldwide. Therefore, human fall detection is becoming an effective technique for assistive living for those people. For assistive living, deep learning and computer vision have been used largely. In this review article, we discuss deep learning (DL)-based state-of-the-art non-intrusive (vision-based) fall detection techniques. We also present a survey on fall detection benchmark datasets. For a clear understanding, we briefly discuss different metrics which are used to evaluate the performance of the fall detection systems. This article also gives a future direction on vision-based human fall detection techniques.

1. Introduction

The review motivates vision-based fall detection as assistive technology for a growing elderly population, emphasizing rapid detection and the limitations of wearable devices. It surveys non-intrusive sensing approaches and outlines the paper’s coverage of metrics, datasets, and deep-learning techniques.

  • 1. Introduction: The global population aged 65 and above is projected to grow from 727 million in 2020 to 1.5 billion by 2050, increasing concern about elder care and falls.The paper links longer life expectancy and limited caretaking capacity to the need for technological assistance.
  • 1. Introduction: Rapid fall detection and medical assistance are important because remaining on the floor for more than one hour is associated with a 50% six-month mortality risk.The review also identifies automatic incident reporting as useful for analyzing and potentially preventing falls.
  • 1. Introduction: Fall detection can use wearable sensors or vision sensors, including RGB, infrared, depth, and 3D camera-array systems.Wearables measure abnormal velocities and angles, whereas vision-based approaches provide non-intrusive alternatives.
  • 1. Introduction: The paper is organized around related work, evaluation metrics, publicly available human-fall datasets, and deep-learning-based vision techniques, followed by future directions.The supplied outline identifies the review’s main sections and its concluding focus on future work.

2. Related Works

The paper distinguishes itself from broader vision-based and multimodal fall-detection reviews by focusing specifically on deep-learning-based, non-intrusive vision methods.

  • 2. Related Works: Unlike broader reviews covering all vision-based algorithms or multimodal sensor fusion, this paper reviews only vision-based deep-learning fall detection in detail.It further organizes the review according to different deep-learning model types.

3. Evaluation Metrics

The review defines metrics for evaluating binary fall-detection outcomes, relating them to correct and incorrect fall or normal-activity classifications. It emphasizes that balanced assessment requires considering both sensitivity and specificity, alongside precision and F-score.

  • 3. Evaluation Metrics: Fall detection is evaluated as a binary problem using sensitivity, specificity, accuracy, F-score, precision, and geometric mean.These metrics are defined using true positives, true negatives, false positives, and false negatives.
  • 3. Evaluation Metrics: Sensitivity measures the proportion of actual falls correctly identified, whereas specificity measures the correct recognition of normal activity as normal.Sensitivity uses TP and FN, while specificity uses TN and FP.
  • 3. Evaluation Metrics: Accuracy requires both sensitivity and specificity to be high because high performance on only one of them is insufficient.The review presents accuracy as combining TP, TN, FP, and FN.
  • 3. Evaluation Metrics: Precision is the ratio of correctly detected falls to all positive fall detections, combining true positives and false positives.The review gives precision as TP relative to TP and FP.
  • 3. Evaluation Metrics: F-score is the harmonic mean of precision and sensitivity, while geometric mean summarizes sensitivity and specificity.The review presents F-score as combining precision and recall and identifies geometric mean as based on sensitivity and specificity.

4. Benchmark Public Human Fall Datasets

Benchmark human-fall datasets support training, evaluation, and comparison of data-hungry deep-learning systems. The review highlights multi-camera recordings and varied activities and environments as dataset characteristics.

  • 4. Benchmark Public Human Fall Datasets: Benchmark human-fall datasets are needed because deep-learning methods require substantial data for training, testing, validation, evaluation, and comparison.Researchers also use models pretrained on human-activity-recognition datasets such as KTH, UCF-101, and MSRDailyActivity3D.
  • 4. Benchmark Public Human Fall Datasets: The Auvinet dataset uses eight IP cameras to record 24 activities from simultaneous viewpoints, with fall and daily-living activities performed by one subject.The recordings are available in AVI format.
  • 4. Benchmark Public Human Fall Datasets: Its listed activities include walking, falling, lying on the ground, crouching, moving, sitting, lying on a sofa, and standing up.The dataset covers both falls and activities of daily living.

4.2. Le2i Fall Detection Dataset (Le2i FDD)

The Le2i Fall Detection Dataset, originally named S, uses a single RGB camera to capture falls and activities of daily living across varied environments and conditions.

  • Le2i FDD contains 143 fall videos and 79 ADL videos from nine subjects performing three fall types and six ADL activities.
  • The dataset spans forward falls, balance loss, falls from sitting, sitting, walking, standing, moving a chair, and housekeeping.
  • Videos were recorded in home, coffee-room, office, and lecture-room environments while varying lighting, clothing, shadows, reflections, and camera view.

4.4. SDUFall

SDUFall is a depth-based dataset recorded with an MS Kinect camera, covering six actions repeated across subjects and varied recording conditions.

  • SDUFall uses an MS Kinect camera to record six actions performed by ten male and female subjects.
  • The six actions are falling down, walking, bending, lying, squatting, and sitting.
  • Each action was repeated 30 times, producing 1800 eight-second video clips.
  • Recordings varied lighting, room layout, camera angle and position, and whether subjects carried an object.

4.6. Thermal Simulated Fall dataset (TSFD)

The Thermal Simulated Fall Dataset uses a single FLIR ONE thermal camera mounted on an Android phone to record simulated falls in a room environment.

  • TSFD was recorded with a single FLIR ONE thermal camera mounted on an Android phone.
  • The dataset contains 35 fall video segments and nine ADL video segments.
  • Fall activities were performed in a room environment.

4.8. UP-Fall Detection Dataset

UP-Fall Detection is a multimodal dataset combining wearable, ambient, and vision sensors, with 17 subjects performing repeated ADL and fall activities.

  • UP-Fall Detection includes wearable, ambient, and vision sensors.
  • Seventeen subjects aged 18–24, including nine males and eight females, participated in the dataset.
  • Participants performed 11 activities consisting of six ADL activities and five fall activities, with each activity repeated three times.
  • The fall activities were falling backward, falling forward using hands, falling forward using knees, falling sideward, and falling sitting in an empty chair.

4.9. Fallen Person Dataset (FPDS)

FPDS is a single-camera fall-detection dataset containing labeled falls and activities of daily living across varied environments and imaging conditions.

  • 4.9. Fallen Person Dataset (FPDS): FPDS contains 1,072 falls and 1,262 manually labeled ADL images captured with a single camera mounted on a robot at 76 cm.More than one subject may appear in an image, and subjects ranged from 1.2 m to 1.8 m tall.
  • 4.9. Fallen Person Dataset (FPDS): The dataset covers eight environments with varying lighting, shadows, reflections, and other visual conditions.

5. Review on Recent State of the art

The review surveys vision-based deep-learning fall-detection methods, organizing them by model family and summarizing their processing pipelines, datasets, and reported results.

  • Review scope and organization: The review covers deep-learning vision-based fall detection since 2014 and classifies methods into CNN, LSTM, auto-encoder, MLP, and hybrid models.It selects papers from Web of Science, Scopus, and Google Scholar using fall-related search-term permutations.
  • CNN-based techniques: CNN methods commonly process visual or motion representations, including RGB frames, optical flow, dynamic images, depth, silhouettes, and pose features.The reviewed CNN variants include 3D CNN, time-delay CNN, and graph convolutional networks.
  • CNN-based techniques: Fan et al. detect a fall when standing, falling, fallen, and not-moving phases occur sequentially, while dynamic images retain appearance and temporal information.The difference scoring method estimates the temporal extent of the fall.
  • CNN-based techniques: The reviewed CNN work also includes optical-flow architectures, trajectory-weighted descriptors, UAV-based detection, and methods intended to reduce false positives or handle changing model conditions.Reported designs include transfer learning, tracking, denoising, camera-height adaptation, and wheelchair-related false-positive filtering.
  • CNN-based techniques: Solback et al. combined 2D CNN pose estimation, depth-based 3D pose and ground-plane reasoning, reporting a 93.3% true positive rate.The system used ZED stereo cameras and was implemented in ROS.

5.2. LSTM or RNN (Recurrent Neural Network) based Techniques

The review describes recurrent, auto-encoder, MLP, and hybrid fall-detection techniques that use pose, motion, thermal, and appearance representations for varied environments and tasks.

  • 5.2. LSTM or RNN (Recurrent Neural Network) based Techniques: LSTM and RNN methods commonly extract human pose or skeleton sequences before classifying falls with recurrent models.Hasan et al. selected eight OpenPose joints and fed 24-frame pose sequences into a two-layer LSTM.
  • 5.2. LSTM or RNN (Recurrent Neural Network) based Techniques: Jeong et al. combined raw skeleton data with the speed of the human centerline coordinate to improve accuracy in manufacturing-industry fall detection.The method used two stacked LSTM layers with 256 hidden-layer features and evaluated URFD and SDUFall.
  • 5.2. LSTM or RNN (Recurrent Neural Network) based Techniques: Attention-guided LSTM methods target complex scenes containing multiple people, using pedestrian detection and tracking before fall classification.The method was evaluated on a purpose-built complex-scene dataset as well as URFD and MCFD.
  • Auto-encoder and geometry-based techniques: Other reviewed approaches use geometric head-to-hip angles and distances, sparse or convolutional auto-encoders, thermal imagery, and spatio-temporal residual representations.Auto-encoder methods include adaptation to abrupt visual changes and normalization or augmentation of thermal data.
  • MLP and hybrid techniques: MLP and hybrid systems combine pose estimation with classification or fuse multiple deep-learning models and representations for fall detection.One MLP prototype sent an SMS to caretakers after detecting a fall, while hybrid reviews group works using more than one DL model.

6. Discussions on Limitations and Future Scope

The review finds CNNs are the most commonly used deep-learning models for fall detection, while emphasizing data, computation, privacy, and deployment challenges.

  • CNNs account for the largest share of reviewed deep-learning fall-detection work, followed by hybrid, LSTM, auto-encoder, and MLP approaches.
  • Deep-learning fall-detection methods are generally data-hungry and computationally demanding, complicating deployment on embedded IoT devices.
  • Cloud-edge data transfers can create security and privacy concerns, while RGB inputs may expose subjects’ identities.
  • Thermal, infrared, and depth cameras are proposed as privacy-preserving alternatives that can also operate in low light.
  • Multiple-camera systems can preserve operation when one camera fails or loses view, but they are more complex and require synchronization.

7. Conclusion

The review surveys vision-based deep-learning fall-detection methods, evaluation metrics, and public datasets, then identifies priorities for more robust and privacy-aware systems.

  • The review covers deep-learning vision-based fall-detection developments since 2014, evaluation metrics, publicly available datasets, and model-based method classifications.
  • Future work should address multiple people, occlusion, privacy, security, and the limited realism of datasets built from simulated falls.
  • Thermal and infrared cameras and edge computing are identified as directions for improving privacy and security.
Loading 2207.10952v1…