Source-linked AI summary

RhythmNet: End-to-end Heart Rate Estimation from Face via Spatial-temporal Representation

Xuesong Niu, Shiguang Shan, Hu Han, Xilin Chen

arXiv:1910.11515v2cs.CV

TL;DR

Remote heart-rate estimation from face videos is needed because contact monitors are inconvenient, while existing methods and datasets offer limited support for less-constrained conditions. The paper introduces RhythmNet, an end-to-end spatial-temporal CNN-GRU approach, and VIPL-HR, a multimodal database covering varied conditions. RhythmNet outperforms prior methods on the evaluated public and VIPL-HR databases, including RMSE values of 8.14 bpm on VIPL-HR and 3.99 bpm after fine-tuning on MAHNOB-HCI.

  • Problem

    Existing remote HR methods have limited known generalization under head movement and illumination variation, while available databases are small and do not adequately cover practical variations.

  • Method

    RhythmNet encodes HR signals in spatial-temporal representations, processes them with a convolutional network, and models adjacent HR estimates with a GRU.

  • Results

    RhythmNet outperforms the state-of-the-art methods, reaching 8.14 bpm RMSE on VIPL-HR and 3.99 bpm RMSE on MAHNOB-HCI after fine-tuning.

  • Takeaways & Limitations

    VIPL-HR's varied pose, illumination, and acquisition conditions support evaluation and training of remote HR estimation for less-constrained scenarios.

Abstract

from arXiv · show

Heart rate (HR) is an important physiological signal that reflects the physical and emotional status of a person. Traditional HR measurements usually rely on contact monitors, which may cause inconvenience and discomfort. Recently, some methods have been proposed for remote HR estimation from face videos; however, most of them focus on well-controlled scenarios, their generalization ability into less-constrained scenarios (e.g., with head movement, and bad illumination) are not known. At the same time, lacking large-scale HR databases has limited the use of deep models for remote HR estimation. In this paper, we propose an end-to-end RhythmNet for remote HR estimation from the face. In RyhthmNet, we use a spatial-temporal representation encoding the HR signals from multiple ROI volumes as its input. Then the spatial-temporal representations are fed into a convolutional network for HR estimation. We also take into account the relationship of adjacent HR measurements from a video sequence via Gated Recurrent Unit (GRU) and achieves efficient HR measurement. In addition, we build a large-scale multi-modal HR database (named as VIPL-HR, available at 'http://vipl.ict.ac.cn/view_database.php?id=15'), which contains 2,378 visible light videos (VIS) and 752 near-infrared (NIR) videos of 107 subjects. Our VIPL-HR database contains various variations such as head movements, illumination variations, and acquisition device changes, replicating a less-constrained scenario for HR estimation. The proposed approach outperforms the state-of-the-art methods on both the public-domain and our VIPL-HR databases.

I. INTRODUCTION

Remote heart-rate estimation aims to avoid inconvenient contact monitors, but existing methods and datasets provide limited evidence for robust performance under varying movement, illumination, and acquisition conditions. The paper addresses these gaps with an end-to-end estimator and the large-scale multimodal VIPL-HR database.

  • Contact monitors are inconvenient, motivating remote heart-rate estimation from face videos without physical contact.Remote methods estimate HR from signals captured in the face area.
  • Existing rPPG methods can work well in constrained settings but may degrade with large head motions, lighting variations, and unsupported skin-reflection assumptions.Prior evaluations also often use private databases, making comparisons difficult.
  • Public HR databases are usually limited to fewer than 50 subjects, restricting large-scale evaluation and deep-model development.The paper identifies the lack of a database covering movement, illumination, and device variation as a community challenge.
  • VIPL-HR contains 2,378 visible face videos and 752 NIR face videos from 107 subjects across three devices and nine situations.The situations cover variations in face pose, scale, and illumination.
  • The proposed approach is evaluated on MAHNOB-HCI, MMSE-HR, and VIPL-HR using intra-database and cross-database protocols.The reported experiments cover visible and near-infrared imaging, video compression, and multiple testing settings.
  • RhythmNet uses spatial-temporal representations, CNN processing, and GRU modeling of adjacent HR measurements for end-to-end remote estimation.The work extends earlier research with adjacent-measurement modeling and broader evaluations across databases, modalities, and acquisition conditions.

II. RELATED WORK

Remote HR estimation from face videos includes rPPG and BCG approaches, but existing methods face robustness and generalization limits under unconstrained conditions. The section motivates data-driven, end-to-end approaches for these challenges.

  • Remote HR estimation approaches: rPPG methods estimate HR from blood-volume-related color changes, whereas BCG methods use subtle head movements caused by cardiovascular circulation.rPPG approaches process optical skin signals; BCG approaches analyze motion associated with periodic blood ejections.
  • Limitations: Hand-crafted rPPG methods may perform poorly with large head movements, lighting variations, noise, or violated skin-reflection assumptions.The cited approaches report good results in constrained settings, but their assumptions may fail in complicated scenarios.
  • rPPG methods: Existing rPPG methods use color transformations, signal decomposition, and frequency analysis to generate and analyze HR signals.Examples include ICA, CHROM, ROI combination, and FFT-based spectral analysis.
  • Limitations: BCG-based methods are highly sensitive to voluntary subject movements, which can reduce HR estimation accuracy and limit practical use.These methods depend on subtle motion analysis rather than skin-color changes.
  • Research gap: Existing approaches commonly rely on assumptions and sophisticated step-by-step pipelines, motivating end-to-end trainable methods for practical remote HR estimation.The paper identifies adaptive data-driven feature learning as a way to address varied application conditions.

B. Public Databases for Remote HR Estimation

Public databases for remote HR estimation are limited in subjects, scenarios, or both, especially for less-constrained conditions. VIPL-HR is introduced to provide larger-scale, more varied data for practical evaluation and model development.

  • Existing databases: Existing public databases contain limited subjects or recording scenarios and are particularly constrained in their acquisition conditions.MAHNOB-HCI, MMSE-HR, PURE, and PFF differ in size and variation but do not collectively provide broad practical coverage.
  • Existing databases: PURE includes 60 videos from 10 subjects with six movements, while PFF includes 104 videos from 10 subjects with small head movements and illumination variations.The cited database descriptions illustrate limited scale and scenario diversity.
  • Motivation: A large-scale database recorded under less-constrained conditions is presented as necessary for studying remote HR estimation in practice.The motivation follows from limitations in both dataset scale and recording scenarios.
  • VIPL-HR: VIPL-HR contains 3,130 face videos from 107 subjects recorded with webcam, RealSense, and smartphone sensors under varying illumination and pose.The dataset was designed to promote research under less-constrained scenarios involving head movement, illumination variation, and device diversity.

A. Device Setup and Data Collection

VIPL-HR data collection varies illumination, pose, sensors, distances, and compression settings to approximate daily application conditions. Compression experiments select a solution that preserves HR signals while reducing dataset size.

  • Data collection: The collection procedure targets diversity in environmental illumination, subject pose, acquisition sensor, and HR distribution to replicate daily application scenarios.The setup includes casual behavior and activities such as talking and looking around.
  • Device setup: Three cameras provide visible-light and near-infrared recordings: Logitech C310, Intel RealSense F200, and Huawei P9 frontal camera.RealSense F200 is also used for NIR recording under dark lighting conditions.
  • Face video compression: The raw VIPL-HR data totals about 1.05 TB, motivating evaluation of video codecs and frame-resizing methods for distribution.The compression study uses Haan2013 as a baseline HR estimator and measures RMSE on compressed videos.
  • Recording conditions: The recording design covers multiple devices and acquisition situations, with frame timestamps recorded because some camera frame rates are unstable.The database specifications and situation details organize these acquisition variations.
  • Face video compression: MJPG best maintains the HR signal while substantially reducing database size, and resolutions below two-thirds damage the HR signal.The selected final solution is MJPG compression with 2/3 frame resizing, producing a dataset of about 48 GB.

C. Database Statistics

VIPL-HR combines visible and NIR videos from 107 participants with broad pose, illumination, and HR variation. Its statistics are intended to represent practical recording diversity and support data-driven HR estimation.

  • Dataset composition: VIPL-HR contains 2,378 color videos and 752 NIR videos from 107 participants, with approximately 30-second videos recorded at about 30 fps.Participants include 79 males and 28 females aged 22 to 41.
  • Pose variation: Maximum head-rotation amplitudes vary across a large range for yaw, pitch, and roll in videos with head movement.The distributions are summarized using OpenFace-estimated head poses and histograms.
  • Illumination variation: Mean face-region gray-scale intensity ranges from 60 to 212 across selected illumination situations.This range quantitatively demonstrates substantial illumination variation in the database.
  • HR distribution: Ground-truth HR ranges from 47 bpm to 146 bpm, covering the typical HR range and reducing the gap with daily-life distributions.The dataset’s size and broad HR distribution support using deep learning for more robust data-driven models.

IV. PROPOSED APPROACH

RhythmNet uses spatial-temporal maps derived from aligned face regions as inputs to an end-to-end deep HR estimator. The representation aggregates weak facial color variations across spatial regions and time while incorporating augmentation for missing detections.

  • Overview: RhythmNet maps face videos to HR through a CNN-RNN model trained on spatial-temporal representations.Face and landmark detection locate facial regions, spatial-temporal maps encode HR signals, and the CNN-RNN learns their mapping to HR.
  • A. Face Detection, Landmark Detection and Segmentation: Face alignment, skin segmentation, and landmark-based cropping define the facial area while retaining the original resolution to limit noise.Faces are aligned using eye centers, non-face regions are removed, and resizing is avoided during subsequent processing.
  • B. Spatial-temporal Map for Representing HR Signals: Average-pooled color values from n ROI blocks form temporal sequences that preserve spatial and temporal information for HR estimation.For each frame, the face is divided into ROI blocks and channel values are averaged; each resulting sequence is min-max normalized.
  • B. Spatial-temporal Map for Representing HR Signals: The spatial-temporal representation has size T × n × c and is constructed by arranging normalized ROI sequences into rows.The representation uses the temporal length, number of ROI blocks, and color-space dimensions as its three axes.
  • B. Spatial-temporal Map for Representing HR Signals: YUV, YCrCb, and HSV are considered as alternative color spaces for representing facial HR signals.The paper evaluates color spaces directly related to color and those commonly used for skin segmentation.
  • Robustness Augmentation: Randomly masking parts of spatial-temporal maps along time simulates missing detections caused by rapid head movement or rotation.The masked maps are used as augmented data to improve RhythmNet’s robustness to missing signal data.

C. Temporal Modeling for HR measurement

RhythmNet models adjacent HR measurements with a GRU after CNN feature extraction and constrains short-term predictions to remain smooth. This temporal design targets stable HR estimation across neighboring video clips.

  • Temporal Modeling: Fixed sliding-window clips provide spatial-temporal maps to ResNet-18 convolutional layers for feature extraction.The clips use a window of w frames and a 0.5-second step before entering the backbone network.
  • Temporal Modeling: A one-layer GRU receives CNN features and passes its output to a fully connected layer that regresses HR for each clip.The GRU models temporal relationships between measurements from succeeding video clips.
  • Smooth Loss: A smooth loss constrains T continuous HR measurements estimated over a short three-second period.The loss uses the mean HR across the continuous measurements to encourage short-term prediction smoothness.
  • Smooth Loss: The final objective combines L1 regression loss with the smoothness term using a balancing parameter λ.The paper identifies Ll1 as the L1 loss and λ as the parameter balancing the two terms.
  • Motivation and Novelty: The paper presents RNN-based temporal modeling as a first known approach for improving temporal stability in physiological HR estimation.It distinguishes this use from established RNN applications such as object tracking and action recognition.

V. EXPERIMENTS

The experiments evaluate RhythmNet across within-database testing, cross-database testing, and key-component analysis. They use three databases and several standard HR estimation metrics under specified temporal-map settings.

  • Evaluation Design: Three evaluation aspects are considered: intra-database testing, cross-database testing, and key components analysis.These aspects define the experimental coverage reported for RhythmNet.
  • Datasets: Evaluations use MAHNOB-HCI, MMSE-HR, and the collected VIPL-HR database.Ground-truth HR for the two public databases is computed from ECG signals using the OSET ECG Toolbox.
  • Metrics: Performance is measured with HR error statistics, MAE, RMSE, MER, and Pearson correlation coefficient r.The reported metric set includes mean and standard deviation of HR error alongside absolute, squared, percentage, and correlation measures.
  • Implementation Settings: Spatial-temporal maps use 300-frame windows for MMSE-HR and VIPL-HR, while MAHNOB-HCI videos are downsampled to 30.5 fps first.All databases use 25 ROI blocks arranged as 5 × 5 grids.

B. Intra-database Testing

Intra-database evaluations show RhythmNet performs strongly on RGB face videos and generalizes across multiple databases and acquisition conditions, while NIR performance is weaker than RGB.

  • Five-fold subject-exclusive cross-validation was used for VIPL-HR, while MAHNOB-HCI and MMSE-HR used three-fold subject-independent cross-validation.
  • 8.14 bpm RMSE was achieved by complete RhythmNet on VIPL-HR RGB videos, improving on the CNN-only result of 8.94 bpm and the best baseline’s 13.8 bpm.
  • 71% of samples had HR estimation errors below 5 bpm with RhythmNet, compared with 41.5% for DeepPhy.
  • Deep learning methods outperformed traditional methods on the challenging VIPL-HR dataset, suggesting that large-scale data support more informative HR representations.
  • RhythmNet achieved the best results on both MAHNOB-HCI and MMSE-HR, indicating generalization across different image acquisition conditions.
  • NIR videos produced weaker results than RGB, partly because one channel conveys less rhythm information and the RGB-trained face detector can miss NIR faces.

C. Cross-database Testing

Cross-database tests evaluate a VIPL-HR-trained RhythmNet on MAHNOB-HCI and MMSE-HR, showing promising direct-transfer and fine-tuned performance.

  • RhythmNet was trained on VIPL-HR color videos and tested directly or after fine-tuning on MAHNOB-HCI and MMSE-HR.
  • 8.28 bpm and 7.33 bpm RMSE were obtained by direct testing on MAHNOB-HCI and MMSE-HR, respectively.
  • Fine-tuning reduced RMSE to 3.99 bpm on MAHNOB-HCI and 5.49 bpm on MMSE-HR.
  • The cross-database results were reported as better than previous methods and as evidence that illumination, movement, and acquisition variations can be handled.
  • The broader analysis used subject-exclusive five-fold cross-validation on VIPL-HR for spatial-temporal maps, color spaces, GRU modeling, and computational cost.

2) Color Space Selection:

Color-space and temporal-modeling analyses identify YUV spatial-temporal maps and GRU as useful design choices, while the model remains computationally efficient.

  • 2) Color Space Selection:: Separating chromaticity from intensity was considered helpful for reducing head-movement artifacts during HR representation learning.
  • 2) Color Space Selection:: YUV color space achieved the best color-space result, with an RMSE of 8.8 bpm.
  • 2) Color Space Selection:: Removing GRU increased RMSE from 8.11 bpm to 8.94 bpm, supporting temporal modeling of adjacent measurements.
  • 2) Color Space Selection:: The HR estimator is approximately 42 MB and takes about 8 ms for inference on a Titan 1080Ti GPU with a 300-frame sliding window.
  • 2) Color Space Selection:: The additional analysis addresses video compression, illumination, and head movement as challenges for less-constrained remote HR estimation.

1) Video Compression:

Robustness analyses show that MJPG compression preserves RhythmNet performance, while illumination and head movement substantially affect estimation accuracy; NIR can help in darkness.

  • 1) Video Compression:: MJPG compression produced the same HR estimation errors before and after compression, indicating that compressed VIPL-HR data preserves the HR signals.
  • 2) Illumination:: Under bright lighting, RhythmNet achieved 5.90 bpm RMSE, compared with 7.94 bpm under dim light.
  • 2) Illumination:: In a dark environment, RGB RMSE increased from 8.14 bpm to 17.7 bpm, while NIR reduced the error to 12 bpm.
  • 3) Head Movement:: Talking increased RMSE to 8.28 bpm, while large-scale head movements increased it to 9.40 bpm.
  • VI. CONCLUSION AND FUTURE WORK: RhythmNet was designed for less-constrained face-video HR estimation using spatial-temporal maps and adjacent-estimate modeling.
  • VI. CONCLUSION AND FUTURE WORK: VIPL-HR covers illumination, head movement, and sensor-diversity variations for within-database and cross-database evaluation.
Loading 1910.11515v2…