Source-linked AI summary

Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions

Timothy Oladunni, Farouk Ganiyu-Adewumi

arXiv:2608.27392v1eess.SPcs.AI

TL;DR

Endpoint accuracy alone cannot show whether camera-derived rPPG preserves other source-PPG properties recording by recording. This study tests property-specific recoverability under a fixed CHROM pathway and finds that preservation varies across properties and observation conditions.

  • Problem

    Existing endpoint-focused rPPG validation does not establish matched, recording-specific preservation of multiple physiological properties from synchronized contact PPG.

  • Method

    The study uses a fixed CHROM camera-rPPG pathway, matched-versus-shuffled property tests, condition-aware discrepancy models, and subject-held-out prediction.

  • Results

    Recoverability was property-specific: population-level plausibility did not guarantee recording-specific preservation, and observation-associated degradation varied across properties.

  • Takeaways & Limitations

    Cross-modal physiological validation should test recording-level preservation of the specific property claimed to be recovered under relevant observation conditions.

  • Takeaways & Limitations

    The condition-aware analysis supports partial MAE reduction rather than complete explanation or correction of PPG–rPPG discrepancy, with maximum out-of-fold R2 of 0.189.

Abstract

from arXiv · show

Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but endpoint performance does not establish whether other physiological properties of source contact photoplethysmography (PPG) remain preserved recording by recording. We evaluated property-specific PPG-to-rPPG recoverability on 655 recordings from the Multi-Domain Mobile Video Physiology Dataset using CHROM as a fixed camera-rPPG observation pathway. The pathway reproduced the published CHROM correlation regime, with heart-rate MAE of 15.26 bpm and Pearson correlation of 0.0801. Matched-versus-shuffled validation revealed modest recording-specific autocorrelation correspondence, while spectral and recurrence-rate measures showed little matched discrimination. Maximal Lyapunov exponents showed essentially no recording-specific PPG-to-rPPG correspondence, with correlation of 0.0231 and permutation p-value of 0.5584, despite population-level overlap. Endpoint discrepancy exhibited Fitzpatrick-associated heterogeneity after adjustment for lighting and motion, including a Fitzpatrick VI versus III contrast of 9.32 bpm, while dynamical discrepancy showed no corresponding gradient. Aggregate RGB signal-to-noise ratio did not materially account for the endpoint contrast. In subject-held-out analysis, adding motion and lighting consistently reduced MAE across linear, ridge, and random-forest learners relative to rPPG-HR-only calibration, with reductions up to 13.32 percent. These findings show that recoverability is property-specific: physiological properties differ in recording-specific preservation and dependence on observation conditions, and population-level plausibility does not establish preservation of individual recordings.

1 Highlights · 3 Introduction

The study frames PPG-to-rPPG validation as property-specific recoverability rather than endpoint fidelity alone. It tests recording-specific correspondence, heterogeneous observation effects, and whether context improves endpoint calibration on unseen subjects.

  • 1 Highlights: The fixed CHROM pathway reproduced the published MMPD correlation regime without post-hoc tuning, with modestly higher absolute error.This establishes the observation pathway used for subsequent recoverability analyses.
  • 1 Highlights: rλ=0.0231 and permutation p=0.5584 showed no detectable recording-specific PPG-rPPG maximal-Lyapunov correspondence.This tests nonlinear dynamical preservation using matched recording identities rather than population-level overlap alone.
  • 1 Highlights: Observation context reduced subject-held-out PPG–rPPG HR MAE by up to 13.32%, with motion and lighting providing the predictive gain.The comparison used rPPG-HR-only calibration as the baseline and evaluated whether context was actionable on unseen subjects.
  • 1 Highlights: Fitzpatrick-associated endpoint heterogeneity persisted after SNR adjustment, while the tested dynamical discrepancy showed a different heterogeneity pattern.Fitzpatrick classification was treated as a measured grouping variable, and aggregate SNR as one possible quality descriptor rather than a complete explanation.
  • 3 Introduction: The unresolved question is which contact-PPG properties retain recording-specific correspondence after camera observation and reconstruction, how loss varies by conditions, and whether context helps on unseen subjects.This distinguishes plausible HR estimation from property-specific recoverability and recording identity preservation.
  • 3.1 Research questions and empirical mapping: The study contributes a unified framework spanning separate property tests, shuffled correspondence controls, heterogeneous-condition analyses, and subject-held-out context calibration.The contributions cover endpoint, conventional structural, and nonlinear dynamical properties across the contact-PPG–to–rPPG observation boundary.
  • 1 Highlights: Recoverability differed across endpoint, temporal, spectral, and nonlinear signal properties.The framework therefore treats these properties as separate validation targets rather than one modality-level fidelity judgment.

4 Methods

The study used a recording-matched, fixed-CHROM analysis of synchronized contact PPG and facial video, treating recoverability as property-specific, recording-specific, and condition-dependent. It evaluated endpoint discrepancy, structure correspondence, heterogeneity, and held-out prediction without assuming that one recovered property implies another.

  • Study design: The retrospective analysis used synchronized contact PPG and facial video, following a fixed hierarchy from CHROM benchmark verification to endpoint, structure, and shuffled-pair analyses.The unit of analysis was the recording, with repeated recordings handled using subject-level random intercepts.
  • Dataset: 655 available recordings were analyzed from an expected 660-recording MMPD design comprising 33 subjects with 20 recordings each.Five source recordings were absent; downstream exclusions and estimator failures followed prespecified validity rules.
  • Observation model: The observation model represented synchronized PPG and facial video through a fixed CHROM pathway under Fitzpatrick group, lighting, and motion/activity conditions.These factors were treated as observed acquisition descriptors rather than manipulated causal exposures.
  • Recoverability framework: Recoverability was evaluated separately for physiological or signal properties using recording-level discrepancies and matched-versus-condition-preserving-permuted correspondence tests.The framework explicitly separated estimability, recording-identity preservation, and variation across observation conditions.
  • Data handling: 645 unique matched records remained in the heterogeneous-analysis table after removing one duplicated matched key and excluding Subject 24 from Fitzpatrick-dependent analyses.Subject 24 was retained for analyses not requiring skin-type classification.
  • Held-out prediction: A subject-held-out follow-up tested whether measured observation variables could reduce contact-PPG reference-HR discrepancy from camera-derived rPPG-HR predictions using out-of-fold MAE.Nested predictor information sets M0–M4 represented different available variables rather than different learning algorithms.

5 Algorithmic specification

This section specifies mathematical procedures for evaluating endpoint discrepancy and recording-specific structural and dynamical correspondence between synchronized contact PPG and camera-derived rPPG. The procedures define reproducible quantities while keeping the fixed CHROM observation pathway unchanged.

  • Observation pathway: The fixed CHROM observation operator remains unchanged, with no post-hoc optimization introduced by the algorithms.Software-specific details remain delegated to the corresponding Methods subsections.
  • Algorithmic scope: The algorithms formalize endpoint discrepancy and recording-specific structural and dynamical correspondence for synchronized contact PPG and camera-derived rPPG.They express the reported analyses through mathematical operations needed for reproduction.
  • Property operators: Each synchronized recording pair is mapped into property space using operators that may return scalar endpoints, structural descriptors, recurrence statistics, or nonlinear dynamical quantities.The property operators are denoted Φ = {ϕj}J, with P representing contact PPG and R camera-derived rPPG.

2. For scalar endpoint HR, define the recording-level loss

Recording-level loss is defined through matched structural comparisons between contact PPG and synchronized rPPG, with shuffled identities serving as correspondence-breaking controls. Statistics that are intrinsically non-discriminative under the implemented parameterization are excluded from interpretation.

  • Matched structural comparison: Matched-pair functionals quantify correspondence for ACF, PSD, and REC structural properties.Sj denotes ACF similarity, cardiac-band PSD distance, or recurrence-rate discrepancy.
  • Matched-versus-shuffled control: Shuffling the rPPG identity breaks recording correspondence while retaining the analyzed signal collection.The shuffled identity is π(i), replacing the synchronized identity i.
  • Matched-versus-shuffled control: Similarity statistics require Smatch to exceed the shuffled counterpart, whereas distance and discrepancy statistics require the reverse direction.This directional comparison determines whether matched recordings show recording-specific correspondence.
  • Statistic validity: Saturated or boundary-concentrated statistics are treated as intrinsically non-discriminative rather than evidence for or against recording-specific correspondence.This rule applies to the saturated determinism statistic, while recurrence-rate discrepancy remains in the structural comparison.

6. Generate the condition-preserving subject-permutation null

This section defines a condition-preserving, property-specific recoverability profile using joint evidence from endpoint, structural, recurrence, and dynamical comparisons. Marginal similarity indicates population-level plausibility but does not alone establish synchronized-recording preservation.

  • Recoverability is assessed jointly through endpoint loss, matched-versus-shuffled structural comparison, recurrence-statistic informativeness, and recording-specific dynamical correspondence.
  • Population-level marginal similarity is treated as plausibility, not as sufficient evidence of synchronized-recording preservation.
  • The output is a property-specific recoverability profile rather than a single aggregate preservation judgment.
  • The dynamical procedure estimates ˆλmax for contact PPG and camera-derived rPPG under identical validity rules before forming matched PPG–rPPG dynamical discrepancy.Inputs undergo specified preprocessing, analysis intervals, SQI gating, and numerical constants; waveforms failing the spectral-quality gate are excluded from Lyapunov estimation.

2. Estimate the delay from the average-mutual-information sequence,

The method estimates delay from the average-mutual-information sequence using a first-local-minimum rule with a prespecified fallback.

  • Delay was estimated from the average-mutual-information sequence using the implemented first-local-minimum rule and its prespecified fallback.

3. Reconstruct the delay-coordinate trajectory in Rm, · 4. Let TW denote the temporal-exclusion duration and let · 5. For forward step k, compute the ensemble mean log-divergence

The method reconstructs delay-coordinate trajectories, excludes temporally proximate neighbors, and estimates maximal-divergence slopes from valid forward log-divergence trajectories. It applies the procedure independently to contact and camera signals before forming paired dynamical discrepancies and evaluating recording-specific correspondence.

  • 3. Reconstruct the delay-coordinate trajectory in Rm,: Delay-coordinate trajectories are reconstructed in Rm using the embedding dimension specified in the Methods.
  • 4. Let TW denote the temporal-exclusion duration and let: Temporally admissible nearest neighbors are selected using an exclusion width defined in samples at the analysis sampling frequency fs.
  • 4. Let TW denote the temporal-exclusion duration and let: TW = 100 ms after resampling to the common analysis time base.
  • 5. For forward step k, compute the ensemble mean log-divergence: For each forward step k, the ensemble mean log-divergence uses neighbor pairs whose trajectories remain defined at that step.
  • 5. For forward step k, compute the ensemble mean log-divergence: Forward divergence time and the study-specific maximal-divergence slope are estimated using the reported fitting region K and implemented linear fit.
  • 5. For forward step k, compute the ensemble mean log-divergence: Failed estimates remain missing rather than being imputed under the operational validity rule used for the reported results.
  • 5. For forward step k, compute the ensemble mean log-divergence: The procedure is applied independently to xP_i and xR_i, with paired dynamical discrepancy formed only when both estimates are valid.
  • 5. For forward step k, compute the ensemble mean log-divergence: Recording-specific correspondence across the common valid set is evaluated by Algorithm 1.

1. Define the two discrepancy variables · 4. Define the descriptive aggregate channel-quality index

The analysis defines model-based endpoint and dynamical discrepancy procedures, including subject-level random intercepts and prespecified Fitzpatrick-by-condition interaction tests. It also specifies descriptive SNR-adjustment sensitivity and subject-disjoint prediction to assess condition-aware endpoint-discrepancy reduction.

  • 1. Define the two discrepancy variables: Fitzpatrick-consistent populations are used for models containing Fi, while dynamical-loss analyses use the common valid Lyapunov set.These population restrictions define the analysis sets for the respective model families.
  • 1. Define the two discrepancy variables: Endpoint and dynamical discrepancy models use a model-specific subject random intercept, bs(i), to represent repeated measurements.The random-intercept variance is estimated separately for each fitted model.
  • 1. Define the two discrepancy variables: Prespecified secondary analyses augment one condition family at a time with Fitzpatrick-by-motion or Fitzpatrick-by-lighting interactions for Di ∈ {DHR,i, Dλ,i}.The full interaction coefficient block is tested jointly rather than interpreting selected cells post hoc.
  • 4. Define the descriptive aggregate channel-quality index: The SNR-adjusted sensitivity model is fitted on exactly the endpoint-model population.Its outputs include fixed-effect contrasts for endpoint and dynamical loss and an SNR-adjusted coefficient-sensitivity measure, AVI.
  • 4. Define the descriptive aggregate channel-quality index: AVI quantifies descriptive attenuation from the Fitzpatrick VI-versus-III coefficient before adjustment, βVI,1, to its SNR-adjusted counterpart, βVI,1b.AVI is interpreted as coefficient sensitivity to aggregate SNR adjustment, not causal mediation.
  • 4. Define the descriptive aggregate channel-quality index: Subject-disjoint prediction tests whether incrementally added observation context reduces held-out PPG–rPPG endpoint discrepancy.The primary performance measure is MAE, using observation variables Mi, Li, Fi, and Si, with Si = SNRRGB,i.

1. Define the nested information sets · 4. Generate exactly one held-out prediction for each recording,

The analysis uses nested observation-information sets to test incremental predictive value for contact-PPG heart rate. Subject-held-out predictions are generated fold-wise, evaluated with endpoint metrics, and interpreted only through uncertainty-calibrated endpoint improvements.

  • 1. Define the nested information sets: X0 = {H}, X1 = {H, M}, and X2 = {H, M, L} define progressively expanded observation-information sets.
  • 1. Define the nested information sets: X3 = {H, M, L, F} and X4 = {H, M, L, F, S} add further observation variables to the nested sets.
  • 1. Define the nested information sets: Each successive model tests whether an added observation variable predicts contact-PPG HR beyond the preceding information set.
  • 1. Define the nested information sets: Five disjoint subject folds, G1, . . . , G5, support subject-held-out evaluation.
  • 4. Generate exactly one held-out prediction for each recording,: For each learner family and information set, fold-specific prediction functions are estimated using training subjects only, with linear, ridge, or random-forest learners.
  • 4. Generate exactly one held-out prediction for each recording,: Predictions are concatenated across the five folds to produce the complete out-of-fold evaluation population, Deval, with Neval = |Deval|.
  • 4. Generate exactly one held-out prediction for each recording,: Performance is summarized using median absolute error, Pearson correlation, and R2 for each nested information set and learner family.
  • 4. Generate exactly one held-out prediction for each recording,: X0 is the learned camera-HR calibration baseline; incremental benefit is ∆MAEk,ℓ, with ∆MAEk,ℓ < 0 indicating lower subject-held-out error after adding context.Bootstrap resampling of subjects recomputes the benefit and its confidence interval; evidence requires the complete interval to lie below zero and concerns endpoint prediction only.

6 Results

Across 655 recordings, camera rPPG preserved source-PPG properties unevenly: temporal autocorrelation showed modest recording-specific correspondence, whereas maximal-Lyapunov dynamics did not. Endpoint discrepancy varied by Fitzpatrick group and motion interactions, but this pattern was not explained by aggregate RGB SNR or mirrored by dynamical discrepancy.

  • Dataset and processing: 655 recordings from 33 participants were processed without technical failure, with dynamics analyses assigning success or failure to every expected subject-recording combination.Five unavailable files reduced the expected 660 subject-recording combinations to 655 available recordings.
  • Recording-specific correspondence: ACF similarity was 0.2863 for matched pairs versus 0.2440 after shuffling, a +0.0422 difference, while cardiac-band PSD and recurrence-rate measures showed little matched discrimination.Cardiac-band PSD distance differed by -0.0009 and recurrence-rate discrepancy by -0.0017 between matched and shuffled pairs.
  • Recording-specific correspondence: rλ=0.0231 across 645 matched recordings, with permutation p=0.5584, indicating essentially absent recording-specific maximal-Lyapunov correspondence.The 95% CI was -0.0542 to 0.1001, and mean λmax was approximately 2.063 for PPG versus 2.256 for rPPG despite strong variance compression.
  • Endpoint–dynamical dissociation: Wald χ2 = 28.789, df=12, p = 0.004234 for the Fitzpatrick×motion interaction on DHR, whereas the corresponding Dλ interaction was nonsignificant.The Fitzpatrick×lighting interaction was also nonsignificant for DHR, and the adjusted motion profiles showed their largest separation in the post-exercise stationary condition.
  • SNR and endpoint discrepancy: Pearson r=0.014, p=0.718 between aggregate RGB SNR and absolute HR discrepancy, while SNR adjustment changed the Fitzpatrick VI-versus-III coefficient from +9.3218 to +9.3678 beats/min.Measured channel SNR varied by approximately 0.67 dB across Fitzpatrick groups, contrasting with the 9.21 beats/min increase in mean DHR from Fitzpatrick III to VI; the attenuation model had RMSE 6.47 dB and was not used for correction.

7 Discussion

The discussion shows that PPG-to-rPPG recoverability is property-specific rather than a single axis: recording-specific correspondence and sensitivity to observation conditions differ across tested properties. The findings support condition-aware endpoint-gap reduction but do not establish universal camera-rPPG limits or source-waveform reconstruction.

  • Property-specific recoverability: PPG-to-rPPG conversion showed different recording-specific recoverability across endpoint heart rate, temporal and spectral structure, and maximal-Lyapunov dynamics.Temporal autocorrelation retained a modest matched-pair advantage, whereas spectral and recurrence-rate measures showed little matched-versus-shuffled separation.
  • Property-specific recoverability: Overlapping population-level λmax ranges did not imply recording-specific preservation, because matched correspondence was essentially indistinguishable from the recording-identity null.The discussion identifies this as a counterexample to inferring individual-recording preservation from population-level physiological plausibility.
  • Heterogeneous observation conditions: Endpoint discrepancy varied with motion/activity and Fitzpatrick group, while the tested dynamical discrepancy showed no corresponding Fitzpatrick- or lighting-associated pattern.Adding aggregate RGB SNR did not materially account for endpoint heterogeneity, and a simplified monotonic attenuation model inadequately reproduced the observed SNR pattern.
  • Condition-aware prediction: Motion and lighting reduced held-out endpoint error beyond the rPPG-HR-only calibration baseline, supporting condition-aware gap reduction rather than reconstruction of the source PPG waveform.Fitzpatrick group and aggregate RGB SNR did not add comparable predictive value once other context was available.
  • Scope and implications: The weak CHROM benchmark characterizes properties accessible under this tested observation pathway, not what every possible camera-rPPG method fundamentally cannot recover.The discussion therefore recommends specifying the recovered property, testing recording-specific correspondence when appropriate, and evaluating stability across relevant observation conditions.

8 Limitations

The findings are limited by reliance on one dataset and one fixed CHROM pathway, representative finite-sample property estimators, associational condition analyses, and a condition-aware experiment focused on MAE reduction rather than explained variance.

  • Dataset scope: MMPD was the sole dataset, with fewer independent participants than recordings, especially within some Fitzpatrick strata, and no independent external validation.Subject-aware mixed-effects models, grouped cross-validation, and subject-level bootstrap resampling address repeated measurements but do not replace external validation.
  • Observation pathway: One fixed CHROM-based camera-rPPG pathway limits the findings to that operator on MMPD rather than establishing a fundamental limit of camera-based physiological sensing.Alternative extraction methods, acquisition systems, and learned observation models may preserve different properties and should be tested with the same recording-specific framework.
  • Property estimation: The tested properties are representative rather than exhaustive, and each is measured through a finite-sample estimator whose behavior may affect observed correspondence.Maximal Lyapunov exponent estimates additionally depend on preprocessing, state-space reconstruction, validity criteria, and estimator behavior.
  • Causal interpretation: Observation-condition analyses are associational: Fitzpatrick category and aggregate RGB SNR do not directly or completely measure the relevant optical properties or camera-domain information quality.Persistence of Fitzpatrick-associated endpoint heterogeneity after RGB-SNR adjustment neither identifies a causal mechanism nor establishes that optical signal quality is irrelevant.
  • Causal interpretation: Direct optical measurements, controlled acquisition, and prospective designs are needed to mechanistically separate subject-associated from acquisition factors.These designs would address the limits of attributing observed condition-related differences to specific optical mechanisms.
  • Condition-aware experiment: 13.32% was the largest reported MAE reduction for M2 relative to M0, while maximum out-of-fold R2 was 0.189 in an experiment designed primarily around MAE reduction.M2 reduced MAE by 13.32%, 13.32%, and 12.77% for linear regression, ridge regression, and random forest, respectively.

9 Conclusion · Supplementary figures · 12 Declarations

The study argues that physiological recoverability is property-specific and recording-specific under heterogeneous observation pathways. Supplementary analyses support the main findings, while the work remains a secondary analysis of an existing de-identified dataset.

  • 9 Conclusion: Population-level plausibility did not guarantee recording-specific preservation, and observation-associated degradation varied across properties beyond what simple signal-quality descriptors captured.The conclusion characterizes the evaluated contact-PPG-to-camera-rPPG pathway rather than a fundamental property of all observation pathways.
  • 9 Conclusion: Cross-modal validation should identify the physiological property claimed to be recovered and establish its recording-level preservation under relevant observation conditions.Endpoint performance alone should not be used to infer signal equivalence.
  • Supplementary figures: Supplementary analyses supplied validation and sensitivity checks covering the Fitzpatrick endpoint gradient, RGB-channel SNR, SNR adjustment, Lyapunov correspondence, and out-of-fold prediction.These analyses correspond to Supplementary Figures S1–S5.
  • Supplementary figures: RGB-channel SNR distributions showed extensive overlap and small central-tendency shifts across Fitzpatrick groups, providing descriptive context for the endpoint analysis.These summaries were not direct measures of melanin concentration or tissue optical absorption.
  • Supplementary figures: Aggregate RGB-SNR adjustment changed displayed Fitzpatrick contrasts by less than 1% relative to Model 1, indicating negligible descriptive attenuation.The sensitivity model retained Fitzpatrick group, lighting, motion/activity, and the subject random intercept; this was not interpreted as causal mediation.
  • Supplementary figures: rλ = 0.0231 and permutation p = 0.5584 indicated no detectable matched recording-level PPG–rPPG maximal-Lyapunov correspondence despite overlap in marginal values.The primary matched analysis used 645 recordings with valid PPG and rPPG λmax estimates.
  • Supplementary figures: The best condition-aware learner reduced MAE relative to the rPPG-only learner, but predictions remained compressed toward central HR ranges and departed from identity at extremes.Supplementary Figure S5 shows out-of-fold predictions from linear M2 using rPPG HR, motion, and lighting.
  • 12 Declarations: The study was a secondary analysis of the existing de-identified MMPD dataset, with no new participants recruited and no new data collected.This statement appears in the ethics declaration.
Loading 2608.27392v1…