Source-linked AI summary

Indicators of retention in remote digital health studies: A cross-study evaluation of 100,000 participants

Abhishek Pratap, Elias Chaibub Neto, Phil Snyder, Carl Stepnowsky, Noémie Elhadad, Daniel Grant, Matthew H. Mohebbi, Sean Mooney, Christine Suver, John Wilbanks, Lara Mangravite, Patrick Heagerty, Pat Arean, Larsson Omberg

arXiv:1910.01165v1stat.APcs.CY

TL;DR

Remote digital health studies can recruit large cohorts, but attrition and open-enrollment patterns raise concerns about retention and representativeness. This paper analyzes eight studies with more than 100,000 participants to identify retention and app-usage patterns, finding associations between longer retention and clinician referral, compensation, clinical conditions, and older age, alongside demographic and regional under-representation.

  • Problem

    Retention and recruitment factors in remote research remain insufficiently understood, threatening the representativeness and generalizability of collected data.

  • Method

    The study evaluates recruitment, retention, and app-engagement data from eight remote digital health studies involving more than 100,000 participants.

  • Results

    Clinician referral was associated with a 40-day increase in median retention, while the studies also showed demographic and regional under-representation.

  • Takeaways & Limitations

    The findings can inform recruitment and retention strategies intended to support more equitable participation in future digital health research.

  • Takeaways & Limitations

    Most studies were unable to recruit representative samples demographically or regionally.

Abstract

from arXiv · show

Digital technologies such as smartphones are transforming the way scientists conduct biomedical research using real-world data. Several remotely-conducted studies have recruited thousands of participants over a span of a few months. Unfortunately, these studies are hampered by substantial participant attrition, calling into question the representativeness of the collected data including generalizability of findings from these studies. We report the challenges in retention and recruitment in eight remote digital health studies comprising over 100,000 participants who participated for more than 850,000 days, completing close to 3.5 million remote health evaluations. Survival modeling surfaced several factors significantly associated(P < 1e-16) with increase in median retention time i) Clinician referral(increase of 40 days), ii) Effect of compensation (22 days), iii) Clinical conditions of interest to the study (7 days) and iv) Older adults(4 days). Additionally, four distinct patterns of daily app usage behavior that were also associated(P < 1e-10) with participant demographics were identified. Most studies were not able to recruit a representative sample, either demographically or regionally. Combined together these findings can help inform recruitment and retention strategies to enable equitable participation of populations in future digital health research.

1 Sage Bionetworks, Seattle, WA, USA

The passage identifies an affiliation with the Department of Biomedical Informatics and Medical Education at the University of Washington in Seattle, Washington, USA.

  • The affiliation is with the Department of Biomedical Informatics and Medical Education.
  • The department is part of the University of Washington.
  • The affiliation is located in Seattle, WA, USA.

8 Department of Biostatistics, University of Washington, Seattle, WA, USA

Remote digital health studies offer large-scale, lower-cost, real-world data collection, but retention, diversity, and representativeness remain unresolved. This study evaluates recruitment and retention drivers across eight studies to identify potential sources of bias.

  • Open enrollment may introduce selection and ascertainment bias, while attrition and variable app usage can produce unrepresentative cohorts.
  • The study examines factors associated with participant retention and long-term app usage in remote research.
  • The analysis compares demographic characteristics and engagement strategies associated with retention, including representational bias in real-world data.

Participant Characteristics

Across the eight studies, participants were concentrated among younger adults and non-Hispanic White individuals, with substantial demographic and regional representation differences from the US population.

  • Female participation had a median of 56.9% but varied from 29.4% to 100% across studies.
  • Hispanic/Latino and African-American/Black participants differed from 2010 census metrics by -8.09% and -9.15%, respectively.
  • Median recruited participation by state showed notable differences from each state's US population proportion.

Participant Retention

Retention varied substantially across studies and participant groups. Older age, clinical conditions of interest, and clinician referral were associated with longer retention, while missing demographic information was associated with variation in retention.

  • Participants engaged during the first 12 weeks had median study engagement of 5.5 days, with in-app tasks performed on 2 days.
  • Median retention varied from 2 to 12 days across studies, with Brighten an outlier at 26 days.
  • The first 8 days of engagement were associated with a 25-day increase in median retention for that sub-cohort.
  • Participants aged 60 years and older had median retention of 7 days and remained significantly longer than the younger sample.
  • Declared gender was not significantly associated with retention, with P = 0.3.
  • Participants with study-relevant clinical conditions had median retention of 13 days versus 6 days for non-disease controls.
  • Participants with missing demographics showed variation in retention compared with participants who shared their demographics.

Participant Daily Engagement Patterns

Daily app engagement formed distinct longitudinal patterns that varied across demographic groups, while early discontinuation was common and associated with participant characteristics and recruitment context. These patterns identify potential targets for improving retention and assessing bias in remote research.

  • Engagement clusters: Four distinct clusters captured overall app-usage behavior: high, moderate, sporadic, and abandoner engagement.Participants who remained in the study for fewer than 7 days were assigned to the abandoner group C5* rather than clustered with longer-term users.
  • Engagement clusters: 54.6% of participants across apps belonged to the abandoner group C5*, with median app usage of just 1 day.C5* represented participants who discontinued immediately, whereas sporadic cluster C4 had a median of 5 days between app uses versus 2 days for C3.
  • Demographic differences: Adults aged 60 years or older comprised 15.1–17.2% of higher-engagement clusters versus 5.1–11.7% of lower-engagement clusters.The difference in proportions across engagement clusters was significant at P=1.38e-12.
  • Recruitment representativeness: Most studies did not recruit demographically or regionally representative samples, with southern, rural, and Midwestern states underrepresented.The authors link this recruitment bias to challenges in reaching racial and ethnic underserved communities and to potential effects on studies of geographically patterned conditions.
  • Implications: Daily usage patterns and retention differences suggest combining recruitment, compensation, co-design, and targeted engagement strategies.A run-in period may increase statistical power by creating a smaller, more engaged cohort, but it does not resolve potential bias.
  • Demographic differences: Only 1 in 10 participants were in high-use clusters C1–2, which were largely Non-Hispanic white and older adults.Minority and younger populations were more represented in the clusters with the lowest daily app usage.
  • Retention factors: Clinician referral was associated with the largest retention difference, exceeding tenfold in the present sample.The referral could be light-touch, such as providing a study flyer during a regular clinic visit.

Data Harmonization

The analysis harmonized user-activity and baseline demographic data across studies to compare engagement and retention consistently. Active tasks were distinguished from passive data, and study-specific recruitment and disease-status information supported subgroup comparisons.

  • Activity data: User activity data across apps were harmonized to enable inter-app comparison of engagement metrics.The harmonized activity data included in-app surveys and sensor-based tasks classified as active tasks.
  • Activity data: Passive data, such as step counts and local weather patterns, were excluded from active user-engagement assessment.Active engagement required explicit user action, whereas passive data were gathered without it.
  • Demographic data: Age, gender, race, and state formed the minimal demographic subset used for recruitment and retention analysis.These characteristics were drawn from baseline demographics collected by each app.
  • Subgroup comparisons: Six studies enrolled participants with and without disease, enabling comparisons between case and control retention.The six studies were mPower, ElevateMS, SleepHealth, Asthma, and MyHeartCounts as listed in the supplied passage.
  • Subgroup comparisons: Two studies included clinician-referred and self-referred participants, allowing retention differences between referral groups to be compared.The clinician-referred subgroup came from mPower and ElevateMS.

Statistical Analysis

Retention and engagement were evaluated with duration-based metrics, survival analysis, and unsupervised clustering of longitudinal activity. Sensitivity analyses addressed censoring and missing demographic data, while demographic and geographic representativeness were assessed against census estimates.

  • Retention metrics: Three metrics assessed retention and long-term engagement: duration in the study, days active, and user activity streak.Duration measured total time active, days active counted days with any active task, and the streak encoded activity across 84 study days.
  • Survival analysis: Survival analysis compared retention across studies, sex, age group, disease status, and clinical referral.Log-rank tests stratified by study and Kaplan–Meier plots summarized pooled effects where applicable.
  • Sensitivity analyses: Two censoring approaches tested retention robustness over the first 84 days, using either no censoring or right-censoring based on 20 weeks of activity.The right-censoring sensitivity approach treated later activity as evidence that participants remained active during the initial 84-day period.
  • Representativeness and robustness: Cluster demographic enrichment was assessed with one-way ANOVA, while recruited state proportions were compared with 2018 US Census population estimates.Additional analyses compared participants who provided demographics with those who opted out to assess missing-data sensitivity.
  • Model assumptions: Cox proportional-hazards modeling was not pursued for some studies because the proportional-hazards assumption was unsupported.The assumption was tested using scaled Schoenfeld residuals.

studies

Table 2 summarizes selected participant demographics and study-app usage across the eight digital health studies.

  • Table 2 presents selected participant demographics and study-app usage measures for eight digital health studies.
Loading 1910.01165v1…