Source-linked AI summary
A Reinforcement Learning System to Encourage Physical Activity in Diabetes Patients
Irit Hochberg, Guy Feraru, Mark Kozdoba, Shie Mannor, Moshe Tennenholtz, Elad Yom-Tov
TL;DR
Sedentary patients with type 2 diabetes often fail to follow physical-activity recommendations, creating a need for more effective adherence support. The study tested smartphone monitoring with reinforcement-learning-selected SMS messages and found improved activity, walking pace, and HbA1c outcomes relative to control policies. The authors conclude that personalized, automated mobile feedback may support exercise adherence, while noting limits from the single model and contextual-bandit design.
Problem
Most patients with type 2 diabetes remain sedentary despite physical-activity recommendations, and personalized SMS reinforcement had not been reported.
Method
The study used smartphone activity monitoring and an online reinforcement-learning policy to select personalized messages based on patient context and prior responses.
Results
RL-selected messages increased walking duration and rate, while learned-policy allocation was associated with superior HbA1c improvement compared with weekly reminders and competing policies.
Takeaways & Limitations
The results suggest that automated, personalized mobile feedback can support physical-activity adherence and personalized care in diabetic patients.
Takeaways & Limitations
The system used one model that ignored sex and age, and its contextual-bandit design did not model the patient’s underlying state because available data were limited.
Abstract
from arXiv · showhide
Regular physical activity is known to be beneficial to people suffering from diabetes type 2. Nevertheless, most such people are sedentary. Smartphones create new possibilities for helping people to adhere to their physical activity goals, through continuous monitoring and communication, coupled with personalized feedback. We provided 27 sedentary diabetes type 2 patients with a smartphone-based pedometer and a personal plan for physical activity. Patients were sent SMS messages to encourage physical activity between once a day and once per week. Messages were personalized through a Reinforcement Learning (RL) algorithm which optimized messages to improve each participant's compliance with the activity regimen. The RL algorithm was compared to a static policy for sending messages and to weekly reminders. Our results show that participants who received messages generated by the RL algorithm increased the amount of activity and pace of walking, while the control group patients did not. Patients assigned to the RL algorithm group experienced a superior reduction in blood glucose levels (HbA1c) compared to control policies, and longer participation caused greater reductions in blood glucose levels. The learning algorithm improved gradually in predicting which messages would lead participants to exercise. Our results suggest that a mobile phone application coupled with a learning algorithm can improve adherence to exercise in diabetic patients. As a learning algorithm is automated, and delivers personalized messages, it could be used in large populations of diabetic patients to improve health and glycemic control. Our results can be expanded to other areas where computer-led health coaching of humans may have a positive impact.
1 Introduction
Most patients with type 2 diabetes do not follow physical-activity recommendations despite known metabolic and quality-of-life benefits. This study evaluates smartphone monitoring and personalized reinforcement-learning feedback to improve adherence.
- Regular physical activity is recommended because it improves glucose control, metabolic risk factors, and quality of life.
- Most diabetic patients remain insufficiently active despite recommendations, motivating better ways to encourage exercise.
- Prior interventions included persuasive communication, financial incentives, and community programs to improve adherence.
- Earlier mobile interventions used random messages or activity displays, but none used a personalized learning algorithm to tailor messages to individuals.
- The study assesses automatically tailored feedback delivered through a smartphone application that measures activity and uses reinforcement learning to select messages.
2 Materials and Methods
The intervention combined smartphone-based activity monitoring with personalized SMS policies selected and updated by a reinforcement-learning system. Patients were randomized against weekly reminders, while the algorithm predicted next-day activity responses from user context and message history.
- The app collected patients’ physical activity in the background and transmitted the data to a central server.
- Each morning, the algorithm selected an SMS message for each patient using demographics, past activity, expected activity, and message history.
- The following morning’s activity served as the reward used to train the reinforcement-learning algorithm.
- The study recruited adults with type 2 diabetes, non-optimal glycemic control, sedentary lifestyles, and Android smartphones with data access for a 26-week study.
- Patients were randomized to unchanging weekly reminders or personalized daily feedback with weekly summaries, while medical staff were blinded to message type.
- The learned policy used exploration and a daily-updated linear regression model with interactions to predict activity changes from patient attributes and message actions.
- The approach functioned mainly as a contextual bandit because it predicted immediate action effects rather than modeling the patient’s underlying state.
3 Results
The results section reports recruitment of 27 patients who successfully installed the app and transmitted data for at least one week.
- 27 patients successfully installed the mobile application and transmitted data for at least one week.
3.2 Application Use and Physical Activity Measured
Participants provided activity data for an average of 20.0 weeks, with interruptions mainly related to phone changes or phone numbers. The analysis included all patients who initiated app use, including those who did not complete 26 weeks.
- 20.0 weeks was the average duration for which the app continued providing activity data.The reported standard error was SEM 1.6.
- 139 ± 62 minutes per week was the average target physical activity.
- Analysis included all participants who successfully initiated application use, including those who did not complete the 26-week experiment.
- Participants reported no regular pre-recruitment physical activity, but objective accelerometer data from before recruitment were unavailable.
3.3 Effect of Different Messages Over Time
Participants responded differently to individual messages and message sequences. The learned policy produced a statistically significant improvement over the initial policy, and message effects depended on temporal context.
- Different messages and consecutive-message sequences produced significantly different changes in participants’ activity.
- The day after a positive-social message produced the greatest activity increase, whereas negative and positive-self messages decreased activity.
- The learned policy differed significantly from the initial policy in its effects on activity change (ANOVA, P = 0.0036).
- Message effects depended on the previous day’s feedback, including negative feedback becoming positive before positive-self feedback and repeated positive-social feedback becoming ineffective.
3.4 Variability in patient response
Patients varied substantially in how they responded to feedback messages. Clustering revealed distinct response patterns and demographic differences, supporting individually tailored feedback.
- Patients were represented by four-dimensional vectors summarizing their average activity change after each daily feedback message.
- Three k-means clusters captured distinct response patterns, including patients reacting negatively to all messages and patients reacting positively, especially to positive-social or positive-self messages.
- The observed response differences were presented as evidence for individually tailored feedback.
- Cluster 3 was dominated by males, whereas cluster 2 consisted mostly of women; age differences across clusters were minor.
3.5 The Learning Process of the Algorithm Over Time
The learning algorithm became more stable and increasingly accurate as participant data accumulated. Its predictions explained a substantial portion of day-to-day variation in exercise, while adverse weather appeared to trigger new learning.
- The algorithm gradually improved its prediction of activity using participants’ responses, previous activity, and demographics.
- The algorithm’s stability increased over time as more participant data were collected.
- Adjusted R2 initially increased to approximately 0.43, indicating that the predictions explained much of the variation in exercise on a given day.
- Stability jumps, including around day 60, appeared to correspond to major adverse weather events and new behavioral patterns.
- The authors noted that longitudinal data across varied circumstances may require additional variables such as weather and calendar events.
3.6 Improvement in Activity Quantity and Walking Rate
The learned policy was associated with increasing activity quantity and walking rate over time, unlike the control and initial policies. The model incorporated activity history, feedback timing, message choice, and their interactions.
- For one participant, the fitted linear slope of activity over the experiment was 0.0016.
- The learned policy showed a positive activity slope, whereas the control population and initial policy showed negative changes in activity over time.
- The learned policy’s activity slope was superior to both the control population and the initial policy.
- Patients receiving personalized messages increased their walking rate over time, while control-condition patients reduced theirs.
- The predictive model included interactions involving previous-day activity, the feedback message, activity performed so far, and time since feedback.
3.7 Change in Glycemic Control
HbA1c improved on average, and personalized-policy allocation and longer participation were associated with greater reductions. Because dietary and medical treatment was not restricted, the change reflects multiple influences.
- 0.28 ± 0.84 improvement in HbA1c was observed across patients from an initial 7.8 ± 1.0.Dietary and medical treatment intensification was permitted, so HbA1c changes reflect exercise alongside treatment changes.
- Allocation to the personalized policy, higher initial HbA1c, and lower activity targets led to superior HbA1c reduction (R2 = 0.405, P < 10^-3).The linear model used HbA1c change as the dependent variable and included measurement interval, initial HbA1c, and activity target as predictors.
- 0.05 treatment-population slope versus -0.06 control-population slope indicated greater blood-glucose reduction with longer personalized-policy participation.The corresponding model fits were R2 = 0.07 for treatment and R2 = 0.03 for control.
- Personal messages were associated with a statistically significant reduction in HbA1c levels.
3.8 Participant Satisfaction
Participants receiving learned-policy messages reported greater help from SMS messages in increasing and maintaining activity than control participants. The satisfaction questionnaire identified only one statistically significant between-group response.
- p < 10^-3 difference favored learned-policy messages for helping participants increase and maintain activity.Both groups reported increasing physical activity, but the learned-policy group rated the messages as more helpful than controls.
- Only the second satisfaction-question response differed significantly between control and personalized messages (chi2 test).
4 Discussion
The discussion presents reinforcement learning as a way to tailor exercise feedback to individuals and message timing. Personalized, changing messages were associated with better activity and HbA1c outcomes, but the pilot’s scale and model scope limit generalization.
- Reinforcement learning selected feedback for each individual and situation, using reactions to messages and message sequences to personalize reminders.
- Changing messages based on performed activity increased walking duration and walking rate, whereas constant weekly reminders did not.The authors also report that the RL algorithm learned to sequence messages to maximize efficiency.
- A single model ignored sex and age; separate contextual models might perform better but would require a larger population and another algorithm.
- Online on-policy learning enabled exploration where outcomes mattered, unlike off-policy learning, which can introduce variance and bias.
- Personalized-policy allocation and longer participation were correlated with superior HbA1c improvement over weekly reminders and context-insensitive policies.
- The system is described as both a predictive tool and a method for personalized care, with potential for economical and efficient implementation.
- The study was small-scale, and larger, longer studies are needed to evaluate effects on health-related behavior and actual health.