Source-linked AI summary
Effective injury forecasting in soccer with GPS training data and machine learning
Alessio Rossi, Luca Pappalardo, Paolo Cintia, Marcello Iaia, Javier Fernandez, Daniel Medina
TL;DR
Professional soccer injury forecasting lacks strong evaluation of statistical models despite the substantial impact of injuries. This paper uses GPS-derived training workload and machine learning to build a multidimensional forecaster, reporting strong injury detection and precision with interpretable rules. The model’s usefulness increases after an initial period of data collection.
Problem
Existing research offers only preliminary understanding of injury-risk variables and limited evaluation of statistical models for forecasting injuries, despite injuries’ substantial impact on professional soccer.
Method
The paper builds a multidimensional injury forecaster from GPS training-workload data and machine learning, using a decision tree to support interpretation.
Results
DT detects about 80% of injuries with about 50% precision, outperforming baselines and state-of-the-art injury-risk estimation techniques.
Takeaways & Limitations
The approach provides interpretable rules that can support coaches and athletic trainers in evaluating injury risk and training decisions.
Takeaways & Limitations
Reliable forecasting requires an initial period of data collection because performance is poor when data are scarce.
Abstract
from arXiv · showhide
Injuries have a great impact on professional soccer, due to their large influence on team performance and the considerable costs of rehabilitation for players. Existing studies in the literature provide just a preliminary understanding of which factors mostly affect injury risk, while an evaluation of the potential of statistical models in forecasting injuries is still missing. In this paper, we propose a multi-dimensional approach to injury forecasting in professional soccer that is based on GPS measurements and machine learning. By using GPS tracking technology, we collect data describing the training workload of players in a professional soccer club during a season. We then construct an injury forecaster and show that it is both accurate and interpretable by providing a set of case studies of interest to soccer practitioners. Our approach opens a novel perspective on injury prevention, providing a set of simple and practical rules for evaluating and interpreting the complex relations between injury risk and training performance in professional soccer.
5 Athletic care department, Philadelphia 76ers, Philadelphia, USA
The paper is associated with sports analytics, data science, machine learning, sports science, and predictive analytics.
- The listed keywords place the work across sports analytics, data science, machine learning, sports science, and predictive analytics.
Introduction
Professional soccer injuries impose substantial costs, while existing research offers limited evidence for forecasting injury risk from available workload data. The paper therefore motivates accurate, interpretable, multidimensional models that can support staff decisions.
- Injuries affect professional soccer through team-performance consequences and substantial rehabilitation and absence costs.In Spain, injuries account for about 16% of professional players’ season absence and around 188 million euros per season.
- Existing soccer injury studies provide preliminary understanding of influential variables, while statistical-model evaluation for injury forecasting remains poor.
- Many existing approaches are mono-dimensional, using one variable at a time without exploiting complex patterns in available data.
- Clubs need injury forecasters that balance high accuracy against interpretability and avoid frequent false alarms or opaque black-box reasoning.
- The proposed contribution is a multidimensional, interpretable, data-driven approach intended to improve precision over existing mono-dimensional methods.The introduction reports precision below 5% for existing approaches and 50% for the proposed approach.
Material and Method
The study collected GPS-based workload and player information from 26 professional male soccer players, then represented training sessions with workload-derived features and injury labels for machine-learning forecasting.
- The study monitored 26 Italian professional male players during the 2013/2014 season, recording GPS and personal information across training sessions.The monitoring covered 23 weeks and included 931 individual training sessions.
- GPS measurements provide kinematic, metabolic, and mechanical workload features, supplemented by personal and previous-injury features.
- The dataset labels whether a player becomes injured in the next game or training session, using features describing personal characteristics and recent workload.
- A decision-tree classifier is trained after RFECV feature selection and ADASYN oversampling to address the highly imbalanced training data.The selected training split contains 279 non-injury examples and 7 injury examples before oversampling.
Results
The decision-tree forecaster achieved strong injury recall and substantially higher precision than baselines and ACWR/MSWR forecasters, while performance improved as season data accumulated. Its extracted rules identify interpretable workload patterns associated with injuries after prior injury.
- Recall=0.80±0.07 and precision=0.50±0.11 on the injury class show that DT detects most injuries while correctly labeling half of predicted injury sessions.
- DT substantially exceeds the baselines and ACWR- and MSWR-based forecasters in precision, whose maximum values are about 6% and below 4%, respectively.
- DT typically reduces false alarms, limiting unnecessary player stoppages before the next game or training session.
- DT performance is initially poor but improves as training and injury examples accumulate, outperforming other models from week w14.
- From w6 through season end, DT detects 9 of 14 injuries with F1-score=0.60 and precision=0.56.
- The decision tree yields six injury rules that coaches and trainers can use to interpret injuries and modify training schedules.Three summarized scenarios concern recently injured players’ high-speed-running distance and total-distance monotony; their cumulative frequency is 28% with mean accuracy 75±5%.
- At season end, RFECV selects PI(EWMA), dHSR(EWMA), and dTOT(MSWR), with decision-tree importances of 0.71, 0.23, and 0.06.
Discussion
The decision-tree injury forecaster performs well, balancing predictive accuracy with interpretability and supporting practical injury-prevention decisions. Its reliability improves after sufficient data collection, while selected features and actionable risk patterns evolve during the season.
- Forecasting performance: Around 80% of injuries are detected, while injury-class cumulative F1-score reaches 0.60, outperforming RF and LR baselines.The decision tree also has a small false-positive rate, while its false-negative rate is moderate.
- Forecasting performance: More than half of injuries could be prevented in the real-world scenario by using the forecasting model throughout the season.The scenario assumes data collection begins for the first time and the forecaster is retrained as the season progresses.
- Data requirements: Classifier performance stabilizes after 14 weeks of data collection, indicating that an initial period is needed before training a reliable model.The required collection period depends on training and match frequency, injury frequency, player availability, and tolerated false alarms.
- Feature dynamics: After 14 weeks, feature selection retains only 3 of 55 features, and this set remains stable during subsequent weeks.The selected features are PI(EWMA), dHSR(EWMA), and dTOT(MSWR); PI(EWMA) is consistently selected and reflects distance from returning to regular training after injury.
- Feature dynamics: The three-feature combination predicts injury better than PI(EWMA) alone, while detected injuries occur both soon after return and long after prior injury.About 60% of detected injuries occurred long after a previous injury and were associated with specific dHSR(EWMA) and dTOT(MSWR) values.
S1 Appendix. Descriptive statistics of the workload features.
The study describes 12 workload features using averages, standard deviations, and a normality test. None of the distributions is normally distributed; visual inspection indicates bimodal and right-skewed patterns.
- The study considers 12 training workload features and reports their average and standard deviation.
- None of the 12 workload-feature distributions is normally distributed according to the Shapiro-Wilks’ Normality test.
- Visual inspection shows that the workload-feature distributions tend to be bimodal and right skewed.
S2 Appendix. The ACWR method
The appendix evaluates ACWR and MSWR as injury-risk indicators and forecasting rules using workload features. Low ACWR is associated with the highest observed injury risk, but ACWR- and MSWR-based predictors have limited practical precision.
- ACWR method: ACWR is defined as the ratio between a player’s acute and chronic workload and is applied across 12 workload features.Training sessions are categorized into five predefined ACWR groups ranging from very low to very high.
- ACWR method: Players with ACWR < 1 have the highest observed injury risk, contrary to the cited finding that ACWR > 2 indicates higher risk.The same low-ACWR pattern is substantially confirmed when ACWR groups are defined by quintiles.
- ACWR method: 0.80 ± 0.08 recall and 0.03 ± 0.003 precision characterize single-feature ACWR forecasters, with injuries wrongly predicted in 97% of cases.The high recall therefore coexists with a very high false-alarm rate.
- ACWR method: ACWR-based predictors significantly outperform baseline classifiers on injury recall, but their low precision makes them unsuitable for practical injury prevention.The stated practical concern is that coaches would stop players without injury in the vast majority of cases.
- MSWR method: MSWR is the ratio of a workload feature’s weekly mean to its standard deviation, and high MSWR values are associated with high injury risk for most features.This relationship substantially confirms earlier literature findings.
- MSWR method: 0.10 ± 0.10 recall and 0.03 ± 0.03 precision describe MSWR predictors, whose combined models have poor injury-class accuracy comparable to ACWR predictors.Quantile-based ACWR predictors show similar classification results and remain limited by low precision.
- Machine-learning forecasting: DT(ADA+RFE) uses three of 55 features while maintaining performance comparable to DT(ADA), producing a more interpretable decision tree.The reported DT(ADA) performance is precision=0.88 and recall=0.92 on the injury class.
- Machine-learning forecasting: Among alternative classifiers, DT(T) reports precision=0.70 and recall=0.47, while RF(T) reaches precision=0.88 and recall=0.60; LR(T) performs worse.With feature selection, DT(RFE) detects 56% of injuries at 74% precision, while RF(RFE) has precision=0.78 and recall 0.58.