Source-linked AI summary
Births are difficult to predict even with rich survey and full-population register data
Elizaveta Sivak, Emily M. Cantrell, Thomas Emery, Javier Garcia-Bernardo, Flavio Hafner, Kasia Karpinska, Malte Lüken, Adrienne Mendrik, Joris Mulder, Hanzhang Ren, Varun Satish, Mark Verhagen, Angelica M. Maineri, Paulina Pankowska, Jasmin Abdel Ghany, Bruno Arpino, Giovanni Cassani, Julia Hellstrand, Katya Ivanova, Sanni Kuikka, Ana Macanovic, Charles Rahal, Felix C. Tropf, Roland J. Veen, Nicole Walasek, Daniël van Wijk, Kelsey Q. Wright, Emilio Zagheni, Henry Abbink, Emanuele Aliverti, Matteo Amestoy, Tilbe Atav, Nicola Barban, Sunnee Billingsley, Goan J. Booij, Louis Boucherie, Yael Broos, Li Ya Chang, Jamie C. Chiu, Chiara Ludovica Comolli, Boris Cule, Qixiang Fang, Dennis M. Feehan, Rachel Ganly, Erwin Gielens, Rolando M. Gonzales Martinez, Andrea Gradassi, Rosember Guerra-Urzola, Mario Guerra-Urzola, Stéphane Guerrier, Enamul Hassan, Vincent A. Haverhoek, Andrew T. Hendrickson, Amber Howard, Yuxuan Jin, Sayash Kapoor, Erik-Jan van Kesteren, Iris ten Klooster, Marie Labussiere, Lydia T. Liu, Tiffany Liu, Adam Maghout, Simone Meneghello, Lasse Mohr, Clara H. Mulder, Saul J. Newman, Jessica Nisén, Janis Norden, Mikkel Odgaard, Riccardo Omenti, Ozancan Ozdemir, Christina Pao, Paige Park, Gaia Penta, Juan C. Perdomo, Tanzir Pial, Alessio Piraccini, Federica Querin, Ziwei Rao, Christian Rellama, Adrien Remund, Frederieke Richert, Arnout van de Rijt, Mojtaba Rostami Kandroodi, Stijn J. Rotman, Lucas Sage, Germans Savcisens, Katrin Schwanitz, Steven Skiena, Alessandro Spata, Yannick Stadtfeld, Benedikt Stroebl, Gaetano Tedesco, Mathilde Theelen, Gianluca Tori, Abigail Tun-Mendicuti, Rishabh Tyagi, Keyon Vafa, Luiz Felipe Vecchietti, Linda Vecgaile, Willem R. J. Vermeulen, Maria-Pia Victoria Feser, Lionel A. Voirol, Thom B. Volker, Xinran Wang, Jiani Yan, Xinyi Zhao, Flora Zhou, Zuzana Zilincikova, Malvina Nissim, Matthew J. Salganik, Gert Stulp
TL;DR
The paper asks why major life events remain difficult to predict, even with rich data and modern algorithms. It combines a large Dutch fertility prediction challenge with simulations of reproductive randomness to distinguish reducible prediction error from chance. Predictions were only moderately accurate, and conception and pregnancy randomness imposed a meaningful ceiling on individual fertility prediction.
Problem
Major life events remain difficult to predict, raising questions about the roles of theory, data, algorithms, and chance.
Method
The study evaluates 147 researchers predicting births within three years from Dutch survey and full-population register data, then simulates reproductive randomness to estimate an upper bound.
Results
Best F1 was 0.59 for registers versus 0.76 for surveys, while the best survey-based model outperformed the best register-based model despite the register’s much larger sample and coverage.
Takeaways & Limitations
Even under highly favorable conditions, fertility was only moderately predictable, and chance in conception and pregnancy set a non-trivial ceiling on prediction.
Takeaways & Limitations
The estimated stochastic variation may partly reflect reducible uncertainty, although key biological data such as cycle-specific and cellular-level events will likely remain unavailable.
Abstract
from arXiv · showhide
Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.
Introduction
The paper tests individual fertility prediction under unusually favorable conditions, using rich Dutch survey and register data in a large interdisciplinary challenge. Fertility is a consequential and measurable life outcome, but prediction may still be limited by incomplete theories, imperfect measurement, algorithms, missing data conditions, and chance.
- Why fertility: Fertility is a useful test case because it is important, widely consequential, measurable with low error, and relatively easy to define as a prediction target.Individual-level fertility predictability remains less studied than fertility correlates and population-level trends.
- Data sources: Administrative records provide larger samples, finer resolution, and longer longitudinal coverage, whereas surveys capture intentions and other subjective factors that registers may miss.Survey weaknesses include smaller samples, selective non-response, and attrition; administrative data may omit self-reported behaviours and perceptions.
- Data sources: The challenge combines nationally representative survey data with detailed Dutch registers covering socioeconomic, demographic, family, residential, and social-network characteristics.The survey additionally measures health, values, personality, fertility intentions, and relationship quality.
- Role of chance: The study treats conception, fetal survival, and contraceptive failure as stochastic processes when estimating how much fertility prediction is inherently constrained by chance.The simulation creates a synthetic population in which age, partnership status, and fertility intentions determine outcomes alongside chance.
Results
Birth prediction was only moderately accurate despite rich survey data, full-population registers, and diverse modeling approaches. Performance did not substantially improve through more advanced models, model combination, or larger samples, while simulated reproductive randomness imposed a non-trivial but incomplete predictive ceiling.
- Predictability: 0.59 register F1 and 0.76 survey F1 were the best scores, with gradient-boosted decision-tree ensembles winning both challenge parts.The advanced transformer and large language model reached F1 = 0.57, below the register tree ensemble's 0.59.
- Predictability: 64% register recall and 66% survey recall show that identifying people who had a child remained difficult.The outcome occurred in 15% of the register sample and 22% of the survey sample.
- Data sources: The full-population register model did not outperform the survey model on a shared holdout, with delta MSE = 0.03 favoring survey predictions.Adjusting survey prevalence lowered its F1 to 0.68 and 0.63, but neither adjusted score fell below the register model's 0.59.
- Predictive ceiling: 0.963 register and 0.94 survey F1 were the predictive upper bounds under a 1% contraceptive-failure assumption, decreasing to 0.881 and 0.86 under 17%.The simulated bounds capture stochastic conception, fetal survival, and contraceptive failure, but the remaining gap also reflects data, methods, or other randomness.
- Model agreement: 21% of register holdout cases involving a child were misclassified by all top-five register models, indicating consistently difficult cases.Survey models also converged on similar performance and poorly predicted cases.
- Model agreement: Model aggregation produced no performance gains for either the register or survey data.The tested strategies included aggregating holdout predictions and combining model predictions or features.
- Data volume: Performance gains became marginal beyond approximately 100,000 register or 1,000 survey training observations, limiting the value of sample-size increases alone.These learning-curve results suggest insufficient sample size was not the most important reason for prediction difficulty.
Discussion
Even under unusually favourable conditions, individual births were only moderately predictable. The remaining error reflects missing or underused information, methodological limits, and stochastic processes in reproduction, while the benchmark supports future tracking of predictability.
- Predictive performance: 147 researchers could not very accurately predict who would have a child within three years using rich survey and full-population register data.The result establishes realistic expectations for individual-level fertility prediction and possibly similar life-course outcomes.
- Constraints: Minor gains from survey-register linkage and limited sensitivity to register-data reductions suggest insufficient sample size was not the main constraint.The authors caution that larger samples, richer features, or other settings could still benefit from additional data.
- Constraints: Missing or underutilised social, genetic, and health information constrained both datasets, although gains may remain limited given typically small effect sizes.The paper identifies richer health and genetics data, higher-resolution longitudinal data, and transfer learning as possible improvement avenues.
- Chance: Stochastic conception, pregnancy progression, and unplanned births impose a meaningful constraint on individual prediction that better data and algorithms cannot entirely eliminate.The simulation quantifies part of the variation attributable to chance in reproduction.
- Chance: The stochastic estimate is cautious because some variation attributed to chance may reflect reducible epistemic uncertainty, yet key biological information is likely unavailable in practice.Cycle-specific biological factors and cellular-level events would be needed to capture some remaining variation.
- Chance: Chance likely contributes more uncertainty than estimated because the analysis covers only reproductive physiology and unplanned births, not stochastic events in other life domains.The authors therefore argue that chance and luck deserve a more central place in fertility and social-science theories.
- Data sources: The best survey model outperformed the best register model despite the registers’ vastly larger sample and detailed longitudinal coverage.Fertility intentions explain much of the survey advantage and may summarise latent factors not captured in either source.
- Future evaluation: PreFer provides a repeatable benchmark for testing whether birth predictability changes across cohorts as data, methods, and social conditions evolve.Updated full-population registers make successive reruns feasible.
Methods
PreFer evaluated individual fertility predictions using Dutch register and LISS survey data, household-level holdouts, diverse algorithms, and a simulation-based upper bound for reproductive randomness.
- Data and sample: The target population comprised Dutch residents aged 18–45 in 2020, with register data covering the full population and survey data drawn from LISS panel participants.Register data included population, tax, education, employment, benefits, housing, and neighbourhood sources; LISS combined Core modules with Background surveys.
- Evaluation design: Participants used data through 2020, while outcomes were withheld for holdout groups to reduce leakage; household-level splitting kept household members together.A potential residual leakage source concerned Statistics Netherlands’ identification of unregistered cohabiting partners.
- Outcome: The outcome was having at least one new biological or adopted child between 2021 and 2023, identified from registers or repeated survey measures.Register outcomes used legal-parent and birthdate records; survey outcomes used reported numbers of children and children’s birthdays when needed.
- Challenge phases: The challenge ran in survey and restricted-register phases, with 41 survey teams making 69 valid submissions and 12 register teams making 11 valid submissions.Register access and secure-computing constraints affected participation and methodological choices.
- Predictive ceiling: The predictive upper bound assumed that age, partnership status, and stable fertility intentions fully determined behavior, leaving conception, fetal survival, and contraceptive failure as random.The simulation treated stochastic fecundability, age at sterility, fetal survival, and contraceptive failure as sources of predictive error.
- Assumptions and limitations: Some modeled reproductive variation may reflect genetic differences or environmental exposures rather than pure randomness, making the upper-bound interpretation dependent on current biological knowledge.The authors also describe the upper bound as conservative because real reproductive intentions and attempts may change over time.
Data Availability
The paper provides access routes for the LISS survey data and prepared survey datasets, while register data remains restricted to approved scientific users.
- Data access: LISS panel data and prepared PreFer survey datasets are available through the LISS archive and ODISSEI’s secure analysis environment.Register data is not publicly available and requires scientific-purpose access for vetted researchers affiliated with authorised institutions.
1. PreFer design
PreFer assembled register and LISS survey resources for a fertility-prediction challenge, with separate training and holdout structures designed to support evaluation and limit leakage.
- Register data: The register resources included socioeconomic, demographic, household, education, employment, housing, geographic, and network information from Dutch administrative sources.The Base dataset added constructed variables such as total children, youngest-child age, and linked partner information.
- Register data: The register target population included Dutch residents aged 18–45 at the end of 2020, after exclusions for residence status and other eligibility conditions.The initial population contained 6,092,379 individuals before exclusions.
- Register evaluation: Register participants received features through 2020 and outcomes for training cases, whereas holdout and supplementary outcomes were withheld for evaluation.Features from holdout and supplementary groups could still support network-characteristic construction.
- Survey evaluation: Survey households were assigned to training or holdout groups, stratified by whether any member had a new child, producing similar outcome, age, and participation distributions.The survey training set contained 6,418 members, while the holdout set contained 395 members with known outcomes.
- Survey evaluation: LISS maintained representativeness through refreshment samples, but its annual attrition rate was approximately 10%.The panel began in 2007 with approximately 5,000 households and 8,000 individuals aged 16 years and older.
- Challenge implementation: Access and computing constraints affected the register phase, including restricted data access, scheduled supercomputer use, and one team’s inability to submit a valid model.These constraints also influenced methodological choices made by teams.
2. Analysis of the predictive performance
Predictive performance was moderate and robust across splits, with survey models outperforming register models despite the latter’s larger dataset. Results indicate that fertility intentions, large samples, and additional register features each influence performance, while leakage appears limited.
- F1 decreased from 0.76 to 0.65 for the best model and from 0.71 to 0.54 for the data-driven model after removing fertility-intention variables.These corresponded to relative decreases of 15% and 24%, respectively.
- Recall of the best register model declined across the January 2021–December 2023 outcome period, indicating that later births were harder to anticipate from pre-2021 data.
- Adding register predictions increased survey-model F1 from 0.63 to 0.68, with the final gain requiring the full register sample and complete feature set.A register model trained on the same individuals reduced performance to F1 = 0.60, whereas comparable smaller or separate subsamples produced no improvement.
- Excluding potentially leaked early-2021 cases changed the best register-model F1 only from 0.59 to 0.58, suggesting minimal impact for that model.The leakage arose because household type in 2020 could partly reflect future births for a small subset of cases.
3. Submissions: scores and brief description
The submissions used diverse classical, ensemble, and advanced modelling approaches, with performance shaped by feature choices, validation, and fertility-related variables.
- Register-based submissions: 0.59 was the score for Stork Oracle, a CatBoost model trained on about 100 manually selected fertility-related variables.The variables were selected using ideas about what should be important for predicting fertility.
- Predictive features: Fertility intentions, household income, and marital status were consistently important predictors in submitted models.Fertility intentions predicted outcomes, household income mattered more than individual income, and marital status repeatedly appeared as important.
- Survey-based submissions: 0.771 F1, 0.720 MCC, 0.895 precision, and 0.679 recall were reported for one survey-based model.The reported intervals were [0.667, 0.861], [0.600, 0.828], [0.786, 0.976], and [0.545, 0.808], respectively.
- Validation and robustness: Validation estimates could be overly optimistic: one model fell from about 0.93 F1 on an internal test set to about 0.61 on the competition holdout.The team changed from five-fold cross-validation to repeated stratified shuffle splits after observing this discrepancy.
4. Additional funding
The research received support from multiple national, university, and consortium funding sources across Europe and the United States.
- Funding: Lucas Sage acknowledged funding from the French Agence Nationale de la Recherche under the Investissement d’Avenir programme.The grant was ANR-17-EURE-0010.
- Funding: Julia Hellstrand, Kelsey Wright, and Ziwei Rao received Strategic Research Council support through Finland’s FLUX consortium.The cited decision numbers were 364374, 364375, 345130, and 345131.
- Funding: Other authors reported support from UK Research and Innovation, the Academy of Finland, Italian research funding, the Dutch Research Council, NSF, and Princeton initiatives.The acknowledgements list FINDME, INVEST, SO-UNFER, NWO VI.Veni.231S.148, NSF grants, Princeton Precision Health, the Princeton AI Lab, and the Princeton Catalysis Initiative.