Source-linked AI summary
A Physiology-Informed Digital Twin Framework for Simulating Liver Health Progression
Sumaiya Afroz Mila, Sandip Ray
TL;DR
HEPATWIN addresses limited availability of intermediate-stage liver data and limited continuous physiology-informed simulation for early disease monitoring. It combines mechanistic liver modeling, patient-specific inputs, and stage-transition calibration, then validates estimated trajectories on NIDDK NAFLD data. The model produces clinically acceptable multi-year biomarker forecasts, and its simulated biomarkers support 65% NASH detection accuracy in a long-horizon setting.
Problem
Intermediate liver-disease stages are underrepresented in datasets, and few studies compare physiology-informed models with AI approaches for continuous early-stage disease simulation.
Method
HEPATWIN models hepatic metabolism, detoxification, biosynthesis, and inter-organ feedback from patient-specific inputs to generate calibrated longitudinal biomarker trajectories.
Results
65% NASH detection accuracy was obtained using HEPATWIN-estimated trajectories generated from baseline lifestyle and demographic information up to five years into the future.
Takeaways & Limitations
HEPATWIN demonstrates the potential of physiology-informed digital twins for personalized, non-invasive liver-health monitoring and prediction.
Abstract
from arXiv · showhide
We present a physiology-informed digital twin of the human liver designed for longitudinal simulation of liver function and early-stage disease progression. The model, referred to as HEPATWIN, integrates key hepatic processes, including carbohydrate, lipid, and protein metabolism, bilirubin conjugation, bile production, and detoxification, within a unified systems-level framework to generate clinically observable biomarker trajectories. Unlike purely data-driven approaches, HEPATWIN incorporates mechanistic representations of liver physiology and patient-specific inputs such as diet, activity, and baseline biomarkers to simulate disease evolution over time. To ensure consistency with clinical progression patterns, we introduce a stage-transition-driven calibration mechanism that aligns simulated outputs with population-level biomarker distributions across disease stages, including NAFLD, fibrosis, and cirrhosis. Validation using the NIDDK NAFLD dataset demonstrates that HEPATWIN produces longitudinal biomarker estimates within clinically acceptable ranges and can forecast trajectories over multi-year horizons. Furthermore, simulated biomarkers retain sufficient clinical signal to support downstream NASH detection with competitive performance relative to models using ground-truth laboratory data. These results highlight the potential of physiology-informed digital twins for personalized, non-invasive diagnosis and prediction of organ health in general and liver health monitoring in particular.
I. INTRODUCTION
HEPATWIN addresses limited support for continuous, physiology-informed simulation of early liver disease by modeling hepatic processes and patient-specific lifestyle inputs. It validates longitudinal biomarker estimation and uses simulated trajectories for personalized disease monitoring and NASH detection.
- Motivation: Early liver disease is difficult to detect because symptoms often appear only after progression to advanced stages.NAFLD can progress through fibrosis and cirrhosis before clinical symptoms manifest.
- Research gap: Existing datasets underrepresent intermediate stages such as NAFLD, NASH, and fibrosis, while physiology-informed continuous simulation remains uncommon.The passage contrasts advanced-stage dataset availability with limited systematic comparisons of AI and physiology-informed approaches.
- Approach: HEPATWIN simulates core hepatic processes and estimates clinically measurable biomarker trajectories from diet, baseline biomarkers, demographics, and physical activity.Modeled outputs include biomarkers such as ALT, ALP, albumin, and bilirubin across multi-year horizons.
- Approach: Unlike purely statistical forecasting, HEPATWIN models how biomarkers evolve over time under patient-specific lifestyle patterns.This produces longitudinal trajectories rather than isolated static estimates.
- Validation: Validation compares simulated biomarkers with longitudinal clinical measurements and tests their usefulness for downstream NASH detection.The two strategies assess estimation accuracy and preservation of clinically meaningful disease-stage signal.
- Implications: The framework supports personalized, scenario-based evaluation of liver-health trends and non-invasive early-stage monitoring.These capabilities are presented as potential uses of HEPATWIN-estimated biomarker trajectories.
II. RELATED WORKS
Related work highlights the need for reliable, less invasive early-stage liver assessment. Existing imaging and digital-twin approaches provide useful capabilities but remain limited for real-time, stage-specific biomarker simulation under data scarcity.
- Disease progression: Liver disease progresses from hepatic fat accumulation through inflammation and fibrosis toward cirrhosis and hepatocellular carcinoma.Early NAFLD is described as reversible, whereas later-stage treatment primarily manages damage.
- Diagnostic limitations: MRI-PDFF and transient elastography measure liver fat more accurately than traditional ultrasound or CAP approaches, but early detection remains challenging.Imaging biomarkers are more commonly used for patients already diagnosed with fibrosis or cirrhosis.
- Diagnostic limitations: Biopsy is invasive and costly, limiting its routine use for early NAFLD detection in the general population.The passage identifies less invasive, reliable early-stage biomarkers as an active research need.
- Digital twins: Digital twins integrate biomarkers, demographics, and lifestyle factors to construct personalized virtual liver profiles for progression or treatment simulation.Their hepatology applications include predictive care and tracking biomarker trends over time.
- Digital twins: Existing liver-focused digital twins mainly address pharmacokinetics, surgical planning, injury, regeneration, or posthepatectomy outcomes.Physiology-informed stage-specific biomarker estimation under data scarcity is described as limited.
- Physiological basis: Liver physiology depends on metabolism, bile production, bilirubin conjugation, detoxification, protein synthesis, and communication with peripheral organs.Pancreas, skeletal muscle, adipose tissue, and gallbladder contribute contextual regulatory signals.
B. Carbohydrate (Glucose) Metabolism Pathway
The carbohydrate pathway describes how dietary glucose and fructose are routed through immediate energy use, hepatic and peripheral glycogen storage, and further metabolic processing. HEPATWIN represents these physiological processes through biomarker-generating computational functions.
- Carbohydrate metabolism: Dietary glucose and fructose reach the liver through portal circulation and are allocated according to immediate and anticipated energy demands.Some energy supports vital organs, while remaining substrate is stored or metabolically processed.
- Glycogen storage: After immediate energy needs are met, glucose is stored as hepatic glycogen and later used during fasting or overnight periods.Finite liver and skeletal-muscle glycogen capacity means sustained surplus is routed toward further processing.
- Pathway illustration: The figure presents selected liver physiological pathways, including glucose metabolism and bilirubin conjugation.Its caption identifies the illustration as a simplified representation rather than a complete cellular model.
- Digital abstraction: HEPATWIN models metabolism, protein synthesis, bilirubin conjugation, bile production, and detoxification to generate updated liver biomarkers.The framework outputs ALT, ALP, albumin, and bilirubin as the physiological state evolves.
- Simulation: Simulation duration is configurable in days and depends on specified dietary and lifestyle conditions.The resulting time-series biomarkers can support disease-stage analysis such as NASH detection.
A. Conceptual Design of HEPATWIN
HEPATWIN abstracts essential liver functions into interacting computational components that transform patient context into evolving physiological states and clinical biomarkers. Its validation combines biomarker comparison with downstream NASH classification.
- Conceptual design: HEPATWIN focuses on clinically observable hepatic functions rather than detailed cellular biochemical pathways.The modeled functions include carbohydrate metabolism, glycogen storage, lipid handling, fat accumulation, protein synthesis, bile production, bilirubin conjugation, and detoxification.
- Conceptual design: The digital representation translates dietary, exercise, physical-profile, and biomarker information into structured model inputs.These inputs drive process variables that track liver physiology over time.
- Inter-organ coupling: Peripheral-organ links preserve contextual physiological influences without implementing full mechanistic models of those organs.This design is intended to retain extra-hepatic effects while keeping the framework computationally tractable.
- Component organization: The framework organizes hepatic functions into metabolism and energy homeostasis, detoxification and waste processing, and biosynthesis and production.Computational submodules collectively drive longitudinal biomarker dynamics.
- State evolution: Patient-specific parameters update internal states such as fat accumulation, glycogen storage, bile processes, and protein synthesis before biomarker generation.The resulting biomarkers correspond to clinical measurements associated with each patient’s health state.
- Validation: Validation Strategy I compares simulated biomarkers with clinical ground truth, while Strategy II uses simulated outputs for independent NASH classification.Together they assess physiological fidelity and clinical relevance.
D. Summary of Methodology
HEPATWIN translates liver physiology into a digital twin that simulates liver function and biomarker trajectories under patient-specific conditions. Its methodology combines physiological modules, longitudinal patient inputs, and biomarker estimation.
- Methodology: HEPATWIN simulates liver function under patient-specific dietary and lifestyle conditions.The framework is validated through direct biomarker comparison and downstream disease detection tasks.
- Methodology: The architecture converts key hepatic processes into computational modules and estimates longitudinal biomarker changes from baseline measurements.The methodology covers physiology-to-digital abstraction and subsequent biomarker trajectory estimation.
- Patient Inputs: Patient-specific dietary information represents macronutrients arriving at the liver, including total calories and carbohydrate, protein, and lipid composition.The model represents physiological states such as fed, fasting, caloric surplus, and caloric deficit.
- Patient Inputs: Baseline liver biomarker values are initialized from screening-visit measurements, with longitudinal assessments approximately 48 weeks apart.The described follow-up schedule includes weeks 48, 96, 144, and 192.
- Simulation Workflow: HEPATWIN integrates dietary intake, physical profile, and inter-organ communication to simulate hepatic metabolism, nutrient storage, redistribution, and detoxification over time.These patient-specific inputs drive subsequent visit-level biomarker values through internal liver modules and process variables.
- Simulation Workflow: The glucose metabolism and bilirubin conjugation pathways illustrate how physiological processes are translated into HEPATWIN’s computational framework.These are presented as detailed examples of the broader physiology-informed design.
B. Digital Abstraction of the Glucose Metabolism Pathway in HEPATWIN
HEPATWIN’s digital abstraction models glucose metabolism through computational submodules that route dietary glucose according to energy demand, storage capacity, and fat conversion. The framework also includes a physiology-informed bilirubin pathway with dynamic pools, disease-sensitive conjugation, and bile-excretion effects.
- Glucose Metabolism: The glucose metabolism module represents immediate energy expenditure, liver and muscle glycogen storage, adipose fat storage, and caloric-deficit compensation.Glucose is modeled as a subset of total carbohydrate intake, while fructose is treated separately for simplicity.
- Glucose Metabolism: Dietary glucose intake, activity level, and demographic and physiological profiles determine caloric state and total daily energy expenditure.TDEE is estimated using the Mifflin–St Jeor equation, with approximately 65-85% initially allocated to immediate energy expenditure.
- Glucose Metabolism: Liver glycogen storage is capped at approximately 120 g, while skeletal muscle glycogen storage is limited to approximately 350–400 g.After storage limits are reached, excess glucose enters fat conversion and contributes to hepatic or adipose fat accumulation.
- Glucose Metabolism: The glucose pathway figure depicts how hepatic glucose physiology is translated into computational modules within HEPATWIN.Its stated purpose is to show the digital implementation of the hepatic glucose metabolism pathway.
- Bilirubin Conjugation: The bilirubin pathway initializes total and direct bilirubin from screening measurements and models unconjugated bilirubin production as a sustained RBC-turnover influx.Separate total and direct bilirubin pools allow the unconjugated component to be inferred dynamically.
- Bilirubin Conjugation: Bilirubin conjugation efficiency is computed from UGT activity and conjugation capacity, then constrained between 0.05 and 0.90.The pathway updates unconjugated bilirubin after combining newly produced and previously retained pools.
- Bilirubin Conjugation: Elevated cholestasis retains more conjugated bilirubin in circulation, increasing direct bilirubin under impaired bile excretion conditions.This state-dependent mechanism supports clinically observed conjugated, unconjugated, and total bilirubin patterns.
D. Brief Description of Other Modules
HEPATWIN extends its physiology-informed abstraction across fructose, protein, lipid, and bile-production modules. These modules represent nutrient handling, hepatic fat dynamics, protein synthesis, and bile-flow regulation under changing functional states.
- Fructose Metabolism: Fructose is modeled as a distinct input when specified, with higher lipogenic weighting than glucose and shared downstream pathways for energy use and lipid storage.When fructose is not specified, total carbohydrate is primarily treated as glucose for modeling simplicity.
- Protein Metabolism: Protein metabolism represents amino acid availability, hepatic protein synthesis including albumin production, and energy production when carbohydrate and lipid metabolism cannot meet demand.The module captures the liver’s roles in amino acid regulation and nutrient-deficit compensation.
- Lipid Metabolism: Lipid metabolism tracks dietary and glucose-derived substrates involved in hepatic fat accumulation, storage, mobilization, and transport.The abstraction represents lipid accumulation and redistribution associated with metabolic dysfunction and fatty liver disease.
- Bile Production, Flow and Excretion: Bile production models bile-acid synthesis, conjugated-bilirubin incorporation, and biliary-flow dynamics under varying physiological states.Bile formation is driven primarily by lipid metabolism and hepatocyte functional capacity.
- Bile Production, Flow and Excretion: Reduced hepatocyte capacity and elevated inflammation increase the cholestasis index, representing impaired bile flow.The model links bile flow to hepatocyte capacity, cholestasis, and inflammation, exposing bile production and flow as interacting process markers.
E. Biomarker Trajectory Estimation
After physiological simulation, HEPATWIN updates internal liver-state variables and applies a stage-transition-driven correction to biomarker estimates. This calibration aligns mechanistic outputs with disease-stage trends while preserving baseline-referenced progression or improvement.
- Trajectory Estimation: HEPATWIN updates inflammation, hepatocyte fraction, damage, cholestasis, hepatic fat, adipose storage, and post-physiology liver-state variables before estimating biomarkers.These variables summarize cumulative metabolic, detoxification, and storage effects during the visit interval.
- Trajectory Estimation: The final biomarker estimate applies a baseline-referenced drift correction to the mechanistic output.The correction uses a stage-transition progress factor to adjust the physiology-derived value relative to baseline.
- Stage Calibration: When disease stage is unchanged, p = 1 and the final estimate equals the mechanistic output.The progress factor is defined from the pre- and post-physiology disease stages.
- Stage Calibration: When disease severity increases, p > 1 produces upward drift relative to baseline, whereas improvement with p < 1 produces a downward adjustment.This stage-transition behavior provides the direction of calibration for biomarker trajectories.
- Stage Calibration: ALT, ALP, and albumin are updated using a similar stage-transition-based calibration framework.The stated purpose is consistency between simulated hepatic physiology and population-level disease-stage transition patterns.
F. Summary of Digital Abstraction of Liver Physiology
HEPATWIN abstracts liver physiology into computational modules that generate observable biomarker outputs while preserving physiological interpretability. Its architecture encodes directional dependencies, feedback coupling, baseline values, stage constraints, and physiology-derived variables.
- Digital abstraction: HEPATWIN decomposes metabolism, detoxification, and biosynthesis into computational modules, submodules, intermediate variables, and biomarker outputs.The architecture includes glucose, protein, and lipid metabolism; bile synthesis and flow; conjugation capacity; and RBC turnover.
- Physiological interpretability: Unlike purely data-driven models, HEPATWIN explicitly encodes physiological mechanisms and causal dependencies within its computational architecture.
- Dynamic feedback: Bidirectional dependencies model feedback coupling, such as glucose intake, hepatic fat accumulation, inflammation, and later glucose metabolic efficiency.The framework distinguishes directional dependencies from feedback interactions between physiological components.
- Biomarker generation: Final biomarker estimates integrate screening-day baselines, stage-specific trend constraints, and physiology-derived process variables.This integration is intended to preserve physiological grounding while maintaining patient-specific personalization.
VI. VALIDATION OF HEPATWIN
HEPATWIN is validated with longitudinal NIDDK NAFLD data by forecasting follow-up biomarkers from earlier patient information and comparing estimates with clinical ground truth. The validation uses repeated patient-level evaluation and selected demographic, lifestyle, laboratory, and histology features.
- Validation design: The NIDDK NAFLD dataset is divided into HEPATWIN inputs and ground-truth validation data for longitudinal forecasting.For a patient with five data points, the first is used as input and the subsequent four are forecast.
- Evaluation metrics: Forecast errors are evaluated with MAE, R², Pearson correlation, Spearman correlation, and estimated-versus-ground-truth boxplots.Performance is assessed across biomarkers and follow-up visits.
- Validation design: Repeated evaluation across patients examines generalizability and robustness using error stability and clinically plausible value ranges.
- Cohort and features: The analysis focuses on dietary intake, physical activity, laboratory tests, central histology, and registration information.
- Cohort and features: Of 1410 initial patients, 825 participants with at least two visits are retained for forecasting and longitudinal modeling.Patients with only one visit are excluded; core biomarkers include ALT, ALP, bilirubin, and albumin.
B. Validation Strategy I: HEPATWIN-Estimated Biomarker Performance across Ground Truth Clinical Biomarker Values
HEPATWIN forecasts longitudinal biomarker trajectories from screening-day information and compares them with NIDDK follow-up measurements. Performance is generally clinically plausible, with stable albumin error, greater long-horizon transaminase variability, and reported bilirubin errors of 0.28 total and 0.08 direct over five steps.
- Estimation setup: HEPATWIN uses screening-day biomarkers and lifestyle information to generate biomarker estimates for future visits.Inputs include diet, physical activity, age, gender, height, and weight.
- Evaluation: Estimated values are evaluated per biomarker and follow-up horizon using MAE, R², Pearson correlation, Spearman correlation, and boxplot comparisons.
- Biomarker results: AST/ALT estimates show greater variability at extended horizons, likely reflecting inflammation-driven nonlinear fluctuations not fully captured by the current parameterization.AST may also vary because it is released from cardiac and skeletal muscle tissue, not only the liver.
- Biomarker results: MAE = 0.28 for total bilirubin and 0.08 for direct bilirubin across five-step predictions spanning approximately five years.The passage compares these longitudinal errors with prior single-time-point TwinScan bilirubin errors of 0.27–0.82 across cirrhosis stages.
- Downstream NASH detection: 65% detection accuracy is achieved for NASH using HEPATWIN-estimated time-series biomarkers generated from baseline information.The estimated trajectories extend up to five years into the future.
D. Validation Experiments Takeaway
Across its validation experiments, HEPATWIN generates clinically plausible longitudinal biomarker estimates and preserves disease-related signal for downstream NASH detection. The findings support the feasibility of physiology-informed digital twins for multi-year liver disease simulation and progression analysis.
- Validation findings: HEPATWIN generates longitudinal biomarker estimates with stable error behavior and clinically plausible ranges relative to ground truth.
- Validation findings: The simulated trajectories preserve clinically relevant disease progression patterns and support downstream NASH detection.
- Validation findings: Physiologically consistent biomarker trajectories extend up to five years using only baseline lifestyle and demographic inputs.
- Conclusion: These results validate the feasibility of physiology-informed digital twin modeling for longitudinal liver disease simulation and progression analysis.
- Conclusion: HEPATWIN integrates metabolic, detoxification, and biosynthetic processes with stage-transition-driven calibration to estimate biomarkers across disease stages.The calibration aligns simulated outputs with population-level progression patterns while preserving physiological interpretability.
- Future work: Future work will add inter-organ interactions and adaptive feedback mechanisms to improve physiological realism and personalized lifestyle guidance.