Source-linked AI summary
What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems
Emanuele Ratti, Lena Zuchowski
TL;DR
Opaque and uncertain ML systems raise questions about when they can be trusted. The paper interprets medicine–ML parallels as generative analogies and develops a reliabilist framework for ML, while noting open questions about construction and evaluation.
Problem
The paper addresses how to establish grounds for trusting opaque and uncertainty-laden machine-learning systems.
Method
The paper uses Hesse-inspired analysis of generative analogies to relate clinical translation’s warrants to the construction of machine-learning systems.
Results
The paper develops a reliabilist framework for machine learning from its philosophical interpretation of medicine–AI parallels.
Takeaways & Limitations
Clinical translation can provide a generative research strategy for articulating epistemic and methodological warrants in machine learning.
Takeaways & Limitations
The paper acknowledges that more work is needed on how machine-learning systems are constructed and evaluated.
Abstract
from arXiv · showhide
In the past few years, machine learning (ML) has been widely (and to an extent, successfully) implemented in medicine. However, uncertainties surrounding ML have made it difficult to establish the bases of its epistemic and methodological warrants. In the literature, a parallel has been drawn between medicine and ML, suggesting that we should model epistemic and methodological standards for ML on the standards of clinical translation. By developing tools from Hesse work, we characterise the nature of this parallel as a generative analogy between the process of clinical translation and the process of building ML systems. We identify more precisely the epistemic and methodological warrants of clinical translation that are typically only mentioned when appealing to the analogy, and we show in which sense such warrants apply analogically to the context of ML. In particular, we interpret warrants of clinical translation in reliabilist terms, and we show how this can inform a new form of ML reliabilism, which is distinct from (though compatible with) existing reliabilist accounts in philosophy of AI.
1 INTRODUCTION
The section frames ML’s opacity and uncertainty as a reliability problem and examines whether clinical translation can provide applicable epistemic and methodological standards. It develops this comparison as a generative analogy, interpreting clinical warrants reliabilistically to inform ML evaluation.
- Motivation: ML systems raise questions about reliance because their essential opacity and uncertainty complicate epistemic and methodological standards.These concerns are posed especially in medical ML, where complex systems must be evaluated despite opaque processes and unknowns.
- Motivation: Medicine sidesteps comparable unknowns through methodologies shown effective for establishing reliable results.Medical knowledge is described as fragmented, atheoretic, associationist, and opaque, with pharmaceuticals exemplifying uncertainty about mechanisms of action and theoretical justification.
- Research gap: Existing proposals draw a parallel between clinical-translation warrants and the epistemic standards used to evaluate ML, but the comparison has remained largely suggestive.Calls for randomized controlled trials in medical AI implicitly recognize this parallel, while the section identifies a lack of concrete guidance beyond a fascinating suggestion.
- Contribution: The article formalizes the comparison as a generative analogy, using Hesse’s framework to show how clinical-translation evaluation can inform ML evaluation.It identifies the analogy’s generative mechanisms as processes for establishing clinical translation’s epistemic and methodological warrants.
- Contribution: The authors interpret clinical-translation warrants in reliabilist terms and investigate their transferability into ML-compatible epistemic and methodological warrants.The resulting framework is presented as compatible with reliabilist accounts already discussed in the philosophy of AI.
2. GENERATIVE ANALOGIES
The paper develops generative analogies as a third purpose of scientific analogy: rather than predicting properties, they generate fruitful research strategies. It argues that the parallel between machine learning and medicine is best understood in this generative sense.
- Hessian analogies: Hesse’s account distinguishes horizontal relations between properties of different analogues from vertical relations between properties of the same analogue.The paper focuses on horizontal relations for its analysis.
- Generative analogies: Beyond prediction and persuasion, the paper identifies a third purpose of reasoning from horizontal relations between positive analogues: generation.This generative purpose is proposed as an important additional function of analogical reasoning in science.
- Application to ML and medicine: The machine-learning–medicine parallel is best viewed as a generative analogy because it generates a research strategy rather than predicting properties.The paper develops this interpretation in section 3 and outlines the resulting strategy in section 4.
MACHINE LEARNING
The section frames clinical translation and ML-system construction as a generative analogy: medicine supplies a research strategy for building reliable ML systems rather than specific predictions. It explains that shared associationism, atheoreticity, and opacity create ML risks, motivating reliability through intervention ensembles.
- Analogy: Clinical translation and ML-system construction are treated as analogues, with shared properties used to generate methodological guidance for ML.The analogy restricts the medical analogue to clinical translation and the ML analogue to constructing ML systems.
- Shared properties: Clinical translation is characterized by associationism, atheoreticity, and opacity, while reliability mechanisms such as RCTs manage their risks without resolving them.These properties concern noncausal associations, limited reliance on theory, and insufficient mechanistic understanding.
- Shared properties: ML exhibits the same three properties: it learns statistical patterns rather than theories or causes, and both trained models and optimization processes are opaque.These features make ML associations uncertain and can leave systems vulnerable to shortcuts, distribution shifts, and adversarial attacks.
- ML risks: Shortcuts, natural distribution shifts, and adversarial attacks show how associationism, atheoreticity, and opacity are connected to distinctive ML failure modes.Opaque systems can misclassify images through nonmedical correlations, degrade when deployment data change, and resist detection of deliberate tampering.
- Generative analogy: The analogy is generative rather than predictive because it produces a fruitful research strategy for investigating how clinical-translation mechanisms might establish ML effectiveness.The proposed strategy concerns constructing reliable intervention ensembles rather than deducing specific outcomes.
- Reliability: Reliability in clinical translation is attributed to assembling intervention ensembles, whose construction becomes the generative strategy for reliable ML.The paper distinguishes use-reliability and notes compatibility with, but distinction from, existing reliabilist accounts.
MACHINE LEARNING
The paper proposes learning ensembles (LEs) as an ML analogue of clinical intervention ensembles, using multiple dimensions and mechanisms to support ensemble- and use-reliability. This framework mitigates, but does not eliminate, ML risks and treats reliability as a holistic assessment beyond quantitative performance metrics.
- Learning ensembles: Learning ensembles adapt clinical intervention-ensemble templates to ML, aiming to increase both ensemble-reliability and use-reliability while managing risks associated with p1-p3.An LE is defined as components operating within specified conditions and boundaries that provide reasons to believe an ML system will remain reliable in a new context when those conditions apply.
- LE dimensions: LEs organize reliability-relevant information into three dimensions: boundaries of reliability, performance, and functional use.These dimensions cover construction circumstances and methodological justifications, performance metrics and their justification, and intended uses, evidence of achievability, and relations to domain knowledge.
- Functional dimension: The functional dimension concerns the ML system’s intended uses and their relation to domain knowledge, including evidence that those uses are achievable.The paper identifies this as the third LE dimension and illustrates reliability concerns such as systems being prone to false positives.
- Reliability mechanisms: Components support reliability through positive mechanisms that indicate adequate outputs and negative mechanisms that show known failures have been avoided.Negative mechanisms establish that a system is more reliable than it would be without the component, rather than proving reliability per se.
- Implications: Together, LE components increase confidence that ML outputs are adequate and known failures have been avoided, while mitigating rather than eliminating p1-p3.The authors argue that establishing adequate outputs is a holistic endeavor, not merely an examination of brute quantitative performance metrics.
5 CONCLUSION
The article interprets AI–medicine parallels as a generative analogy and develops a reliabilist framework for assessing machine-learning systems through learning ensembles. It distinguishes key dimensions and limitations of learning ensembles while positioning the account relative to computational reliabilism.
- Framework: Clinical translation is understood as assembling intervention ensembles whose ensemble- and use-reliability are assessed, motivating equivalent learning ensembles for ML systems.The analogy is interpreted as generative, using learning ensembles to assess both ensemble- and use-reliability.
- Framework: Learning ensembles have three dimensions: boundaries of reliability, performance, and functionality.
- Relation to computational reliabilism: The account treats transparency as redundant for risk management because reliability-based approaches need not solve p1-3, while remaining compatible with computational reliabilism.Both accounts reject transparency as necessary for evaluating ML systems.
- Relation to computational reliabilism: Unlike computational reliabilism’s broad algorithmic scope, this account is restricted to ML, emphasizes how ensemble components confer reliability, and distinguishes ensemble-reliability from use-reliability.Its learning-ensemble dimensions and components reflect ML specificities, and its reliability indicators overlap with but do not coincide with Duran’s.
- Limitations and contribution: The article discusses only some relevant learning-ensemble components, while robust ML construction requires trial and error, external out-of-distribution evaluation, augmentation, explainability, and continuous retraining.Despite these limitations, it advances an epistemology of learning ensembles and provides guidance for their construction and evaluation.