Source-linked AI summary

FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare

Karim Lekadir, Aasa Feragen, Abdul Joseph Fofanah, Alejandro F Frangi, Alena Buyx, Anais Emelie, Andrea Lara, Antonio R Porras, An-Wen Chan, Arcadi Navarro, Ben Glocker, Benard O Botwe, Bishesh Khanal, Brigit Beger, Carol C Wu, Celia Cintas, Curtis P Langlotz, Daniel Rueckert, Deogratias Mzurikwao, Dimitrios I Fotiadis, Doszhan Zhussupov, Enzo Ferrante, Erik Meijering, Eva Weicken, Fabio A González, Folkert W Asselbergs, Fred Prior, Gabriel P Krestin, Gary Collins, Geletaw S Tegenaw, Georgios Kaissis, Gianluca Misuraca, Gianna Tsakou, Girish Dwivedi, Haridimos Kondylakis, Harsha Jayakody, Henry C Woodruf, Horst Joachim Mayer, Hugo JWL Aerts, Ian Walsh, Ioanna Chouvarda, Irène Buvat, Isabell Tributsch, Islem Rekik, James Duncan, Jayashree Kalpathy-Cramer, Jihad Zahir, Jinah Park, John Mongan, Judy W Gichoya, Julia A Schnabel, Kaisar Kushibar, Katrine Riklund, Kensaku Mori, Kostas Marias, Lameck M Amugongo, Lauren A Fromont, Lena Maier-Hein, Leonor Cerdá Alberich, Leticia Rittner, Lighton Phiri, Linda Marrakchi-Kacem, Lluís Donoso-Bach, Luis Martí-Bonmatí, M Jorge Cardoso, Maciej Bobowicz, Mahsa Shabani, Manolis Tsiknakis, Maria A Zuluaga, Maria Bielikova, Marie-Christine Fritzsche, Marina Camacho, Marius George Linguraru, Markus Wenzel, Marleen De Bruijne, Martin G Tolsgaard, Marzyeh Ghassemi, Md Ashrafuzzaman, Melanie Goisauf, Mohammad Yaqub, Mónica Cano Abadía, Mukhtar M E Mahmoud, Mustafa Elattar, Nicola Rieke, Nikolaos Papanikolaou, Noussair Lazrak, Oliver Díaz, Olivier Salvado, Oriol Pujol, Ousmane Sall, Pamela Guevara, Peter Gordebeke, Philippe Lambin, Pieta Brown, Purang Abolmaesumi, Qi Dou, Qinghua Lu, Richard Osuala, Rose Nakasi, S Kevin Zhou, Sandy Napel, Sara Colantonio, Shadi Albarqouni, Smriti Joshi, Stacy Carter, Stefan Klein, Steffen E Petersen, Susanna Aussó, Suyash Awate, Tammy Riklin Raviv, Tessa Cook, Tinashe E M Mutsvangwa, Wendy A Rogers, Wiro J Niessen, Xènia Puig-Bosch, Yi Zeng, Yunusa G Mohammed, Yves Saint James Aquino, Zohaib Salahuddin, Martijn P A Starmans

arXiv:2309.12325v3cs.CYcs.AIcs.CVcs.LG

TL;DR

Healthcare AI remains limited in real-world clinical practice amid technical, clinical, ethical and legal risks. FUTURE-AI presents an international consensus guideline developed through a 24-month process, and its recommendations achieved less than 5% disagreement among members. The framework is risk-informed, but recommendation applicability varies across medical use cases and further work is needed for regulation.

  • Problem

    Healthcare AI deployment remains limited despite research progress, while technical, clinical, ethical and legal risks complicate trustworthy real-world use.

  • Method

    The consortium developed a risk-informed framework through a 24-month international consensus process involving iterative recommendations and expert feedback.

  • Results

    Less than 5% disagreement among FUTURE-AI members accompanied approval of the recommendations and establishment of the consensus guideline.

  • Takeaways & Limitations

    FUTURE-AI provides a consensus guideline for trustworthy healthcare AI that addresses application-specific risks across the AI lifecycle.

  • Takeaways & Limitations

    The applicability of many recommendations varies across medical use cases, and further work is needed on regulating medical AI.

Abstract

from arXiv · show

Despite major advances in artificial intelligence (AI) for medicine and healthcare, the deployment and adoption of AI technologies remain limited in real-world clinical practice. In recent years, concerns have been raised about the technical, clinical, ethical and legal risks associated with medical AI. To increase real world adoption, it is essential that medical AI tools are trusted and accepted by patients, clinicians, health organisations and authorities. This work describes the FUTURE-AI guideline as the first international consensus framework for guiding the development and deployment of trustworthy AI tools in healthcare. The FUTURE-AI consortium was founded in 2021 and currently comprises 118 inter-disciplinary experts from 51 countries representing all continents, including AI scientists, clinicians, ethicists, and social scientists. Over a two-year period, the consortium defined guiding principles and best practices for trustworthy AI through an iterative process comprising an in-depth literature review, a modified Delphi survey, and online consensus meetings. The FUTURE-AI framework was established based on 6 guiding principles for trustworthy AI in healthcare, i.e. Fairness, Universality, Traceability, Usability, Robustness and Explainability. Through consensus, a set of 28 best practices were defined, addressing technical, clinical, legal and socio-ethical dimensions. The recommendations cover the entire lifecycle of medical AI, from design, development and validation to regulation, deployment, and monitoring. FUTURE-AI is a risk-informed, assumption-free guideline which provides a structured approach for constructing medical AI tools that will be trusted, deployed and adopted in real-world practice. Researchers are encouraged to take the recommendations into account in proof-of-concept stages to facilitate future translation towards clinical practice of medical AI.

2 Institució Catalana de Recerca i Estudis Avançats (ICREA), Barcelona, Spain

The section lists affiliations for contributors from universities, hospitals, research institutes and companies across multiple countries.

  • Affiliations include institutions in Europe, North America, South America, Asia, Africa and Oceania.
  • Contributors are affiliated with medical, engineering, computer science, law and research organisations.
  • The listed affiliations include universities, hospitals, government-linked research bodies and private-sector organisations.

INTRODUCTION

Healthcare AI remains limited in real-world clinical practice despite research progress and faces technical, clinical, ethical and legal risks. FUTURE-AI addresses this gap with an internationally developed framework organised around six principles and intended to cover the AI lifecycle.

  • Healthcare AI deployment and adoption remain limited in real-world clinical practice despite major advances.
  • Medical AI raises risks involving patient harm, bias, health inequalities, transparency, accountability, privacy and security.
  • Existing reporting guidelines standardise development and evaluation reporting but do not provide best practices for developing and deploying AI tools.
  • Prior proposals lacked wide international consensus and did not cover healthcare AI’s full lifecycle from design through monitoring.
  • FUTURE-AI is presented as a structured, holistic guideline established through international consensus and covering the entire AI lifecycle.
  • Its six guiding principles are Fairness, Universality, Traceability, Usability, Robustness and Explainability.

FUTURE-AI GUIDELINE

The FUTURE-AI guideline organises recommendations for trustworthy healthcare AI across evaluation, implementation, governance and lifecycle activities. It includes compliance expectations for research and deployable tools.

  • The framework summarises recommendations with proposed compliance levels for research and deployable AI tools.
  • Recommendations address risk management, documentation, quality control, auditing, logging and AI governance.
  • The guideline includes intended-use definition, human-AI oversight, user training, user-experience assessment and clinical utility and safety evaluation.
  • It recommends representative real-world data, evaluation across external datasets or multiple sites, and optimisation against real-world variation.
  • Additional recommendations cover privacy and security, AI risks, regulatory requirements, ethical issues and social and societal issues.

Fairness

The Fairness principle seeks comparable AI performance across individuals and groups, including under-represented and disadvantaged populations. It recommends identifying, measuring, mitigating and reporting bias while balancing non-discrimination with re-identification risks.

  • Fairness requires AI tools to maintain the same performance across individuals and groups, including under-represented and disadvantaged groups.
  • The framework defines three Fairness recommendations: define sources of bias, collect relevant attributes, and evaluate fairness.
  • Potential bias sources include individual attributes, medical profiles, data acquisition, labelling, curation and input-feature selection.
  • Relevant individual and dataset attributes should be systematically collected to support bias identification and address technical and human biases.
  • Attribute collection requires informed consent and ethics approval to balance non-discrimination benefits against re-identification risks.
  • Bias mitigation measures should be tested for effects on fairness and model accuracy, with remaining bias documented and reported.

Universality

Universality requires healthcare AI tools to generalise beyond controlled development settings and remain interoperable and transferable across clinical environments. FUTURE-AI defines four recommendations covering setting definition, standards, external-data evaluation, and local clinical validity.

  • Universality: Universality requires AI tools to generalise outside the controlled environment where they were built and across relevant clinical settings.Settings may differ in populations, equipment, workflows, end-users, and IT infrastructures.
  • Universality: FUTURE-AI defines four Universality recommendations: define clinical settings, use existing standards, evaluate using external data, and evaluate local clinical validity.The recommendations address transferability, interoperability, generalisability, and site-specific performance.
  • Universality: Development teams should specify intended settings and anticipate obstacles such as differing end-users, clinical definitions, equipment, and IT infrastructures.The intended settings may include primary healthcare centres, hospitals, home care, and different resource levels.
  • Universality: AI tools should use community-defined standards, including medical ontologies, data models, interface standards, annotation protocols, evaluation criteria, and technical standards.Examples include SNOMED CT, OMOP, DICOM, FHIR HL7, IEEE, and ISO standards.
  • Universality: Generalisability should be assessed with distinct external datasets and, except for single-centre tools, clinical evaluations should involve multiple sites.If generalisability is limited, transfer learning or domain adaptation should be applied and tested.
  • Universality: Local clinical validity should be evaluated on new patients, users, and applicable sites, with recalibration tested when local performance decreases.The tool should fit local clinical workflows and populations.

Usability

Usability means enabling end-users to use medical AI safely and efficiently in real-world clinical environments. FUTURE-AI addresses user requirements, human-AI oversight, training, clinical usability, and clinical utility.

  • Usability: Usability requires end-users to use AI tool functionalities and interfaces easily, efficiently, and safely in real-world clinical environments.The framework also considers digital literacy, age, ergonomics, and automation bias.
  • Usability: FUTURE-AI defines five Usability recommendations: user requirements, human-AI interactions and oversight, training, clinical usability, and clinical utility.These recommendations target adoption, reduced errors and harm, and beneficial clinical use.
  • Usability: Developers should engage clinical experts, end-users, and relevant stakeholders early to define intended use, interfaces, and human factors.Interfaces should support standardised input annotation and verification of AI inputs and results.
  • Usability: Human-in-the-loop mechanisms should perform quality checks and allow users to overrule AI predictions when necessary.Checks may flag biases, errors, or implausible explanations.
  • Usability: Accessible training materials or activities should account for diverse end-users, including specialists, nurses, technicians, citizens, and administrators.Training is intended to facilitate best usage, minimise errors and harm, and increase AI literacy.
  • Usability: Clinical usability and utility should be evaluated in representative real-world workflows, including user satisfaction, productivity, patient or clinician benefits, and safety.Comparisons should use the current standard of care; randomised clinical trials are one possible safety evaluation.

Explainability

Explainability is intended to make medical AI decisions clinically meaningful to end-users, while recognising that explanations can be difficult to evaluate and may have limitations. FUTURE-AI recommends defining explainability needs and evaluating explanations quantitatively and qualitatively.

  • Explainability: Explainability can help end-users interpret AI outputs, understand tool capacities and limitations, and intervene when necessary.Its value is described across technological, medical, ethical, legal, and patient perspectives.
  • Explainability: Explainability is a complex task requiring careful attention during development and evaluation so explanations remain clinically meaningful and beneficial.The paper identifies clinically incoherent explanations and unreasonable increases in confidence as limitations to assess.
  • Explainability: FUTURE-AI defines two Explainability recommendations: define explainability needs and evaluate explainability.Requirements should be established with representative experts and end-users.
  • Explainability: Explainability requirements should specify the explanation goal, suitable approach, and limitations such as potential end-user over-reliance.Goals may involve global model behaviour or local explanations of individual decisions.
  • Explainability: Explainable AI methods should be evaluated quantitatively for computational correctness and qualitatively for effects on user satisfaction, confidence, and clinical performance.Evaluations should also identify clinically incoherent, noise-sensitive, adversarially sensitive, or confidence-inflating explanations.

General recommendations

FUTURE-AI adds seven general recommendations applying across all trustworthy-AI principles. They address stakeholder engagement, data protection, risk mitigation, evaluation, regulation, and application-specific ethical and societal issues.

  • General recommendations: FUTURE-AI defines seven general recommendations that apply across all trustworthy-AI principles.They extend beyond system performance to broader deployment requirements.
  • General recommendations: Stakeholders should be engaged continuously to understand and anticipate needs, obstacles, and pathways toward acceptance and adoption.Suggested methods include working groups, advisory boards, interviews, co-creation meetings, and surveys.
  • General recommendations: Data privacy and security measures should operate throughout the AI lifecycle, including privacy-enhancing techniques, impact assessments, governance, and cybersecurity defences.De-identification requires balancing health benefits against re-identification risks.
  • General recommendations: AI modelling plans should address identified risks through tested measures for robustness, generalisability, and subgroup bias.Examples include data augmentation, domain adaptation, knowledge distillation, re-sampling, and equalised-odds post-processing.
  • General recommendations: Evaluation plans should define separated test data, suitable metrics, reference methods, and benchmarks against AI tools or standard practice.The plan should assess each trustworthy-AI dimension while preventing data leakage.
  • General recommendations: Development teams should identify jurisdiction- and time-dependent regulations early to anticipate obligations based on intended classification and risks.The paper gives the EU AI Act as an example of a healthcare AI regulatory context.
  • General recommendations: Application-specific ethical, social, societal, and environmental issues should be addressed with relevant experts as part of development and deployment.Issues include working conditions, power relations, skills, future interactions, and carbon footprint.

OPERATIONALISATION OF FUTURE-AI

FUTURE-AI operationalises its recommendations as step-by-step guidance embedded across the AI lifecycle. The process addresses stakeholder engagement, heterogeneity, bias, explainability, ethics, risk management, validation and deployment.

  • Lifecycle process: The framework embeds best practices chronologically across key AI lifecycle stages, from design and development through validation and deployment.The approach is presented as an agile process with practical steps and methods for operationalising recommendations.
  • Design: The design phase engages relevant stakeholders to specify clinical, technical, ethical and social requirements and identify associated risks.Stakeholders include patients, clinicians, epidemiologists, ethicists and technical personnel; the analysis produces specifications and risks to monitor.
  • Validation and deployment: Validation examines performance, robustness, fairness, generalisability and explainability, while deployment addresses local validity, training, monitoring and regulatory compliance.These stages are linked to adoption in real-world healthcare practice.
  • Design: Design guidance addresses cross-setting data heterogeneity, fairness risks, explainability requirements and application-specific ethical and societal issues.Examples include variation in equipment, protocols, operators and populations, as well as bias sources and end-user needs for explanations.
  • Implementation tools: The implementation guidance provides concrete examples including documentation, human oversight, standard ontologies, technical standards, risk prioritisation and mitigation measures.Risk management includes assessing likelihood and consequences, prioritising risks, defining mitigations and monitoring them over time.

DISCUSSION

FUTURE-AI presents an internationally developed, risk-informed framework for trustworthy and deployable medical AI, addressing the full AI lifecycle. Its recommendations are intended to support adoption across healthcare settings while recognizing regulatory and use-case-specific limitations.

  • Motivation: The framework responds to limited clinical translation amid persistent technical, clinical, socio-ethical and legal challenges.The paper links these challenges to the limited number of AI tools transitioning into clinical practice.
  • Framework and scope: 30 recommendations organized under six guiding principles cover the whole lifecycle of medical AI.The principles are Fairness, Universality, Traceability, Usability, Robustness and Explainability.
  • Limitations and future work: The framework achieved broad international consensus, but recommendation applicability varies across medical use cases and important regulatory issues remain unresolved.The paper specifically identifies unresolved liability and tensions between continuous model modification and regulations that can invalidate initial validation.
  • Risk-informed design: FUTURE-AI uses early, continuous risk management and stakeholder engagement to identify application-specific risks and tailor mitigation measures.Examples include discrimination, poor generalisability, data drift, patient harm, lack of transparency and security vulnerabilities.
  • Assumption-free implementation: The guideline is assumption-free and flexible about implementation techniques, allowing developers to select methods according to the domain, use case and data.Possible privacy measures include de-identification, federated learning, differential privacy and encryption, each with advantages and limitations.
  • Deployment emphasis: 26 of 30 recommendations were rated highly recommended for deployable tools, compared with 12 for research and proof-of-concept tools.Researchers are nevertheless advised to use as many guideline elements as possible to facilitate later transition to real-world practice.

COMPETING INTERESTS

The authors report competing interests including equity, consultancy, advisory, board, shareholding, patent, royalty, and grant relationships with multiple AI and healthcare companies.

  • Competing interests: Several authors disclose equity interests, shareholdings, consultancy, advisory, board, or founder roles in healthcare and AI companies.Disclosed relationships include Artrya, Radiomics SA, Quantib, Bunker Hill Health, and other companies.
  • Competing interests: Some authors report patents, royalties, research agreements, grants, or gifts involving medical-imaging and technology companies.The disclosures include radiomics patents licensed to Radiomics SA and institutional support from multiple companies and foundations.
  • Competing interests: Other disclosures include service on radiology society AI committees and consultancy or funding relationships with GE, Genentech, and Siloam Vision.One author also reports committee service, while another reports funding and consultancy relationships.
  • Competing interests: The authors state that all other authors declare no competing interests.

FUNDING

The work received partial support from European Union Horizon 2020 grants associated with several medical-imaging and healthcare AI projects.

  • Funding: Partial support came from the European Union’s Horizon 2020 programme under grants for EuCanImage, ProCAncer-I, CHAIMELEON, PRIMAGE, and INCISIVE.The listed grant numbers are 952103, 952159, 952172, 826494, and 952179, respectively.

Appendix Table 1 – A glossary of main terms used in the FUTURE-AI guideline (ranked alphabetically).

The appendix glossary defines core FUTURE-AI concepts, while the guideline identifies stakeholder groups that can use it to support trustworthy medical AI across clinical practice.

  • Glossary: The glossary defines AI auditing as periodic evaluation of an AI tool’s performance and working conditions to identify potential problems over time.
  • Glossary: It defines Explainability, Fairness, Traceability, Universality, and Usability as capabilities related to decisions, equal treatment, lifecycle monitoring, generalisation, and clinical fit.Usability concerns fitness for end-users in the intended clinical setting.
  • Glossary: Trustworthy AI is described as AI with proven characteristics that enable stakeholders to rely on and adopt it in real-world practice.The listed characteristics include efficacy, safety, fairness, robustness, and transparency.
  • Stakeholders: The guideline can benefit data managers, educators, funders, health organisations, healthcare professionals, IT managers, legal experts, manufacturers, public authorities, researchers, societies, social scientists, and standardisation bodies.These groups are assigned roles spanning governance, education, evaluation, deployment, monitoring, regulation, research, and standardisation.
  • Stakeholders: Public authorities and regulatory bodies are identified as users who can adapt policies and improve evaluation, certification, and monitoring procedures for medical AI.
  • Stakeholders: Researchers and developers are encouraged to investigate trustworthy-AI methods and create proof-of-concepts that transition more easily into clinical practice.
Loading 2309.12325v3…