Search arXivSearch

arXiv · 2309.12325

FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare

Karim Lekadir·Aasa Feragen·Abdul Joseph Fofanah·Alejandro F Frangi·Alena Buyx·Anais Emelie·Andrea Lara·Antonio R Porras·An-Wen Chan·Arcadi Navarro·Ben Glocker·Benard O Botwe·Bishesh Khanal·Brigit Beger·Carol C Wu·Celia Cintas·Curtis P Langlotz·Daniel Rueckert·Deogratias Mzurikwao·Dimitrios I Fotiadis·Doszhan Zhussupov·Enzo Ferrante·Erik Meijering·Eva Weicken·Fabio A González·Folkert W Asselbergs·Fred Prior·Gabriel P Krestin·Gary Collins·Geletaw S Tegenaw·Georgios Kaissis·Gianluca Misuraca·Gianna Tsakou·Girish Dwivedi·Haridimos Kondylakis·Harsha Jayakody·Henry C Woodruf·Horst Joachim Mayer·Hugo JWL Aerts·Ian Walsh·Ioanna Chouvarda·Irène Buvat·Isabell Tributsch·Islem Rekik·James Duncan·Jayashree Kalpathy-Cramer·Jihad Zahir·Jinah Park·John Mongan·Judy W Gichoya·Julia A Schnabel·Kaisar Kushibar·Katrine Riklund·Kensaku Mori·Kostas Marias·Lameck M Amugongo·Lauren A Fromont·Lena Maier-Hein·Leonor Cerdá Alberich·Leticia Rittner·Lighton Phiri·Linda Marrakchi-Kacem·Lluís Donoso-Bach·Luis Martí-Bonmatí·M Jorge Cardoso·Maciej Bobowicz·Mahsa Shabani·Manolis Tsiknakis·Maria A Zuluaga·Maria Bielikova·Marie-Christine Fritzsche·Marina Camacho·Marius George Linguraru·Markus Wenzel·Marleen De Bruijne·Martin G Tolsgaard·Marzyeh Ghassemi·Md Ashrafuzzaman·Melanie Goisauf·Mohammad Yaqub·Mónica Cano Abadía·Mukhtar M E Mahmoud·Mustafa Elattar·Nicola Rieke·Nikolaos Papanikolaou·Noussair Lazrak·Oliver Díaz·Olivier Salvado·Oriol Pujol·Ousmane Sall·Pamela Guevara·Peter Gordebeke·Philippe Lambin·Pieta Brown·Purang Abolmaesumi·Qi Dou·Qinghua Lu·Richard Osuala·Rose Nakasi·S Kevin Zhou

Abstract

Despite major advances in artificial intelligence (AI) for medicine and healthcare, the deployment and adoption of AI technologies remain limited in real-world clinical practice. In recent years, concerns have been raised about the technical, clinical, ethical and legal risks associated with medical AI. To increase real world adoption, it is essential that medical AI tools are trusted and accepted by patients, clinicians, health organisations and authorities. This work describes the FUTURE-AI guideline as the first international consensus framework for guiding the development and deployment of trustworthy AI tools in healthcare. The FUTURE-AI consortium was founded in 2021 and currently comprises 118 inter-disciplinary experts from 51 countries representing all continents, including AI scientists, clinicians, ethicists, and social scientists. Over a two-year period, the consortium defined guiding principles and best practices for trustworthy AI through an iterative process comprising an in-depth literature review, a modified Delphi survey, and online consensus meetings. The FUTURE-AI framework was established based on 6 guiding principles for trustworthy AI in healthcare, i.e. Fairness, Universality, Traceability, Usability, Robustness and Explainability. Through consensus, a set of 28 best practices were defined, addressing technical, clinical, legal and socio-ethical dimensions. The recommendations cover the entire lifecycle of medical AI, from design, development and validation to regulation, deployment, and monitoring. FUTURE-AI is a risk-informed, assumption-free guideline which provides a structured approach for constructing medical AI tools that will be trusted, deployed and adopted in real-world practice. Researchers are encouraged to take the recommendations into account in proof-of-concept stages to facilitate future translation towards clinical practice of medical AI.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Karim Lekadir, Aasa Feragen, Abdul Joseph Fofanah, Alejandro F Frangi, Alena Buyx, Anais Emelie, Andrea Lara, Antonio R Porras, An-Wen Chan, Arcadi Navarro, Ben Glocker, Benard O Botwe, Bishesh Khanal, Brigit Beger, Carol C Wu, Celia Cintas, Curtis P Langlotz, Daniel Rueckert, Deogratias Mzurikwao, Dimitrios I Fotiadis, Doszhan Zhussupov, Enzo Ferrante, Erik Meijering, Eva Weicken, Fabio A González, Folkert W Asselbergs, Fred Prior, Gabriel P Krestin, Gary Collins, Geletaw S Tegenaw, Georgios Kaissis, Gianluca Misuraca, Gianna Tsakou, Girish Dwivedi, Haridimos Kondylakis, Harsha Jayakody, Henry C Woodruf, Horst Joachim Mayer, Hugo JWL Aerts, Ian Walsh, Ioanna Chouvarda, Irène Buvat, Isabell Tributsch, Islem Rekik, James Duncan, Jayashree Kalpathy-Cramer, Jihad Zahir, Jinah Park, John Mongan, Judy W Gichoya, Julia A Schnabel, Kaisar Kushibar, Katrine Riklund, Kensaku Mori, Kostas Marias, Lameck M Amugongo, Lauren A Fromont, Lena Maier-Hein, Leonor Cerdá Alberich, Leticia Rittner, Lighton Phiri, Linda Marrakchi-Kacem, Lluís Donoso-Bach, Luis Martí-Bonmatí, M Jorge Cardoso, Maciej Bobowicz, Mahsa Shabani, Manolis Tsiknakis, Maria A Zuluaga, Maria Bielikova, Marie-Christine Fritzsche, Marina Camacho, Marius George Linguraru, Markus Wenzel, Marleen De Bruijne, Martin G Tolsgaard, Marzyeh Ghassemi, Md Ashrafuzzaman, Melanie Goisauf, Mohammad Yaqub, Mónica Cano Abadía, Mukhtar M E Mahmoud, Mustafa Elattar, Nicola Rieke, Nikolaos Papanikolaou, Noussair Lazrak, Oliver Díaz, Olivier Salvado, Oriol Pujol, Ousmane Sall, Pamela Guevara, Peter Gordebeke, Philippe Lambin, Pieta Brown, Purang Abolmaesumi, Qi Dou, Qinghua Lu, Richard Osuala, Rose Nakasi, S Kevin Zhou. 2024-07-08. FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. https://arxiv.org/abs/2309.12325

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An Investigation Into Secondary School Students' Debugging Behaviour in Python

Background and context: Debugging is a significant and often frustrating challenge for beginner programmers. Understanding students' debugging behaviours and strategies can help to identify common difficulties and inform approaches for alleviating these. Currently, there are limited studies of school students' debugging behaviour in a text-based programming language, a medium through which millions are learning to program. Objectives: In this paper, we investigate the debugging behaviour of 12-14-year-old students learning Python through a lesson-long classroom study. Method: We collected program snapshots from 73 students' attempts at a set of Python debugging exercises in an online code editor. Through qualitative content analysis of these snapshots, we developed a granular categorisation of the changes students made when debugging. Findings: A range of debugging behaviours were exhibited by students, many of which were ineffective. Students added errors through small-scale changes, reverted corrective changes, and repeatedly ran identical programs in quick succession. From the results, we identify four barriers to successful and reliable debugging for students learning a text-based programming language: fragile knowledge, a lack of systematicity and reflection, the syntax barrier, and dynamics of emotions and attitudes. Implications: This paper highlights some of the difficulties that secondary school students have when debugging in Python and the challenges of analysing program snapshots through manual inspection. We recommend that school teachers explicitly teach a systematic approach to debugging and discourage the use of ineffective debugging behaviours, and that programming environments should contain features that facilitate successful debugging.

cs.CY

Printed but not benchmarkable: most building-decarbonisation disclosure cannot be matched to the pathways that stranding regulation assumes

Cities are beginning to enforce carbon limits on existing buildings. Science-based decarbonisation pathways set those limits one asset type and one jurisdiction at a time. Owners, however, report for the whole firm. We measure what that mismatch costs on two sets of public corporate reports: a census of 502 reports from the 119 listed built-environment firms with a collected report inside a 2,246-firm panel (2003-2023), and 519 real-estate reports from 101 firms (2007-2024). BeDA, a multimodal language-model tool whose reliability we test first, read them. Running the pathway frameworks' own entry tests over published disclosure: 16.5% of census reports (43.7% of real-estate reports) print an operational carbon intensity per square metre; 6.6% (25.0%) can be matched to a pathway for their property type in a covered jurisdiction; and only 5.0% (16.4%) disclose the floor area they divided by. Of the failures at the pathway test, 82-84% follow from reports lumping the portfolio together and 16-18% from a missing curve in the pathway library. The obstacle is the reporting unit, not missing data. The rate is roughly twice as high for European as for US listings (65-71% versus 35% in listed real estate). We also show that a US portfolio's carbon verdict cannot be worked out from disclosure at all. Within one climate zone, the pathway's carbon limit varies by up to 2.79-fold with the electricity subregion, which no report names; its energy limit does not move. Extraction is checked against the source PDFs (97.3% of extracted intensities appear verbatim) and repeats on a second extractor (kappa = 0.97). Recall of the non-disclosing class was 95.1% in a blinded hand audit of 122 reports. The fix follows from the measurement: split intensity by asset type and jurisdiction, and report floor area.

cs.CY

Incipit: Axiom-Grounded Scaffolding for Human-AI Literary Creation

Large language models can produce fluent prose from short prompts, but a direct interaction gives writers little access to the assumptions that shape a long narrative. We present Incipit, an implemented research prototype that inserts an explicit planning layer between writer intent and generated prose. The layer is grounded in literary axioms, which are curated and reusable propositions about human experience and narrative craft. The prototype connects a knowledge base of 1455 axioms and 472 typed relationships to a five round direction dialogue, a retrieval and selection pipeline, and a three level blueprint covering creative premises, story beats, character arcs, and chapter outlines. Writers can inspect and edit the resulting structures before using them as context for scene generation. Additional modules support real event abstraction and five dimensional diagnostic feedback. We describe the design rationale, data flow, implementation boundaries, and a worked design example. Because no controlled user study or independently rated output study has yet been completed, we do not claim that the system improves literary quality. Instead, we outline a future preregistered comparison designed to distinguish the contribution of axiom grounding from that of hierarchical planning. The paper contributes a concrete architecture for making literary knowledge an inspectable coordination object in human AI writing.

cs.CY