Science

Penn State study finds AI models rank kidney transplant candidates differently than doctors

Chatbots fixate on single traits like alcohol use while humans weigh age and ambiguity, allocation ethics risk becoming a software default

Images

Who deserves a transplant? AI disagrees with human doctors Who deserves a transplant? AI disagrees with human doctors euronews.com

AI chatbots asked to choose who receives a kidney transplant often make different picks than human doctors, according to a Penn State University study reported by Euronews. In head-to-head hypothetical cases, the models tended to commit quickly to a single answer, even when researchers added an explicit “flip a coin” option to detect indecision.

The study set up paired scenarios in which two eligible patients competed for one available kidney, varying traits such as age, health and drinking habits. Researchers then tested large language models by changing one trait at a time, mixing several traits, and watching which detail dominated the output. Human respondents, Euronews reports, generally leaned toward younger patients over older ones—an uncomfortable but familiar priority in allocation debates. Many AI systems instead elevated lower alcohol consumption over age, and often did so by latching onto a single attribute rather than balancing several.

That mismatch matters because transplant allocation is not a quiz with a correct answer; it is a formal process designed to ration scarcity while remaining publicly defensible. When a model treats a single behavioural marker as decisive, it effectively rewrites the moral arithmetic of the system without announcing that it has done so. The same “confidence” that makes chatbots attractive in clinical workflows—clear prose, no hesitation—also makes their blind spots harder to catch in real time, especially for non-specialists who may assume a machine’s certainty reflects deeper calculation.

The paper arrives as hospitals and health systems integrate AI into scheduling, documentation, triage and decision support, often through vendor tools that sit upstream of clinician judgment. Allocation decisions already depend on structured criteria and datasets; adding a model that is trained to produce plausible-sounding answers can turn a value dispute into a software default. Euronews notes the researchers say they are not encouraging AI to replace professional judgment in high-stakes settings, but the study’s premise reflects where the market is heading: automation first, governance later.

For now, the gap the authors highlight is narrower than “AI versus doctors” and closer to “which assumptions get to be baked into the form.” In the study’s scenarios, the machines did not discover a new ethic; they amplified whichever variable their prompting and training made easiest to treat as a rule.

The experiment ends with two fictional patients and one fictional kidney. Real transplant committees still have to explain their choices to families.