Human + AI: Winning Team?

In 2009, Charles Friedman articulated what he called the “Fundamental Theorem of Informatics”: a person working in partnership with an information resource is better than that same person unassisted. For decades, this has been a bedrock principle of biomedical informatics. It’s a simple idea: computers augment human capability. A calculator improves accuracy and saves time. A drug–drug interaction checker reduces medication errors. UpToDate expands a clinician’s knowledge at the point of care.
These were, and still are, powerful tools. But generative AI is different: it’s becoming a collaborator, maybe even a competitor.
As AI systems progress from decision support to autonomous reasoning, it’s reasonable to ask whether the Fundamental Theorem is still true.
The Teammate Outperforms the Team
Recent studies have shown AI outperforming clinicians in defined tasks.
For example, a 2024 Stanford-led randomised trial compared physicians with a large language model (GPT-4) on challenging diagnostic cases. The model, working alone, identified correct diagnoses roughly 92% of the time. Physicians using conventional online resources scored around 74%.
Strikingly, when physicians were given access to the AI, their performance barely changed. They reached only 76% – not significantly better than working unassisted, and still well below the AI’s performance alone.
The Stanford trial got a lot of attention but it was an early and artificial test of human–AI interaction, using short vignettes, a generic chat interface, and a small number of participants (50) with little experience prompting LLMs, in conditions very different from the clinic. This matters because how we evaluate human–AI collaboration affects what we conclude about its value. Benchmarking on written scenarios and narrow tasks isn’t enough; we need study designs that reflect real work in real clinical settings, and, most importantly that can illuminate how humans and machines might collaborate effectively.
Other studies in more realistic embedded workflows have shown a different pattern. The 2023 Swedish “MASAI” trial of screening mammography compared single radiologists using AI for triage and cancer detection support against screened two-radiologists, 40,000 patients in each group. With AI support, 6.1 cancers per 1,000 participants were detected versus 5.1 in the control group, with almost the same false-positive rate (1.5 %) and 44% less reading workload.
These findings indicate solid human-AI synergy (one radiologist plus AI matched or improved on two radiologists) but give no evidence that human radiologists will become redundant or that AI should operate without human oversight.
Human + AI Doesn’t Automatically Add Up
Why doesn’t adding a human always make things better? Three explanations stand out.
1. The Partnership Isn’t Designed Well
Most clinical AI tools, especially large language models, are not (yet) intentionally designed into real workflows. In the Stanford diagnostic trial, for example, many physicians had never used an AI chatbot before. They were given an unfamiliar interface, no training in effective prompting, and limited guidance on interpreting outputs. Not surprisingly, the combination didn’t create value.
Medicine has seen this pattern before. Early electronic prescribing tools showed little benefit until workflows, user training and interface design improved. Simply handing clinicians a new technology is unlikely to improve performance. Successful human–machine teams require co-design, training and role clarity.
2. Automation Bias and Erosion of Vigilance
A more concerning effect is that AI can sometime worsen human performance. Automation bias – over-reliance on a machine’s recommendation – has long been recognised in clinical decision support. If the AI is wrong, a clinician may anchor on the incorrect suggestion and override their own judgement.
The risk increases when clinicians don’t verify AI outputs. Modern large language models produce fluent, confident output; the confidence itself can be misleading. Without training in “AI literacy” – knowing when a model is likely to err and how to cross-check – human oversight may be unreliable.
A 2025 study tested augmentative AI tools in a safety-critical nursing task and found that when the AI was highly accurate, teams improved, but when the AI erred, the combined human–AI performance was sometimes worse. The study points out that you can’t assume human + AI is better than AI alone without empirically testing the combined system in real-world conditions.
3. Superhuman AI Leaves Little Room for Human Value-Add
Nevertheless, in some domains, AI may simply be better. The chess analogy is instructive. A decade ago, “centaur” teams (human + chess engine) could beat either alone. Today, top engines are so strong that even world champions cannot meaningfully improve their moves. Adding a human now might weaken the system.
The same pattern may emerge in well-defined clinical tasks like image triage, ECG classification, dermoscopy, or generating a differential diagnoses. If the AI has absorbed millions of examples, it can easily produce a more complete or more consistent assessment than a clinician with time to review only a few cases per week. When the AI is right, the human can only dilute accuracy. When the AI is wrong, the human must detect the error, but this is precisely where humans struggle.
The challenge is about combining strengths but also avoiding new weaknesses.
Is the Theorem Dead?
Not quite. But it needs updating.
The core idea – that humans and computers can outperform either alone – still holds. But it depends on designing systems where the contributions are genuinely complementary.
Where AI excels:
- Speed: scanning vast databases or images instantly.
- Scale: absorbing guidelines, textbooks, trial evidence, patterns from millions of cases.
- Consistency: avoiding fatigue, distraction and variability.
- Pattern recognition: detecting subtle signals in imaging or longitudinal data.
Where humans excel:
- Context: integrating social, cultural and behavioural factors.
- Ethics and trust: navigating uncertainty with patients.
- Values-based decisions: weighing risk tolerance and preferences.
- Complex reasoning: understanding atypical presentations, rare exceptions, or conflicting information.
- Relationship-building: empathy, counselling, shared decision-making.
The question is not so much whether humans or machines are “better,” but how to design a workflow where each partner augments the other.
One promising model is delegation: AI handles the first pass (triage, risk estimation, summarisation), and humans focus on ambiguity, exceptions and decisions requiring judgement. This is how the Swedish mammography trial achieved success. It is also how aviation has balanced automation and pilot oversight for decades.
But this model still requires careful planning. If junior clinicians rarely engage with straightforward diagnostic tasks because AI handles them, how will future clinicians develop expertise? If clinicians only see highly complex cases, does that accelerate burnout? We must design systems that preserve training, protect wellbeing and avoid creating a “hollowed-out” workforce.
The Real Goal: Better Patient Care
Ultimately, the measure of success isn’t technical benchmark scores but whether systems help clinicians deliver safer, more compassionate care.
If AI flags a subtle abnormality that a human might miss, the patient wins.
If AI reduces documentation load, allowing a clinician to spend more time listening, the patient wins.
If AI handles routine tasks so clinical teams can focus on high-value, human-centred work, the patient wins.
Friedman’s theorem should be seen as an aspiration. The goal is to build human–AI partnerships where the machine’s strengths augment the clinician’s and the clinician’s judgement keeps the machine safe.
Achieving this will require:
- New training focused on AI literacy, uncertainty and oversight.
- New workflows where AI is embedded into the clinical environment, not bolted on as an afterthought.
- New roles that emphasise patient relationships, complex decision-making and system stewardship.
- New safeguards against bias, hallucination, and over-reliance.
This is an evolving partnership. When computers do what they do best, clinicians have more room to do what only humans can. That, arguably, is the definition of a “winning team” – one where the patient benefits from the best of both.
Further reading
- Friedman CP. A “Fundamental Theorem” of Biomedical Informatics. J Am Med Inform Assoc. 2009;16(2):169–170.
- Goh E, Gallo R, Hom J, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw Open. 2024;7(10):e2440969. doi:10.1001/jamanetworkopen.2024.40969
- Lång K, Eriksson M, Sahlén P, et al. AI-supported mammography screening versus double reading: a prospective, population-based, paired-reader, non-inferiority study. Lancet Oncol. 2023;24(8):936–944.
- Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review. J Am Med Inform Assoc. 2012;19(1):121-127.
- Morey, D.A., Rayo, M.F. & Woods, D.D. Empirically derived evaluation requirements for responsible deployments of AI in safety-critical settings. npj Digit. Med. 8, 374 (2025).