AI
Doctors Trusted Faulty AI Even When Recovery Data Showed It Was Wrong
A 2026 study of 223 physicians found that doctors trusted faulty AI treatment recommendations even when patient recovery data showed the algorithm was wrong.
A new study of 223 physicians, published July 9, 2026 in PLOS Digital Health, found that physicians trusted faulty AI treatment recommendations even when patient recovery data contradicted the algorithm’s advice. The pattern held across two online experiments the researchers had specifically designed so the AI’s classification and the actual treatment outcomes would not align.
The physicians worked through a series of hypothetical cases for a rare disease treated with an unproven therapy. An AI system told the doctors which patients would benefit from the treatment and which would not; the doctors then saw whether each patient recovered. In both experiments, physicians rated the AI as reliable and continued to defer to its classification in the face of contradicting outcomes.
Two Experiments, 223 Doctors, One Hard Conclusion
The study, led by Aranzazu Vinas of the University of the Basque Country, ran two experiments that together enrolled 223 physicians as anonymous online participants. Each doctor walked through a series of cases for a hypothetical rare disease and an unproven treatment, with an AI patient classification system flagging which patients would benefit more or less from the therapy. After treating each case, the physician received the patient’s recovery outcome and then rated how trustworthy the AI had been.
The experimental design was deliberately rigged against the AI. In Experiment 1, the treatment was equally moderately effective for every patient, regardless of what the AI’s sensitivity score predicted. In Experiment 2, the treatment was completely ineffective for every patient; no patient got better at all. The two datasets gave doctors two separate opportunities to notice that the AI’s guidance was misaligned with the visible outcomes. In both cases, physicians routinely kept trusting the AI and continued to defer to its classification even when the data in front of them said otherwise.
- Two online experiments, both with deliberately misleading AI patient classifications.
- Experiment 1: the treatment was moderately effective for every patient, regardless of AI classification.
- Experiment 2: the treatment was completely ineffective across the board; physicians did not pick up on that.
- Both experiments: physicians rated the AI as reliable after seeing recovery outcomes.
- Both experiments: physicians did not adjust their trust in the AI despite contradicting outcomes.
How the Researchers Built a Test the AI Could Not Pass
Both experiments used an identical setup: an AI patient classification system sorted each patient into a high-sensitivity or low-sensitivity bucket, and the physician treated (or did not treat) the patient before seeing the actual recovery outcome. The doctors were told the AI’s classifications were predictions, not ground truth, leaving them free to override them. The recovery data was available throughout, giving each physician multiple chances to update their model of what the algorithm could and could not do.
In Experiment 1, recovery was moderate and identical across both sensitivity categories, and the data implied that the AI’s split added nothing useful. In Experiment 2, no patient recovered at all, and the data implied the same thing and then some: the treatment did not work. Across the two experiments, the recovery data uniformly contradicted the AI’s guidance. None of this was hidden or ambiguous; it was just outcomes, listed patient by patient. The setup, the authors wrote, gave physicians a direct test of whether they would treat the AI’s output as a starting hypothesis or as an authority.
In most cases, the doctors treated the AI as authority. The recovery outcomes, which were right there in the interface, did not get pulled into the decision-making in the way the design assumed. The result was the same in both experiments.
| Scenario Element | Experiment 1 | Experiment 2 |
|---|---|---|
| AI’s patient classification | Highly vs. lowly sensitive (the split was wrong) | Highly vs. lowly sensitive (the split was wrong) |
| Actual patient response to treatment | Moderately effective for every patient | Completely ineffective for every patient |
| What the recovery data implied | That the AI’s split added nothing | That the AI’s split added nothing and the treatment did not work at all |
| How physicians responded | Rated the AI as reliable | Rated the AI as reliable; missed that the treatment was ineffective |
Why Even Trained Doctors Could Not Overrule the Algorithm
The authors point to three documented processes that, in their reading, explain why physicians missed what was right in front of them. The first is automation bias, a default of treating the algorithmic output as a trusted answer rather than a hypothesis to test. Past healthcare studies have flagged the same pattern in clinical decision support systems. The PLOS team treats automation bias as the entry point; the others compound it once the algorithmic signal feels credible.
The second process is confirmation bias. Once a doctor accepts the AI’s patient classification, decisions follow: the patient gets the treatment if the AI says they will benefit, and skips it if the AI says they won’t. Each treated patient then becomes a chance to see recovery; each skipped patient does not. The feedback loop stops including the cases where the AI would have predicted failure, so the doctor’s experience with the system looks better than the visible recovery data would justify. The doctor is not collecting unbiased evidence; they are collecting evidence shaped by the AI’s prior classification.
The third process is what the paper calls causal illusions. When a treatment is given often and recovery happens often, both the physician and the patient can read that pairing as cause and effect, even when the underlying data shows no link. The pattern persists because the feedback loop never includes the contrasting cases that would disprove it.
In both experiments, physicians mostly trusted the AI’s classifications and had trouble learning from the feedback. Furthermore, in the second experiment, professionals did not notice that the treatment was completely ineffective.
That summary comes from lead author Aranzazu Vinas, a researcher at the University of the Basque Country in Spain, released with the open-access paper in PLOS Digital Health. Her co-authors are Fernando Blanco of the University of Granada’s Mind, Brain and Behavior Research Center (CIMCYC) and Helena Matute, a psychology professor at the University of Deusto. The paper’s editor was Po-Chih Kuo of National Tsing Hua University in Taiwan. The journal article, DOI 10.1371/journal.pdig.0001490, was published on July 9, 2026, with the underlying data openly posted on the Open Science Framework.
Earlier Work With Non-Doctors Pointed the Same Way
The pattern was already visible a year earlier, in a 2025 paper from the same lab on anonymous internet users. That earlier experiment used the same patient classification setup and found that ordinary participants kept trusting the AI even when the outcomes refused to confirm it. The point of returning to the design in 2026 with doctors was to find out whether medical training would interrupt the same loop.
It carried through. A separate 2025 study of 300 readers, the 300-reader study of AI-generated medical advice trust that the MIT Media Lab tracked, found that participants could not reliably tell AI-written medical responses from physician-written ones and reported a higher tendency to follow inaccurate AI advice than accurate doctor advice. Together, the two streams of evidence point at the same place: trained judgment and untrained judgment both defer to a confident AI, and neither reliably uses the contradicting evidence in front of them. The wider trust failure is not limited to clinicians. Related coverage at this site looks at how confidently wrong online material fools medical AI on the patient side.
The European AI Act Puts Doctors in the Loop on Purpose
Healthcare policy already assumes this is the one place where the loop will hold. The EU AI Act, passed in 2024, classifies medical AI as a high-risk domain and prohibits autonomous operation; a human actor must validate or reject each AI output. The Act treats the supervising physician as the safety layer that catches what the algorithm misses. That assumption is built into the regulation’s core.
The PLOS study calls that assumption into question. Vinas and her team note that the EU rule, and the clinical practice that follows from it, depend on doctors being able to override AI errors when they see them. The experiments suggest this override is harder than regulators have assumed, even for doctors with active training and with outcome data right in front of them. The PLOS paper leaves the policy gap to further research, but the underlying direction is the same: human oversight of healthcare AI is positioned to be harder than the Act presumes.
People tend to say that there is always a human controlling the algorithm, but our experiments show that doctors (as well as anyone else) have problems in learning from the available evidence when it contradicts the suggestions of an algorithm.
Co-author Helena Matute, a psychology professor at the University of Deusto in Spain, framed the broader takeaway that way in a statement accompanying the paper. Automation bias in clinical decision support has been documented in CDSS research, with a Bowtie analysis of automation bias in clinical decision support systems covering how alert fatigue and deference to alerts are recurring findings. What the Vinas study adds is a controlled demonstration of the problem in working physicians, with a clean before-and-after signal in the recovery data they were given.
Paths Out of Automation Bias for Working Physicians
The authors stop short of prescribing specific tools, but they call for testing protocols that train doctors to expect an AI to be wrong sometimes, and to read the visible data first. The paper suggests the field could build on this design to compare different safeguards head-to-head. Several approaches already being explored line up with the levers the PLOS team calls out.
- Training programs that build in deliberately wrong AI cases, so clinicians get practice catching errors.
- Better feedback mechanisms after each AI-assisted decision, so doctors can update on what the algorithm is getting wrong.
- Enhanced transparency in how AI systems reach their recommendations, so a low-confidence case is harder to defer to blindly.
All three approaches already exist in scattered form across the field. The PLOS paper and adjacent commentary point to them as the most promising next steps. None of them has been tested head-to-head against a control group of physicians using the same simulated setup, which is the next study the authors say the field needs. Existing research on automation bias in clinical decision support, including the Bowtie analysis from ScienceDirect, frames the same set of levers in similar terms. The goal, in the authors’ words, is to maximize the benefits of the human-AI collaboration while minimizing the errors that, on present evidence, are easier to make than the policy assumes.
Frequently Asked Questions
What did the PLOS Digital Health study find about doctors and AI?
The 2026 study ran two online experiments with 223 anonymous physicians and found that the doctors consistently trusted the AI’s patient classification even when recovery data was visibly inconsistent with it. In Experiment 2, physicians did not realize that the treatment was completely ineffective across both patient categories.
How was the experiment set up, and why did recovery data matter?
Physicians worked through cases for a rare disease treated with an experimental therapy, and an AI classified each patient as either highly or lowly sensitive to the treatment. After each decision, the doctor saw whether the patient recovered. The data gave physicians a direct opportunity to notice when the AI’s classification did not match the outcome, which the study found most of them did not.
Is this only a lab problem, or does it apply to real hospitals?
The study used online simulations, not live clinical care. The authors describe the result as flagging a potential challenge for healthcare AI deployment, particularly given that regulations such as the EU AI Act already assume doctors will override algorithmic errors in high-risk medical domains. They call for further research in clinical settings.
Why do even experienced doctors defer to AI recommendations?
The paper lists three contributing processes: automation bias, a default of trusting the algorithmic output; confirmation bias, in which selective use of the AI’s classification skews the feedback doctors collect; and causal illusions, the habit of inferring that a treatment works when prescriptions and recoveries are both frequent. Past research has documented all three in other healthcare and decision-making settings.
What could reduce automation bias in clinical practice?
The authors point to training programs and protocols that explicitly test doctors’ ability to spot AI errors, alongside stronger feedback loops after AI-assisted decisions and greater transparency from AI systems about their confidence in each prediction. They stop short of endorsing any one tool or workflow.
Disclaimer: This article reports research findings and is for informational purposes only. It does not constitute medical advice. The cited study used simulated patient scenarios, and its findings may not generalize to all clinical settings. Consult a qualified healthcare professional before making medical decisions. Figures and findings are accurate as of publication.
-
AI1 month agoFable 5 and Mythos 5 Return as US Lifts Anthropic Export Controls
-
AI2 months agoSpaceX’s Google Deal Turns a Rocket Company Into a Cloud Landlord
-
AI1 month agoOracle Cuts 21,000 Jobs in a Year, Cites AI in 10-K Filing
-
GAMING1 month agoCD Projekt Red Co-CEO: Redemption Arc Isn’t Done, Witcher 4 in 2027
-
CRYPTO2 months agoXPL Rallies 30% Ahead of Plasma One Card Tier Launch
-
APPS2 months agoDGO App Brings Rs 549 Mobile Pass for FIFA World Cup 2026 in Nepal
-
NEWS2 months agoGoogle Search Profiles Build a Follow Graph Inside Discover
-
AI2 months agoMoonshot AI Targets $30 Billion in China’s Fastest AI Funding Sprint
