The Illusion of Expertise: When AI Sounds Authoritative Without the Clinical Context

Large language models (LLMs) can generate medical information that is fluent, coherent, and highly persuasive—even when the underlying response is incomplete or incorrect. This creates a distinctive challenge in clinical care: patients may have difficulty distinguishing an answer that sounds authoritative from one that has been appropriately evaluated in the context of their individual circumstances.
Research has documented inaccurate information, hallucinations, incomplete responses, and clinically problematic recommendations from generative AI systems.1,2,3 The concern is not simply that an AI system can be wrong. It is that a wrong or incomplete answer can be delivered with the confidence and accessibility of a good one.
This distinction matters particularly in oncology, where treatment decisions can depend on details that may be absent from a patient’s prompt: tumor biology, stage, prior therapies, comorbidities, organ function, drug interactions, performance status, and patient preferences.

Why Fluency Can Be Mistaken for Expertise
One reason AI-generated medical information can be persuasive is that LLMs can produce communication that people perceive as supportive and empathetic. In a widely cited study comparing responses to patient questions posted on a public social media forum, blinded physician evaluators rated chatbot responses as higher in both quality and empathy than physician responses.4
This does not mean that AI possesses empathy in the human sense. It demonstrates something clinically important: AI-generated communication can create a strong perception of empathy. Clear formatting, conversational language, and confident explanations can make information easier to understand and can also make it harder to recognize uncertainty or error.
For patients, fluency can become a proxy for credibility. A polished answer may feel more trustworthy than a cautious clinical explanation, even when the latter reflects a fuller assessment of the patient’s circumstances.4
The Context Gap
“Hallucination” has become shorthand for inaccurate information generated by AI. In the context of generative AI, a “hallucination” generally refers to content generated by a model that is presented as plausible or factual but is unsupported, fabricated, or inconsistent with the available information.6 It is an important problem, but focusing exclusively on hallucinations misses a larger clinical issue: an LLM can provide generally accurate medical information and still give a patient an inappropriate impression of what that information means for them.
A recent scoping review of AI chatbots in oncology identified persistent concerns regarding accuracy, hallucinations, incomplete responses, and the reliability of general-purpose chatbots as sources of cancer information.1 Research examining cancer-related generative AI has also documented hallucinations and trade-offs involved in attempts to reduce them.2 Broader reviews of ChatGPT in healthcare similarly identify concerns about factual accuracy, fabricated information, and the need for human oversight.3
The clinical risk therefore extends beyond a fabricated fact. A response can be broadly correct yet clinically incomplete because the model does not know what it has not been told. It may describe a treatment that is appropriate for a population while failing to identify the patient-specific factor that changes whether it is appropriate. The broader healthcare literature likewise emphasizes that LLM outputs require human evaluation rather than being treated as inherently reliable clinical guidance.3

Why Overreliance Matters
This problem also intersects with a well-established phenomenon in human interaction with automated systems: automation bias. People may accept an automated recommendation too readily or fail to seek contradictory information when a system appears capable and authoritative.5 As generative AI becomes more fluent and conversational, the risk is not necessarily that patients will believe everything it says. It is that they may have difficulty knowing which parts deserve skepticism.
In oncology, where uncertainty is common and decisions often involve competing risks and preferences, the more persuasive the presentation, the more important it becomes to evaluate the underlying information rather than the confidence of its delivery.
The Clinician’s Role: Contextualize, Don’t Compete
The response to AI-generated medical information should not be to create a competition between clinicians and chatbots. Nor should clinicians dismiss patients who use AI. Patients may be seeking information, reassurance, help understanding unfamiliar terminology, or more time to process a diagnosis.
Instead, AI-generated material can become a starting point for a clinically useful conversation: “What did the AI tell you? What information did you give it? What did it not know about your health or your cancer? Let’s look at which parts apply to your situation.”
This shifts the discussion from whether AI is simply “right” or “wrong” to a more useful question: What information is missing, and how does that missing context change the answer?
Patients can increasingly access enormous amounts of medical information themselves. The clinician’s value is not simply knowing more facts. It is determining which information is relevant, identifying what is missing, evaluating competing possibilities, integrating evidence with patient preferences, and recognizing when uncertainty requires additional evaluation.
As AI-generated health information becomes more sophisticated, clinicians may need to spend less time policing whether patients use these tools and more time teaching patients how to evaluate what the tools produce. The future of the exam room may include AI. It will still require clinicians who can provide what AI cannot: context, judgment, and care tailored to the person sitting in front of them.
References
- David Chen, Kate Avison, Saif Alnassar, Ryan S Huang, Srinivas Raman, Medical accuracy of artificial intelligence chatbots in oncology: a scoping review, The Oncologist, Volume 30, Issue 4, April 2025, oyaf038, https://doi.org/10.1093/oncolo/oyaf038
- Nishisako S, Higashi T, Wakao F. Reducing Hallucinations and Trade-Offs in Responses in Generative AI Chatbots for Cancer Information: Development and Evaluation Study. JMIR Cancer. 2025;11:e70176. doi:10.2196/70176.
- Draelos RL, Afreen S, Blasko B, et al. Large language models provide unsafe answers to patient-posed medical questions. npj Digit Med. 2026;9:241. doi:10.1038/s41746-026-02428-5.
- Chen D, Parsa R, Hope A, Hannon B, Mak E, Eng L, Liu FF, Fallah-Rad N, Heesters AM, Raman S. Physician and Artificial Intelligence Chatbot Responses to Cancer Questions From Social Media. JAMA Oncol. 2024;10(7):956–960. doi:10.1001/jamaoncol.2024.0836.
- Al-Anezi FM. Generative Artificial Intelligence in Healthcare: Automation Bias, Deskilling, and Cognitive Implications—A Systematic Review. J Healthc Leadersh. 2026;18:590498. doi:10.2147/JHL.S590498.
- Ji Z, Lee N, Frieske R, et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys. 2023;55(12):1–38. doi:10.1145/3571730.
Medical disclaimer: This article is intended for general educational and informational purposes for healthcare professionals and does not constitute personalized medical advice, diagnosis, or treatment recommendations.

