Accurate diagnosis is one of the foundations of high-quality healthcare. Every clinical decision, from selecting an appropriate treatment to predicting patient outcomes, depends on identifying the right condition at the right time. However diagnostic uncertainty and error remain persistent challenges, contributing to delayed treatment, unnecessary procedures, increased costs, and preventable patient harm.
Artificial intelligence is beginning to change how clinicians approach this challenge. By analyzing complex clinical information and identifying patterns that may be difficult to detect consistently, AI systems can support diagnostic reasoning across a growing range of specialties. What was once primarily a research field is rapidly becoming one of the most consequential applications of AI in healthcare.
The current evidence is promising, but it also requires careful interpretation. AI models have achieved high levels of diagnostic accuracy in selected tasks, while performance remains highly variable across specialties, models, and study designs. The most useful question is therefore not whether AI can diagnose disease independently, but how it can be used safely to strengthen clinical decision-making.
Key Findings
What the diagnostic evidence shows
Studies in which an LLM achieved higher accuracy than clinicians
Highest primary diagnostic accuracy reported in a selected clinical task
Clinical cases evaluated across 30 studies and 19 language models
High Performance, but Wide Variation
A 2025 systematic review and meta-analysis published in JMIR Medical Informatics evaluated 30 studies involving 19 large language models and 4,762 clinical cases across multiple medical specialties1.
The highest-performing models achieved primary diagnostic accuracy ranging from 25% to 97.8%, while reported triage accuracy ranged from 66.5% to 98%1. This breadth is important: it demonstrates that AI can perform exceptionally well in selected applications, but also that diagnostic performance is not consistent across every clinical context.
Several factors can explain this variation. Different studies evaluated different specialties, model versions, case formats, prompts, and outcome measures. Some models received structured case descriptions, while others were evaluated using patient records or published clinical cases. As a result, a headline accuracy figure cannot be generalized automatically to every organization, specialty, or patient population.
How AI Compares With Clinicians
The comparison between AI and healthcare professionals was mixed. In 10 of the 30 included studies, ChatGPT-based models achieved higher diagnostic accuracy than the human comparison group. Several studies also reported performance approaching or matching clinicians in specific diagnostic tasks, particularly within ophthalmology and selected general medicine applications1.
However, these individual findings should be interpreted alongside the pooled analysis. When the authors combined 18 studies that used primary diagnostic accuracy as a common outcome, clinical professionals generally outperformed the language models1.
This distinction matters. The evidence supports the diagnostic capabilities of AI, but it does not support replacing clinicians with autonomous general-purpose models. Instead, the findings suggest a more practical opportunity: using AI as a clinical assistant that expands the information available to professionals while preserving human judgment and accountability.
Evidence Summary
A promising but still evolving evidence base
The study demonstrates meaningful diagnostic capabilities while also identifying substantial variation and methodological limitations.
| Outcome | Observed result | Interpretation |
|---|---|---|
| Primary diagnostic accuracy | 25–97.8% | |
| Triage accuracy | 66.5–98% | |
| LLMs outperforming clinicians | 10 of 30 studies | |
| Pooled diagnostic comparison | Favored clinicians | |
| Studies with high risk of bias | 20 of 30 studies |
The reported figures describe the studies included in the review and should not be interpreted as guaranteed performance in routine clinical practice.
Why Validation Matters
The quality assessment found that 20 of the 30 included studies had a high risk of bias1. Many were retrospective, relied on relatively small test sets, or used previously documented cases rather than prospective real-world clinical encounters.
All of the included studies evaluated models through data testing rather than using them for real-time diagnosis of patients. The authors also identified significant heterogeneity between studies, including differences in clinical departments, diagnostic methods, prompts, input formats, and evaluation measures1.
These limitations do not eliminate the value of the findings. They clarify what healthcare organizations must do before implementation: validate model performance using representative local data, define appropriate human oversight, monitor errors, and establish clear boundaries for clinical use.
The Opportunity for Healthcare Organizations
The most immediate opportunity is not autonomous diagnosis. It is the use of AI to augment clinical expertise, reduce diagnostic uncertainty, and help professionals consider relevant possibilities more consistently.
Depending on the use case, AI systems may help organize complex clinical information, generate differential diagnoses, identify potentially overlooked patterns, prioritize high-risk cases, or provide a second layer of review. When appropriately designed, these tools can support more informed decisions without removing clinicians from the diagnostic process.
The key is to begin with a specific clinical problem and a measurable objective. Organizations should determine what type of diagnostic error or uncertainty they are trying to reduce, which information the model will receive, how its recommendations will be presented, and where responsibility for the final decision will remain.
Conclusion
Large language models have demonstrated meaningful diagnostic capabilities across a diverse range of clinical applications. In selected tasks, they achieved very high levels of accuracy and sometimes outperformed the clinicians included in individual studies.
At the same time, performance varied widely, the pooled evidence continued to favor clinical professionals, and many studies presented a substantial risk of bias. The current evidence therefore supports AI-assisted diagnosis with careful clinical oversight, not unsupervised replacement of professional judgment.
At Argenticare, we help healthcare organizations design and implement practical data and AI solutions that deliver measurable clinical and operational value. From clinical decision support and analytics to workflow automation and AI assistants, we combine medical and technical expertise to turn promising ideas into safe, effective solutions.
Ready to Explore AI for Your Organization?
Explore how Argenticare can support your goals with AI, data, and automation.