Generalist Artificial Intelligence Models Outperform Specialized Tools in Medicine

Generalist Artificial Intelligence Models Outperform Specialized Tools in Medicine

Artificial intelligence tools designed specifically for medical use are beginning to establish themselves in clinical practice, but independent evaluation remains rare. A recent analysis compared two such tools, OpenEvidence and UpToDate Expert AI, with three cutting-edge generalist artificial intelligence models: GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6. The evaluation took place in three distinct stages.

The first stage involved submitting 500 US medical exam-style questions to the models to test their knowledge. The generalist models performed better, with scores exceeding 90% correct answers. Gemini achieved 97.4% accuracy, closely followed by GPT with 94.2%, while the specialized tools scored lower, around 88 to 89%.

The second stage assessed alignment with clinical judgments through 500 scenarios drawn from the HealthBench database. Once again, the generalist models stood out. GPT achieved the highest score with 88%, ahead of Gemini at 79.3% and Claude at 77%. The specialized tools, however, produced far less convincing results, with scores below 63%.

Finally, the third stage involved 100 real clinical queries posed by doctors in a hospital setting. Twelve American clinicians blindly evaluated the responses from the six models, producing 1,800 annotations. The generalist models dominated again, with average ratings above 3.5 out of 4. The specialized tools and Google’s analysis tool received lower scores, around 3.2. The differences were particularly notable in terms of clarity, where OpenEvidence showed significant weaknesses, with responses that were sometimes incomplete, disorganized, or poorly suited to the medical audience.

The specialized tools also had a higher refusal rate, with UpToDate declining to answer 19% of questions, compared to just 1 to 3% for the generalist models. No model produced more dangerous content or factual errors than the others, suggesting that safety is not a major differentiating factor among them.

This study highlights a surprising finding: generalist models, thanks to their extensive training and advanced alignment, outperform specialized tools in a variety of medical tasks. Their ability to reason and integrate general knowledge seems to compensate for, or even surpass, the assumed advantage of medical specialization. Dedicated tools, although designed for clinical use, struggle to compete with the versatility and precision of the most advanced models.

The results also emphasize the importance of rigorous and independent evaluation before adopting these technologies in medical settings. Traditional benchmarks, often created by the same actors developing the models, can introduce biases. The study therefore prioritized real-world queries, evaluated by clinicians without knowing the origin of the responses, to ensure an objective comparison.

Generalist models could thus represent a more effective solution for common clinical tasks, while specialized tools might find their use in highly specific areas or contexts requiring deep integration with specific hospital systems. However, their current superiority is not guaranteed to last, as the rapid evolution of these technologies could reshuffle the deck.


Sources and Credits

Source Study

DOI: https://doi.org/10.1038/s41591-026-04431-5

Title: General-purpose large language models outperform specialized clinical AI tools on medical benchmarks

Journal: Nature Medicine

Publisher: Springer Science and Business Media LLC

Authors: Krithik Vishwanath; Anton Alyakin; Mrigayu Ghosh; Ali Hage; Sean N. Neifert; Cordelia Orillac; Nataniel J. Mandelberg; Hammad A. Khan; Jin Vivian Lee; Jie J. Yao; William Robert Small; Aakaash Varma; D. Brock Hewitt; Yindalon Aphinyanaphongs; Daniel Alexander Alber; Eric Karl Oermann

Speed Reader

Ready
500