A Nature Medicine benchmark published in June 2026 found that general-purpose chatbots — not FDA-cleared clinical AI — gave the more accurate answers to real physician questions. For a movement founded by a physician, that finding is not a curiosity. It is the whole regulatory gap in one sentence. This is the fourth article in LIWARSE’s five-part series on the AI landscape by type of use.
Purpose-Built Clinical Models: Med-Gemini and the On-Premise Alternatives
Google DeepMind’s Med-Gemini is the clearest example of a model built specifically for clinical use rather than adapted from a general chatbot. It is designed to process radiology images, pathology slides, electronic health records, lab results, and genomic data within a single system, and it has scored above 91% on standardized medical exam questions (MedQA) — a record-setting benchmark result. Critically, Med-Gemini is not sold as an off-the-shelf hospital product; it is accessed through research partnerships and Google Cloud’s Vertex AI platform, and any hospital deploying it must independently build HIPAA-compliant configuration and pursue regulatory clearance.
For institutions with strict data-governance requirements that rule out any cloud-hosted model, Meditron-3 has become the leading on-premise clinical alternative — an open-weight model a hospital’s own IT department can deploy entirely inside its firewall. Meanwhile, narrower fine-tuned models like BioMedLM 2 handle specific administrative tasks — ICD-10, CPT, and HCPCS medical coding — with measurably higher accuracy than general-purpose models asked to do the same narrow job.
The General-Purpose Models Doctors Are Actually Using
Here is the uncomfortable part of the picture. Multiple independent evaluations through 2026 — including the Nature Medicine study, and consumer tools that run several models head-to-head against PubMed and FDA drug labels on the same clinical question — have found that general-purpose frontier models (GPT-5.x, Gemini 3.1 Pro, Claude Opus) frequently outperform purpose-built, FDA-cleared clinical decision support tools on real-world physician queries. This is not an argument that clinicians should abandon regulated tools for a chatbot. It is evidence that the regulatory category “cleared medical device” and the practical category “most accurate answer” have quietly come apart, and no framework has caught up to that gap yet.
The regulatory picture is, at least, moving. The FDA’s January 2026 guidance relaxed oversight of certain digital-health and clinical-decision-support software, while its Predetermined Change Control Plan (PCCP) framework — finalized in December 2024 — lets AI/ML-based devices update post-market without a fresh submission each time, provided the sponsor pre-specifies exactly what may change and how it will be tested. The EU AI Act’s high-risk provisions, covering most medical-device AI, took effect for most systems in August 2026, with full compliance obligations following in August 2027.
Risks and Benefits Through the LIWARSE Lens
Benefits
- Multimodal clinical models like Med-Gemini offer a genuine path toward ending the “diagnostic odyssey” many rare-disease patients face, by linking clinical phenotypes to their likely genetic drivers far faster than manual review.
- On-premise options like Meditron-3 let hospitals under strict data-governance rules adopt capable clinical AI without sending patient data to any outside server — directly serving patient privacy.
- Regulatory frameworks like the FDA’s PCCP allow safety-relevant model updates to reach patients faster than a full resubmission cycle would permit, without abandoning oversight.
Risks
- The finding that unregulated general-purpose models can outperform FDA-cleared tools on real clinical queries is precisely the validation gap LIWARSE warns about: regulatory clearance is being treated, informally, as a proxy for accuracy it does not actually guarantee.
- No LLM — general-purpose or purpose-built — currently functions as a standalone FDA-cleared diagnostic system; every credible clinical use case still requires a human clinician as the accountable decision-maker, a boundary that is easy to erode under time pressure.
- Fine-tuned narrow models (BioMedLM 2 and similar) reduce error on the specific task they were built for, but that narrowness is also a limit — the same model cannot be trusted to generalize to a clinical judgment outside its training scope.
The LIWARSE Assessment
Medicine is where LIWARSE’s No Harm to Life principle meets the most immediate, individual stakes of any category in this series. The honest clinical read, from a physician’s chair, is this: the most capable model is not always the most regulated one, and the most regulated one is not always the most capable — which means the clinician in the room, not the badge on the software, remains the last and necessary safeguard. That will not change until regulation measures real-world diagnostic accuracy directly, rather than certifying a development process and assuming accuracy follows.
Under the 3 Absolute Laws, a clinical AI’s value is measured only by the patient outcome it protects — not by its benchmark score, and not by which agency’s seal it carries.
— The LIWARSE Movement | liwarse.org
Safety of Life · Advancement of Life · Together.