By Dr. Jay Anders, Chief Medical Officer, Medicomp Systems
LinkedIn: Jay Anders MD
LinkedIn: Medicomp Systems, Inc.
Host of Tell Me Where IT Hurts – #TellMeWhereITHurts
Behind the polished demos and bold promises, a large language model remains a fast, powerful next-word predictor. It reviews an enormous matrix of possibilities and generates the word it calculates should come next. It does this so well that it can appear to reason. But it does not reason as clinicians do. It predicts.
That distinction matters because healthcare needs more than just faster answers. It needs clinical understanding.
Clinicians do not base decisions on a single data point. They weigh evidence, context, risk, patient history, and what may be missing from the record. A probabilistic model can summarize a chart, draft a note, or explain a recommendation in plain language. These are valuable uses. But when the work shifts from communication to clinical logic, the standard changes. The system must be able to connect symptoms, history, medications, tests, diagnoses, and treatments in an evidence-based, reproducible, and clinically relevant way.
To be clear: AI belongs in healthcare. The documentation burden is real, the data fragmentation problem is real, and the opportunity to make clinical information more usable is real. I believe AI will have a positive impact on healthcare over the next decade. But we need to be clear-eyed about what AI can and cannot do when deployed. Too many organizations feel pressured to move quickly, often faster than they can secure, govern, or validate these tools. That should be a warning sign.
The problem with “usually right”
Outside of medicine, if an automated trading system repeatedly gets the math wrong, it costs money. In a clinical setting, the stakes are higher. Medicine does not always operate within clean, well-defined parameters. The rule set is complex, variable, and far less forgiving.
A trading error can wipe out capital. A clinical error can harm a patient.
Sometimes the case is simple: the bone is broken, and you set it. Far more often, the patient has comorbidities, genetic factors, medication history, prior procedures, and years of clinical context that no amount of pattern inference can safely shortcut. A confident answer that is wrong is more dangerous in clinical care than silence, because it can move treatment forward even when it is inappropriate or harmful.
What LLMs are genuinely good at
This does not mean LLMs do not belong in healthcare. It means they belong in the right place.
Their core strength is translation: turning computer logic into human language and human language back into structured queries. That is why their most natural applications are communication-facing: summarizing an encounter, drafting discharge notes, generating a prior authorization rationale, helping a clinician query a large data repository in plain English, or connecting systems that were never designed to communicate with each other.
That is the right job for a probabilistic tool. It can be exceptional at the human-facing surface. It should not be treated as the clinical logic beneath it.
Where determinism is non-negotiable
Clinical decision support is the logic behind it, and it must be solid. It must be built on evidence and return the same answer every time.
Think about a drug-drug interaction. It either fires or it does not. You cannot have a system that catches the interaction one time but misses it the next. The same standard applies across clinical decision support: a lab value crossing a threshold that should trigger a diagnostic flag, a hallmark symptom that points toward or away from a diagnosis, or evidence-based criteria applied against what is actually documented in the chart.
These are not places where “maybe” is acceptable. The rule must hold every time, or it is not a rule.
This is the failure mode I see most often: an LLM dropped into the middle of a clinical stack, hoping it will simply handle everything. It will not. You end up fighting a system that may not give you the same answer twice. In healthcare, that is not a quality-assurance nuisance. It is a patient safety problem.
The architecture that actually works
The answer is not less AI. It is AI in the right architecture, with each layer doing what it is built to do.
The deterministic layer contains the clinical logic and serves as the source of truth: rule-based, evidence-based, and reproducible. The probabilistic layer communicates.
The deterministic engine triggers the drug-interaction alert; the LLM explains it clearly to the clinician. The rule-based system flags the diagnostic gap; the LLM surfaces it in a clear prior authorization narrative. The clinical-logic engine evaluates the chart; the LLM helps answer the clinician’s natural-language question about it.
Newer connector standards make this pairing easier than ever, exposing trustworthy, deterministic tools through the conversational surfaces clinicians actually want to use.
That pairing, not the LLM alone, is what makes AI useful in clinical workflows. The probabilistic model makes the system easier to use, and the deterministic layer makes it safer to trust.
Questions worth asking before you deploy
For leaders making these decisions, the practical test begins with a single question: how does the system handle evidence?
Not whether it sounds confident, but whether it verifies its outputs against what is actually documented in the record.
A system that passes that threshold earns the next question: is that verification deterministic? Will it return the same answer tomorrow as it does today? If so, it should also be able to show its work with an audit trail explaining why a given clinical assertion was accepted or flagged.
Underneath it all sits the question that matters most yet gets asked least: when the AI is wrong, what actually happens, and where does the liability lie?
Organizations that work through this chain of reasoning now will be better positioned as regulatory scrutiny tightens and as the downstream consequences of AI-generated documentation begin to appear in quality metrics, risk-adjustment audits, and patient outcomes.
AI will not be removed from clinical workflows, nor should it be. The documentation burden is real, and the efficiency gains are too. But speed without accuracy is a liability, not a benefit.
A diagnosis is either supported by the record or it is not. That determination is not probabilistic and never will be. The systems we can trust at the bedside will be those that pair clear communication with clinical logic that distinguishes what sounds right from what is supported by the record.