In healthcare, large language models (LLMs) are used for many cognitive tasks and are good at a lot of them – but not all.
Using one kind of AI to handle every kind of task would be like staffing a hospital entirely with internal medicine doctors. They may be great physicians, but eventually someone will ask why they’re also running the pharmacy, changing wound dressings and servicing the MRI machine.
A different kind of AI
A new AI model called Jev, launched in September by TypeSafe AI, is one reason to think through this issue. Jev is unusual because unlike a standard chatbot it doesn’t write an essay, engage in a friendly chat or issue a cheery “Certainly! Here are seven things to consider.”
Instead, you give it some information and ask for a judgement. For example:
Does this patient need urgent review?
Jev: Yes – estimated probability 0.91.
or:
Which category best describes this incident?
Jev: Medication 0.62 / equipment 0.21 / communication 0.11 / other 0.06.
or:
How concerning is this report?
Jev: 8.2 out of 10.
TypeSafe calls Jev a “System One Model”: an AI designed for rapid judgement rather than extended reasoning or language generation.
Returning to the starting point of this article, we can ask:
Should one AI do everything?
One patient, several tasks
Imagine an AI-assisted online pre-surgery interview in which a 72-year-old patient says:
“I had one of those heart tubes put in about eight months ago and I take aspirin and another blood thinner. I can’t remember the name.”
First, we need to understand what the patient means. That’s the job for an LLM, which copes well with messy human language, can recognise that “heart tube” probably means a coronary stent and can suggest which medications the patient might mean.
Once the medication has been verified as clopidogrel, ideally by checking pharmacy records or the prescription, it’s not necessary for an LLM-based AI to speculate on what it is. Knowledge is needed, and that’s what an ontology is for.
An ontology is a standardised map of clinical concepts and their relationships. It can represent that clopidogrel is a P2Y12 inhibitor, that P2Y12 inhibitors are antiplatelet drugs, that coronary stents are associated with antiplatelet therapy, and so forth.
An ontology is like a slightly obsessive colleague who gets annoyed when someone treats “heart attack”, “MI” and “myocardial infarction” as three different conditions.
If a particular combination of procedure, medication or condition mandates a particular action, a language model’s creativity is not needed. We need rules:
IF X, THEN Y.
Rules are boring but in safety-critical systems boring is desirable, provided the rules are current and match local policy.
Unfortunately, healthcare is often not so tidy. Suppose the patient adds:
“I don’t get chest pain exactly. Sometimes my chest feels heavy when I walk uphill, but it goes away when I stop.”
This is territory where judgement is needed. It may be possible to write 400 rules covering “heaviness”, “pressure”, “tightness”, walking, hills, stairs, exertion and relief with rest, but along the way the rule book may start to look like the Income Tax Act.
Here is where a model like Jev could help. Instead of asking it to write an assessment, we could say:
Does this history warrant clinician review?
Jev returns a yes or no with an estimated probability for a triage judgement. The model doesn’t decide whether the patient is safe but can help determine whether clinical review is needed.
The probability is meaningful if it has been tested in the intended patient population, setting and workflow, and should be monitored after deployment.
Finally, the patient may ask:
“Why are you making such a fuss about this? I feel perfectly well.”
At this point we’re back to the LLM. What’s needed now is language, context and explanation rather than a numeric probability.

How a pre-surgery conversation moves through the AI clinical team.
Meet the AI clinical team
This points to a division of labour for healthcare AI:
- Ontologies for how clinical concepts relate
- Rules for what must happen
- System One AI models for rapid judgement
- LLMs for interpretation, reasoning or communication
Let’s not forget one last member of the team:
- Humans for clinical responsibility, expertise and managing exceptions
The five members must work together. A medication mentioned by a patient is flagged by a language model, verified against records, mapped to a standard clinical concept by an ontology, checked against a deterministic safety rule, incorporated into a rapid risk judgement and referred to a clinician if the situation is consequential or uncertain.
This looks less like a single “AI doctor” and more like a well-run clinical team.
Why split it up?
LLMs tempt us to solve every problem with an elaborate prompt. But healthcare problems aren’t all the same.
Some involve facts, some require rules, others call for pattern recognition, deliberation or clear communication. Certain things need a professional to take responsibility for decisions.
The aim is to apply the simplest validated tool that can be governed for the task.
As we build healthcare and other high-stakes AI systems, separating tasks could make them cheaper and faster as well as easier to understand, test and govern.
TypeSafe says Jev is roughly two orders of magnitude faster and cheaper than LLMs on the same tasks, so using it for quick judgements and keeping LLMs for language could cut the cost of an entire process, although these figures are company reported and have not been tested in healthcare.
When the AI gets it wrong the separation makes it easier to fix compared with trying to work out which part of a vast neural network was at fault.
If an ontology supplies the wrong relationship, attend to the knowledge stored in it. If a rule is wrong, change the rule. If a System One model is poorly calibrated, validate or retrain it. If the LLM explains something badly, improve the communication layer. If there is uncertainty, human judgement, experience and contextual knowledge is needed.
One caveat is that more components also mean more hand-offs and more ways to fail.
Jev is brand new and may or may not turn out to be important. But the idea of Jev deserves attention.

The AI clinical team in a huddle: the clinician weighs what the LLM, ontology, rules engine and Jev contribute, and the patient is part of the conversation.
Conclusion
Ever-larger LLMs taking on ever more tasks may not be the best route to better healthcare AI. A more promising one is a well-designed team of AI tools and humans, with clear roles, careful hand-offs and a validated clinical workflow.