The debate about artificial intelligence in medical education is usually conducted at the wrong altitude — for it or against it, as though it were one thing.
It is several things doing several jobs, and the jobs have very different risk profiles. Generating practice volume is not the same as certifying competence. Simulating a patient's answers is not the same as deciding whether a student's management was safe.
A useful position requires separating them, and the separation is not subtle once made.
Where it genuinely helps
Producing practice volume. The binding constraint in clinical education is the number of encounters a student can have under observation, and that constraint is faculty time. A system that lets a student conduct many encounters, at any hour, without consuming a supervisor, changes something real — the observation bottleneck that limits every competency-based programme.
Presenting unfamiliar cases. Practising with peers means practising with cases both parties already know. Encountering a presentation cold, without knowing the answer, is a different exercise and it is the one that resembles practice.
Immediate structured feedback. A student who receives specific feedback at the end of an encounter, rather than a mark a fortnight later, can act on it while the encounter is still in memory.
Aggregating patterns. Across many encounters, patterns emerge that no individual observation shows — including the quiet weakness in a student who is passing comfortably, discussed in what faculty can see that exam scores hide.
The common thread: all four are about volume and structure. None of them is about judgement.
Where it does not belong
Certifying competence to practise. Licensure is a judgement about whether someone may treat patients unsupervised. It carries consequences that a system cannot bear and should not be delegated one.
Replacing patient contact. A simulated encounter cannot produce the things that make clinical work difficult — a patient who is frightened, a family in the room, a presentation that does not resemble any description, the physical examination itself. Simulation prepares for those; it does not substitute for them.
Rendering verdicts in contested areas. Where guidelines genuinely disagree, presenting one answer as correct teaches false certainty. The handling is to name the divergence, as argued in guideline alignment.
Standing in for faculty relationships. A significant part of medical training is professional formation — how to conduct oneself, how to break bad news, what it looks like to be a good doctor. That transmits through people.
The claim worth being sceptical of
The strongest marketing claim in this category is some version of the system knows whether the student is right.
It is worth being precise about what such a system actually does: it compares a student's actions against an expected path defined by clinicians in advance. That is genuinely useful and it is not the same as clinical judgement. It cannot recognise a defensible alternative approach that its authors did not anticipate, and it cannot weigh the contextual factors a supervisor would.
Which means the feedback is best treated as structured, consistent and bounded — valuable precisely because it is consistent, limited precisely because it is bounded. A tool that presents it as more than that is overstating itself, and the general form of that failure is described in what "AI-powered" should mean.
The specific risks
Learning to satisfy the system. Anything scored gets optimised toward. If a system rewards a particular question sequence, students will produce that sequence — including where it is not what the patient in front of them needs.
False confidence. A student who has performed well across many simulated encounters may over-estimate their readiness for a real one. Framing matters, and the framing should be explicit rather than left to inference.
Narrowing. A case library represents the conditions someone chose to include. Students trained heavily on it develop expectations shaped by that selection, which is a reason for breadth and for saying plainly what is and is not covered.
Substitution by budget. The risk that concerns us most: an institution under resource pressure treating simulation as a reason to reduce clinical exposure. Nothing in the technology justifies that, and any vendor implying it should be distrusted.
Our position, with the interest disclosed
Truffaire builds SYNTAX — 500 cases across five specialties and 43 sub-specialties, delivered as voice-based encounters where information must be elicited rather than presented, with structured review across nine dimensions. We have a commercial interest in institutions adopting it.
Given that, the position we hold to:
It provides practice volume and structured evidence about reasoning. It does not certify competence, replace rotations, or substitute for faculty teaching. Where it says a management approach was incorrect, that judgement traces to a stated guideline reviewed by clinicians, not to a system forming a clinical opinion.
We would rather state the boundary plainly than let an institution infer something broader — because an inference of that kind, in this domain, eventually reaches a patient. The reasoning behind the voice-based design specifically is in voice-first AI in clinical training, and the assessment framework in how clinical reasoning is actually assessed.
Frequently asked questions
Will AI replace clinical rotations?
No, and an institution that reduces rotations on that basis has made an error the technology does not support.
Can it assess communication skills?
It can assess structural aspects — whether information was explained, whether the patient was oriented, whether safety-netting occurred. Empathy as experienced by a patient is not something it observes.
Is student data from these systems safe?
That depends entirely on the vendor and the contract, and it is worth settling before deployment rather than after. The questions to ask are in security and data ownership in client systems.
Should AI be used in examinations of record?
That is a regulatory and institutional judgement with invigilation and identity implications. Formative use raises far fewer questions and is where most institutions should establish value first.
How do we evaluate a vendor's claims?
Ask what it does not do, and how clinical content is reviewed and when. Vendors who answer both plainly are describing a real system; those who answer neither are describing a demonstration.
Where to start
When any platform is proposed, ask for the list of things it explicitly does not do.
A vendor who can produce that list has thought about the boundary. One who cannot is selling something whose limits your institution will discover on its own — which is a more expensive way to find them.
If you are evaluating simulation for a college or teaching hospital, get in touch.