Two students reach the same diagnosis. One worked systematically — took a structured history, examined the right systems, ordered investigations in a defensible order, and narrowed a differential as evidence arrived. The other recognised a pattern from a textbook case and guessed correctly.
A written examination cannot tell them apart. Both wrote the same answer. Both score the same mark.
This is the central difficulty in medical education assessment: the thing that determines whether a graduate is safe is not what they concluded, but how they got there — and the reasoning is invisible in the output. A correct diagnosis reached by luck is indistinguishable on paper from one reached by competence, right up until the patient presents atypically and luck runs out.
What clinical reasoning actually consists of
Reasoning is not a single skill that can be scored once. It is a sequence of decisions, each of which can be done well or badly independently of the others.
Broken into its parts, a clinical encounter requires a student to:
- Elicit a history, including information the patient will not volunteer
- Decide which systems to examine, and in what order
- Interpret findings against a forming hypothesis
- Select investigations proportionate to the question
- Build and narrow a differential as evidence arrives
- Commit to a working diagnosis under time pressure
- Plan management, and communicate it
- Recognise when something is beyond their competence
A student can be excellent at examination and poor at investigation selection. Strong at pattern recognition and weak at safety-netting. These are separable competencies, and a single mark averages them into a number that hides exactly the variation a supervisor needs to see.
Why the differential matters more than the diagnosis
The most diagnostic signal in a student's reasoning is not their final answer. It is what they considered and rejected, and why.
A student who lists three plausible conditions and eliminates two on specific evidence is demonstrating reasoning. A student who names the correct condition immediately and considers nothing else may be demonstrating recall — or may have been lucky. On the answer sheet, the second looks better.
This is why assessment that captures the path is more informative than assessment that captures the endpoint.
What a written exam can and cannot see
Written examinations are efficient, standardised and defensible. They are also structurally blind to several things.
They cannot see elicitation. In a written vignette, the history is provided. The student never has to ask. But in practice, the quality of a history depends entirely on the quality of the questions — a patient does not volunteer that the pain radiates unless asked.
They cannot see sequence. Ordering a troponin before an ECG, or after, produces the same list of investigations on paper. In a real presentation the order matters.
They cannot see what was not done. Omission is often the failure that matters — the examination not performed, the red flag not asked about, the safety-net not given. An exam scores what was written, not what was missed.
They cannot see behaviour under pressure. Time pressure changes reasoning. Assessment conducted calmly measures a different capability than the one used at 3am.
None of this makes written examination worthless. It makes it partial — good at knowledge, blind to application.
What assessment that sees the path looks like
If reasoning is a sequence, assessment has to observe the sequence. In practice that means recording what a student did, in what order, and evaluating each dimension separately rather than collapsing them.
SYNTAX assesses across nine dimensions: history taking, communication, physical examination, investigation strategy, differential diagnosis, clinical reasoning, management planning, patient safety, and professionalism.
The reason for nine rather than one is the point above — these fail independently. A student with strong reasoning and weak safety behaviour is a specific, addressable problem. A student with a score of 68% is not.
The decision timeline
Recording the order in which decisions were made turns assessment into something teachable. A chronological replay — history started, symptom recognised, investigation ordered, result reviewed, diagnosis committed — lets a supervisor see not just that a student reached the right answer but when, and on what basis.
That is the artefact that makes a debrief specific. "You ordered the right investigation late, and here is where the evidence was already sufficient" is teaching. "Your score was 68" is not.
Assessment against current guidance
A defensible assessment needs a reference standard. Decisions in SYNTAX are compared against current clinical recommendations — NICE, WHO, ICMR, ACC/AHA and ESC — so that feedback explains what is recommended, why it is recommended, and where the student's reasoning diverged.
Comparing against guidance rather than an answer key matters because it distinguishes a defensible alternative from an error. Medicine frequently permits more than one reasonable path.
Truffaire's position on this
We built SYNTAX because the assessment problem and the practice problem are the same problem, and neither is solved by more content.
A medical student can spend years studying and graduate having independently managed a limited number of patients — because exposure depends on which cases happen to be admitted during a rotation. The gap is not knowledge. It is repetitions of the decision itself.
Three principles follow, and they shaped what gets assessed:
Information must be elicited, not provided. Inside an encounter the patient does not volunteer what is not asked. This is the single most important design decision, because it makes history taking assessable at all.
Software does not grade the student. Scores are computed by logic, not by a language model. The simulated patient can never reveal the diagnosis. A model writes the debrief; it does not decide the mark.
Failure has to be safe and repeatable. Students hesitate in real settings because the cost of error is a person. Removing that cost is what makes deliberate practice possible — the same case, repeated, until the reasoning is reliable.
What we do not claim
SYNTAX does not replace clinical rotations, and is not intended to. Simulation prepares a student to use real exposure better; it does not substitute for patients. Any claim otherwise would be a claim about medicine we are not entitled to make.
Where SYNTAX fits
SYNTAX is Truffaire's clinical reasoning platform — 500 MD-reviewed cases across five specialties and 43 sub-specialties, each encounter closing with a structured review across the nine dimensions above.
The related arguments are covered elsewhere: why patient simulation matters, why students need to practise failure, and what changes when the interface is voice. The broader case is in the future of clinical education in India.
SYNTAX is live here, and the systems page covers where it sits in Truffaire's work.
Frequently asked questions
Is this meant to replace OSCEs?
No. OSCEs assess performance in a standardised setting with examiners present, and that has value simulation does not replicate. Simulation adds volume and repetition — a student can attempt far more encounters than an OSCE schedule allows, which is where deliberate practice comes from.
How is a score defensible if AI is involved?
Because the score is not produced by the AI. Scoring is computed by logic against defined criteria. The language model generates the patient's behaviour and writes the narrative debrief. Separating those two things is what makes the assessment auditable.
Why nine dimensions rather than an overall mark?
Because they fail independently, and an average conceals the failure that matters. A student weak in patient safety and strong elsewhere is a specific problem with a specific remedy. One number hides that.
Can this identify students who are struggling early?
That is the main institutional use. Patterns — repeated omissions, low safety scores, consistent errors in one dimension — appear across encounters long before they appear in an examination result.
Does assessing reasoning mean there is one correct path?
No, and this is why comparison is against current guidance rather than a single answer key. Medicine permits defensible alternatives. Feedback explains what is recommended and where reasoning diverged, which is different from marking a deviation wrong.
Where to start
The distinction worth holding onto is between the answer and the path. Written assessment measures the first efficiently and cannot see the second at all. Competence lives in the second.
For an institution, the practical question is not whether to replace examinations — it is whether anything currently makes reasoning visible before graduation. If the only evidence of how a student thinks is a mark, the answer is no.
If you are evaluating this for a college or teaching hospital, get in touch.