How the AI marks your MMI and CASPer responses
A plain-language explanation of what the AI actually scores, why the criteria were chosen, and what separates a 4/5 from a 5/5.
The short version
The AI reads your transcribed or typed response and assigns scores against the six Key2MD practice criteria, the same lens Dan uses when he marks a student himself. It is not listening for polish or filler words. It is checking whether you actually addressed what the scenario was testing: what you would do, why it matters, and whether you noticed the person it affects.
Every score comes with written feedback explaining what was present, what was absent and what a stronger version of the same response could look like. Use it to identify patterns quickly and decide what to practise next.
The AI cannot reproduce the subjective reaction of a specific interviewer. It looks for evidence in the response, such as whether you noticed the person involved, explained your reasoning and reflected on uncertainty.
How MMI responses are marked
We assess empathy, communication and reasoning, not specialist knowledge. You do not need named policies, legislation, clinical facts, institutional procedures or memorised frameworks. Use the facts in the question, explain your thinking and recognise the limits of your role. Pushback invites you to consider another perspective: you can maintain or change your view if you explain why.
MMI feedback uses six coaching criteria, scored where the question gives them scope. Scores run from 1 to 5 per criterion per prompt. The AI marks every prompt in the station separately, then produces a per-criterion average and an overall station score.
What a 1/5 looks like vs a 5/5
Empathy 1/5: The candidate jumps immediately to policy or procedure. The people in the scenario are treated as a problem to manage rather than people to understand. There is no pause for feelings, no acknowledgement that the situation is difficult.
Empathy 5/5: The candidate names the emotional weight of the situation in specific terms. Their understanding shapes the response. They can explore possible feelings without claiming to know them, and there is no required order or script.
Reasoning 1/5: The candidate states a conclusion with no visible logic. "I would tell the patient" with no explanation of what competing values were weighed.
Reasoning 5/5: The candidate explicitly names the tension: respecting someone's choices versus concern for their wellbeing, for example, or the colleague's wellbeing versus patient safety. They state which value they are weighing more heavily and say why, not just that they are weighing them.
The AI marks each follow-up question in a station as a separate unit. If you gave a strong answer to the first prompt and then failed to engage with the second, the AI will flag it. That is deliberate: a strong opening should not carry a halo over a prompt you did not really answer.
Specialist mode
Specialist mode applies a more demanding Key2MD coaching lens for applicants with clinical experience. It considers the responsibilities stated in the scenario while assessing empathy, communication and reasoning. It does not require clinical governance, policy or procedural recall. It remains a practice benchmark, not a programme-specific score.
How CASPer responses are marked
Key2MD marks CASPer practice across nine competencies. Casper itself is run by Acuity Insights, not by Key2MD, and only some Australian and New Zealand programmes use it, so check the admissions page for your own intake year.
Scores are given on a 1 to 9 scale. The AI also produces a short summary of the strongest element of the response and the single most useful improvement that would move the score.
The model and its constraints
AI-assisted feedback can vary between runs. Treat small score differences cautiously and focus on themes that recur across several attempts.
The model reads the full transcript of your spoken response, not just keywords. It can identify that you named empathy in an abstract sense while giving a response that contains no concrete acknowledgement of the person in front of you. It will flag this gap.
What the feedback tells you
Every marked response includes:
- A score for each criterion (1-5 for MMI, 1-9 for CASPer)
- An overall station or response score
- A written explanation for each score: what was present and what was absent
- A practical note on what a stronger version of the same response would contain
- For MMI premium tier: voice quality metrics including pacing, filler word frequency, and clarity
The delta view
When you attempt the same station more than once, your history page shows a criterion-by-criterion comparison between attempts. This lets you confirm that a specific improvement you worked on actually moved the score, rather than inferring improvement from the overall number alone.
How the coach learns your recurring mistakes
Marking one answer is useful. Marking a hundred and noticing the same mistake in most of them is what actually changes your result. The adaptive coach is an opt-in layer that reads across all of your own marked MMI and CASPer answers and identifies the mistakes that keep recurring, rather than treating each station as a fresh start.
The taxonomy is a fixed set of well-defined patterns: reasoning-based traps such as committing to a position before weighing both sides, or solving a problem before acknowledging the person in it; and structural traps such as leaving out the people a decision affects. When a pattern shows up repeatedly in your feedback, the coach does three things:
- It shows you the pattern and how often it has appeared, so you can see it as a habit rather than a one-off comment.
- It surfaces your most frequent trap as a short reminder immediately before your next station, while you can still act on it.
- It passes your recurring mistakes to the marker as an explicit instruction to check them, so your next piece of feedback confirms whether you actually closed the gap.
The coach reads only your own marked answers and is off unless you turn it on. It uses recurring patterns from your previous feedback to choose what to remind you about before the next station.
What the AI does not penalise
Filler words and hesitations. "Um," "ah," and brief pauses do not reduce your score. In fact, a response with genuine pauses for thought often scores higher on Reasoning and Reflection than a response that rushes through without them.
Imperfect sentence structure. The AI marks spoken transcripts. Spoken language is not syntactically clean and the model knows this. You will not lose marks for a sentence that trails off and restarts.
Agreement with any particular ethical position. The AI does not have a preferred answer to ethical dilemmas. It marks on whether you reasoned through the dilemma, not on which side you landed.
What the AI does penalise
Skipping empathy entirely. If the scenario involves a person in distress and your response goes straight to action without acknowledgement, that is a Empathy score of 1 or 2 regardless of how good the action is.
Generic responses. A response that could be pasted under any scenario in the same category will score poorly on Communication and Real-world Awareness. The AI checks whether your response is specific to the scenario or is essentially a template.
No visible reasoning structure. Naming competing values and concluding is not enough. The response needs to show that you weighed them, meaning it should say something like "I am giving more weight to X because in this context Y is less at risk" rather than just "there is a tension between X and Y."