Skip to content
Key2MD
Marking methodology

How the AI marks your MMI and CASPer responses

A plain-language explanation of what the AI actually scores, why the criteria were chosen, and what separates a 4/5 from a 5/5.

By Dan Brittain | Updated June 2026

The short version

The AI reads your transcribed or typed response and assigns scores against the six Key2MD practice criteria, the same lens Dan uses when he marks a student himself. It is not listening for polish or filler words. It is checking whether you actually addressed what the scenario was testing: what you would do, why it matters, and whether you noticed the person it affects.

Every score comes with written feedback explaining what was present, what was absent and what a stronger version of the same response could look like. Use it to identify patterns quickly and decide what to practise next.

What the AI is not

The AI cannot reproduce the subjective reaction of a specific interviewer. It looks for evidence in the response, such as whether you noticed the person involved, explained your reasoning and reflected on uncertainty.

MMI marking

How MMI responses are marked

We assess empathy, communication and reasoning, not specialist knowledge. You do not need named policies, legislation, clinical facts, institutional procedures or memorised frameworks. Use the facts in the question, explain your thinking and recognise the limits of your role. Pushback invites you to consider another perspective: you can maintain or change your view if you explain why.

MMI feedback uses six coaching criteria, scored where the question gives them scope. Scores run from 1 to 5 per criterion per prompt. The AI marks every prompt in the station separately, then produces a per-criterion average and an overall station score.

Empathy
Does the candidate consider what people might be feeling, check rather than assume, and let that understanding shape their response?
Communication
Is the response structured? Does it speak to the person, not past them? Is it free of rehearsed script delivery?
Reasoning
Does the response name and weigh competing values? Is there a clear decision logic, or does it just list considerations?
Reflection
Does the response acknowledge what the candidate does not yet know, or what they would do differently? Does it avoid false certainty?
Real-world Awareness
Does the response account for everyday constraints in the scenario, such as trust, privacy, time and the limits of their role? No healthcare-system knowledge is required.
Problem Solving
Does the candidate understand the problem and explain a sensible way forward? A reasoned next step can be enough; exact procedures and named contacts are not required.

What a 1/5 looks like vs a 5/5

Empathy 1/5: The candidate jumps immediately to policy or procedure. The people in the scenario are treated as a problem to manage rather than people to understand. There is no pause for feelings, no acknowledgement that the situation is difficult.

Empathy 5/5: The candidate names the emotional weight of the situation in specific terms. Their understanding shapes the response. They can explore possible feelings without claiming to know them, and there is no required order or script.

Reasoning 1/5: The candidate states a conclusion with no visible logic. "I would tell the patient" with no explanation of what competing values were weighed.

Reasoning 5/5: The candidate explicitly names the tension: respecting someone's choices versus concern for their wellbeing, for example, or the colleague's wellbeing versus patient safety. They state which value they are weighing more heavily and say why, not just that they are weighing them.

Per-prompt scoring

The AI marks each follow-up question in a station as a separate unit. If you gave a strong answer to the first prompt and then failed to engage with the second, the AI will flag it. That is deliberate: a strong opening should not carry a halo over a prompt you did not really answer.

Specialist mode

Specialist mode applies a more demanding Key2MD coaching lens for applicants with clinical experience. It considers the responsibilities stated in the scenario while assessing empathy, communication and reasoning. It does not require clinical governance, policy or procedural recall. It remains a practice benchmark, not a programme-specific score.

CASPer marking

How CASPer responses are marked

Key2MD marks CASPer practice across nine competencies. Casper itself is run by Acuity Insights, not by Key2MD, and only some Australian and New Zealand programmes use it, so check the admissions page for your own intake year.

Collaboration
Working constructively with others; sharing credit; avoiding conflict escalation.
Communication
Clarity, structure, and appropriate tone for the audience described in the scenario.
Empathy
Genuine understanding of the perspective of affected parties, not just token acknowledgement.
Ethics
Application of ethical principles with appropriate nuance; avoidance of black-and-white thinking.
Fairness
Equitable treatment of all parties; awareness of bias or conflicting interests.
Motivation
Demonstration of genuine engagement with the scenario rather than performance of expected answers.
Problem Solving
Practical, realistic steps rather than generic or impossibly idealistic resolutions.
Professionalism
Appropriate handling of hierarchy, institutional constraints, and boundaries.
Service Orientation
Consistent orientation toward the needs of patients and communities over personal benefit.

Scores are given on a 1 to 9 scale. The AI also produces a short summary of the strongest element of the response and the single most useful improvement that would move the score.

How the model works

The model and its constraints

AI-assisted feedback can vary between runs. Treat small score differences cautiously and focus on themes that recur across several attempts.

The model reads the full transcript of your spoken response, not just keywords. It can identify that you named empathy in an abstract sense while giving a response that contains no concrete acknowledgement of the person in front of you. It will flag this gap.

Use the feedback to find repeated gaps in your reasoning. Treat the score as a Key2MD practice benchmark, not a prediction of how a university panel will mark you.

What the feedback tells you

Every marked response includes:

The delta view

When you attempt the same station more than once, your history page shows a criterion-by-criterion comparison between attempts. This lets you confirm that a specific improvement you worked on actually moved the score, rather than inferring improvement from the overall number alone.

Adaptive AI coach

How the coach learns your recurring mistakes

Marking one answer is useful. Marking a hundred and noticing the same mistake in most of them is what actually changes your result. The adaptive coach is an opt-in layer that reads across all of your own marked MMI and CASPer answers and identifies the mistakes that keep recurring, rather than treating each station as a fresh start.

The taxonomy is a fixed set of well-defined patterns: reasoning-based traps such as committing to a position before weighing both sides, or solving a problem before acknowledging the person in it; and structural traps such as leaving out the people a decision affects. When a pattern shows up repeatedly in your feedback, the coach does three things:

The coach reads only your own marked answers and is off unless you turn it on. It uses recurring patterns from your previous feedback to choose what to remind you about before the next station.

Common misunderstandings

What the AI does not penalise

Filler words and hesitations. "Um," "ah," and brief pauses do not reduce your score. In fact, a response with genuine pauses for thought often scores higher on Reasoning and Reflection than a response that rushes through without them.

Imperfect sentence structure. The AI marks spoken transcripts. Spoken language is not syntactically clean and the model knows this. You will not lose marks for a sentence that trails off and restarts.

Agreement with any particular ethical position. The AI does not have a preferred answer to ethical dilemmas. It marks on whether you reasoned through the dilemma, not on which side you landed.

What the AI does penalise

Skipping empathy entirely. If the scenario involves a person in distress and your response goes straight to action without acknowledgement, that is a Empathy score of 1 or 2 regardless of how good the action is.

Generic responses. A response that could be pasted under any scenario in the same category will score poorly on Communication and Real-world Awareness. The AI checks whether your response is specific to the scenario or is essentially a template.

No visible reasoning structure. Naming competing values and concluding is not enough. The response needs to show that you weighed them, meaning it should say something like "I am giving more weight to X because in this context Y is less at risk" rather than just "there is a tension between X and Y."

See it mark your answer now

Practise a station, read the AI feedback, and use the delta view to track what actually changes between attempts.

Try an MMI station CASPer practice