Back to Resources
How Is Casper GradedCasper Scoring SystemCasper QuartilesCasper Z-ScoreCasper 1-9 ScaleCasper Test Results

How Is the Casper Test Graded? Raters and the 1-9 Scale

How Casper's 1-9 scale, quartiles, and z-scores actually work, what raters look for in each response, and how to avoid a red flag on test day.

By Rajani Katta, MD Published September 13, 2025 9 min read
How Is the Casper Test Graded? Raters and the 1-9 Scale

Quick answer

Trained human raters, not AI, score each of your 22 responses (11 scenarios x 2 questions) on a 1-9 scale, with most test takers landing between 4 and 7. No rater sees more than one of your scenarios, so no single impression follows you through the test. Your combined score becomes a quartile, from 1st (bottom 25%) to 4th (top 25%), and that quartile, not a raw number, is what you actually see.

Casper doesn't have right answers, which makes "how is it graded" the question almost every applicant asks right before test day. The short version: real people read what you wrote or watch what you recorded, score it on a simple scale, and your combined performance gets converted into a quartile. No algorithm decides whether your ethics are good enough. A trained human does, twenty-two separate times.

This guide walks through exactly how that process works, what raters are trained to look for, how quartiles get calculated, and what actually moves your score.

How Many Responses Get Scored

Casper's current format has 11 scenarios, each with 2 questions, for 22 total responses across the test. Every single one is graded independently by a different rater, so no single person's impression of you follows you through the whole exam. A rough answer on scenario 4 has no bearing on how scenario 5 gets scored.

Responses are either typed or video-recorded depending on the section. Raters evaluate the same core qualities either way: empathy, ethical reasoning, communication, and professionalism, using the 10 competencies Casper measures as their rubric.

The 1-9 Scoring Scale

Each response gets a score from 1 to 9. Here's what separates the bands:

Score Band What it typically looks like
8-9 Excellent Sophisticated analysis, specific and practical solutions, strong command of the competencies
5-7 Solid to strong Reasonable reasoning, workable solutions, room for more depth or specificity
3-4 Fair Surface-level competency, vague or generic solutions
1-2 Poor Minimal reasoning, impractical or concerning solutions

Most test takers land between 4 and 7. Scores of 8 or 9 are genuinely rare, and scores of 1 to 2 usually signal a real problem with the response, not just a weak one. Your responses are also converted into a z-score, which tells schools how far your performance sits from the average, in standard deviations, rather than just a raw number.

Who Actually Reads Your Answers

Raters come from a mix of professional backgrounds: healthcare, education, human resources, and current medical students among them. They go through structured training and regular calibration sessions so that a 6 from one rater means roughly the same thing as a 6 from another. Each rater is blinded to your identity and to your other answers, grades a single response, and never sees more than one scenario from the same applicant. That structure is deliberate: it's what keeps one bad five minutes from sinking your entire test.

What Raters Are Actually Looking For

Beyond matching your answer to a competency, raters weigh a few things consistently:

  • Depth of analysis. Do you go past the obvious problem and consider other perspectives, consequences, and stakeholders, or stop at the first thing you noticed?
  • Practical solutions. Is your proposed action specific and realistic, or a vague placeholder like "I would gather more information" with nothing after it?
  • Communication quality. Is the answer organized and easy to follow, even under a tight word count?
  • Range across competencies. Strong answers often touch more than one competency naturally, without forcing it.

Reading the commentary in the sample questions and answers, including the annotated weak answers, is the fastest way to see these standards applied to real answers instead of just described in the abstract.

How Quartiles Get Calculated

Your 22 individual scores get combined into one overall result, then compared against everyone else who took the same test type around your test date. That comparison produces your quartile:

Quartile Percentile range What it means
4th 75th-100th Top 25% of test takers
3rd 50th-74th Above average
2nd 25th-49th Below average
1st 0-24th Bottom 25%

You only ever see your quartile. Schools receive your exact standardized score and percentile, which gives them more resolution than you get yourself. That gap is intentional; Acuity's stated goal is to give applicants a general signal without letting anyone reverse-engineer the scoring model from their own result. What a given quartile actually means for your application, school by school, is covered in our Casper quartiles guide

Red Flags: A Separate Risk From a Low Score

Independent of your numerical score, a response can be flagged if it demonstrates something raters consider genuinely concerning: unsafe advice, unethical reasoning, discriminatory attitudes, or unprofessional conduct. A red flag gets communicated to your schools and can hurt your application regardless of how the rest of your test scored.

Avoiding one is mostly about instinct rather than technique: prioritize safety and professionalism, avoid extreme or dismissive positions, and default to the conservative, ethical option when a scenario is ambiguous. Almost nobody gets flagged by accident while genuinely trying to act ethically. It tends to happen when applicants try to sound clever or decisive at the expense of sound judgment.

How to Actually Raise Your Score

Two response frameworks show up often in strong Casper answers:

  • CARE, for situational questions: Clarify the problem, Ask what information is missing, Recognize other perspectives, Explore ethical solutions.
  • ARC, for personal questions: Anecdote, Reflection, Connection to your future profession.

Beyond structure, the habits that separate 5s from 7s and 8s are consistent:

  1. Get specific. Replace "I would talk to them about it" with what you'd actually say and do first.
  2. Show more than one competency. A single strong answer often demonstrates empathy and problem-solving together, not just one in isolation.
  3. Use your full time. A thorough answer to fewer points beats a rushed answer that tries to touch everything.
  4. Watch for common pitfalls: assuming facts the scenario didn't give you, offering a generic response that could fit any scenario, or proposing something unrealistic or unethical to sound decisive.

The question-type guide breaks down how these frameworks apply to situational, personal, and policy questions specifically, and StudyCasper's timed practice with AI feedback will flag exactly which of these habits is costing you points.

Frequently Asked Questions

Is Casper graded by AI?

No. Every response is scored by a trained human rater. AI grading tools (including StudyCasper's own practice feedback) exist to help you prepare, but they aren't part of how your actual test gets scored.

How many people grade my Casper test?

Twenty-two, one per response. Each response is routed to a different rater, and no rater sees more than one scenario from the same test taker, so no single person's impression carries across your test.

Do typos or grammar mistakes hurt my Casper score?

No. Raters are explicitly instructed to ignore spelling, grammar, and syntax and to focus on the substance of your reasoning. Bullet points are officially acceptable in typed responses.

Can I see my raw Casper score?

No. Applicants only receive a quartile. Schools receive your full standardized score and percentile, which is a meaningfully more detailed result than what you get to see.

What's a good Casper score?

Generally, 3rd or 4th quartile is considered strong, though how much that matters depends heavily on which schools are on your list. The full breakdown is in What Is a Good Casper Score?

The Bottom Line

Casper scoring is more structured, and more human, than it feels from the test taker's side. Twenty-two independent raters, a simple 1-9 scale, and a quartile system built to compare you fairly against people who tested around the same time. You can't game it, but you can prepare for it: learn what raters reward, practice under real time pressure, and give specific, well-reasoned answers instead of safe, vague ones. Start with the Casper study plan, drill with the practice scenarios, and check your test date against the current Casper calendar so your results reach schools in time. If you want to know which of the four rating standards you're actually losing points on, the AI-scored prep course grades your responses against them one by one.