A clinical case review has exactly one question in it. Did the care this patient received meet the written standard, judged on this patient's own record. Not whether it was better than the case you read before lunch. Not whether it was about average for the site. The unit of assessment is one record against one standard, and the review methods that survive contact with real reviewers are built to protect that. The moment cases start being ranked against each other, a fixed standard has been swapped for a moving one, and a moving standard cannot detect anything.
01
What a case review is actually asking
There are only two things you can hold a case up against. A standard, or another case.
Only the first is stable. A standard is written down before the case exists, it does not move while you are reading, and it can be quoted back to the clinician you are asking to change something. Another case is none of those things. It varies by who was on shift, what the patient presented with, and how bad the last record you opened happened to be.
The best documented method in this area works exactly this way. The structured judgement review, used across the National Health Service for mortality case record review and described in Clinical Medicine, breaks a record into phases: admission and the first twenty four hours, ongoing care, care during a procedure, end of life or discharge care, and an overall assessment. For each phase the reviewer writes a narrative judgement in plain words, then scores that phase from one to five, very poor through to excellent. The Royal College of Physicians, which ran the national programme that spread the method, reports it being taken up across well over a hundred acute trusts.
Notice what is absent from that method. There is no field for how this case compares with the last one. No ranking, no forced distribution, no quota of poor scores per hundred records. Each phase is judged against what good care in that phase looks like, and the words come before the number.
02
Why comparison feels rigorous and is not
Reviewers reach for comparison when the standard is thin. If nobody has written down what an adequate pre-operative assessment contains, the only way to grade one is to hold it next to other pre-operative assessments. That feels like rigour. It is the opposite. It is a group of people agreeing on a house average and then measuring themselves against it, which means a network can drift a long way and never produce a bad score.
The drift runs both ways, and the second direction gets less attention. Comparison punishes cases unfairly too. Read a competent, unremarkable record straight after an exceptional one and it will feel thin, because you are still carrying the exceptional one. Same care, different verdict, decided by reading order. I have watched a reviewer mark down a perfectly sound case for no reason other than what they had open twenty minutes earlier, and nothing anywhere in the record shows that this is what happened.
Relative grading has a third cost that is easy to miss. A rank tells nobody what to fix. If a case comes sixth out of ten, there is no action in that. If a case is judged against a clause saying the consent discussion must record the alternatives offered, and this record does not, there is an action, a named owner, and something to check next month.
A case review is not a competition between patients. It is a comparison between one patient's care and one written standard.
03
The outcome is the loudest thing in the room
Once you accept that a case has to be judged on its own record, the next problem arrives immediately. The record has an ending, and the ending contaminates the judgement of everything before it.
A 2024 study in the European Journal of Emergency Medicine put numbers on this. Physicians were given case vignettes that were identical apart from how the patient ended up. In an abdominal pain case, the care was rated adequate by 44 percent of reviewers when the outcome was bad and by 88 percent when it was good. A trauma case ran at 9 percent against 43 percent. Same decisions, same information available at the time, and roughly double the criticism or worse once the reviewer knew the patient had done badly.
That is not reviewer sloppiness. It is how people read stories. Which means it has to be designed out rather than trained away. The fix is unglamorous. Give the reviewer the record up to the decision point, ask what a reasonable clinician would do with what was known at that moment, get the judgement written down, and only then reveal how it ended.
04
Agreement collapses at the point of blame
There is a related finding worth knowing before you design any review form. A 2018 study in PLOS ONE re-examined the records of patients who had died and looked at how well two committees agreed. On whether an adverse event had occurred, agreement was reasonable, a kappa of 0.66. On whether that event had been preventable, agreement was 0.03, which is to say none at all. The authors attributed it to the absence of an agreed definition of preventability.
Read that as an instruction. Reviewers can agree about facts and about specific criteria. They cannot agree about global verdicts, and a comparison between cases is a global verdict wearing a number. If your form asks whether this case was acceptable, you will collect noise. If it asks whether the operative note records the graft count, whether escalation happened inside the stated window, whether the alternatives were documented, you will collect signal, and you will collect it consistently across reviewers, sites and countries.
05
The single case check
Here is the checklist I use. It takes about a minute to set up before each review.
Name the standard first. Before you open the record, write down the clause you are assessing against. If you cannot find one, stop, because you are about to invent it after the fact.
Read it forward. Follow the record in the order the events happened, not in the order the notes were filed.
Words before numbers. Write the judgement as a sentence for each phase, in explicit language, then assign the score. A score written first will bend the sentence to fit it.
Hide the ending. Judge the decisions on what was known at the time, and reveal the outcome only once that judgement is on the page.
Score the phase, not the person. The unit is a phase of care against a standard. Nobody is being ranked.
Record good care in the same detail as poor care. This is the part everyone skips, and it is what stops a review programme turning into a hunt.
06
Where comparison actually belongs
None of this means comparison has no place. It has a very specific one, and it comes later.
Compare after the reviews, never during them. Once you hold a hundred cases each judged independently against the same standard, patterns are exactly what you want: which phase scores lowest across the network, which criterion fails most often, whether one site's consent documentation is slipping. That is aggregate analysis of independent judgements, and it is both legitimate and useful. Ranking cases against each other while you are still reviewing them is a different thing, and it corrupts the data the aggregate is built from.
So here is the split. You do not control who walks through the door, how sick they are, how a case turns out, or whether a reviewer had a bad morning. You do control what the standard says, whether it exists in writing before the case does, whether the reviewer sees the ending too early, whether the judgement is written in words a clinician can act on, and whether good care gets recorded as carefully as poor care. All of that sits inside your process. Get it right and the outcome stops being the thing you are grading, which is the only condition under which a quality review is worth running.
Questions people ask
What is the right method for reviewing a clinical case?
Judge one record against one written standard, phase by phase. Read the notes in the order the events happened, write a narrative judgement in words for each phase, then score that phase. Never grade the case relative to other cases you have reviewed. The comparison that produces learning is case against standard, not case against case.
Why should cases not be compared with each other in a quality review?
Because a relative grade moves with whatever you reviewed last. A weak case looks acceptable next to a worse one and a sound case looks thin next to an exceptional one, so the same care earns different verdicts on different days. Ranking also hides the specific failure, because a position in an order tells nobody which step went wrong or who should fix it.
How do you stop the outcome influencing a case review?
Withhold the ending until the reviewer has judged the process. Assess what was known to the clinician at each decision point, using the record as it stood at that moment, and get that judgement written down before the outcome is revealed. Published vignette studies show identical care is rated far more harshly once reviewers know the patient did badly.