Updated on
August 29, 2026
Checking for Understanding: What the Evidence Really Says
The 0.4 to 0.7 effect size for formative assessment does not survive scrutiny, and over a third of feedback interventions make performance worse.

What is checking for understanding?
Checking for understanding is eliciting evidence of what learners currently know during the lesson, and then changing what you do next in response. It is the fastest cycle of formative assessment, running minute by minute rather than lesson by lesson. On Black and Wiliam’s own definition the evidence has to be used, so a check that does not change the teaching is not formative assessment at all.
The number everybody quotes for formative assessment is wrong, and the person most associated with it has said so himself, in print, without anybody much noticing.
Checking for understanding is what a teacher does inside a lesson to find out what learners currently know, so that the next few minutes can change. It is the fastest cycle of formative assessment: elicit evidence, read it, act on it, all within the lesson rather than after it.
The famous claim is that formative assessment is worth an effect size of 0.4 to 0.7. That range comes from Black and Wiliam's (1998) review, which was a narrative review of roughly 250 studies rather than a meta-analysis, and the range was a summary of what different studies reported rather than a pooled figure. When Kingston and Nash (2011) went looking for studies solid enough to pool, they screened more than 300 and found 13 usable. The median effect was 0.25 and the weighted mean was 0.20. Briggs and colleagues (2012) then criticised that meta-analysis in turn, concluding that considerable uncertainty remains.
If you want the checks themselves, skip to the 34, grouped by what each one actually tells you.
So the honest position is not "0.20 replaces 0.7". It is that nobody currently knows the size of the effect, because the research base is too thin to say. That is a much better thing for a teacher to know than a number, because it changes what you do: you check for understanding because a lesson without it is flying blind, not because a meta-analysis promised you seven months.
Checking for understanding is eliciting evidence of what learners know at the point of teaching, and then changing what you do next in response. It runs on the shortest cycle in the classroom, minute by minute rather than lesson by lesson. Doug Lemov gives it a whole chapter in Teach Like a Champion 3.0, and the moves in it are all versions of the same question: who actually has this, and how would I know?
Mostly it is branding, with one distinction that survives. The difference is cycle length, not kind.
Formative assessment is the whole system: eliciting evidence, interpreting it, and acting on it. Checking for understanding is that system's first step run at speed, inside the lesson. There is no study comparing them and no evidence that they have different effects, so anyone claiming a difference in size is asserting a taxonomy rather than reporting a finding.
The distinction collapses entirely in one situation, and it is the common one. A check that elicits evidence and does not change what the teacher does next fails Black and Wiliam's own definition of formative. The whole thing turns on the second step. Our guide to formative assessment strategies covers the longer cycles.
The direction of the evidence is good and the magnitude is genuinely unknown. Three findings matter more to a teacher than any effect size: the headline number is not defensible, feedback frequently backfires, and the reputable sources disagree with each other in ways nobody hides.

Dylan Wiliam set the competing estimates side by side himself in 2011, which is more honest than most of the people citing him. Kluger and DeNisi (1996) found an average of 0.41 for feedback interventions. Black and Wiliam (1998) estimated 0.4 to 0.7. Hattie and Timperley (2007) proposed around 0.95. Wiliam's own classroom field trial, run over a year with ordinary teachers and measured on externally mandated standardised tests, produced 0.32.
His explanation is worth having, because it is not "somebody was wrong". The estimates diverge because outcome measures differ in how sensitive they are to instruction, and because the populations differ in variance. A test built around the thing you taught will always show a bigger effect than a national exam.
The Education Endowment Foundation currently rates feedback at +6 months additional progress, very low cost, on high-strength evidence from 155 studies. That is the most useful public summary for a school, and it is a different question from "what is the effect size of formative assessment", which is the question that has no reliable answer.
Kluger and DeNisi's (1996) meta-analysis is the largest of feedback interventions: 131 usable papers, 607 effect sizes, 23,663 observations. The weighted mean was 0.41, which is the number people quote.
The number people do not quote sits in the same results section. Over 38 per cent of the effects were negative. More than a third of feedback interventions made performance worse. The authors checked whether that was driven by one prolific researcher's unusually negative studies; excluding them, it was still 33 per cent. They conclude the negative effects are robust rather than an artefact.
That is the single most useful fact in this article. Feedback is not a safe operation with a variable dose. It has a real failure mode, it fires often, and the general shape of the failure is feedback that directs attention to the person rather than the task.
Hinge questions, the diagnostic multiple-choice item placed at the pivot of a lesson, are treated in British CPD as an evidence-based technique. Searching the largest education research database by title returns one record, and it is a practitioner magazine article by Wiliam, who popularised the term.
What is genuinely evidenced is the adjacent practice under a different name. Peer Instruction in undergraduate physics and biology uses exactly this move, and it has been tested with control conditions. Smith and colleagues (2009), in Science, found that peer discussion improved performance on in-class concept questions, and, strikingly, that it did so even when nobody in the discussion group originally knew the correct answer. Crouch and Mazur (2001) report ten years of the same structure.
So the move is well evidenced in undergraduate science teaching and the school-facing label has not itself been studied. Both sentences are true, and using the second to dismiss the first would be as wrong as using the first to dress up the second.

The design problem is always the same: make it impossible for the check to be answered by the learners who were going to be fine anyway. Four rules do that, and the first one rules out the commonest classroom practice in the world.
Never let learners choose whether to be counted. Hands up, "does everyone understand", and "any questions" all sample volunteers. They tell you about the fastest three learners and nothing about the twenty-seven you needed to hear from. Replace them with something that samples everyone.
Ask for the answer and the reason. A correct answer with wrong reasoning is the most expensive thing you can fail to notice, because it looks exactly like success. Requiring a because clause is the cheapest fix available.
Design the wrong answers. A diagnostic question is only diagnostic if each wrong option corresponds to a specific misconception. Four plausible distractors tell you what to teach next. Three obviously silly ones tell you nothing.
Decide in advance what you will do about it. Before you ask, know what you will do if a third of the class gets it wrong. Without that, the check produces information at exactly the moment you have no time to think, and the natural response is to carry on as planned. Which is not formative assessment.
The techniques that make these work are the ones that get everyone answering: cold calling, mini whiteboards, exit tickets at the end, and show call for the written work in front of you.
Grouped by the kind of information each one gives you, because that is the decision a teacher is actually making. A whole-class signal and a written check answer different questions, and the commonest mistake is using the fast one when you needed the slow one.
Jump to: Whole-Class Signals · Written Checks · Talk-Based Checks · Checks That Surface Misconceptions, Not Confidence.
Everyone answers at once and you read the room. Fast, and the weakest at telling you what an individual actually thinks.
Slower, and they give you the reasoning rather than the answer. This is where misconceptions actually surface.
You hear reasoning as it forms. Powerful and slow, and it samples very few learners unless you design against that.
The category the competing lists do not have. These are designed so a wrong answer tells you exactly what the learner believes.
The question 62 undifferentiated strategies leave open is which one to reach for. Cycle length answers it. The framing is Fisher and Frey’s, and the mapping below is ours.
| Cycle | What it is for | Use | What you do with the result |
|---|---|---|---|
| Short, inside a minute | Can we move on from this sentence? | Mini whiteboards, fingers, cold call | Re-explain now, or carry on. No record kept. |
| Medium, inside the lesson | Did the explanation land? | Diagnostic multiple choice, because clause, turn and talk with report back | Change the next activity, or reteach to a group while others practise. |
| Long, lesson to lesson | What do I teach tomorrow? | Exit ticket, two-minute paper, entrance ticket | Rewrite tomorrow’s starter. This is the only cycle where reading the answers later is any use. |
Four common checks measure something other than understanding, and all four feel productive while doing it.
Thumbs up and traffic lights measure confidence. A learner who has misunderstood cleanly and completely is often the most confident person in the room. The hypercorrection literature says that learner is also the most worth catching, and a confidence check will not find them.
"Does everyone understand?" samples volunteers. Worse, it asks learners to assess their own understanding, which is the thing they are least able to do about material they have just met.
A whole-class choral answer hides the individual. A learner who does not know can mouth along a fraction behind and be undetectable, and the confident half of the room carries the sound.
A correct answer with wrong reasoning looks exactly like success. This is the most expensive miss available, and the only defence is asking how they knew.
Whole-class checks are designed to reach everyone, and the format itself can exclude the learners it most needs to hear from. A public answer raises the stakes for an anxious learner, a fast cue samples only the fast, and a written check measures writing as well as understanding. Each of those is a format problem with a cheap fix, and none of them requires lowering what you are checking for.
A public check is a public test. For an anxious learner, a whole-class response method with a visible answer raises the stakes of being wrong at exactly the moment you want honesty. Mini whiteboards held to the chest, or a written check you read while circulating, get you the same information privately.
Speed excludes. A check that runs on a five-second cue samples the learners who can produce an answer in five seconds. For a learner with slower processing or a language disorder, extending to fifteen seconds changes who is in the data. This is the same adjustment as wait time, and it costs nothing.
Writing is not neutral. A written check measures writing as well as understanding. For a learner with dysgraphia or as an EAL learner early in acquisition, a diagram, a pointed-to option or a spoken answer gets you the understanding without the confound.
Wrong for a reason you did not design for. Distractor analysis assumes the learner picked the option for the reason you intended. Learners with additional needs frequently arrive at an answer by a route nobody modelled. If a pattern of answers makes no sense, ask rather than infer.
Checking for understanding is close to unarguable as a practice and its research base is far weaker than its status. Four qualifications are worth carrying into any conversation about it.
The effect size is not known. Not 0.7, and not confidently 0.20 either. Two credible sources disagree, and the second says the uncertainty is the finding. Do not build a policy on any specific number.
The famous review was not a meta-analysis. Black and Wiliam (1998) was a narrative review of around 250 studies, with the authors hand-examining ten years of 76 journals. It is a serious piece of work and it is not what it is usually described as.
Feedback has a real failure mode. Over a third of interventions in the largest meta-analysis reduced performance. Any technique that increases the amount of feedback in a classroom increases exposure to that too.
The sources disagree about subjects. Kingston and Nash found the largest effects in English and the smallest in science; the EEF reports slightly higher effects in mathematics and science. Two credible sources point opposite ways, and there is no basis for choosing between them.
It is eliciting evidence of what learners know during the lesson and changing what you do next in response. It is the fastest cycle of formative assessment, running minute by minute rather than lesson by lesson.
No. That range came from a narrative review rather than a meta-analysis. When someone did meta-analyse it, only 13 studies were usable and the weighted mean was 0.20, and even that figure has been contested. The honest answer is that the size of the effect is not known.
Yes, and often. In Kluger and DeNisi's meta-analysis of 607 effect sizes, over 38 per cent of feedback interventions reduced performance. The authors treated that as a robust finding rather than an artefact.
It samples volunteers. The learners who answer are the ones who already know, and a learner who does not understand usually cannot tell that they do not. Use a method where every learner has to produce something.
The term is not. A title search of the largest education research database returns one record, a practitioner magazine article. The underlying move, a diagnostic question at the pivot of a lesson, is well evidenced in undergraduate physics and biology teaching under the name Peer Instruction.
Every reference below was checked twice: once to confirm it exists, and again to confirm the record it resolves to is the study being cited. Where a figure could not be established at source, it is not printed here.