Checking for Understanding: What the Evidence Really Says

Updated on  

August 29, 2026

Checking for Understanding: What the Evidence Really Says

|

August 28, 2026

The 0.4 to 0.7 effect size for formative assessment does not survive scrutiny, and over a third of feedback interventions make performance worse.

Start your metacognitive learning plan
Copy citation

Main, P. (2026, August 28). Checking for Understanding: What the Evidence Really Says. Structural Learning. https://www.structural-learning.com/post/checking-for-understanding

What is checking for understanding?

Checking for understanding is eliciting evidence of what learners currently know during the lesson, and then changing what you do next in response. It is the fastest cycle of formative assessment, running minute by minute rather than lesson by lesson. On Black and Wiliam’s own definition the evidence has to be used, so a check that does not change the teaching is not formative assessment at all.

The number everybody quotes for formative assessment is wrong, and the person most associated with it has said so himself, in print, without anybody much noticing.

Checking for understanding is what a teacher does inside a lesson to find out what learners currently know, so that the next few minutes can change. It is the fastest cycle of formative assessment: elicit evidence, read it, act on it, all within the lesson rather than after it.

The famous claim is that formative assessment is worth an effect size of 0.4 to 0.7. That range comes from Black and Wiliam's (1998) review, which was a narrative review of roughly 250 studies rather than a meta-analysis, and the range was a summary of what different studies reported rather than a pooled figure. When Kingston and Nash (2011) went looking for studies solid enough to pool, they screened more than 300 and found 13 usable. The median effect was 0.25 and the weighted mean was 0.20. Briggs and colleagues (2012) then criticised that meta-analysis in turn, concluding that considerable uncertainty remains.

If you want the checks themselves, skip to the 34, grouped by what each one actually tells you.

So the honest position is not "0.20 replaces 0.7". It is that nobody currently knows the size of the effect, because the research base is too thin to say. That is a much better thing for a teacher to know than a number, because it changes what you do: you check for understanding because a lesson without it is flying blind, not because a meta-analysis promised you seven months.

What You Need to Know

  1. A check that changes nothing is not a check: on Black and Wiliam's own definition, evidence has to be used, so if your teaching does not move, no formative assessment happened.
  2. The 0.4 to 0.7 figure does not survive scrutiny: 13 usable studies gave a weighted mean of 0.20, and even that is contested.
  3. Feedback can make things worse: in the largest meta-analysis of feedback interventions, over 38 per cent of effects were negative.
  4. Hands up tells you about the fastest three learners: any check where learners choose whether to be counted is a check on volunteers.

What Is Checking for Understanding?

Checking for understanding is eliciting evidence of what learners know at the point of teaching, and then changing what you do next in response. It runs on the shortest cycle in the classroom, minute by minute rather than lesson by lesson. Doug Lemov gives it a whole chapter in Teach Like a Champion 3.0, and the moves in it are all versions of the same question: who actually has this, and how would I know?

Is It Different from Formative Assessment?

Mostly it is branding, with one distinction that survives. The difference is cycle length, not kind.

Formative assessment is the whole system: eliciting evidence, interpreting it, and acting on it. Checking for understanding is that system's first step run at speed, inside the lesson. There is no study comparing them and no evidence that they have different effects, so anyone claiming a difference in size is asserting a taxonomy rather than reporting a finding.

The distinction collapses entirely in one situation, and it is the common one. A check that elicits evidence and does not change what the teacher does next fails Black and Wiliam's own definition of formative. The whole thing turns on the second step. Our guide to formative assessment strategies covers the longer cycles.

What the Evidence Actually Shows

The direction of the evidence is good and the magnitude is genuinely unknown. Three findings matter more to a teacher than any effect size: the headline number is not defensible, feedback frequently backfires, and the reputable sources disagree with each other in ways nobody hides.

Evidence summary card with 3 figures: +6 months progress, the eef feedback rating, 155 studies; average 0.41, across 607 effect sizes; over 38% went negative, feedback that made performance worse.
What the evidence says

The Numbers, and Why They Disagree

Dylan Wiliam set the competing estimates side by side himself in 2011, which is more honest than most of the people citing him. Kluger and DeNisi (1996) found an average of 0.41 for feedback interventions. Black and Wiliam (1998) estimated 0.4 to 0.7. Hattie and Timperley (2007) proposed around 0.95. Wiliam's own classroom field trial, run over a year with ordinary teachers and measured on externally mandated standardised tests, produced 0.32.

His explanation is worth having, because it is not "somebody was wrong". The estimates diverge because outcome measures differ in how sensitive they are to instruction, and because the populations differ in variance. A test built around the thing you taught will always show a bigger effect than a national exam.

The Education Endowment Foundation currently rates feedback at +6 months additional progress, very low cost, on high-strength evidence from 155 studies. That is the most useful public summary for a school, and it is a different question from "what is the effect size of formative assessment", which is the question that has no reliable answer.

The Finding Teachers Are Rarely Told

Kluger and DeNisi's (1996) meta-analysis is the largest of feedback interventions: 131 usable papers, 607 effect sizes, 23,663 observations. The weighted mean was 0.41, which is the number people quote.

The number people do not quote sits in the same results section. Over 38 per cent of the effects were negative. More than a third of feedback interventions made performance worse. The authors checked whether that was driven by one prolific researcher's unusually negative studies; excluding them, it was still 33 per cent. They conclude the negative effects are robust rather than an artefact.

That is the single most useful fact in this article. Feedback is not a safe operation with a variable dose. It has a real failure mode, it fires often, and the general shape of the failure is feedback that directs attention to the person rather than the task.

Hinge Questions Are a Convention, Not a Literature

Hinge questions, the diagnostic multiple-choice item placed at the pivot of a lesson, are treated in British CPD as an evidence-based technique. Searching the largest education research database by title returns one record, and it is a practitioner magazine article by Wiliam, who popularised the term.

What is genuinely evidenced is the adjacent practice under a different name. Peer Instruction in undergraduate physics and biology uses exactly this move, and it has been tested with control conditions. Smith and colleagues (2009), in Science, found that peer discussion improved performance on in-class concept questions, and, strikingly, that it did so even when nobody in the discussion group originally knew the correct answer. Crouch and Mazur (2001) report ten years of the same structure.

So the move is well evidenced in undergraduate science teaching and the school-facing label has not itself been studied. Both sentences are true, and using the second to dismiss the first would be as wrong as using the first to dress up the second.

Three hand-drawn teal panels: a teacher and a learner with quotation marks in their speech bubbles, labelled elicit; a teacher holding a written sheet beside a clock, labelled read; two people pointing at a ticked checklist, labelled act.
Elicit, read, act

How to Check for Understanding Properly

The design problem is always the same: make it impossible for the check to be answered by the learners who were going to be fine anyway. Four rules do that, and the first one rules out the commonest classroom practice in the world.

Never let learners choose whether to be counted. Hands up, "does everyone understand", and "any questions" all sample volunteers. They tell you about the fastest three learners and nothing about the twenty-seven you needed to hear from. Replace them with something that samples everyone.

Ask for the answer and the reason. A correct answer with wrong reasoning is the most expensive thing you can fail to notice, because it looks exactly like success. Requiring a because clause is the cheapest fix available.

Design the wrong answers. A diagnostic question is only diagnostic if each wrong option corresponds to a specific misconception. Four plausible distractors tell you what to teach next. Three obviously silly ones tell you nothing.

Decide in advance what you will do about it. Before you ask, know what you will do if a third of the class gets it wrong. Without that, the check produces information at exactly the moment you have no time to think, and the natural response is to carry on as planned. Which is not formative assessment.

The techniques that make these work are the ones that get everyone answering: cold calling, mini whiteboards, exit tickets at the end, and show call for the written work in front of you.

34 Checks, Grouped by What They Actually Tell You

Grouped by the kind of information each one gives you, because that is the decision a teacher is actually making. A whole-class signal and a written check answer different questions, and the commonest mistake is using the fast one when you needed the slow one.

Jump to: Whole-Class Signals · Written Checks · Talk-Based Checks · Checks That Surface Misconceptions, Not Confidence.

Whole-Class Signals

Everyone answers at once and you read the room. Fast, and the weakest at telling you what an individual actually thinks.

  1. Mini whiteboards. Every learner writes and holds up. The single best whole-class check, because thirty answers become visible at once.
  2. Fingers 1 to 4. Learners show a number for the multiple-choice option they chose, held at the chest so neighbours cannot copy.
  3. Show me the answer. Written on a whiteboard, revealed on a count of three so nobody adjusts.
  4. Heads down, hands up. Removes the social cost of admitting you are unsure. Useful once, then it becomes a game.
  5. Response cards. Pre-printed A, B, C, D cards. The best-evidenced version of this family.
  6. Traffic lights. Red, amber, green cups or cards. Measures confidence, not understanding, which is a real limitation.
  7. Stand up if you agree. Physical, quick, and only worth using where the disagreement is genuine.
  8. Silent thumbs at the chest. Same information as thumbs up, without the class seeing who is struggling.

Written Checks

Slower, and they give you the reasoning rather than the answer. This is where misconceptions actually surface.

  1. Exit ticket. One question at the end, collected. Read them before you plan tomorrow, or do not set one.
  2. Entrance ticket. The same question at the start of the next lesson, which is retrieval as well as a check.
  3. Two-minute paper. Write everything you know about X. Ungraded. Reveals gaps a question would not have asked about.
  4. One sentence summary. Force the whole idea into fifteen words. Impossible to do without understanding it.
  5. Because clause. The answer plus "because...". Converts a guess into a claim you can inspect.
  6. Annotate the model. Give a worked answer and ask learners to mark where the thinking happens.
  7. Predict then check. Written prediction before a demonstration, then compare. The gap is the diagnostic.
  8. Corrections column. Learners rewrite a wrong answer with what they now understand, which is the part that teaches.
  9. 3-2-1. Three things learned, two questions, one thing still unclear. The third is the useful one.
  10. Draw it. A diagram instead of prose, which unmasks a learner whose writing hides their understanding.

Talk-Based Checks

You hear reasoning as it forms. Powerful and slow, and it samples very few learners unless you design against that.

  1. Cold call. Named individual, no hands. The only reliable way to sample someone who would not volunteer.
  2. Turn and talk, then report back. Everyone rehearses, then one pair is accountable for reporting.
  3. Say it again, better. A learner restates their own answer in the subject register after you model it.
  4. Revoicing. "So you are saying that..." Puts the claim back for the learner to accept or correct.
  5. Explain it to a partner who was away. Forces completeness rather than shorthand.
  6. Ask for a non-example. Knowing what something is not is a harder and more revealing test than defining it.
  7. Probe the right answer. Ask a learner who was correct how they knew. A correct answer with wrong reasoning is the most expensive thing to miss.
  8. Chain the answer. One learner starts, the next continues, a third judges. Three learners checked in a minute.

Checks That Surface Misconceptions, Not Confidence

The category the competing lists do not have. These are designed so a wrong answer tells you exactly what the learner believes.

  1. Diagnostic multiple choice. Every wrong option corresponds to one specific misconception, so the distribution names what to teach.
  2. Hinge question. One question at the pivot of the lesson that decides whether you move on. A convention rather than an evidenced technique, but a good one.
  3. Which is wrong and why. Present four answers, one flawed. Identifying the flaw is a harder test than producing an answer.
  4. Odd one out. Three items, one different. The justification is the check, not the choice.
  5. Always, sometimes, never. A statement to classify. Ruthless at exposing an over-generalised rule.
  6. Sort the examples. Cards into categories. The disputes between learners are where the misconceptions live.
  7. Confidence plus answer. Ask for the answer and how sure they are. A confident wrong answer is the one worth your next five minutes.
  8. The deliberate error. Put a mistake on the board and wait. Silence tells you as much as a correction.

Which Check, and When

The question 62 undifferentiated strategies leave open is which one to reach for. Cycle length answers it. The framing is Fisher and Frey’s, and the mapping below is ours.

CycleWhat it is forUseWhat you do with the result
Short, inside a minuteCan we move on from this sentence?Mini whiteboards, fingers, cold callRe-explain now, or carry on. No record kept.
Medium, inside the lessonDid the explanation land?Diagnostic multiple choice, because clause, turn and talk with report backChange the next activity, or reteach to a group while others practise.
Long, lesson to lessonWhat do I teach tomorrow?Exit ticket, two-minute paper, entrance ticketRewrite tomorrow’s starter. This is the only cycle where reading the answers later is any use.

The Check That Tells You Nothing

Four common checks measure something other than understanding, and all four feel productive while doing it.

Thumbs up and traffic lights measure confidence. A learner who has misunderstood cleanly and completely is often the most confident person in the room. The hypercorrection literature says that learner is also the most worth catching, and a confidence check will not find them.

"Does everyone understand?" samples volunteers. Worse, it asks learners to assess their own understanding, which is the thing they are least able to do about material they have just met.

A whole-class choral answer hides the individual. A learner who does not know can mouth along a fraction behind and be undetectable, and the confident half of the room carries the sound.

A correct answer with wrong reasoning looks exactly like success. This is the most expensive miss available, and the only defence is asking how they knew.

Checking for Understanding and Learners with SEND

Whole-class checks are designed to reach everyone, and the format itself can exclude the learners it most needs to hear from. A public answer raises the stakes for an anxious learner, a fast cue samples only the fast, and a written check measures writing as well as understanding. Each of those is a format problem with a cheap fix, and none of them requires lowering what you are checking for.

A public check is a public test. For an anxious learner, a whole-class response method with a visible answer raises the stakes of being wrong at exactly the moment you want honesty. Mini whiteboards held to the chest, or a written check you read while circulating, get you the same information privately.

Speed excludes. A check that runs on a five-second cue samples the learners who can produce an answer in five seconds. For a learner with slower processing or a language disorder, extending to fifteen seconds changes who is in the data. This is the same adjustment as wait time, and it costs nothing.

Writing is not neutral. A written check measures writing as well as understanding. For a learner with dysgraphia or as an EAL learner early in acquisition, a diagram, a pointed-to option or a spoken answer gets you the understanding without the confound.

Wrong for a reason you did not design for. Distractor analysis assumes the learner picked the option for the reason you intended. Learners with additional needs frequently arrive at an answer by a route nobody modelled. If a pattern of answers makes no sense, ask rather than infer.

Limitations and Critiques

Checking for understanding is close to unarguable as a practice and its research base is far weaker than its status. Four qualifications are worth carrying into any conversation about it.

The effect size is not known. Not 0.7, and not confidently 0.20 either. Two credible sources disagree, and the second says the uncertainty is the finding. Do not build a policy on any specific number.

The famous review was not a meta-analysis. Black and Wiliam (1998) was a narrative review of around 250 studies, with the authors hand-examining ten years of 76 journals. It is a serious piece of work and it is not what it is usually described as.

Feedback has a real failure mode. Over a third of interventions in the largest meta-analysis reduced performance. Any technique that increases the amount of feedback in a classroom increases exposure to that too.

The sources disagree about subjects. Kingston and Nash found the largest effects in English and the smallest in science; the EEF reports slightly higher effects in mathematics and science. Two credible sources point opposite ways, and there is no basis for choosing between them.

Frequently Asked Questions

What is checking for understanding?

It is eliciting evidence of what learners know during the lesson and changing what you do next in response. It is the fastest cycle of formative assessment, running minute by minute rather than lesson by lesson.

Is formative assessment really worth 0.4 to 0.7?

No. That range came from a narrative review rather than a meta-analysis. When someone did meta-analyse it, only 13 studies were usable and the weighted mean was 0.20, and even that figure has been contested. The honest answer is that the size of the effect is not known.

Can feedback make learning worse?

Yes, and often. In Kluger and DeNisi's meta-analysis of 607 effect sizes, over 38 per cent of feedback interventions reduced performance. The authors treated that as a robust finding rather than an artefact.

What is wrong with asking "does everyone understand?"

It samples volunteers. The learners who answer are the ones who already know, and a learner who does not understand usually cannot tell that they do not. Use a method where every learner has to produce something.

Are hinge questions evidence-based?

The term is not. A title search of the largest education research database returns one record, a practitioner magazine article. The underlying move, a diagnostic question at the pivot of a lesson, is well evidenced in undergraduate physics and biology teaching under the name Peer Instruction.

References and Further Reading

Every reference below was checked twice: once to confirm it exists, and again to confirm the record it resolves to is the study being cited. Where a figure could not be established at source, it is not printed here.

  1. Assessment and classroom learning. Black, P., and Wiliam, D. (1998). Assessment in Education: Principles, Policy and Practice, 5(1), 7-74. The origin of the 0.4 to 0.7 range. A narrative review of around 250 studies, built by hand-examining every issue of 76 journals across ten years. Not a meta-analysis, and it never claimed to be.
  2. Formative assessment: A meta-analysis and a call for research. Kingston, N., and Nash, B. (2011). Educational Measurement: Issues and Practice, 30(4), 28-37. Screened over 300 studies, found 13 usable, and reported a weighted mean of 0.20. The paper that broke the famous number.
  3. Meta-analytic methodology and inferences about the efficacy of formative assessment. Briggs, D. C., Ruiz-Primo, M. A., Furtak, E., Shepard, L., and Yin, Y. (2012). Educational Measurement: Issues and Practice, 31(4), 13-17. The critique of the critique, and the reason the honest answer is "nobody knows" rather than "0.20".
  4. The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Kluger, A. N., and DeNisi, A. (1996). Psychological Bulletin, 119(2), 254-284. 607 effect sizes, mean 0.41, and over 38 per cent of them negative. The most important paper here for anyone about to increase the amount of feedback in their classroom.
  5. The power of feedback. Hattie, J., and Timperley, H. (2007). Review of Educational Research, 77(1), 81-112. The high estimate, at around 0.95 on overlapping literature, and half of the reason the numbers in this field cannot all be right.
  6. What is assessment for learning? Wiliam, D. (2011). Studies in Educational Evaluation, 37(1), 3-14. Open access. Wiliam sets the incompatible estimates against each other himself and explains why they diverge. Read this before quoting any effect size for formative assessment.
  7. Formative assessment: A critical review. Bennett, R. E. (2011). Assessment in Education: Principles, Policy and Practice, 18(1), 5-25. The formal critical review of the field.
  8. Why peer discussion improves student performance on in-class concept questions. Smith, M. K., Wood, W. B., Adams, W. K., Wieman, C., Knight, J. K., Guild, N., and Su, T. T. (2009). Science, 323(5910), 122-124. The diagnostic-question move, tested properly, with the striking finding that discussion helped even when nobody in the group started with the right answer.
  9. Peer Instruction: Ten years of experience and results. Crouch, C. H., and Mazur, E. (2001). American Journal of Physics, 69(9), 970-977. A decade of the individual-commit-then-discuss structure that a hinge question borrows.
  10. Inside the black box: Raising standards through classroom assessment. Black, P., and Wiliam, D. (2010). Phi Delta Kappan, 92(1), 81-90. The 2010 reprint of the 1998 booklet that took this argument into schools. Cite the reprint if you need a DOI, and name the booklet as the original.
Paul Main, Founder of Structural Learning
About the Author
Paul Main
Founder & Metacognition Researcher

Paul Main is an educator and metacognition researcher who founded Structural Learning in 2002. With a psychology degree from the University of Sunderland and 22+ years helping schools embed thinking skills, he bridges the gap between educational research and classroom practice. Fellow of the RSA and Chartered College of Teaching, with 128+ Google Scholar citations.

More →

Oracy

Back to Blog