Visualizing and Verbalizing: concept imagery for comprehension

Updated on  

July 2, 2026

Visualizing and Verbalizing: concept imagery for comprehension

|

June 13, 2026

A comprehensive, research-grounded guide to using concept imagery and verbalisation strategies to support language and reading comprehension for learners in UK and international classrooms.

Key Takeaways

  • Concept imagery is the sensory-cognitive ability to construct an imaged gestalt, or whole picture, from written and oral language.
  • The approach systematically targets the connection between nonverbal mental imagery and verbal language, moving beyond traditional questioning routines.
  • Evidence suggests positive impacts on reading comprehension, but the high dosage requirements and lack of automatic transfer require careful school planning.
  • Integrating these principles with structural frameworks like Map It, Say It, and Build It helps learners anchor, articulate, and write down their mental models.

What is Visualizing and Verbalizing?

Visualizing and verbalizing is a sensory-cognitive intervention designed to develop concept imagery, the ability to create a mental picture from spoken or written language. Developed originally by Nanci Bell, the approach is built on the premise that language comprehension is fundamentally a sensory process of constructing a whole mental image, or gestalt, in the mind. Many learners who read fluently but struggle to understand or recall what they have read suffer from a weakness in this visual-verbal connection. Without a mental image to organize and anchor the details, language remains a series of disconnected words.

The primary goal of this intervention is to make sensory processing automatic, allowing learners to translate language into mental movies and then translate those mental movies back into verbal or written language. Traditional comprehension strategies often test comprehension by asking questions, rather than teaching the cognitive processes required to comprehend. This intervention reverses that dynamic by explicitly tutoring the sensory system. It provides learners with a structured toolkit to retrieve, organize, and synthesize details into a coherent mental representation.

In the classroom, this is achieved through structured Socratic dialogue. The teacher uses specific language to guide the learner's mind to construct, expand, and monitor their mental pictures. Rather than telling learners what to see, the teacher asks choice-and-contrast questions that prompt them to actively create their own mental models. Over time, this dual-coding process becomes self-generative and supports independent reading, listening, and writing.

Who is this Intervention For?

This intervention is primarily designed for learners who present a specific profile of reading difficulty: adequate word recognition and decoding skills, but significantly depressed language and reading comprehension. These individuals are often described as "hyperlexic" or "poor comprehenders" (Johnson-Glenberg, 2000). They can decode complex vocabulary and read aloud with speed and accuracy, yet they cannot summarize the text, answer inferential questions, or follow multi-step oral directions. This gap occurs because their sensory systems fail to automatically generate the mental representations that support working memory and recall.

The strategy is also highly beneficial for learners with specific speech, language, and communication needs (SLCN), including those on the autism spectrum. These learners frequently struggle with figurative language, abstract concepts, and inference, because they process language in a literal, word-by-word manner. By explicitly training them to construct a mental gestalt, teachers help them make abstract concepts concrete. Additionally, learners with working memory difficulties benefit because an image reduces cognitive load, allowing them to chunk information more efficiently.

Conversely, this intervention is not appropriate for learners whose primary barrier to reading is weak decoding, word recognition, or phonological awareness. If a learner cannot read the words on the page accurately, their cognitive energy is entirely consumed by the mechanics of decoding, leaving no capacity for concept imagery. For these learners, structured synthetic phonics must remain the primary intervention. It is also not a substitute for general vocabulary instruction, as a learner cannot visualize a concept for which they have no underlying semantic knowledge.

The Sensory-Cognitive Mechanism of Comprehension

The cognitive foundation of this approach relies heavily on Dual Coding Theory (Paivio, 1971, 1986). This theory posits that human cognition is governed by two independent but highly interacting systems: a verbal system for language and a nonverbal system for mental imagery. When both systems are activated simultaneously, memory and comprehension are significantly enhanced. Standard reading instruction in many schools over-emphasises the verbal system, treating comprehension as a linguistic puzzle solved by identifying synonyms, grammatical clues, or text structures.

Without concept imagery, the verbal system is forced to work in isolation, which quickly overloads the working memory. When a learner reads a paragraph, they must hold each word in their mind while waiting for the next, hoping to synthesize them. An imaged gestalt acts as a cognitive filing cabinet, organising individual details into a single, cohesive mental picture (Bell, 1991). This imagery-language connection is not a separate learning style; rather, it is a universal cognitive architecture that must be systematically trained to reach automaticity.

When a learner successfully visualizes a scene, they can easily make inferences because they are reading information directly from their own mental image. For example, if a text states that "the dry grass crackled underfoot," a learner with strong concept imagery will picture a sun-drenched, arid landscape, automatically inferring the season and weather without the author stating them explicitly. If a learner does not visualize, the phrase remains a literal, abstract statement, and the opportunity for deep inference is lost.

Step-by-Step Classroom Implementation Sequence

Implementing this strategy in the classroom requires a structured, cumulative sequence that gradually increases the cognitive demand on the learner. Each step must be mastered before moving to the next, ensuring that sensory processing becomes secure and automatic.

Step 1: Setting the Climate and Introduction

Before introducing the structured steps, the teacher must explain the purpose of the strategy to the learners. The teacher explains that the brain works like a movie screen, and that to understand what we read, we must turn words into pictures and pictures into words. A simple diagram of a head with a thought bubble can be used as a visual anchor. The teacher models this by describing a simple, high-imagery object, showing how the words allow them to build a picture in their mind.

Step 2: Picture-to-Picture Imaging

The teacher displays a highly colourful, detailed picture to the group and asks a learner to describe it using specific structure words. These structure cards are the foundational vocabulary of the intervention, comprising twelve keys: What, Size, Colour, Number, Shape, Where, Movement, Mood, Background, Perspective, When, Sound. The learner points to elements of the picture and describes them, using the cards as prompts to ensure completeness. The teacher asks choice-and-contrast questions to refine the description, such as: "Is that dog big or small?" This step establishes the shared vocabulary and the expectation of detail.

Step 3: Word-by-Word and Single-Sentence Imaging

The teacher transitions from a physical picture to a single word, asking the learner to create a mental picture of a concrete noun, such as "castle" or "tractor". The learner describes their image using the structure cards. Once single words are consistently imaged, the teacher introduces a single, descriptive sentence. For example: "The large, rusty tractor sat in the muddy field." The learner places a coloured square on the table to represent the sentence, visualizes it, and describes their mental image, systematically matching their picture to the details in the sentence.

Step 4: Sentence-by-Sentence Imaging

The teacher presents a short paragraph of three to five sentences, introducing them one at a time. For each sentence, the learner places a coloured square of a different colour on the table to anchor the image. As each new sentence is read, the learner must integrate the new details into their existing mental image, expanding the gestalt. The teacher uses Socratic questioning after each sentence to ensure the picture is updated. At the end of the paragraph, the learner points to each square and summarizes the overall image, verbalising the complete story.

Step 5: Multiple Sentence and Whole Paragraph Imaging

Once sentence-by-sentence processing is secure, the teacher reads two or three sentences before pausing for visualization and verbalisation. This step increases the demand on working memory, as the learner must process larger chunks of language before anchoring them. Eventually, the learner reads or hears a whole paragraph, visualizes the entire scene, and describes the complete mental movie. The focus shifts from tracking individual words to synthesizing a global understanding of the text.

Structural Learning Practice: Map It, Say It, Build It

Concept imagery becomes a powerful classroom tool when integrated with the Structural Learning framework. By connecting mental visualization to physical manipulation, oral language, and structured writing, teachers can ensure that concept imagery translates into robust academic achievements.

Map It: Anchoring the Mental Model

The "Map It" stage uses graphic organizers to give a physical structure to the mental gestalt. When a learner is visualising a text, they can translate their mental movie directly onto a spatial organizer, such as a storyboard, a character map, or a sequencing timeline. This physical representation reduces the cognitive load on the working memory, allowing the learner to organize the elements of their image visually.

Visualising and Verbalising framework infographic explaining What, How, and Why concept imagery supports reading comprehension.
Visualising and Verbalising Framework

For example, when visualising a history passage about the Roman invasion of Britain, learners can place physical symbols or keywords on a map of Britain. This process forces them to clarify the "Where," "Size," and "Background" of their mental image. It ensures that the spatial relationships in their visualization are accurate and organised, turning a fleeting mental picture into a permanent, structured model.

Say It: Scaffolding Verbalisation and Peer Dialogue

The "Say It" stage leverages exploratory and explanatory talk to refine the mental gestalt. Once a learner has visualized a sentence or paragraph, they must verbalise their image to a peer or teacher. This verbalisation is critical because it forces the brain to translate the nonverbal image code back into the verbal language code, reinforcing the dual-coding pathway.

Teachers can scaffold this dialogue using specific role cards to structure the conversation:

  • The Starter: "I see a large, dark forest under a stormy sky."
  • The Builder: "I want to add to that picture. In the middle of the forest, I see a small wooden cabin with smoke coming from the chimney."
  • The Challenger: "How does that match the text? The story said it was summer, so would there be smoke coming from the chimney?"

This collaborative visualization ensures that learners actively monitor their comprehension, using peer feedback to correct and expand their mental images.

Build It: From Mental Imagery to Structured Writing

The "Build It" stage bridges the gap between visualization and written comprehension, particularly for learners who experience "writer's block". Writing requires a massive amount of cognitive coordination, from spelling and grammar to content generation. By separating content generation from transcription, teachers can help learners write more fluently and descriptively.

In this stage, learners use their mental gestalt to build physical representations of their ideas using modular blocks, cards, or plastic figures before writing.

  • Level 1 (Concept): Learners build the physical scene of their story or argument on their desk, arranging the blocks to represent key details.
  • Level 2 (Sentence): Learners use the "Hockman method" of sentence expansion, using their physical model to generate sentences containing "because," "but," and "so".
  • Level 3 (Word): Learners label their physical models with precise vocabulary cards, focusing on rich adjectives and verbs derived from their visualization.

Once the physical model is complete, the learner simply describes their model in writing, reducing the cognitive load of transcription and resulting in rich, detailed prose.

Research Evidence, Practical Limitations, and Critiques

While many commercial providers heavily promote the efficacy of visualizing and verbalizing, schools must evaluate the empirical evidence base objectively to plan effective interventions. The National Center on Intensive Intervention (NCII) taxonomy report (2021 review cycle) noted that the programme drives a statistically significant reading effect size of 0.51 across paragraph reading comprehension and oral language measures. However, the academic studies supporting this effect size often suffer from methodological limitations, with the US Department of Education's What Works Clearinghouse (WWC) frequently rating primary studies as ineligible for review due to a lack of rigorous comparison groups or randomized controlled designs.

Furthermore, some clinical claims regarding the programme have been disputed. For instance, the British Society of Audiology (2018) clinical guidance on Auditory Processing Disorder (APD) states that there is currently no empirical research evidence supporting the use of the programme as an effective treatment for APD. Schools should therefore be cautious of marketing materials that position the programme as a clinical cure for auditory or language disorders, and instead view it as a structured educational scaffold for reading comprehension.

Another critical limitation is the practical "dosage barrier". The official programme recommends highly intensive instruction, up to four hours per day, five days a week, for ten weeks, totaling up to two hundred hours of instruction. In resource-constrained school environments, where special educational needs teams face significant constraints on time and funding, such a high dosage is operationally impossible to deliver. Research suggests that reading group sizes larger than four or five are significantly less effective, meaning schools cannot easily dilute the intervention into large-class formats without compromising its efficacy.

Finally, educators must address the "transfer gap". Research has consistently shown that learners do not automatically generalize mental imagery skills to their broader academic coursework (Bell, 1991). A learner may excel at visualizing short, highly descriptive paragraphs in a structured intervention room, but struggle to apply those same skills to a dense, abstract science textbook in a busy secondary classroom. To overcome this, schools must explicitly plan "bridging" sessions that teach learners how to actively apply visualization and verbalisation routines to standard curriculum materials.

Comparative Analysis of Comprehension Approaches

To make informed decisions, schools should understand how concept imagery strategies compare to other established reading comprehension interventions.

Visualizing and Verbalizing vs. Reciprocal Teaching

Reciprocal teaching (Palincsar & Brown, 1984) is a highly verbal, dialogue-based approach where learners work in groups to predict, clarify, question, and summarize texts. It is primarily verbally coded, relying on linguistic analysis and metacognitive strategies. In contrast, visualizing and verbalizing is visually coded, focusing on the sensory-cognitive transition from words to images.

Learner Profile Suitability Intervention Focus
High decoding, low comprehension Highly Suitable Sensory-cognitive concept imagery and verbalisation
Weak decoding, weak comprehension Not Suitable (Initially) Structured synthetic phonics and phonological awareness
Speech, language, and communication needs Highly Suitable Socratic language training and mental gestalt building
Severe vocabulary deficits Partially Suitable Combine with explicit vocabulary instruction and concrete objects

A seminal study by Johnson-Glenberg (2000) compared these two approaches with poor comprehenders in grades three to five. The study found that while both interventions improved reading comprehension compared to a control group, the visually based visualization programme was particularly effective for learners with weaker verbal memory, as the mental images provided an alternative cognitive pathway to store and retrieve information.

Visualizing and Verbalizing vs. Direct Instruction of Graphic Organizers

Direct instruction of graphic organizers involves teaching learners to map text structures, such as cause-and-effect or comparison grids, directly from the page. While highly effective for structural understanding, it operates at a conceptual level rather than a sensory level. Visualizing and verbalizing, however, operates at a pre-conceptual, sensory level, training the brain to build the raw mental images that are later organised by graphic templates.

For maximum impact, schools should not view these as mutually exclusive, but rather as complementary parts of a comprehensive reading strategy. Sensory-cognitive visualization builds the raw mental content, and graphic mapping organizes that content into logical academic structures.

Classroom Case Study: A Socratic Dialogue in Action

The following scenario illustrates a Year 5 (Key Stage 2) small-group intervention session, focusing on a descriptive paragraph about a historic castle. The group consists of three learners who decode words accurately but have weak comprehension.

The Target Text

"The massive grey castle stood on top of a steep, grassy hill. A deep, muddy moat surrounded the high stone walls. A heavy wooden bridge was lowered across the water."

Socratic Walkthrough

Teacher: Let's read our first sentence: "The massive grey castle stood on top of a steep, grassy hill." Let's place a blue square on our table to anchor this image. Alfie, what do those words make you picture on your movie screen?

Alfie: A castle. It's on a hill.

Teacher: Let's look at our structure cards. I see the "Size" and "Colour" cards. Alfie, what size is your castle, and what colour are the walls?

Alfie: It's massive and grey.

Teacher: Excellent. Now look at the "Shape" and "Background" cards. Priya, what shape is the hill, and what is in the background behind the castle?

Priya: The hill is really steep, like a triangle pointing up. In the background, the sky is grey and there are some dark clouds.

Teacher: Perfect. You've matched the mood of that grey castle. Now let's read the second sentence: "A deep, muddy moat surrounded the high stone walls." Let's place a green square next to our blue square. Priya, how does this sentence change your mental movie?

Priya: Now there's water around the castle.

Step Phase Visual Anchor Cognitive Demand
Step 1 Setting Climate Thought bubble diagram Metacognitive awareness of mental movies
Step 2 Picture-to-Picture Physical picture card Translating physical details to structured language
Step 3 Single-Sentence One coloured square Translating written details to individual images
Step 4 Sentence-by-Sentence Multiple coloured squares Integrating sequential details into an expanding gestalt
Step 5 Whole Paragraph Single square for gestalt Synthesizing whole text chunks independently

Teacher: Let's use the "Colour" and "Where" cards. What colour is that water, and where exactly does it go?

Priya: It's brown and muddy, and it goes all the way around the bottom of the high stone walls, like a big circle.

Teacher: Excellent tracking. Alfie, does your castle have windows, and can you see anyone on the walls? Look at the "Perspective" and "What" cards.

Alfie: I'm standing close to the bottom of the hill looking up. The walls are really high and have tiny square windows, but there's nobody there. It looks empty.

Teacher: Now let's read our third sentence: "A heavy wooden bridge was lowered across the water." Let's place a red square down. Sam, what do you see happening in this part of our movie? Look at the "Movement" and "Sound" cards.

Sam: The bridge is coming down. It's making a loud creaking sound as the chains let it go.

Teacher: Wonderful sound detail. What is the bridge made of, and where does it land?

Sam: It's made of thick, dark wood. It lands with a heavy thud on the grass on our side of the muddy water.

Teacher: Fantastic. Now, look at all three squares on our table. Sam, point to the blue square and tell us the first part of your movie. Then the green square, then the red square.

Sam: [Points to blue] First, I see the massive grey castle standing high up on that steep, grassy hill. [Points to green] Then, I see the deep, muddy moat going all the way around the stone walls. [Points to red] Finally, the heavy wooden bridge lowers down creaking and lands with a thud so we can cross.

Analysis of the Socratic Method

This dialogue illustrates several critical principles of sensory-cognitive tutoring:

  • The teacher never tells the learner what to see: The teacher uses Socratic prompts to elicit details from the learner's own mind.
  • The teacher uses choice-and-contrast prompts: Questions like "Is that dog big or small?" force the learner to make a conscious cognitive decision, which solidifies the mental image.
  • The structure cards prevent cognitive overload: Instead of asking vague questions, the teacher points to a specific card (e.g., "Size" or "Movement") to direct the learner's attention to one detail at a time.
  • The coloured squares act as memory anchors: The squares provide a physical representation of the text sequence, allowing the learner to chunk and retrieve information easily.

The SENCO and Classroom Teacher Toolkit

To implement this approach effectively, SENCOs and classroom teachers can use the following planning frameworks and checklists to structure their sessions.

One-Page Implementation Checklist

  • Establish the Climate: Explain the "mental movie screen" concept to learners, showing how words and pictures connect.
  • Display Structure Words: Ensure the twelve structure words (What, Size, Colour, Number, Shape, Where, Movement, Mood, Background, Perspective, When, Sound) are visible to all learners.
  • Use Visual Anchors: Keep a supply of coloured squares on hand to represent sentences during paragraph-level work.
  • Maintain Socratic Language: Avoid telling learners what to visualize; instead, ask choice-and-contrast questions to guide their processing.
  • Check for Understanding: Ensure learners can summarize their mental pictures by pointing to each coloured square in sequence.
  • Plan Bridging Sessions: Allocate time to explicitly teach learners how to transfer their visualization skills to curriculum textbooks.
  • Monitor Progress: Keep detailed records of each learner's ability to retain details and make inferences over time.

Classroom Prompt and Planning Template

LESSON PLAN: CONCEPT IMAGERY BRIDGING SESSION

Phase Educational Focus Classroom Action Structural Tool
Map It Spatial organisation Draw or sequence the mental image Graphic organizers and storyboards
Say It Socratic verbalisation Talk through the image with a peer Role cards and structured prompts
Build It Concept translation Build a physical model before writing Modular blocks and vocabulary labels

Target Text: ___________________________________________________________ Reading Level: ______________ Interest Level: ______________

  1. SETTING THE CLIMATE

    • Prompt: "Today, we are going to build a mental picture of this passage on our movie screens to help us understand it."
  2. STRUCTURE WORD FOCUS (Select 3-4 for this session) [ ] What [ ] Size [ ] Colour [ ] Where [ ] Movement [ ] Background

  3. SENTENCE-BY-SENTENCE SOCRATIC PROMPTS

    • Sentence 1: _______________________________________________________
      • Prompt: "What do you picture for this sentence? What size is it? Where is it located?"
    • Sentence 2: _______________________________________________________
      • Prompt: "How does this sentence change your movie? What is happening now?"
    • Sentence 3: _______________________________________________________
      • Prompt: "What new details are we adding? What is in the background?"
  4. VERBALISATION AND BRIDGING

    • Prompt: "Point to each square and tell me the story of your movie. How does this help us answer our comprehension question?"

Intervention Fit and Caution Box

Frequently Asked Questions (FAQs)

Is this intervention based on "learning styles" theory?

No. While the name sounds visual, the cognitive framework behind this strategy is Dual Coding Theory, which is a universally validated model of cognitive architecture (Paivio, 1986). It does not categorize children into "visual" or "verbal" learners. Instead, it asserts that all human brains process information more effectively when they connect nonverbal mental imagery with verbal language. The intervention aims to strengthen the connection between these two systems for all learners, particularly those with a specific deficit in visual-verbal integration.

How do we adapt this strategy for secondary school subjects?

In secondary schools, texts become dense and abstract, making automatic visualization more difficult. Teachers can adapt this strategy by focusing on key concepts rather than entire chapters. In a science lesson on cell structure, for example, the teacher can use the structure cards to help learners visualize the organelles within a cell, asking: "What shape is the mitochondrion, and where is it located relative to the nucleus?" In geography, learners can visualize tectonic plate movements, focusing on the "Movement," "Shape," and "Background" cards.

How long should an intervention session last?

While commercial clinics run intensive sessions for several hours a day, a realistic and effective school-based adaptation is to run short, targeted sessions of twenty to thirty minutes, three to four times a week. These sessions must be highly focused, keeping group sizes small (maximum of four learners) to ensure every learner has ample opportunity to verbalise their images. Consistency and high-quality Socratic dialogue are far more important than long, exhausting sessions that exceed the learners' attention spans.

What should I do if a learner says they "cannot see anything" on their movie screen?

If a learner struggles to visualize, the teacher should temporarily step back to concrete objects. Display a physical object, such as a colourful toy or a cup, and have the learner describe it using the structure words while looking at it. Then, cover the object with a cloth and ask the learner to describe what they saw from memory, using the structure words as prompts. This process scaffolds the brain to move from direct physical perception to mental visualization, building the confidence and cognitive pathways needed for text-based imaging.

Next-Lesson Action

Next lesson, select a short, descriptive paragraph from your curriculum text, read the first sentence aloud, and ask your learners to describe exactly what they visualize on their mental movie screens using the "What," "Size," and "Colour" structure words before they write anything down.

Feature Visualizing & Verbalizing Reciprocal Teaching Graphic Organizers
Primary Cognitive Code Visual and sensory (Nonverbal) Linguistic and analytical (Verbal) Conceptual and spatial
Core Strategy Socratic imagery prompts Predict, clarify, question, summarize Mapping text structures
Best Suited For Learners with weak visual-verbal integration Learners with strong oral language but weak strategies Learners struggling to organize information
Classroom Delivery Small group or individual intervention Collaborative peer-led groups Whole-class or small-group instruction

Research sources

Further reading from peer-reviewed research

These 5 studies give source context for the classroom guidance in this article on Visualizing and Verbalizing: concept imagery for comprehension. They are included as starting points for deeper reading, not as a substitute for local professional judgement.

Meta Analysis 77 citations link.springer.com

The Impact of Motivational Reading Instruction on the Reading Achievement and Motivation of Students: a Systematic Review and Meta-Analysis

Miriam McBreen et al. (2020) | Educational Psychology Review

Translate the finding into explicit modelling, guided practice and progress monitoring rather than relying on one-off exposure.

View study

Meta Analysis 97 citations journals.sagepub.com

A Review of Summarizing and Main Idea Interventions for Struggling Readers in Grades 3 Through 12: 1978–2016

E. Stevens et al. (2019) | Remedial and Special Education

Translate the finding into explicit modelling, guided practice and progress monitoring rather than relying on one-off exposure.

View study

Systematic Review linkinghub.elsevier.com

A systematic review and narrative synthesis of mental imagery tasks in people with an intellectual disability: Implications for psychological therapies.

O. Hewitt et al. (2022) | Clinical Psychology Review

Translate the finding into explicit modelling, guided practice and progress monitoring rather than relying on one-off exposure.

View study

Meta Analysis 114 citations idp.springer.com

The effectiveness of visual-based interventions on health literacy in health care: a systematic review and meta-analysis

Elisa Galmarini et al. (2024) | BMC Health Services Research

Use this as a caution: check learner fit, delivery quality and progress data before treating the approach as settled practice.

View study

Peer Reviewed Study 99 citations doi.apa.org

Training reading comprehension in adequate decoders/poor comprehenders: Verbal versus visual strategies

Johnson-Glenberg (2000) | Journal of Educational Psychology

Translate the finding into explicit modelling, guided practice and progress monitoring rather than relying on one-off exposure.

View study

Paul Main, Founder of Structural Learning
About the Author
Paul Main
Founder & Metacognition Researcher

Paul Main is an educator and metacognition researcher who founded Structural Learning in 2002. With a psychology degree from the University of Sunderland and 22+ years helping schools embed thinking skills, he bridges the gap between educational research and classroom practice. Fellow of the RSA and Chartered College of Teaching, with 128+ Google Scholar citations.

More →

SEND

Back to Blog