Updated on
September 24, 2026
B. F. Skinner's Theory: Operant Conditioning Explained
B. F. Skinner's theory explained: operant conditioning, the four consequences, schedules of reinforcement, KS2 and KS4 examples and a function check.

Updated on
September 24, 2026
B. F. Skinner's theory explained: operant conditioning, the four consequences, schedules of reinforcement, KS2 and KS4 examples and a function check.
What is B. F. Skinner's theory?
B. F. Skinner's theory explains behaviour through its relation to consequences. He defined reinforcement by an increase in future behaviour. In 1953, Skinner treated punishment as presenting a negative reinforcer or removing a positive one; the later functional convention defines punishment by a decrease in future behaviour. Radical behaviourism also includes private events such as thoughts and feelings.
B. F. Skinner's theory, usually called operant conditioning, is the idea that what happens straight after an action changes how often that action happens again. His behaviourism treats learning as a change in what a person does. A teacher can test it in a week: pick one action, watch what follows it, and see whether it happens more or less.
Skinner (1904 to 1990) built the theory from laboratory work with rats and pigeons. He then applied it to teaching machines, language and society. His wider philosophy, radical behaviourism, counts thoughts and feelings as behaviour too, rather than denying them.
For example, a teacher offers praise after a learner starts independent work. Does starting work become more likely later? The teacher's intent does not decide whether the praise reinforced the action.
This guide covers Skinner's main ideas, the four consequences, his schedules of reinforcement and what they look like in a KS2 and a KS4 classroom. It sits within our guide to the fundamental theories of learning. The operant conditioning guide adds more everyday examples from outside school.
Skinner's theory says that behaviour changes through its relation to what follows it. An action that produces reinforcement becomes more likely in a similar setting.
Skinner's account of punishment covered two procedures. One presented a negative reinforcer; the other removed a positive one. He stressed that their effects differed from reinforcement (Skinner, 1953).
Later behaviour analysts adopted a functional definition of punishment. Under this convention, a consequence counts as punishment when it reduces the future probability of an action (Azrin and Holz, 1966). This later convention is useful. It should not be presented as Skinner's own 1953 definition.
Skinner called such actions operants because they operate on the world. Their exact movement can vary while serving the same function. A learner can ask for help by speaking, writing or using AAC. Each form can belong to one operant class when it has the same effect.
This approach differs from Pavlovian conditioning. Pavlov studied how signals come to elicit responses by predicting other events. Skinner studied how the results of an action alter its future occurrence. The Pavlov guide explains the conditioned-reflex evidence.
Burrhus Frederic Skinner was an American psychologist who lived from 1904 to 1990. He studied English at Hamilton College before turning to psychology. He completed his doctorate at Harvard in 1931. He later worked at the universities of Minnesota and Indiana before returning to Harvard.
Skinner built on Edward Thorndike's work on the consequences of action. Yet he rejected mental bonds or “satisfaction” as final explanations. He aimed to measure how behaviour varied with the conditions in which it occurred. Our Thorndike guide covers puzzle boxes, connectionism and the changing Law of Effect.
His major books include The Behavior of Organisms (1938), Science and Human Behavior (1953), Verbal Behavior (1957) and About Behaviorism (1974). He also wrote fiction and social criticism, including Walden Two.
Skinner joined an experimental method to a broad account of behaviour. Operant selection, environmental context, reinforcement, punishment and the analysis of language all belong to that programme. His view was more than a set of classroom reward techniques.
| Idea | Account | Boundary |
|---|---|---|
| Operant behaviour | Consequences select actions from the variation a person or animal produces. | The same visible act can have different functions in different settings. |
| Reinforcement | A consequence increases the future chance of an operant. | An intended reward is not a reinforcer unless later behaviour increases. |
| Punishment | Skinner's 1953 account was procedural; later functional use defines a decrease in future behaviour. | Suppression may be brief and does not teach a useful alternative. |
| Radical behaviourism | Private and public behaviour both belong to the natural world. | A thought or feeling is not denied, but naming it does not complete the explanation. |
| Selection by consequences | Selection occurs across biological, individual and cultural time. | The three levels use different processes and should not be collapsed into one. |
Radical behaviourism includes events that only one person can observe directly. Skinner treated thinking, feeling and sensing as behaviour within the world. He rejected an inner copy of an act as its explanation (Skinner, 1945).
A learner may say, “I avoided the task because I felt anxious.” The feeling matters. A fuller account still asks what the task meant and what happened before. It also asks what followed avoidance and how this history developed.
The aim is not to deny experience. It is to avoid treating one label as the whole cause.
This differs from methodological behaviourism, which restricted scientific evidence to public observation. The behaviourism overview compares Watson's programme with Skinner's later view.
Skinner studied operant behaviour in a controlled chamber, commonly called an operant chamber or Skinner box. The apparatus recorded repeated actions and their consequences across time. A rat could press a lever; a pigeon could peck a key.
The controlled setting let Skinner change one condition at a time. He could then study patterns across time rather than rely on a few trials.
Skinner also used the cumulative recorder. Its pen moved across paper. The vertical line rose with each response.
A steep line showed a high response rate; a flat line showed no responses. Changes in the line made the effects of different conditions visible.
The chamber did not prove that all human learning is simple or mechanical. It isolated a relation so researchers could test it. Claims about language, school or society require extra evidence and careful limits.
A full operant analysis often examines an antecedent, an action and a consequence. An antecedent sets the occasion for an action; it does not force it. A consequence has a reinforcing function only when later behaviour shows an increase. Under the later functional convention, punishment is identified by a decrease.
Operant conditioning sorts consequences with two questions: was something added or taken away, and did the action then rise or fall? Crossing the two answers gives four consequences, often called the four quadrants. Positive and negative describe adding and removing. They do not mean pleasant and unpleasant.
The quadrant teachers most often get wrong is negative reinforcement. It is not a sanction. It makes an action more likely by removing something unwelcome, such as lifting a silent-work rule once the whole class has settled.
The same logic explains a common backfire. A learner who disrupts may be sent out of a task they find too hard. The adult sees a sanction, but the learner has escaped the task. If disruption then rises, the removal reinforced it.
Extinction is the idea that sits beside the grid. When reinforcement that used to follow an action stops, the action usually fades. It often gets worse first, a short rise known as an extinction burst. Plans that rely on withholding attention therefore need agreement across staff and careful oversight.
A schedule of reinforcement is the rule that decides which responses get reinforced. Charles Ferster and Skinner mapped these rules with pigeons and showed that each one produces its own pattern of responding (Ferster and Skinner, 1957). Every classroom routine already runs on one, whether or not the teacher chose it.
| Schedule | The rule | Where it shows up in class | Pattern in Skinner's research |
|---|---|---|---|
| Continuous | Every correct response is reinforced. | A new routine in September, such as lining up at the door, noticed every time. | Fast learning, but the action fades quickly once reinforcement stops. |
| Fixed ratio | Reinforcement after a set number of responses. | A sticker for every fifth piece of homework handed in. | Bursts of work, then a pause after each reward. |
| Fixed interval | The first response after a set time is reinforced. | A test every Friday, or merits counted at the end of each week. | Little effort early on, then a rush just before the deadline. |
| Variable ratio | Reinforcement after an unpredictable number of responses. | Praising some correct answers but not all. | Steady responding that is slow to fade. |
| Variable interval | The first response after an unpredictable time is reinforced. | Scanning the room and praising on-task work at random moments. | Steady, moderate responding. |
The last column comes from laboratory work, so treat it as a prediction to check, not a promise. The practical rule is to start dense and then thin. Reinforce a new routine every time until it is secure, then move to an unpredictable schedule so it survives the busy weeks. In one primary class, a team points game kept working when the teacher stretched the gap between points from two minutes to five (Bohan and Smyth, 2023).
Fixed intervals also explain the Friday-test scramble. If effort only pays off on test day, most of it arrives the night before. Short quizzes on unpredictable days spread that effort across the week, which also suits retrieval practice in secondary lessons.
Praise is only reinforcement when the praised action becomes more likely afterwards. Skinner's definition turns a kind word into a claim you can check. Most of the time praise does reinforce, but four common patterns stop it working.
First, general praise is weaker than specific praise. In four Year 4 classes, praise that named the behaviour produced more on-task work than general positive comments (Chalk and Bizo, 2004). A review of classroom studies rated behaviour-specific praise a potentially evidence-based practice (Royer et al., 2019). "Well done" tells a learner they pleased you; precise praise tells them what to repeat.
Second, public praise can punish. A Year 10 learner who hates being singled out may stop answering after being praised in front of the class. By Skinner's definition that praise worked as punishment, whatever the teacher meant.
Third, expected rewards can crowd out interest. Deci, Koestner and Ryan found that expected tangible rewards reduced later free-choice interest in tasks learners already enjoyed (Deci, Koestner and Ryan, 1999). Cameron and Pierce reached a less negative overall conclusion, and both reviews found verbal praise less risky than tangible rewards (Cameron and Pierce, 1994). The overjustification effect explains this pattern in more detail.
Fourth, praise that is too rare does little. In four secondary classrooms, specific praise about once every two minutes brought large gains, while once every four minutes had mixed effects (O'Handley et al., 2023). Observations of primary classrooms suggest that teachers' natural rate of specific praise is often low (Floress et al., 2018).
Try the function check below on one action you are working on this week. It names what your consequence actually did, whatever you meant it to do.
Skinner's ideas are most useful as a way of looking at routines, not as a rewards system. The two worked examples below apply shaping, the four consequences and schedules to one primary routine and one secondary routine. They are illustrations, not case studies.
A Year 4 class takes six minutes to move from the carpet to their tables for maths. The teacher defines the target action: seated, book open and date written within one minute. She does not wait for the whole routine before she responds. Shaping means reinforcing closer and closer steps towards the final action.
In week one she names every table seated within two minutes: "Blue table, seated with books open." In week two only tables with the date written earn the comment. Once most days hit the target, she stops naming tables every time and praises at random points instead. That is the move from a continuous schedule to a variable one.
She also checks function. One learner wanders more after being sent back to the carpet to try again. The walk back was an escape from the maths starter, so she changes it: he sits first, and she reads him the first question. If wandering falls over the next week, the escape was the reinforcer.
A Year 10 chemistry teacher sets a quiz every Friday. Revision happens on Thursday night and hardly at all on Monday, which is the fixed-interval pattern from the table. She switches to three short quizzes a week on unpredictable days, each worth nothing except feedback.
She also drops the whole-class merit count, which kept rewarding the same six learners. Instead she writes one quiet, specific comment on each set of corrections: "You fixed the balancing error in question 3 by counting atoms on both sides." The action she wants to see more of is correcting, so that is what now earns a response.
Both examples keep the check Skinner insisted on. Neither teacher assumes a consequence worked because it felt like a reward or a sanction. Each looks at what the action did over the following week.
Teachers describe the same routines on social media. The table below sets three of their posts beside what Skinner's research and the classroom studies say about each one.
Three routines teachers described on social media this term. The routine and notice columns are our summary of each post; the evidence column is Skinner's research and the classroom studies, not the post.
| Routine, as teachers describe it | What you would notice | What the evidence says |
|---|---|---|
| Say the credit out loud as you give specific praise, and log it after the lesson | ||
You can just narrate that you are giving a 'tick for you, David!' after the specific praise. |
The praise names the action and the tick is spoken at once, so nobody waits for the rewards platform. It answers a head of physics who found logging credits live too slow. | Specific praise produced more on-task behaviour than general praise in four Year 4 classes (Chalk and Bizo, 2004). The consequence the learner meets is the spoken praise straight after the action; the record is for the teacher. |
| Golden time earned in short daily slices, banked towards Friday | ||
Currently focusing on positives and rewarding with ‘golden time’ (short amounts each day until Friday). |
A Year 5 teacher with an unruly class pairs it with a focus on positives. Each day carries its own small reward instead of one payout at the end of the week. | A single Friday payout is a fixed-interval schedule, which in Skinner's work produced effort bunched near the deadline (Ferster and Skinner, 1957). Daily slices shorten the delay between the action and its consequence. |
| A sticker on every quiz that scores 100 | ||
They love stickers on their papers for 100s, I had students looking for 10 mins after the quiz today |
Learners spent ten minutes after the quiz looking for the sticker. The reward is visible and wanted, and it follows the score rather than the checking that produced it. | Whether a sticker reinforces anything is shown by what learners do on later quizzes, not by how excited they are (Skinner, 1953). Expected tangible rewards can reduce interest in tasks learners already enjoy (Deci, Koestner and Ryan, 1999). |
The routine teachers argue about most is the public points score. Read the posts below with the four consequences in mind: the same public rating can reinforce effort for one learner and punish it for another who dreads being watched. RJ in the last post means restorative justice.
Four short posts from social media on public points and praise. A parent and a school leader disagree about a points app; two teachers describe where praise and charts stop working.
By the time he came home all he cared about was his rating.
I guarantee you your child's teacher hates using it for behavior management. It's a great communication tool, however.
Those who received praise from me, are those who are really exceptional
going from clip chart to talk to kids over minor issues and trying to solve them went well but for totally off the rails behavior RJ as practiced is a disaster.
None of these posts is evidence that a routine works. Read them as the questions to test in your own room, using the function check above.
Skinner compared operant selection with selection in biology. Actions vary. Some consequences make particular forms more likely to occur again. Across repeated events, a person's behavioural repertoire changes over time.
No plan inside the person has to choose every response.
In 1981, Skinner described three levels of selection. Natural selection works across generations. Operant conditioning works across an individual's life. Cultural practices are selected across groups. The levels affect one another, but they are not the same process (Skinner, 1981).
This idea helps explain why one reward has no fixed meaning. Praise can increase one learner's participation. It can have no effect on another or make a third learner avoid public attention. The observed change, in context, decides how the consequence functioned.
Skinner applied his theory to instruction through teaching machines and programmed learning. He wanted each learner to respond actively. Material would appear in a careful order, with prompt knowledge of results. His 1954 paper argued that classroom feedback often came too late. Learners also had too few chances to respond (Skinner, 1954).
A teaching machine displayed a small item and required an answer. It then showed the learner whether the answer was correct. Later items depended on earlier steps.
Skinner saw this as a way to let learners work at a suitable pace. It also freed teachers for work that a machine could not do.
The design was not merely a reward dispenser. Correct responses and progress through a well-built sequence were central. Yet a sequence can still teach fragments without deep understanding. Good instruction also needs explanation, dialogue, varied examples and checks for transfer.
Modern adaptive software sometimes resembles programmed instruction. Giving points alone does not make it Skinnerian. The quality of the task, feedback and model of learning matter more than the presence of a score.
Skinner's 1957 book Verbal Behavior offered a functional account of language. It grouped verbal acts by the conditions that control them. It also considered their effects on a listener. A request and a label can sound alike while serving different functions.
Noam Chomsky's 1959 review argued that laboratory terms became too loose when applied to human language. He also questioned how the account explained new sentences. People can produce and understand them from limited experience (Chomsky, 1959).
Chomsky's review did not experimentally disprove all operant learning. Nor did one review make behaviourism disappear. Kenneth MacCorquodale later argued that parts of the critique misread Skinner's definitions. He also defended the purpose of a functional analysis (MacCorquodale, 1970).
The debate exposed a real boundary. Laboratory control can show how a consequence changes responding. A full theory of language must also explain structure, novelty and meaning. It must account for the social setting in which people speak.
Skinner's theory is strongest when behaviour and its setting can be defined and measured. It directs attention to what a person can do. It also asks about the conditions and how the pattern changes. This can replace a vague trait label with a testable question.
For example, “the learner lacks motivation” does not yet explain anything. A teacher can check whether the task is clear and whether the learner can start it. The teacher can also check what follows an attempt or delay. This may reveal a barrier in the task rather than a defect in the learner.
The method also corrects adult assumptions. If praise does not increase the target action, it was not reinforcement in that case. Under the later functional convention, a sanction is not punishment unless the target action decreases.
Skinner's work produced tools for repeated measurement and single-case experimental design. These methods can show change that a broad group average may hide. They still require care about consent, goals and generalisation. The measured action must also matter to the person.
Skinner's theory does not by itself explain every form of learning or settle which goals schools should pursue. A change in visible performance can have many causes. It may reflect understanding, imitation, recall, fear, guessing or a wish to please. The same response can have more than one history.
Bandura's social learning theory showed that people can learn through watching models and forming expectations, not only through direct consequences. Cognitivist learning theories examine attention, memory and mental representation. Social and cultural theories examine language, tools, identity and participation with others. Our comparison of behaviourism and constructivism sets out the other side of that debate.
Skinner warned that punishment can produce escape, avoidance and emotional effects. It may also fail to teach a useful alternative. This is not a reason to turn extinction or planned ignoring into a casual teacher technique. Such procedures depend on function, safety, consent and skilled oversight.
There is also an ethical question about control. All teaching shapes the setting in some way. Schools still need to ask who chose the goal and whether the learner benefits. They must protect dignity, communication and agency.
Skinner's educational legacy lies in active responding, prompt feedback, careful sequencing and the study of consequences. These ideas influenced programmed instruction. They also influenced behaviour analysis and debate about rewards, sanctions and learner autonomy.
Teachers can use the analytic questions without turning class into a token system. What can the learner do now? What is the next response?
Is feedback clear? Does the skill transfer? What barrier in the setting can be changed?
Detailed rewards, sanctions, routines and whole-school policy belong in the classroom behaviour management guide. Communication, sensory access and individual SEND support belong with the school's inclusive and safeguarding routes.
Choose one routine that is not working this week. Write down the exact action you want, watch the ten seconds after it for five lessons, and use the function check above to decide what to change. Then set Skinner beside the theories that explain what he could not, in our overview of learning theories for teachers.
These answers separate Skinner's broad theory from common classroom shorthand. They cover his main claim, radical behaviourism, the Skinner box, punishment and education. They also answer the questions teachers ask most about the four consequences, schedules and praise.
Skinner argued that what follows an action changes how often that action happens again. A consequence that makes the action more frequent is reinforcement; one that makes it rarer is punishment. Most of his research tracked how different patterns of consequences shape behaviour over time.
Radical behaviourism is Skinner's philosophy of a science of behaviour. It includes thoughts and feelings as private events within the natural world. It does not use them as inner agents that end the search for a fuller account.
The operant chamber let researchers measure how actions changed when their consequences changed. It provided controlled evidence about response rates and schedules. It did not show that all human learning can be reduced to a lever press.
Skinner studied punishment but stressed its limits. It can suppress behaviour without teaching an alternative and can produce escape, avoidance and emotional effects. This makes a punishment-led reading of Skinner misleading.
Skinner designed teaching machines and promoted programmed instruction. He stressed active responses, small steps, prompt feedback and work at a suitable pace. His ideas influenced behaviour analysis and later debates about educational technology.
Pavlov's dogs learned that a sound came before food, so the sound alone began to trigger salivation. Skinner's rats and pigeons learned that pressing or pecking produced food, so they pressed and pecked more. One is learning what predicts what; the other is learning what an action produces.
They are positive reinforcement, negative reinforcement, positive punishment and negative punishment. Positive means something is added and negative means something is taken away. Reinforcement makes the action more frequent afterwards and punishment makes it less frequent.
Skinner and Ferster described continuous reinforcement and four intermittent schedules: fixed ratio, variable ratio, fixed interval and variable interval. Ratio schedules count responses and interval schedules count time. Unpredictable schedules tend to keep an action going longest once reinforcement becomes rare.
Skinner's behaviourism holds that behaviour, including thinking and feeling, can be explained by a person's history of consequences in their surroundings. He called it radical behaviourism because it leaves nothing out as purely mental. In a classroom it becomes a practical habit: explain an action by what reliably follows it.
Praise may be too general, too rare, unwelcome in public or tied to a reward the learner does not value. In Skinner's terms it then fails to reinforce, or it even punishes. Check whether the praised action actually becomes more frequent, and change the praise if it does not.
The evidence on praise is summarised below: what the classroom studies found, how strong each one is and what a teacher can take from them.
These primary and scholarly sources support the account of Skinner's theory, research, educational work and major criticisms.