# Operant Conditioning — full text > The complete guide to how consequences shape behavior. Generated 2026-09-13 from https://operantconditioning.com. Each section below is one page; the URL is given under its title. # Operant Conditioning: The Complete Guide > Operant conditioning is learning through consequences. The definition, Skinner's four quadrants, schedules of reinforcement, examples, and how to apply it. - Source: https://operantconditioning.com/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *The complete guide* Operant conditioning is learning from consequences: a behavior becomes more likely if it’s followed by reinforcement and less likely if it’s followed by punishment. - **350 BC** — [Aristotle](https://operantconditioning.com/history/#aristotle-on-habit): "Pleasure and pain are also the standards by which we all, in a greater or less degree, regulate our actions." - **1898** — [Thorndike](https://operantconditioning.com/history/)'s law of effect - **1937** — [Skinner](https://operantconditioning.com/bf-skinner/) coins the term "operant conditioning" - **Now** — [Reinforcement learning](https://operantconditioning.com/applications/#economics) is used to improve AI ## What is operant conditioning? > **Definition** > > **Operant conditioning** is a form of learning in which the future frequency of a behavior is changed by the consequences that follow it. Behaviors followed by [reinforcement](https://operantconditioning.com/positive-reinforcement/) become more likely; behaviors followed by [punishment](https://operantconditioning.com/positive-punishment/) become less likely. > > The term was introduced by American psychologist [B. F. Skinner](https://operantconditioning.com/bf-skinner/) in 1937, building on Edward Thorndike's law of effect (1898). It is also called *instrumental conditioning*, and it is the foundation of applied behavior analysis, modern animal training, and most evidence-based habit-change methods.[1][2] **Pick a door.** - [I can't put my phone down.](https://operantconditioning.com/#the-machine-you-are-in) — Checking and scrolling run on two different schedules. One of them is the strongest of the five. - [I want a habit that actually sticks.](https://operantconditioning.com/#how-to-use-operant-conditioning-on-yourself) — How motivated you feel is the part you cannot arrange. Three other things you can. - [My kid keeps doing the thing.](https://operantconditioning.com/parenting/) — What the evidence says about time-out, rewards, and why consistency beats severity. - [My dog is ignoring me.](https://operantconditioning.com/dog-training/) — Marker timing, clicker mechanics, and the method that replaced dominance theory. - [Just explain it properly.](https://operantconditioning.com/#operant-conditioning-in-one-paragraph) — The definition, the four quadrants, the evidence, the five schedules, and where the theory runs out. Teaching this, or studying it? The quiz, glossary, diagrams, citation formats and discussion questions all live on [one page](https://operantconditioning.com/for-teachers/). ## The machine you're already in Sometime in the last hour you picked up your phone without deciding to. You weren't looking for anything in particular. You just checked. Then, probably, you kept scrolling. Those are two different mechanisms, and behavior analysts have a name for each. **Checking** pays off on a time basis — either something arrived since you last looked or it didn't, and pressing the button twice as often will not make messages arrive twice as fast. That is a [variable-interval schedule](https://operantconditioning.com/schedules-of-reinforcement/), and it produces a moderate, stubbornly steady rate of responding. It is why you check again four minutes later. **Scrolling** is the other kind. Each swipe is a response, and whether it pays depends on how many swipes you make rather than on how long you wait. That is a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/), and of the five basic schedules — catalogued, with dozens of combinations, in a 700-page book Ferster and Skinner published in 1957 — it is the one that produces the highest and steadiest rate of responding, and the one that keeps behavior going longest after the payoffs stop. It is what a slot machine is. It is what a loot box is. It is what an infinite feed is. None of which is a character flaw, and in most cases it is not addiction in any clinical sense. It is a schedule doing what seventy years of data say a schedule will do. Which is the genuinely useful part. A mechanism strong enough to keep you swiping against your own stated wishes is strong enough to aim somewhere else, and aiming it is a skill rather than a personality trait. The rest of this page is that mechanism from the ground up: what a consequence has to do to count as one, the four things it can be, the five ways to time it, and what happens when it stops. > **One honest caveat before you go further.** Schedules explain a great deal about behavior and not all of it. People also learn by watching someone else, by building a map of a situation before any reward arrives, and — most awkwardly for the theory — through language. Where this runs out, the page says so rather than papering over it. ## Watch: operant vs. classical conditioning in four minutes Start here if you are new to the topic. This TED-Ed lesson by Peggy Andover walks through Pavlov's dogs, then shows how Skinner's operant conditioning differs — behavior first, consequence second — and how reinforcement and punishment change what an animal (or a person) does next. *Figure: "The difference between classical and operant conditioning," a TED-Ed lesson by Peggy Andover (2013). Embedded from YouTube's privacy-enhanced player.* > **What to watch for** > > Two things the video makes vivid: in classical conditioning the animal is *passive* — the bell and the food arrive whether or not it does anything — while in operant conditioning the animal's own action is what produces the consequence. And "negative" never means "bad": it means something was *taken away*. The rest of this page builds on both ideas. ## Operant conditioning in one paragraph Every organism that can learn is constantly running the same experiment: *do something, notice what happens next, adjust.* A rat presses a lever and a food pellet drops, so it presses again. A toddler says "please" and gets the cookie, so "please" becomes a habit. You check your phone, a notification rewards you, and the checking becomes automatic. In each case the behavior **operates** on the environment (hence "operant") and the environment answers back with a consequence. Operant conditioning is the study of how those consequences select which behaviors survive and which fade away.[3] Three ideas do most of the work: - **Consequences are defined by their effect, not their appearance.** A "reward" that doesn't increase behavior is not a [reinforcer](https://operantconditioning.com/glossary/#reinforcer). A "punishment" that doesn't decrease behavior is not a punisher. This is a functional definition, and it is the single most important thing to understand about the whole field. - **Behavior is selected over time, the way evolution selects traits.** Skinner called this "selection by consequences." Reinforcement doesn't teach a rule; it shifts probabilities.[4] - **The unit of analysis is the three-term contingency:** an antecedent sets the occasion, a behavior occurs, a consequence follows. Learn to see the A-B-C pattern and you can read almost any behavior. ## How operant conditioning works Skinner's central insight was that behavior is not just triggered by what comes *before* it (as in Pavlov's reflexes), it is shaped by what comes *after* it. He formalized this as the **[three-term contingency](https://operantconditioning.com/glossary/#three-term-contingency)**, often written A → B → C:[3] Two more variables determine how much a consequence matters: - **[Contiguity](https://operantconditioning.com/glossary/#contiguity) (timing).** Consequences that follow within seconds are far more effective than delayed ones. In laboratory studies the strength of learning drops off steeply as the delay grows to even a few seconds — one reason a paycheck at the end of the month is a poor reinforcer for any specific behavior on a Tuesday morning.[5] - **[Contingency](https://operantconditioning.com/glossary/#contingency) (dependability).** The consequence must actually depend on the behavior. If food arrives whether or not the rat presses the lever, lever-pressing doesn't get learned — and, as Skinner showed in his famous "superstition" experiment, pigeons fed on a timer developed odd ritual behaviors that happened to precede the food by chance.[6] > **Operant vs. respondent behavior** > > Skinner distinguished **[respondent](https://operantconditioning.com/glossary/#respondent-behavior)** behavior (reflexes elicited by a prior stimulus — salivation, the eye-blink, the startle response) from **[operant](https://operantconditioning.com/glossary/#operant)** behavior (actions emitted by the organism and controlled by their consequences). Pavlov studied the first; Skinner studied the second. Most of what people mean by "behavior" — walking, talking, working, scrolling — is operant. [Full comparison ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## Inside the Skinner box The apparatus that made all of this measurable was Skinner's **operant conditioning chamber** — the "Skinner box." Thorndike had timed cats escaping from puzzle boxes one trial at a time; Skinner's innovation was a box the animal never had to leave, so it could respond whenever it liked and the *rate* of responding could be recorded continuously.[2] [The Skinner box in full: every part, and why it was built that way ›](https://operantconditioning.com/skinner-box/)  *An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a food pellet on any schedule; the light signals when pressing will pay off; the recorder draws responses over time.* A typical experiment runs in four steps. The rat is kept mildly hungry so food works as a reinforcer. It first learns that the click of the food dispenser means a pellet has arrived (a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer)). Then the experimenter [shapes](https://operantconditioning.com/shaping/) lever-pressing, reinforcing closer and closer approximations until the rat presses on its own. Finally the schedule is thinned — every press, then every fifth, then an unpredictable number — and the recorder shows how the pattern of responding changes. Pigeons peck a lit key instead of pressing a lever; the logic is identical. [More on Skinner and the box ›](https://operantconditioning.com/bf-skinner/) ## Run your own Skinner box At the top of this page you were the rat. Here you hold the pellet button. Reading about shaping is one thing; doing it is another. This lab puts a hungry, untrained rat in a chamber and hands you the pellet button. Magazine-train it, shape a lever press one approximation at a time, put the press on a schedule, then take the food away and watch extinction — all in about three minutes. The counters under the chamber track every pellet you deliver and every press the rat makes, from the first step on. The rat is a stylized model — its tendencies shift with what you reinforce, drift back when you don't, and follow the schedule patterns Ferster and Skinner documented — not a replay of real data. Real shaping takes longer and real rats are more surprising. ## The four quadrants: reinforcement and punishment Every consequence can be sorted along two questions. **Did the behavior increase or decrease?** (That tells you whether it was reinforcement or punishment.) **Was a stimulus added or removed?** (That tells you whether it was "positive" or "negative.") Crucially, in this vocabulary *positive* and *negative* mean plus and minus — added and removed — not good and bad. - **Positive reinforcement** — A pleasant stimulus is added after the behavior. The dog sits; the dog gets a treat. You finish a task; you feel a hit of satisfaction. (https://operantconditioning.com/positive-reinforcement/) - **Negative reinforcement** — An aversive stimulus is removed after the behavior. You buckle up; the seat-belt chime stops. You take an aspirin; the headache goes away. (https://operantconditioning.com/negative-reinforcement/) - **Positive punishment** — An aversive stimulus is added after the behavior. You touch a hot pan; it burns. You speed; you get a ticket. (https://operantconditioning.com/positive-punishment/) - **Negative punishment** — A pleasant stimulus is removed after the behavior. A teenager breaks curfew; the car keys are gone for a week. A player fouls; they sit out. (https://operantconditioning.com/negative-punishment/) > **The two-question test** > > Ask, in this order: **(1) Did the behavior become more or less likely?** More likely means reinforcement; less likely means punishment. **(2) Was something added or removed?** Added means positive; removed means negative. Those two answers name the quadrant every time. ### The four types of operant conditioning at a glance | Type | What happens after the behavior | Effect on the behavior | Example | | --- | --- | --- | --- | | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | A stimulus is added | Increases | A dog sits and gets a treat; sitting becomes more frequent | | [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) | A stimulus is removed | Increases | You buckle up and the seat-belt chime stops; buckling up becomes faster and more reliable | | [Positive punishment](https://operantconditioning.com/positive-punishment/) | A stimulus is added | Decreases | You touch a hot pan and get burned; touching hot pans becomes rarer | | [Negative punishment](https://operantconditioning.com/negative-punishment/) | A stimulus is removed | Decreases | A teenager breaks curfew and loses the car keys; breaking curfew becomes rarer | A fifth process, [extinction](https://operantconditioning.com/extinction/), is not a quadrant: the reinforcer that used to follow the behavior simply stops arriving, and the behavior fades. ### Which quadrant is it? An interactive check Almost everyone gets one pair backwards the first time, and it is nearly always *negative reinforcement* mistaken for *punishment*. Answer the two questions about any scenario and the tool will classify it. ### Five mistakes almost everyone makes 1. **Treating "negative reinforcement" as a polite word for punishment.** It is the opposite: negative reinforcement makes a behavior *more* likely by taking something unpleasant away. Taking an aspirin to end a headache is negative reinforcement of aspirin-taking. 2. **Reading "positive" as good and "negative" as bad.** They mean added and removed. A spray of water in a dog's face is *positive* punishment. 3. **Calling something a reinforcer because it seems nice.** A reinforcer is anything that increases the behavior it follows — it is defined by its effect. Scolding that a bored child finds attention-worthy is a reinforcer; a sticker a teenager finds embarrassing is not. 4. **Mixing up operant and classical conditioning.** If the key event comes *before* the response and the response is a reflex (salivating, flinching), it is classical. If the key event comes *after* a voluntary behavior, it is operant. 5. **Confusing extinction with punishment.** Extinction means the reinforcer simply stops arriving; nothing is added or taken away as a consequence. The behavior fades — usually after a brief burst — rather than being suppressed. ### Try three ## Reinforcement versus punishment: what the evidence says Skinner believed punishment was a poor way to change behavior, in part because an early experiment by his student W. K. Estes suggested that punishment only temporarily suppressed responding.[7] Later research complicated that picture: punishment *can* produce lasting decreases when it is immediate, consistent, and sufficiently intense from the outset.[8] But those same studies documented why practitioners still prefer reinforcement: - Punishment teaches what *not* to do without teaching what to do instead. Reinforcement builds a replacement. - Punishment tends to produce escape and avoidance — of the punisher as much as the behavior. (The child learns not to get caught.) - It can elicit aggression and emotional side effects, and it models the use of aversive control. - It works best at intensities and consistencies that are ethically or practically unavailable in most human settings. In parenting specifically, a large body of research links corporal punishment to worse, not better, long-term behavioral outcomes.[9] Modern applied behavior analysis therefore treats reinforcement-based procedures as the default and reserves punishment for narrow, supervised cases where reinforcement alone has failed and the behavior is dangerous.[10] [Reinforcement: the two types, kinds of reinforcers, and what makes it work ›](https://operantconditioning.com/reinforcement/) [Punishment: the two types, the side effects, and the alternatives ›](https://operantconditioning.com/punishment/) ## Schedules of reinforcement Once a behavior is learned, *how often* it gets reinforced changes both how fast the organism responds and how long the behavior persists when reinforcement stops. Ferster and Skinner catalogued these patterns in a 700-page 1957 volume, and the basic findings have held up for seventy years.[11] | Schedule | Reinforcer delivered… | Typical response pattern | Everyday example | | --- | --- | --- | --- | | **Continuous (CRF)** | after every response | Fast learning; fast extinction | A vending machine | | **Fixed ratio (FR)** | after a set number of responses | High rate with a pause after each reinforcer | Paid per piece; "buy 10, get 1 free" | | **Variable ratio (VR)** | after an unpredictable number of responses | Highest, steadiest rate; most resistant to extinction | Slot machines; social-media feeds | | **Fixed interval (FI)** | for the first response after a set time | "Scallop": slow after a reinforcer, accelerating as the interval ends | Checking the oven as the timer nears zero | | **Variable interval (VI)** | for the first response after an unpredictable time | Moderate, steady rate | Checking email | The practical rule: **use continuous reinforcement to build a behavior, then thin to an [intermittent schedule](https://operantconditioning.com/glossary/#intermittent-reinforcement) to make it durable.** The variable-ratio schedule is why gambling and infinite-scroll apps are so hard to quit — and why a behavior you reinforce only sometimes can end up stronger than one you reinforce every time. [Run the interactive schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) · [How organisms choose between schedules: the matching law ›](https://operantconditioning.com/matching-law/) ## Extinction, shaping, and stimulus control ### Extinction When a previously reinforced behavior stops producing reinforcement, it gradually declines. But not immediately: there is usually an **[extinction burst](https://operantconditioning.com/glossary/#extinction-burst)** — a temporary spike in the frequency, intensity, and variability of the behavior — before it fades. (Push the elevator button; nothing happens; you push it harder and faster before giving up.) Behavior that has been extinguished can also show **[spontaneous recovery](https://operantconditioning.com/glossary/#spontaneous-recovery)** after a rest period. [More on extinction ›](https://operantconditioning.com/extinction/) ### Shaping Complex behavior is rarely emitted fully formed, so it can't simply be reinforced. **Shaping** solves this by reinforcing *[successive approximations](https://operantconditioning.com/glossary/#successive-approximations)* — first any movement toward the lever, then touching it, then pressing it. Skinner used shaping to teach pigeons to play ping-pong; trainers use it to teach dolphins to jump through hoops; speech therapists use it to build words from sounds. [How shaping works ›](https://operantconditioning.com/shaping/) ### Stimulus control and the antecedent A behavior reinforced in one context and not in another comes under **[stimulus control](https://operantconditioning.com/glossary/#stimulus-control)**: it appears when the "discriminative stimulus" (S^D) is present and not otherwise. The rat presses only when the light is on. You swear with friends and not with your grandmother. This is the "A" in A-B-C, and it is the most under-used lever in self-improvement — changing the cue is often easier than willing a new response. [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) · [The ABC model ›](https://operantconditioning.com/abc-model/) ## Operant vs. classical conditioning The two great forms of associative learning are often confused. The clean distinction is *what gets associated with what*: | Aspect | Classical (Pavlovian) conditioning | Operant (instrumental) conditioning | | --- | --- | --- | | **Association** | Stimulus ↔ stimulus (bell → food) | Behavior ↔ consequence (press → food) | | **Behavior type** | Involuntary, reflexive (salivation, fear, nausea) | Voluntary, "emitted" (pressing, speaking, working) | | **Organism's role** | Passive; the stimulus is presented regardless | Active; the consequence depends on what it does | | **Timing of key event** | Stimulus comes *before* the response | Consequence comes *after* the response | | **Founders** | Ivan Pavlov (1890s–1927) | Edward Thorndike (1898); B. F. Skinner (1937–38) | In real life the two run together. The sound of the treat bag classically conditions excitement in a dog *and* operantly reinforces running to the kitchen. [Full comparison with examples ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## Can you spot it? Eight scenarios, instant explanations, no score kept. When you can call all eight without hesitating, you have it — and there is a [longer set of twenty](https://operantconditioning.com/quiz/). ## What happens in the brain Reinforcement has a physical address. In 1953 James Olds and Peter Milner found that a rat would press a lever thousands of times an hour for a pulse of electricity to its own brain, and in 1997 Wolfram Schultz and colleagues showed what the relevant neurons are doing: midbrain **dopamine** cells fire when a reinforcer is *better than expected*, fall silent when it is exactly as expected, and dip when an expected one fails to arrive.[21][16] That **reward prediction error** is the brain's teaching signal, and it explains why unpredictable reinforcers hold behavior so well and why a fully predictable one stops teaching. It is not a pleasure signal: Kent Berridge and Terry Robinson showed that animals without dopamine still *like* sugar but no longer *want* it.[22] [The neuroscience of operant conditioning, in depth ›](https://operantconditioning.com/neuroscience/) ## A brief history - **1898** — Thorndike's puzzle boxes Edward Thorndike times cats escaping from latched boxes. Escapes get faster, and he proposes the **[law of effect](https://operantconditioning.com/glossary/#law-of-effect)**: responses followed by satisfaction are "stamped in"; those followed by discomfort are "stamped out."[1] - **1930–1938** — Skinner's operant chamber At Harvard, as a graduate student and then a junior fellow, Skinner builds the apparatus later nicknamed the "Skinner box" and invents the cumulative recorder; at Minnesota he coins "operant" (1937) and lays out the science of operant behavior in *The Behavior of Organisms* (1938).[12][2] - **1948–1957** — Superstition, schedules, and Verbal Behavior The pigeon "superstition" study (1948), *Science and Human Behavior* (1953), *Schedules of Reinforcement* with Ferster (1957), and *Verbal Behavior* (1957) extend the analysis to society and language. - **1968** — Applied behavior analysis is born Baer, Wolf, and Risley publish the founding paper of ABA in the first issue of the *Journal of Applied Behavior Analysis*.[15] - **1997** — The dopamine connection Schultz, Dayan, and Montague show that midbrain dopamine neurons encode a *reward prediction error* — a biological implementation of the learning signal operant theory had assumed.[16] [The full history of operant conditioning ›](https://operantconditioning.com/history/) · [B. F. Skinner: life, work, and the Skinner box ›](https://operantconditioning.com/bf-skinner/) · [The original books, full text, in the library ›](https://operantconditioning.com/library/) ## Examples of operant conditioning in everyday life Once you know the pattern you see it everywhere: - **Positive reinforcement:** A barista is thanked warmly for a latte-art heart and starts making them on every cup. A student's essay earns praise and she writes more. Your phone lights up with a like. - **Negative reinforcement:** You clean the kitchen to end your partner's nagging. A student finishes homework early to escape the anxiety of a deadline. A driver takes the side street to avoid the traffic jam. - **Positive punishment:** A dog gets sprayed with water for jumping on the couch. A late invoice draws a fee. You bite into a moldy strawberry. - **Negative punishment:** A child loses screen time for hitting a sibling. A driver loses points from their license. A team member loses a project after missing deadlines. - **Schedules in the wild:** Slot machines (variable ratio). A bakery loyalty card (fixed ratio). Watching the clock in the last minutes of a shift (fixed interval). Fishing (variable interval). [50+ examples, sorted by quadrant and setting ›](https://operantconditioning.com/examples/) ## Operant conditioning in pop culture Two scenes worth watching a second time. Both are embedded from YouTube and load only when you press play. *Figure: **What to look for:** — every time Penny does something Sheldon likes, a chocolate appears — — [positive reinforcement](https://operantconditioning.com/positive-reinforcement/) — , delivered immediately, on a continuous schedule. When Leonard objects, Sheldon reaches for a spray bottle — that would be — [positive punishment](https://operantconditioning.com/positive-punishment/) — . Sheldon even names the procedure.* *Figure: **What almost everyone gets wrong here:** — Jim calls it Pavlov, but is it? Dwight's hand reaching out is a voluntary behavior that has been reinforced with a mint whenever the chime sounds — the chime is working as a — [discriminative stimulus](https://operantconditioning.com/glossary/#discriminative-stimulus) — , which makes the reaching — *operant* — . The dry mouth he notices is the — [classical](https://operantconditioning.com/operant-vs-classical-conditioning/) — part. Most real learning is both at once.* Clips are uploaded by third parties and may disappear; NBC hosts [the official Office clip](https://www.nbc.com/the-office/video/jims-pavlovian-prank-on-dwight-the-office/4141507). ## Key terms to know The twelve terms that do most of the work, defined the way behavior analysts define them. Hover or tap any underlined term anywhere on this page for its definition; the [full glossary](https://operantconditioning.com/glossary/) has 128 entries. - **Operant**: A class of behavior defined by its effect on the environment (what it accomplishes), not by its exact form. Lever-pressing with the left paw or the right paw is the same operant. - **Reinforcer**: Any consequence that increases the future frequency of the behavior it follows. *Positive* reinforcers are added; *negative* reinforcers are removed. - **Punisher**: Any consequence that decreases the future frequency of the behavior it follows. - **Primary vs. secondary reinforcer**: Primary reinforcers work without learning (food, water, warmth). Secondary (conditioned) reinforcers acquire their power by being paired with primary ones — money, praise, grades, a clicker. - **Three-term contingency**: Antecedent → Behavior → Consequence: the basic unit of analysis. [The ABC model ›](https://operantconditioning.com/abc-model/) - **Discriminative stimulus (S^D)**: A cue that signals a behavior will be reinforced. When a behavior reliably occurs in its presence and not otherwise, the behavior is under **stimulus control**. - **Schedule of reinforcement**: The rule for which responses get reinforced: continuous, fixed ratio, variable ratio, fixed interval, or variable interval. [Schedules ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Extinction**: Withholding the reinforcer that maintained a behavior, so the behavior declines — often after an **extinction burst**, a temporary increase. [Extinction ›](https://operantconditioning.com/extinction/) - **Shaping**: Building a new behavior by reinforcing successive approximations of it. [Shaping ›](https://operantconditioning.com/shaping/) - **Motivating operation**: A condition such as deprivation or satiation that changes how effective a reinforcer is. Food reinforces a hungry rat, not a full one. - **Premack principle**: A more probable behavior can reinforce a less probable one — "finish your homework, then you can play." [Premack ›](https://operantconditioning.com/premack-principle/) - **Law of effect**: Thorndike's 1898 principle that responses followed by satisfying consequences are strengthened and those followed by discomfort are weakened — the ancestor of operant conditioning. [The law of effect in depth ›](https://operantconditioning.com/law-of-effect/) *In practice* ## Where operant conditioning is used. The same three-term contingency runs a therapy session, a classroom, a family dinner, a factory floor, and a habit tracker. - [Applied behavior analysis](https://operantconditioning.com/applications/#aba): The clinical discipline built on operant principles, used in autism intervention, developmental disability support, and behavioral medicine. - [Education](https://operantconditioning.com/classroom/): Token economies, positive behavioral supports, immediate feedback, and the teaching machines Skinner pioneered in the 1950s. - [Parenting](https://operantconditioning.com/parenting/): Catch them being good, planned ignoring, time-out done correctly, and why consistency beats severity. - [Animal training](https://operantconditioning.com/dog-training/): Clicker training, marker signals, and the reinforcement-based methods that replaced dominance theory. - [Workplace & management](https://operantconditioning.com/applications/#workplace): Organizational behavior management, safety programs, and why annual reviews fail to change daily behavior. - [Habits & self-management](https://operantconditioning.com/habits/): Using antecedents, tiny behaviors, and immediate consequences to build habits that stick — on yourself. ## Criticisms and limitations An honest account includes what operant conditioning does *not* explain well. - **Biological constraints.** Organisms are not blank slates. The Brelands found that raccoons trained to deposit coins would instead "wash" them — an instinctive food-handling behavior that intruded even though it delayed reinforcement, a drift toward species-typical behavior they called [instinctive drift](https://operantconditioning.com/glossary/#instinctive-drift).[14] Reinforcement works with an animal's evolved tendencies, not against them. - **Cognition and language.** Chomsky argued that reinforcement cannot explain how children acquire grammar from limited input;[13] Tolman's rats learned mazes without obvious reinforcement, and Bandura's children learned by watching.[17] Operant learning is one powerful process among several, not a complete theory of mind. - **Rewards and intrinsic motivation.** Expected, tangible rewards for an activity someone already enjoys can reduce their interest once the rewards stop — the overjustification effect. A large meta-analysis found it; a rival meta-analysis found it small and narrow.[18][19] The fair reading: praise and feedback rarely undermine motivation; paying people for things they already love sometimes does. - **Ethics of control.** Skinner's *Beyond Freedom and Dignity* (1971) argued that since behavior is always controlled by its environment, we should design that environment deliberately. Critics saw a road to manipulation; the ethics codes of behavior analysis answer with consent, least-restrictive procedures, and the client's own goals.[10] What survived every critique is the core: the law of effect, schedule effects, extinction bursts, stimulus control, and shaping replicate across species and remain the working toolkit of clinicians, teachers, and trainers. Reinforcement learning — the branch of AI behind game-playing systems and the tuning of language models — is a mathematical descendant of the same ideas.[20] ## How to use operant conditioning on yourself The same contingency that trains a pigeon can be turned inward, and most self-improvement advice ignores two-thirds of it: it obsesses over motivation, which is not a term in the equation, and neglects antecedents and consequences, which are. The protocol is short. Attach the new behavior to a cue that already happens every day; shrink the behavior until it is almost embarrassing; deliver a small reinforcer within seconds, every time at first; then thin the schedule so the habit survives missed days. [The full seven-step protocol, with the evidence on how long habits take ›](https://operantconditioning.com/habits/) ### Design your own contingency Fill in the three terms for a habit you actually want. The tool checks each against the science — is the cue stable, is the behavior really a behavior, is the consequence immediate — and gives you a card to print. Nothing you type leaves your browser. [Full guide to building habits ›](https://operantconditioning.com/habits/) > **The app: Operant runs this loop for you.** A habit app for iPhone and Apple Watch from the publisher of this site. Each habit is set up as an antecedent, a behavior and a consequence — the three terms, not just the middle one. Free to download and try; a subscription unlocks the full app. [About the app](https://operantconditioning.com/app/) · [Download on the App Store](https://apps.apple.com/us/app/operant-behavior-change-app/id6802081776) ## Before you leave: can you answer these without opening them? **What is operant conditioning in simple terms?** Operant conditioning is learning from consequences. When a behavior is followed by something good (or the removal of something bad), it happens more often. When it is followed by something bad (or the loss of something good), it happens less often. The organism "operates" on its environment and the results shape what it does next. **Who discovered operant conditioning?** The underlying principle — the law of effect — was discovered by Edward Thorndike in 1898 through his puzzle-box experiments with cats. B. F. Skinner named it "operant" conditioning in 1937, developed the experimental methods to study it, and built the field around it beginning with *The Behavior of Organisms* in 1938. **What are the four types of operant conditioning?** Positive reinforcement (add something, behavior increases), negative reinforcement (remove something, behavior increases), positive punishment (add something, behavior decreases), and negative punishment (remove something, behavior decreases). "Positive" and "negative" refer to adding and removing a stimulus, not to whether the outcome is good or bad. **What is Skinner's theory of operant conditioning?** Skinner's theory is that behavior is selected by its consequences, much as species are selected by their environments. A behavior that is followed by reinforcement becomes more frequent; one followed by punishment or by no reinforcement at all becomes less frequent. He distinguished this *operant* behavior, which acts on the environment, from *respondent* behavior, the reflexes studied by Pavlov, and he showed that the schedule on which reinforcement arrives controls how fast and how persistently an organism responds. The theory deliberately explains behavior by its history of consequences rather than by inner states such as wants or intentions, which Skinner treated as behavior to be explained rather than as causes. [B. F. Skinner: the theory, the experiments, and the critiques ›](https://operantconditioning.com/bf-skinner/) **Why is it called "operant" conditioning?** Skinner chose the word in 1937 because the behavior *operates* on the environment to produce a consequence: the rat's press operates the lever, the lever delivers food, and the food changes future pressing. He contrasted operant behavior with respondent behavior, which is elicited by a stimulus that comes before it, as a puff of air elicits a blink. The older name, *instrumental* conditioning, makes the same point from the other side: the behavior is instrumental in producing the outcome. **What are the three components of operant conditioning?** The antecedent, the behavior, and the consequence, usually written A-B-C and called the three-term contingency. The antecedent is the situation or cue that sets the occasion for the behavior; the behavior is what the organism does; the consequence is what follows, and it is the consequence that changes how likely the behavior is next time. Every example on this site can be broken into those three parts. [The ABC model in depth ›](https://operantconditioning.com/abc-model/) **Is negative reinforcement the same as punishment?** No — this is the most common mistake in the whole subject. Negative reinforcement *increases* a behavior by removing something unpleasant (taking a painkiller to end a headache). Punishment *decreases* a behavior. "Negative" only means something was taken away. The two forms of negative reinforcement, escape and avoidance, have [their own page](https://operantconditioning.com/avoidance-learning/). **Which schedule of reinforcement is most resistant to extinction?** The variable-ratio schedule, in which reinforcement follows an unpredictable number of responses. Because the organism can never tell whether the next response will pay off, responding persists long after reinforcement has stopped. This is why gambling and social-media checking are hard to extinguish. ## References 1. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. See also Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 4. Skinner, B. F. (1981). Selection by consequences. *Science, 213*(4507), 501–504. 5. Grice, G. R. (1948). The relative effects of delay of reinforcement in the discrimination learning of rats. *Journal of Experimental Psychology, 38*(1), 1–16. See also Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139. 6. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 7. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 8. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 9. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 10. Behavior Analyst Certification Board. (2020). *Ethics Code for Behavior Analysts*. BACB. See also Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 11. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 12. Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279. 13. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58. 14. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 15. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97. 16. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 17. Tolman, E. C. (1948). Cognitive maps in rats and men. *Psychological Review, 55*(4), 189–208; Bandura, A. (1977). *Social Learning Theory*. Prentice-Hall. 18. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 19. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423. 20. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. 21. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427; Olds, J. (1958). Self-stimulation of the brain. *Science, 127*(3294), 315–324. 22. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? *Brain Research Reviews, 28*(3), 309–369. *Keep going* ## Go deeper. - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The most-used tool in the kit — and the most misunderstood word in it. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Fixed, variable, ratio, interval — with a live cumulative-record simulator. - [B. F. Skinner](https://operantconditioning.com/bf-skinner/): The man, the box, the pigeons, the controversies. - [50+ examples](https://operantconditioning.com/examples/): Everyday, classroom, workplace, and animal examples for every quadrant. - [Glossary](https://operantconditioning.com/glossary/): Every term from "abolishing operation" to "variable interval," defined. - [Can you spot it?](https://operantconditioning.com/quiz/): A 20-question quiz with instant explanations. Can you tell the quadrants apart? - [The library](https://operantconditioning.com/library/): Thorndike, Morgan, James, Yerkes, Darwin: the public-domain sources in full text, for readers and for AI. - [The Operant app](https://operantconditioning.com/app/): A habit app for iPhone and Apple Watch that runs the A-B-C loop for each habit. --- # Reinforcement: Definition, the Two Types, and What Makes It Work > Reinforcement is any consequence that makes a behavior more likely. Positive vs. negative reinforcement, types of reinforcers, what makes it work, and more. - Source: https://operantconditioning.com/reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · The hub* Half of operant conditioning is about making behavior more likely. This page defines reinforcement, sorts its two types and its kinds of reinforcers, and points you to the deep pages on each. > **Definition** > > **Reinforcement** is the process in which a consequence that follows a behavior makes that behavior *more likely* in the future. The consequence is called a **reinforcer**. Whether something is a reinforcer is decided by its effect on behavior, never by how it looks or feels.[1] > > A "reward" that does not increase the behavior it follows is not a reinforcer. A scolding that does increase it is. **In brief** - Reinforcement is any consequence that makes a behavior more likely; it is defined by that effect, never by how it looks or feels. - Positive reinforcement adds a stimulus and negative reinforcement removes one; both increase behavior, and the signs are arithmetic, not judgments. - Most failures of reinforcement are failures of implementation: late, non-contingent, too small, delivered to a satiated person, or on the wrong [schedule](https://operantconditioning.com/schedules-of-reinforcement/). ## The two types of reinforcement Both types make behavior more likely. They differ only in what happens to the stimulus: it is **added** (positive, +) or **removed** (negative, −). "Positive" and "negative" are arithmetic signs, not judgments. - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): A stimulus is **added** after the behavior and the behavior increases. The dog sits and gets a treat; you finish a task and feel the satisfaction. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): A stimulus is **removed** after the behavior and the behavior increases. You buckle up and the chime stops; you take an aspirin and the headache goes. > **The two-question test** > > Ask, in this order: **(1) Did the behavior become more or less likely?** More likely means reinforcement. **(2) Was something added or removed?** Added means positive; removed means negative. [Try it on any scenario with the quadrant checker ›](https://operantconditioning.com/#the-four-quadrants-reinforcement-and-punishment) ## Kinds of reinforcers | Kind | What it is | Examples | | --- | --- | --- | | Primary (unconditioned) | Reinforcing without any learning, because of biology | Food, water, warmth, sleep, sex, relief from pain | | Secondary (conditioned) | Acquires its power by being paired with other reinforcers | Praise, grades, a clicker, a "like," a check mark | | Generalized | A conditioned reinforcer paired with many others, so it works almost regardless of the person's state | Money, tokens, attention, approval | | Activity | The opportunity to do something more probable than the target behavior | Play after homework, the walk after the coffee — the [Premack principle](https://operantconditioning.com/premack-principle/) | | Social | Delivered by other people | Smiles, thanks, being listened to | | Automatic | Produced by the behavior itself, with no one delivering it | The feel of scratching an itch, the sound of your own humming | ## What makes reinforcement work - **Immediacy.** Reinforcers that arrive within seconds teach; delayed ones mostly strengthen whatever happened just before they arrived. Bridge a delay with a conditioned reinforcer (a word, a click, a check-off).[2] - **Contingency.** The reinforcer has to depend on the behavior. Reinforcement that arrives anyway teaches nothing — or teaches superstition.[3] - **Magnitude and quality.** Bigger and better reinforcers work better, with diminishing returns; a cut in magnitude is felt as a loss.[4] - **Motivating operations.** Deprivation makes a reinforcer stronger; satiation makes it weaker. Food does not reinforce a full rat. - **Schedule.** Reinforce every occurrence while a behavior is being learned, then thin to an intermittent schedule to make it durable. [Schedules of reinforcement ›](https://operantconditioning.com/schedules-of-reinforcement/) - **The alternatives.** A behavior's strength depends on what everything else pays. Enrich the alternatives and the behavior weakens without any punishment. [The matching law ›](https://operantconditioning.com/matching-law/) ## Reinforcement is not bribery, and not a reward A bribe is offered *before* a behavior to induce it; a reinforcer follows the behavior. A reward is something a person gives because it seems nice; a reinforcer is defined afterward, by the fact that the behavior went up. Most failures of "positive reinforcement" in classrooms and homes are failures of one of the five factors above — the reinforcer was late, non-contingent, too small, delivered to a satiated person, or on the wrong schedule — rather than failures of the principle. ## Key takeaways - A reinforcer is defined afterward, by the fact that the behavior went up. A "reward" that does not increase the behavior is not a reinforcer; a scolding that does increase it is. - Ask two questions, in order: did the behavior become more or less likely, and was something added or removed? More likely means reinforcement; added means positive and removed means negative. - Reinforcers come in kinds: primary (biological), conditioned (learned by pairing), generalized (money, tokens, approval), activity (a more probable behavior), social, and automatic (produced by the behavior itself). - Reinforcement depends on immediacy, contingency, magnitude, motivating operations, and schedule. Reinforce every occurrence while a behavior is being learned, then thin to an intermittent schedule to make it durable. - A bribe is offered before a behavior; a reinforcer follows it. A behavior's strength also depends on what everything else pays, so enriching the alternatives weakens it without any punishment. ### Check yourself **A teacher scolds a student every time he calls out, and calling out becomes more frequent. Was the scolding a punisher?** No. A consequence is classified by its effect on behavior, never by how it looks or feels. The behavior became more likely, so the scolding was a reinforcer: attention added after the behavior, which makes it positive reinforcement. **Food is delivered to a pigeon on a timer, regardless of what the pigeon does. Why might it end up repeating some odd movement?** Because reinforcement that arrives anyway strengthens whatever happened just before it arrived. Without contingency, the reinforcer does not teach the intended behavior; it teaches superstition. **You take an aspirin and your headache fades, and you reach for aspirin sooner next time. Is this reinforcement or punishment, positive or negative?** Reinforcement, because the behavior became more likely; negative, because a stimulus (the headache) was removed. Negative reinforcement is not punishment: the sign only says whether something was added or taken away. **A trainer wants to reinforce a dog's recall, but the treats are in a bag across the yard. What should happen the instant the dog arrives?** A conditioned reinforcer, such as a word or a click, should mark the behavior immediately. Reinforcers that arrive within seconds teach, while delayed ones mostly strengthen whatever happened just before they arrived; the conditioned reinforcer bridges the delay to the treat. **Explain it to a friend.** Explain what makes a consequence a reinforcer, using one example involving a person and one involving an animal. ## Go deeper - [Schedules](https://operantconditioning.com/schedules-of-reinforcement/): Fixed and variable, ratio and interval — with a live simulator. - [Shaping](https://operantconditioning.com/shaping/): Building behavior that does not yet exist by reinforcing approximations. - [Premack principle](https://operantconditioning.com/premack-principle/): When a behavior is the reinforcer. - [Avoidance learning](https://operantconditioning.com/avoidance-learning/): Negative reinforcement's strangest and most persistent form. - [The matching law](https://operantconditioning.com/matching-law/): How reinforcement divides behavior between options. - [Examples](https://operantconditioning.com/examples/): Fifty-plus scenarios sorted by quadrant. ## Frequently asked questions **What is reinforcement in psychology?** Any consequence that makes the behavior it follows more likely in the future. It is defined by its effect: if the behavior does not increase, no reinforcement occurred, whatever the consequence looked like. **What is the difference between positive and negative reinforcement?** Both increase behavior. Positive reinforcement adds a stimulus (a treat, praise); negative reinforcement removes one (a chime stops, a headache ends). "Negative" does not mean bad, and negative reinforcement is not punishment. **What are the types of reinforcers?** Primary (biological: food, water), secondary or conditioned (learned: praise, money, a clicker), generalized (paired with many reinforcers: money, tokens), activity (a preferred behavior), social, and automatic (produced by the behavior itself). **Why does reinforcement sometimes fail?** Usually because of timing (too late), contingency (delivered regardless of behavior), magnitude (too small), motivating operations (the person is satiated), or schedule (thinned too fast). The principle rarely fails; its implementation often does. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 3. Skinner, B. F. (1948). "Superstition" in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 4. Hutt, P. J. (1954). Rate of bar pressing as a function of quality and quantity of food reward. *Journal of Comparative and Physiological Psychology, 47*(3), 235–239. ## Related - [Punishment](https://operantconditioning.com/punishment/): The other half: making behavior less likely, and what the evidence says. - [Extinction](https://operantconditioning.com/extinction/): What happens when reinforcement stops. - [The complete guide](https://operantconditioning.com/): Everything on one page, in order. --- # Punishment: Definition, the Two Types, and What the Evidence Says > Punishment is any consequence that makes a behavior less likely. Positive vs. negative punishment, why it is not extinction, side effects, and alternatives. - Source: https://operantconditioning.com/punishment/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Punishment · The hub* The other half of operant conditioning: making behavior less likely. This page defines punishment precisely, separates its two types from extinction and from negative reinforcement, and gives an honest reading of whether it works. > **Definition** > > **Punishment** is the process in which a consequence that follows a behavior makes that behavior *less likely* in the future. The consequence is called a **punisher**. As with reinforcement, the definition is functional: a consequence that does not reduce the behavior is not punishment, however unpleasant it was meant to be.[1] > > In everyday speech "punishment" means a penalty someone intended. In behavior analysis it means a consequence that worked. **In brief** - Punishment is any consequence that makes the behavior it follows less likely; a consequence that does not reduce the behavior is not punishment. - Positive punishment adds a stimulus, negative punishment removes one; both decrease behavior, unlike [negative reinforcement](https://operantconditioning.com/negative-reinforcement/), which increases it. - It works only when immediate, consistent, intense from the outset, and paired with a reinforced alternative — conditions rarely met outside a laboratory. ## The two types of punishment - [Positive punishment](https://operantconditioning.com/positive-punishment/): A stimulus is **added** after the behavior and the behavior decreases. You touch the hot pan and it burns; you speed and get a ticket. - [Negative punishment](https://operantconditioning.com/negative-punishment/): A stimulus is **removed** after the behavior and the behavior decreases. A teenager misses curfew and loses the car keys; a player fouls and sits out. ## Three things punishment is not - **It is not negative reinforcement.** Negative reinforcement *increases* a behavior by removing something aversive. Punishment decreases behavior. The word "negative" is shared; the direction is opposite. [Why students mix these up ›](https://operantconditioning.com/negative-reinforcement/) - **It is not extinction.** In extinction nothing is added or removed as a consequence; the reinforcer that used to follow the behavior simply stops arriving. Ignoring a tantrum is extinction (if attention was the reinforcer); taking away the tablet for the tantrum is negative punishment. [Extinction ›](https://operantconditioning.com/extinction/) - **It is not defined by intent.** A parent who "punishes" a child by yelling, and finds the behavior increasing, has reinforced it — attention was the reinforcer. The behavior, not the parent, decides what the consequence was. ## Does punishment work? Skinner thought not. An early experiment by his student W. K. Estes found that punishing rats' lever pressing suppressed it only temporarily; when the punishment stopped, the pressing returned, and the total number of responses to extinction was about the same as for unpunished rats.[2] Skinner concluded that punishment merely suppresses, and argued against it for the rest of his career. Later research complicated the picture. Azrin and Holz's 1966 review, still the reference work, found that punishment can produce large and lasting decreases when it is **immediate**, **delivered every time**, **intense enough from the outset** (rather than escalated gradually), and combined with reinforcement of an alternative behavior.[3] A review of the applied literature three decades later reached the same conclusion for clinical settings.[4] Punishment, in other words, works — under conditions that are rarely met outside a laboratory, and at a price. ## The side effects - **It teaches nothing new.** Punishment says what not to do. The gap is filled by whatever else the person can do, which may be worse. Reinforcement of an alternative fills the gap deliberately. - **Escape and avoidance.** The punished organism learns to avoid the punisher as much as the behavior — the child learns not to get caught, the employee stops reporting mistakes. [Avoidance learning ›](https://operantconditioning.com/avoidance-learning/) - **Emotional and aggressive responding.** Aversive stimulation produces fear and aggression, sometimes toward bystanders.[3] - **It models aversive control.** People treated with punishment use punishment; the punisher's own behavior is negatively reinforced by the brief pause in the problem, which is why punishing tends to escalate. - **Corporal punishment specifically.** A 2016 meta-analysis of 75 studies covering more than 160,000 children found spanking associated with worse, not better, outcomes — more aggression, more antisocial behavior, more mental-health problems — with no evidence of benefit.[5] ## What to do instead | Alternative | How it works | Where it is explained | | --- | --- | --- | | Differential reinforcement of an alternative | Reinforce a behavior that does the same job for the person; the problem behavior loses its function | [Glossary: DRA](https://operantconditioning.com/glossary/#differential-reinforcement-of-alternative-behavior) | | Extinction | Identify the reinforcer maintaining the behavior and stop it arriving; expect a burst first | [Extinction](https://operantconditioning.com/extinction/) | | Antecedent change | Remove the cue, add a prompt, change the setting so the behavior is never triggered | [The ABC model](https://operantconditioning.com/abc-model/) | | Time-out done correctly | Brief, calm removal from reinforcement — which only works if the situation left was reinforcing | [Negative punishment](https://operantconditioning.com/negative-punishment/) | | Response cost | A predictable, proportionate loss of a token or privilege inside a system that also pays for good behavior | [Negative punishment](https://operantconditioning.com/negative-punishment/) | Modern applied behavior analysis treats reinforcement-based procedures as the default and reserves punishment for narrow, supervised cases where reinforcement alone has failed and the behavior is dangerous, with consent and the least restrictive procedure as ethical requirements.[6] ## Key takeaways - In behavior analysis, punishment means a consequence that worked: the behavior became less likely. The behavior, not the punisher's intent, decides what the consequence was. - Punishment is not negative reinforcement, which increases behavior by removing something aversive, and it is not extinction, in which the reinforcer simply stops arriving. Ignoring a tantrum is extinction if attention was the reinforcer; taking away the tablet for it is negative punishment. - Punishment can produce large, lasting decreases when it is immediate, delivered every time, intense enough from the outset, and combined with reinforcement of an alternative. Those conditions are rarely met outside a laboratory. - Its side effects: it teaches nothing new, it produces escape and avoidance of the punisher, it elicits fear and aggression, and it models aversive control. Spanking is associated with worse outcomes and no evidence of benefit. - Reinforcement-based procedures are the default. Differential reinforcement of an alternative, extinction, antecedent change, and time-out or response cost done correctly come first; punishment is reserved for narrow, supervised cases where the behavior is dangerous. ### Check yourself **A parent yells at a child for whining, and the whining increases over the following weeks. What kind of consequence was the yelling?** A reinforcer. Punishment is not defined by intent: the behavior, not the parent, decides what the consequence was. Attention was added and the behavior increased, so the yelling was positive reinforcement. **A child throws tantrums for attention. One parent stops responding to tantrums entirely; the other takes away the tablet each time. Which one is punishment?** Taking the tablet is negative punishment: a stimulus is removed as a consequence. Ignoring the tantrum is extinction, because nothing is added or removed; the reinforcer that used to follow the behavior simply stops arriving. Expect a burst before extinction works. **A manager waits until the monthly review to reprimand an employee for skipping safety steps, and mentions it only sometimes. Why is this unlikely to reduce the behavior?** Punishment produces lasting decreases only when it is immediate and delivered every time. A delayed, occasional reprimand meets neither condition, and it teaches nothing about what to do instead. **A teenager loses her phone for a day each time she misses curfew, and missed curfews decrease. A friend calls this negative reinforcement, because something was removed. What is wrong with that?** The direction. Negative reinforcement increases a behavior by removing something aversive; here a behavior decreased when something wanted was removed, which is negative punishment. The two share the word "negative" and nothing else. **Explain it to a friend.** Explain why behavior analysts reach for reinforcement before punishment, using an example from your own life in which a penalty failed to change what someone did. ## Go deeper - [Positive punishment](https://operantconditioning.com/positive-punishment/): Examples, the spanking evidence, alternatives. - [Negative punishment](https://operantconditioning.com/negative-punishment/): Time-out and response cost, done right. - [Extinction](https://operantconditioning.com/extinction/): The burst, spontaneous recovery, and how to use it. ## Frequently asked questions **What is punishment in psychology?** Any consequence that makes the behavior it follows less likely in the future. Positive punishment adds a stimulus (a burn, a fine); negative punishment removes one (losing privileges). It is defined by its effect on behavior, not by anyone's intention. **Is negative reinforcement a type of punishment?** No. Negative reinforcement increases behavior by removing something unpleasant; punishment decreases behavior. They share the word "negative" and nothing else. **Does punishment work?** It can reduce behavior when it is immediate, consistent, sufficiently intense from the start, and paired with reinforcement of an alternative. Those conditions are rarely met in daily life, and punishment carries side effects — avoidance, aggression, no new learning — which is why behavior analysts use reinforcement first. **What is the difference between punishment and extinction?** Punishment adds or removes a stimulus as a consequence. Extinction changes nothing except that the reinforcer stops coming. Both reduce behavior; extinction usually produces a temporary burst first. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), 1–40. 3. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 4. Lerman, D. C., & Vorndran, C. M. (2002). On the status of knowledge for using punishment: Implications for treating behavior disorders. *Journal of Applied Behavior Analysis, 35*(4), 431–464. 5. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 6. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. ## Related - [Reinforcement](https://operantconditioning.com/reinforcement/): The other half: making behavior more likely. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): The idea most often confused with punishment. - [The complete guide](https://operantconditioning.com/): Everything on one page, in order. --- # Positive Reinforcement: Definition, Examples, and How to Use It > Positive reinforcement adds a stimulus after a behavior to make it more likely. Definition, examples, types of reinforcers, what makes it work, and mistakes. - Source: https://operantconditioning.com/positive-reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Adding something* The workhorse of operant conditioning: add something after a behavior, and the behavior grows. Here is exactly how it works, what counts as a reinforcer, and why so many attempts at it fail. > **Definition** > > **Positive reinforcement** is the process in which a behavior is followed by the *addition* of a stimulus, and as a result the behavior becomes more frequent, intense, or likely in the future. The added stimulus is called a **positive reinforcer**. > > "Positive" means something is added (think plus sign), not that the outcome is pleasant — although positive reinforcers usually are. Whether a stimulus is a reinforcer is determined only by its effect on behavior.[1] **In brief** - Positive reinforcement adds a stimulus after a behavior, and the behavior becomes more frequent, intense, or likely in the future. - A reinforcer is defined by its effect on behavior, not by whether the giver thinks it is nice. - Reinforcement must be immediate and contingent; reinforce every occurrence while learning, then thin to an [intermittent schedule](https://operantconditioning.com/schedules-of-reinforcement/) so the behavior persists. ## How positive reinforcement works The sequence is always the same: an [antecedent](https://operantconditioning.com/abc-model/) sets the occasion, a behavior occurs, and immediately afterward a stimulus appears that was not there before. If the behavior then happens more often under similar conditions, positive reinforcement has taken place. - A rat presses a lever → a food pellet drops → lever-pressing increases. - A child says "thank you" → a parent smiles and says "you're welcome" → "thank you" increases. - You post a photo → likes appear → posting increases. Notice that in each case the consequence is delivered *because of* the behavior (contingency) and *right after* it (contiguity). Remove either and the effect weakens sharply. A bonus paid in December for effort in March reinforces very little of that effort; the pellet that arrives ten seconds after the press teaches the rat almost nothing about pressing.[2] > **The functional definition, again** > > A reinforcer is not a "reward." A reward is something the giver thinks is nice. A reinforcer is something that demonstrably increases the behavior it follows. Praise that a teenager finds embarrassing is not a reinforcer for them. Attention — even scolding — often *is* a reinforcer for a child who gets little of it. You find out what reinforces a behavior by watching what happens to the behavior. ### Check any scenario in two questions Not sure whether a consequence is this quadrant or a neighbor? Answer the two questions and the checker names it. ## Examples of positive reinforcement | Setting | Behavior | Stimulus added | Result | | --- | --- | --- | --- | | Home | Toddler uses the potty | Sticker and enthusiastic praise | Uses the potty more often | | Home | Child clears the table without being asked | "That was really helpful — thank you." | Clears the table more often | | Classroom | Student raises hand instead of calling out | Teacher calls on them | Hand-raising increases | | Classroom | Class transitions quietly | Marble added to the class jar (token) | Quiet transitions increase | | Workplace | Employee submits a report early | Public recognition in the team meeting | Early submissions increase | | Workplace | Salesperson closes a deal | Commission | Closing behavior increases | | Dog training | Dog sits on cue | Click, then a treat | Sitting on cue increases | | Dog training | Dog returns when called | Play with a favorite toy | Recall improves | | Technology | User opens the app | New content, likes, badges | App-opening increases | | Health | Patient attends a treatment session | Voucher (contingency management) | Attendance increases | | Self-management | You put on running shoes at 6:30 a.m. | Marking the habit done; a satisfying check | Shoes go on more mornings | | Nature | Bee visits a flower | Nectar | Visits to that flower type increase | Not every example of "adding something nice" is positive reinforcement. If a parent gives a child candy to *stop* a tantrum in the store, the candy positively reinforces the tantrum (and the end of the tantrum negatively reinforces the parent's candy-giving). Both people learned something; neither learned what they intended. ## Types of positive reinforcers ### Primary (unconditioned) reinforcers Stimuli that reinforce without any learning history because they relate to biological needs: food, water, warmth, sexual contact, relief from pain, and — for social species — physical contact. Their effectiveness depends on deprivation: food is a powerful reinforcer for a hungry rat and a weak one for a full one. ### Secondary (conditioned) reinforcers Neutral stimuli that acquire reinforcing power by being paired with existing reinforcers. Money is the classic example; so are grades, praise, a clicker's sound, and a green checkmark. Conditioned reinforcers are what make delayed real-world consequences workable: the click bridges the gap between the dog's sit and the treat that follows a second later.[3] ### Generalized reinforcers Conditioned reinforcers paired with *many* different reinforcers, so they work regardless of the organism's current state. Money, tokens, and social approval are generalized reinforcers; that is why token economies work in classrooms and psychiatric wards.[4] ### Social, tangible, activity, and sensory reinforcers Practitioners often sort reinforcers by category: **social** (attention, praise, a smile), **tangible** (a sticker, a toy, a bonus), **activity** (getting to play, choosing the music), and **sensory** (a pleasant sound or texture). The **Premack principle** is the rule for activity reinforcers: a higher-probability behavior can reinforce a lower-probability one. "You can play video games after you finish your homework" is Premack in action.[5] [The Premack principle in depth ›](https://operantconditioning.com/premack-principle/) ## What makes positive reinforcement effective 1. **Immediacy.** Deliver the reinforcer within seconds. If you can't, deliver a conditioned reinforcer (a click, a word, a checkmark) immediately and the real one later. 2. **Contingency.** The reinforcer follows the behavior and only the behavior. Non-contingent goodies are nice, but they don't teach. 3. **Magnitude and quality.** Bigger and better reinforcers work better up to a point: rats press faster for larger and for sweeter food rewards, with diminishing returns as amount grows.[8] But the effect is relative. Crespi found that rats switched from a large reward to a small one ran *slower* than rats that had only ever had the small one — a negative contrast effect — and rats shifted upward briefly ran faster than rats always given the large amount.[9] Cutting a reinforcer is felt as a loss, not as a smaller gain. And small, immediate reinforcers usually beat large, delayed ones. 4. **Motivating operations.** Deprivation makes a reinforcer stronger; satiation makes it weaker. Training a dog after dinner with kibble is a losing game. 5. **Schedule.** Reinforce every occurrence while the behavior is being learned, then shift to an intermittent schedule so it persists. [Schedules of reinforcement ›](https://operantconditioning.com/schedules-of-reinforcement/) 6. **Individualization.** Reinforcers are personal. What works is discovered by preference assessment and by watching the data, not assumed. ## Positive reinforcement vs. negative reinforcement vs. bribery Three things get confused constantly: | Aspect | Positive reinforcement | Negative reinforcement | Bribery | | --- | --- | --- | --- | | **What happens** | Stimulus is added after the behavior | Stimulus is removed after the behavior | Reward is offered *before* the behavior, often to stop misbehavior | | **Effect** | Behavior increases | Behavior increases | Usually reinforces the misbehavior that prompted the bribe | | **Example** | Praise after homework is done | Nagging stops when homework is done | "If you stop screaming, I'll buy you the toy" | The difference between reinforcement and bribery is timing and target. Reinforcement follows the desired behavior. Bribery precedes it and is typically triggered by an undesired behavior, so the undesired behavior is what gets strengthened. [More on negative reinforcement ›](https://operantconditioning.com/negative-reinforcement/) ## Common mistakes - **Reinforcing too late.** "Great job on the presentation" three days later is pleasant, not reinforcing. - **Reinforcing the wrong behavior.** Attention given during a tantrum reinforces the tantrum. Reinforcement must be contingent on the behavior you want. - **Assuming the reinforcer.** Stickers do not reinforce every child; public praise does not reinforce every employee. - **Reinforcing every time forever.** Continuous reinforcement builds behavior quickly but leaves it fragile. Thin the schedule. - **Reinforcing outcomes you can't control.** "Lose weight" can't be reinforced in the moment; "walked after lunch" can. - **Pairing praise with criticism.** "Good, but…" turns the praise into a warning signal. ## Does positive reinforcement undermine intrinsic motivation? This is the most serious research-based objection, and the honest answer is "sometimes, under specific conditions." A meta-analysis of 128 studies found that *expected, tangible* rewards for doing an already-interesting activity reduced free-choice engagement with it afterward; verbal praise and unexpected rewards did not, and informational feedback generally increased intrinsic motivation.[6] A competing meta-analysis found the negative effect small and limited to a narrow set of conditions.[7] The practical guidance that follows from both: - Use social and informational reinforcers (specific praise, feedback on progress) freely. - Use tangible rewards for behaviors the person would not otherwise do, and fade them as the behavior contacts natural reinforcers. - Avoid paying people for things they already love doing, especially with rewards contingent merely on engaging rather than on quality. ## How to use positive reinforcement: a six-step protocol 1. **Define the behavior so a camera could see it.** "Raises her hand before speaking," not "is respectful." If you cannot count it, you cannot reinforce it. 2. **Find a reinforcer that works for this individual.** Watch what the person chooses when free, ask, or test. A reinforcer is proven by its effect, not by your intentions; a sticker that a teenager finds embarrassing is not one. 3. **Deliver it immediately and every time — at first.** Within seconds, and after every occurrence, until the behavior is reliable. If the real reinforcer must wait, mark the behavior instantly with a conditioned reinforcer: a word, a click, a check. 4. **Make it contingent and only contingent.** The reinforcer follows the behavior and nothing else. Free access to the same reinforcer at other times drains its power. 5. **Thin the schedule.** Once the behavior is steady, reinforce it only sometimes, unpredictably. Intermittent reinforcement is what makes the behavior survive days when no one is watching. [How to thin a schedule ›](https://operantconditioning.com/schedules-of-reinforcement/) 6. **Measure, and fade to natural reinforcers.** Count the behavior before and after. Then hand the job to consequences the world already provides — the finished essay, the dog's walk, the satisfaction of the check mark — so the behavior no longer depends on you. ### Praise, done properly Praise is the cheapest positive reinforcer available, and most of it is wasted. Jere Brophy's analysis of teacher praise found that it changed behavior only when it was *contingent* (delivered for the behavior, not as a reflex), *specific* (named what was done well), and *credible* (sincere, and not inflated).[10] A later review added a fourth condition: praise for effort and strategy supports motivation, while praise for ability or for merely finishing can undermine it.[11] "You kept the paragraph to one idea — that made it clear" reinforces; "great job" does not. [Praise, token economies, and the Good Behavior Game in the classroom ›](https://operantconditioning.com/classroom/) ### A worked example A second-grade teacher wants a student, Maya, to start her worksheet without a reminder. **Behavior:** pencil on paper within one minute of the worksheet landing on her desk. **Reinforcer:** the teacher notices that Maya lights up when asked to hand out materials, so "hand out the next set" becomes the consequence. **Delivery:** the moment the pencil moves, a quiet "you started on your own — you're handing out the readers at ten." Every time, for a week. **Thinning:** in week two, the job goes to two starts out of three, then to unpredictable ones. **Result:** starts rose from one in five worksheets to almost all of them, and by the end of the month the teacher's occasional nod was enough. Every step is on this page; none of it required a sticker chart. ### Quick check ## Positive reinforcement in practice Positive reinforcement is the default procedure of [applied behavior analysis](https://operantconditioning.com/applications/#aba), the core of reward-based [dog training](https://operantconditioning.com/dog-training/), the basis of classroom systems like PBIS and token economies, the engine of contingency management in addiction treatment, and the mechanism behind most successful [habit-formation](https://operantconditioning.com/habits/) methods. The through-line: identify the behavior, find a reinforcer that actually works for that individual, deliver it immediately and contingently, and thin the schedule over time. ## Key takeaways - "Positive" means a stimulus is added, not that the outcome is pleasant. A reinforcer is not a reward: it is defined only by its effect on the behavior it follows. - Reinforcers are personal and are discovered by watching the behavior, not assumed. Praise a teenager finds embarrassing is not a reinforcer; attention during a tantrum, even scolding, often is. - Contingency and contiguity are both required: the reinforcer arrives because of the behavior and right after it. Bribery is offered before the behavior, usually to stop misbehavior, so the misbehavior is what gets strengthened. - Reinforce every occurrence while the behavior is being learned, then thin to an intermittent schedule and fade to the natural reinforcers the world already provides. - Expected, tangible rewards for an activity a person already enjoys can reduce intrinsic motivation. Specific praise, unexpected rewards, and informational feedback do not. **Explain it to a friend.** Explain why a reinforcer is not the same thing as a reward, using an example from your own week and without using the word "reward" itself. ## Frequently asked questions **What is positive reinforcement in simple terms?** Adding something after a behavior so the behavior happens more often. A dog sits, gets a treat, and sits more often. The treat is the positive reinforcer. **What is an example of positive reinforcement?** A child puts their toys away and a parent says, "Thank you for cleaning up — that was a big help." If the child cleans up more often afterward, the praise was a positive reinforcer. Other examples: a paycheck for hours worked, a like on a post, a treat for a dog's trick. **Is positive reinforcement the same as a reward?** Not exactly. A reward is something intended to be pleasant. A positive reinforcer is defined by its effect: it must actually increase the behavior it follows. Many rewards fail to reinforce, and some unpleasant things (like scolding, which is attention) do reinforce. **What is the difference between positive and negative reinforcement?** Both increase behavior. Positive reinforcement adds a stimulus (a treat appears). Negative reinforcement removes one (an alarm stops). "Positive" and "negative" mean added and removed, not good and bad. **Who developed positive reinforcement?** The principle traces to Edward Thorndike's law of effect (1898). B. F. Skinner developed the concept of reinforcement, the term "positive reinforcement," and the experimental science around it, beginning with *The Behavior of Organisms* (1938). **Is positive reinforcement effective for adults?** Yes. Organizational behavior management uses it to improve safety and productivity, contingency management uses it in addiction treatment with strong evidence, and it is the mechanism behind effective self-management and habit apps. Adults simply have more complex reinforcers (money, status, autonomy, feedback) than treats. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139. 3. Williams, B. A. (1994). Conditioned reinforcement: Experimental and theoretical issues. *The Behavior Analyst, 17*(2), 261–285. 4. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 5. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 6. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 7. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423. 8. Hutt, P. J. (1954). Rate of bar pressing as a function of quality and quantity of food reward. *Journal of Comparative and Physiological Psychology, 47*(3), 235–239. 9. Crespi, L. P. (1942). Quantitative variation of incentive and performance in the white rat. *American Journal of Psychology, 55*(4), 467–517. 10. Brophy, J. (1981). Teacher praise: A functional analysis. *Review of Educational Research, 51*(1), 5–32. 11. Henderlong, J., & Lepper, M. R. (2002). The effects of praise on children's intrinsic motivation: A review and synthesis. *Psychological Bulletin, 128*(5), 774–795. ## Related - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): Removing something to strengthen behavior — and why it isn't punishment. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): How often to reinforce, with a live simulator. - [Reinforcement: the hub](https://operantconditioning.com/reinforcement/): Both types, the kinds of reinforcers, and what makes reinforcement work. --- # Negative Reinforcement: Definition, Examples, and Why It Isn't Punishment > Negative reinforcement removes an aversive stimulus so a behavior increases. Definition, 17 examples, escape vs. avoidance, and why it is not punishment. - Source: https://operantconditioning.com/negative-reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Taking something away* Not punishment — relief. Negative reinforcement is what keeps you buckling your seat belt, hitting snooze, and avoiding what you fear, and it is the most misread idea in operant conditioning. > **Definition** > > **Negative reinforcement** is the process in which a behavior is followed by the *removal* (or prevention) of an aversive stimulus, and as a result the behavior becomes more frequent or more likely in the future. The stimulus that is removed is called a **negative reinforcer**. > > "Negative" means something is subtracted (think minus sign), not that the outcome is bad. Like all reinforcement, negative reinforcement *strengthens* behavior; that is what separates it from punishment, which weakens it.[1] **In brief** - Negative reinforcement removes or prevents an aversive stimulus after a behavior, and the behavior becomes more frequent or more likely. - "Negative" means subtracted, not bad: negative reinforcement strengthens behavior, which is exactly what separates it from punishment. - Escape ends an aversive stimulus that is present; [avoidance](https://operantconditioning.com/avoidance-learning/) prevents one that would have come, and avoidance is extremely persistent. ## How negative reinforcement works In [positive reinforcement](https://operantconditioning.com/positive-reinforcement/), the world is quiet before the behavior and something good appears after it. In negative reinforcement the order is reversed: something unpleasant is already present (or about to arrive), the behavior makes it go away, and the going-away is what teaches. The reinforcer is the *change* — from aversive to not-aversive. - A rat is receiving a mild shock → it presses a lever → the shock stops → pressing increases. - The seat-belt chime is beeping → you buckle up → the chime stops → you buckle faster next time. - Your head is pounding → you take an aspirin → the pain fades → aspirin-taking increases. The same rules apply as for any operant in the [antecedent–behavior–consequence](https://operantconditioning.com/abc-model/) loop: the relief must be *contingent* on the behavior and *immediate*. A chime that stops on its own after 30 seconds teaches you to wait, not to buckle. Skinner devoted whole chapters of *Science and Human Behavior* to aversive control, and argued that governments, schools, and other institutions lean on it far too much.[1] > **The two-question test** > > To classify any consequence, ask only two things. **Did the behavior go up or down?** Up is reinforcement; down is punishment. **Was something added or taken away?** Added is "positive"; taken away is "negative." Negative reinforcement is "behavior went up, something was taken away." How unpleasant the stimulus felt is not part of the test — if removing it increased the behavior, it was aversive by definition. ### Check any scenario in two questions Not sure whether a consequence is this quadrant or a neighbor? Answer the two questions and the checker names it. ## Escape and avoidance: the two kinds of negative reinforcement In **escape**, the aversive stimulus is already present and the behavior ends it: the shock is on, the rat presses, the shock goes off; the baby is crying, the parent picks her up, the crying stops. Escape is learned quickly because the relief is felt directly.  *Escape ends an aversive stimulus that is present; avoidance prevents one that would have come. Both increase the behavior, so both are negative reinforcement.* In **avoidance**, the behavior comes first and prevents or postpones the aversive stimulus. Dogs trained with a warning signal in a shuttle box learned to jump the barrier within seconds of the signal and kept jumping for hundreds of trials without ever being shocked again.[2] That persistence is the puzzle of avoidance — how can the *absence* of a shock reinforce anything? Mowrer's two-factor answer was that the signal first acquires fear by classical conditioning and the response is then reinforced by escape from that fear; later work showed that avoidance can be learned with no signal at all (Sidman avoidance) and that a reduction in the overall rate of aversive events can itself be the reinforcer.[3][4][5][6] [Avoidance learning in depth: the paradox, the theories, and learned helplessness ›](https://operantconditioning.com/avoidance-learning/) | Aspect | Escape | Avoidance | | --- | --- | --- | | **Aversive stimulus** | Already present | Not yet present; the behavior prevents or postpones it | | **What the behavior does** | Ends it | Prevents it (or turns off its signal) | | **Example** | Taking a painkiller for a headache | Leaving early to miss the traffic | | **How it is learned** | Directly; the relief is felt | Usually grows out of escape as the response comes earlier | | **Persistence** | Fades when the behavior stops working | Extremely persistent; the threat is never tested | ## Examples of negative reinforcement In every row, an aversive stimulus is present or imminent, a behavior removes or prevents it, and the behavior becomes more likely. | Setting | Aversive stimulus | Behavior | Result | | --- | --- | --- | --- | | Home | Partner nagging about the dishes | Do the dishes; nagging stops | Dishes get done sooner (and the nagging is reinforced too) | | Home | Baby crying | Parent picks the baby up; crying stops | Parent picks up faster next time | | Classroom | Looming pop quiz | Class works quietly all week; quiz is cancelled | Quiet work increases | | Classroom | Hard math worksheet | Student acts out; is sent to the hall | Acting out during math increases (escape-maintained) | | Workplace | Manager's reminder emails | Submit the report; emails stop | Reports go in sooner | | Workplace | Machine noise | Put on ear protection | Ear protection worn more reliably | | Driving | Seat-belt chime | Buckle up; chime stops | Buckling is faster and more reliable | | Driving | Predicted traffic jam | Take the side street | Side street becomes habitual (avoidance) | | Technology | Alarm blaring | Hit snooze; alarm stops | Snoozing increases | | Health | Headache | Take an aspirin; pain fades | Aspirin at the first twinge increases | | Health | Social anxiety at a party | Leave early; anxiety drops | Leaving (then not going) increases | | Dog training | Leash pressure | Dog steps toward handler; pressure released | Dog follows pressure more readily | | Dog training | Mail carrier at the door | Dog barks; carrier leaves (as always) | Barking at the door increases | | Weather | Rain | Open an umbrella | Umbrella use increases; carrying one becomes avoidance | | Self-management | Dread of writing the essay | Clean the apartment instead; dread fades | Procrastination increases | | Self-management | Dread of writing the essay | Write one sentence; dread lifts | Starting increases — the healthy version of the same loop | | Nature | Midday heat | Lizard moves into shade | Shade-seeking increases | Notice how many rows involve *two* people reinforcing each other. When a baby cries and a parent picks her up, the parent's behavior is negatively reinforced (the crying stops) and the baby's crying is positively reinforced (contact arrives). These reciprocal traps explain an enormous amount of family life, and they are covered below. [More examples for every quadrant ›](https://operantconditioning.com/examples/) ## Negative reinforcement vs. punishment This is the confusion the quadrant vocabulary exists to prevent. "Negative" sounds like "bad," so people hear "negative reinforcement" and picture a scolding. But a scolding that reduces a behavior is *positive punishment* (something added, behavior down). Negative reinforcement always *increases* behavior. | Aspect | Positive reinforcement | Negative reinforcement | Positive punishment | Negative punishment | | --- | --- | --- | --- | --- | | **Stimulus** | Added | Removed | Added | Removed | | **Effect on behavior** | Increases | Increases | Decreases | Decreases | | **Kind of stimulus** | Wanted | Unwanted (aversive) | Unwanted (aversive) | Wanted | | **Before the behavior** | Reinforcer absent | Aversive present or imminent | Aversive absent | Reinforcer present | | **Example** | Dog sits → treat | Dog sits → leash pressure released | Dog jumps → sprayed with water | Dog jumps → person turns away | | **Page** | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | This page | [Positive punishment](https://operantconditioning.com/positive-punishment/) | [Negative punishment](https://operantconditioning.com/negative-punishment/) | The two often use the *same* aversive stimulus for opposite jobs. A shock that stops when the rat presses negatively reinforces pressing; a shock that starts when the rat presses punishes it. A parent who nags until the room is clean is using negative reinforcement; one who starts nagging on seeing the mess is using positive punishment. ## The negative reinforcement trap: how it maintains problem behavior Reinforcement does not care whether anyone wants the behavior it strengthens. Negative reinforcement quietly maintains a large share of the behaviors people go to therapists, behavior analysts, and trainers to get rid of. ### Escape-maintained behavior In applied behavior analysis, a **functional analysis** tests which consequence maintains a problem behavior by presenting each candidate in turn.[7] In the *demand* condition, tasks are presented and briefly withdrawn if the problem behavior occurs; behavior that peaks there is **escape-maintained**. In the largest published series of functional analyses of self-injury, escape from demands was the single most common function identified.[8] The everyday version needs no clinic. A child is asked to put on shoes; she screams; the parent, late for work, drops the demand and carries her to the car. Screaming was negatively reinforced (the demand vanished) and so was giving up (the screaming stopped). Gerald Patterson called this the **coercive family process**: parent and child train each other, in escalating rounds, to use aversive behavior to get their way.[9] The remedy is not more punishment but **escape extinction** — following through so the behavior no longer works — plus reinforcement for compliance and easier, better-prompted demands.[10] [How escape extinction differs from ignoring ›](https://operantconditioning.com/extinction/#extinction-is-not-ignoring) ### Avoidance and anxiety Two-factor theory turned out to describe human anxiety well. A person who fears elevators takes the stairs; the anxiety drops; stair-taking is negatively reinforced. Because they never ride the elevator, they never learn that nothing bad happens, so the fear never extinguishes. The same loop maintains compulsions (checking relieves doubt), social avoidance (leaving relieves dread), and panic-related avoidance.[11] Exposure-based treatment is built on *blocking the avoidance response* long enough for fear to fade — extinction applied to a behavior that negative reinforcement has protected for years. ### Procrastination Procrastination is escape, not laziness. The thought of the task is aversive, doing something else makes the feeling go away, and that relief reinforces the diversion — researchers describe it as short-term mood repair at the expense of long-term goals.[12] The fix follows from the analysis: make the first step so small that the relief of having started outweighs the dread of starting, so that *starting* becomes the escape response. That is why "open the document and write one sentence" works when "write the essay" does not, and it is the logic behind the [tiny-behavior approach to habits](https://operantconditioning.com/habits/). ### Nagging, whining, and the snooze button Nagging persists because it works intermittently, and intermittent reinforcement builds the most durable behavior of all (see [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/)). Whining persists the same way: hold out four times, give in on the fifth, and you have trained a five-whine behavior. The snooze button is the purest case — snooze removes the alarm instantly, every morning, which is why the alarm across the room is the only one that reliably works. ## How to use negative reinforcement ethically Negative reinforcement is not inherently unethical — the seat-belt chime saves lives. But it requires an aversive stimulus to be present, so it carries the side effects of aversive control (stress, escape, avoidance), and if you have to *create* the aversive yourself, you become the thing the learner wants to get away from.[13] 1. **Prefer positive reinforcement.** Use negative reinforcement when an aversive is already unavoidably present — pain, noise, a deadline, a task that must be done. 2. **Use natural aversives, not manufactured ones.** Finishing homework ending the homework is fine. Inventing an unpleasant condition so you can remove it is coercion. 3. **Make the desired behavior the fastest route to relief.** If tantrums end demands faster than compliance does, tantrums win. 4. **Close the other exits.** If the learner can also escape by lying, hiding, or quitting, those get reinforced too. 5. **Pair it with positive reinforcement, then fade it.** The goal is behavior maintained by natural positive consequences, not by relief from something you hold over the learner. 6. **Watch the side effects.** If the learner starts avoiding *you*, the procedure has failed whatever the target behavior is doing. In [dog training](https://operantconditioning.com/dog-training/), methods built on negative reinforcement — leash corrections that stop when the dog complies, electronic collars whose stimulation ends on recall — are associated with more stress behaviors and no better obedience than reward-based methods, which is why veterinary behavior organizations recommend against them.[14] ## Common mistakes - **Calling it punishment.** If the behavior increased, it was reinforcement, whatever it felt like. - **Missing your own reinforcement.** Parents, teachers, and managers are negatively reinforced every time giving in makes an unpleasant situation stop. Ask what *your* behavior is escaping from. - **Reinforcing the escape instead of the task.** Sending a disruptive student out of class may be exactly the consequence the student was working for. - **Waiting for avoidance to extinguish on its own.** It won't; successful avoidance never contacts the change. Something has to block the response. - **Confusing it with negative punishment.** Both remove something; one strengthens behavior (removing an aversive), the other weakens it (removing a reinforcer). [Negative punishment explained ›](https://operantconditioning.com/negative-punishment/) ## Is the positive/negative distinction even real? Some behavior analysts think not. Jack Michael argued in 1975 that the two cannot be told apart in many real cases — is a warm room after a cold one the addition of warmth or the removal of cold? — and that the field should simply say "reinforcement."[15] The distinction survived because it maps onto clearly different laboratory procedures and forces students to notice what was present *before* the behavior.[16] Don't agonize over borderline cases; ask what the behavior accomplishes and whether that makes it more likely. ## Key takeaways - Negative reinforcement is "behavior went up, something was taken away." How unpleasant the stimulus felt is not part of the test; if removing it increased the behavior, it was aversive by definition. - It is not punishment: a scolding that reduces a behavior is positive punishment, while nagging that stops when the room is cleaned is negative reinforcement. The same aversive stimulus can do either job, depending on whether the behavior ends it or starts it. - Escape ends an aversive stimulus that is already present and is learned directly. Avoidance prevents one that would have come, and it persists because the threat is never tested. - Negative reinforcement quietly maintains much of the behavior people try to get rid of: escape-maintained tantrums, procrastination, and anxious avoidance. The fix is to stop the escape from working and make the desired behavior the fastest route to relief, not to add punishment. - Prefer positive reinforcement. Use negative reinforcement only when an aversive is already unavoidably present, and never invent an unpleasant condition so you can remove it. ### Check yourself **A teacher sends a student to the hallway whenever he acts out during math. Acting out during math becomes more frequent. Is the teacher punishing the behavior?** No. The behavior increased, so the consequence was reinforcement, whatever the teacher intended. Being sent out removed the hard worksheet, so the acting out is escape-maintained: negatively reinforced by the removal of the demand. **A parent starts nagging the moment she sees a messy room, and messes become less frequent. A friend calls this negative reinforcement. Is the friend right?** No. The behavior went down and something was added, so this is positive punishment. Negative reinforcement would be nagging that stops once the room is cleaned, which makes cleaning more likely. **Someone who fears elevators has taken the stairs for years, and the fear has never faded. Why not?** Taking the stairs is avoidance: the anxiety drops, so stair-taking is negatively reinforced, and the elevator is never ridden. Because the threat is never tested, the fear never gets the chance to extinguish; something has to block the avoidance response. **You open an umbrella when rain starts. Later you begin carrying one whenever rain is forecast. Which is escape, and which is avoidance?** Opening the umbrella in the rain is escape: the aversive stimulus is present and the behavior ends it. Carrying one so you never get wet is avoidance: the behavior comes first and prevents the stimulus. Both increase the behavior, so both are negative reinforcement. **Explain it to a friend.** Explain why negative reinforcement is not punishment, without using the words "positive" or "negative." ## Frequently asked questions **What is negative reinforcement in simple terms?** Removing something unpleasant after a behavior so the behavior happens more often. The seat-belt chime stops when you buckle up, so you buckle up faster. The unpleasant thing that goes away is the negative reinforcer. **What is an example of negative reinforcement?** Taking an aspirin to end a headache: the headache is present, you take the pill, the pain fades, and you take aspirin sooner next time. Others: hitting snooze, doing chores to stop nagging, opening an umbrella, leaving a party to escape anxiety. **Is negative reinforcement the same as punishment?** No. Negative reinforcement increases a behavior by removing something unpleasant; punishment decreases a behavior. "Negative" only means something was taken away. A scolding that reduces a behavior is positive punishment; nagging that stops when the room is cleaned is negative reinforcement. **What is the difference between escape and avoidance?** In escape, the unpleasant stimulus is already present and the behavior ends it (a painkiller for a headache). In avoidance, the behavior comes first and prevents the stimulus (leaving early to miss traffic). Avoidance is persistent because the person never finds out whether the threat is still real. **Is negative reinforcement bad?** Not inherently — pain relief and seat-belt chimes depend on it. But it requires an aversive stimulus, so it shares punishment's side effects, and it silently maintains tantrums, procrastination, and anxiety. Positive reinforcement is usually the better tool when you have a choice. **What is negative reinforcement in the classroom?** Any case where a behavior removes something a student finds aversive. Useful: finishing work early ends the work period. Problematic: a student acts out during hard tasks and is sent out of the room, which removes the task and strengthens acting out — escape-maintained behavior. **Who came up with negative reinforcement?** B. F. Skinner introduced the reinforcement vocabulary in *The Behavior of Organisms* (1938) and elaborated it in *Science and Human Behavior* (1953). Avoidance research was advanced by O. H. Mowrer's two-factor theory (1947) and Murray Sidman's unsignaled avoidance experiments (1953). ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Solomon, R. L., & Wynne, L. C. (1953). Traumatic avoidance learning: Acquisition in normal dogs. *Psychological Monographs, 67*(4), 1–19. 3. Mowrer, O. H. (1947). On the dual nature of learning — a re-interpretation of "conditioning" and "problem-solving." *Harvard Educational Review, 17*, 102–148. 4. Mowrer, O. H. (1960). *Learning Theory and Behavior*. Wiley. 5. Sidman, M. (1953). Avoidance conditioning with brief shock and no exteroceptive warning signal. *Science, 118*(3058), 157–158. 6. Herrnstein, R. J. (1969). Method and theory in the study of avoidance. *Psychological Review, 76*(1), 49–69. 7. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 8. Iwata, B. A., Pace, G. M., Dorsey, M. F., Zarcone, J. R., Vollmer, T. R., Smith, R. G., et al. (1994). The functions of self-injurious behavior: An experimental-epidemiological analysis. *Journal of Applied Behavior Analysis, 27*(2), 215–240. 9. Patterson, G. R. (1982). *Coercive Family Process*. Castalia. 10. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 11. Barlow, D. H. (2002). *Anxiety and Its Disorders: The Nature and Treatment of Anxiety and Panic* (2nd ed.). Guilford Press. 12. Sirois, F., & Pychyl, T. (2013). Procrastination and the priority of short-term mood regulation: Consequences for future self. *Social and Personality Psychology Compass, 7*(2), 115–127. 13. Iwata, B. A. (1987). Negative reinforcement in applied behavior analysis: An emerging technology. *Journal of Applied Behavior Analysis, 20*(4), 361–378. 14. Ziv, G. (2017). The effects of using aversive training methods in dogs — A review. *Journal of Veterinary Behavior, 19*, 50–60. See also Vieira de Castro, A. C., Fuchs, D., Morello, G. M., Pastur, S., de Sousa, L., & Olsson, I. A. S. (2020). Does training method matter? Evidence for the negative impact of aversive-based methods on companion dog welfare. *PLoS ONE, 15*(12), e0225023. 15. Michael, J. (1975). Positive and negative reinforcement, a distinction that is no longer necessary; or a better way to talk about bad things. *Behaviorism, 3*(1), 33–44. 16. Baron, A., & Galizio, M. (2005). Positive and negative reinforcement: Should the distinction be preserved? *The Behavior Analyst, 28*(2), 85–98. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): Adding something to strengthen behavior — the tool to reach for first. - [Positive punishment](https://operantconditioning.com/positive-punishment/): Adding an aversive to weaken behavior, and what the evidence says about it. - [Reinforcement: the hub](https://operantconditioning.com/reinforcement/): Both types, the kinds of reinforcers, and what makes reinforcement work. --- # Positive Punishment: Definition, Examples, Side Effects, and What the Evidence Says > Positive punishment adds an aversive stimulus after a behavior to reduce it. Definition, 16 examples, side effects, the spanking evidence, and alternatives. - Source: https://operantconditioning.com/positive-punishment/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Punishment · Adding something* Add something unpleasant after a behavior and the behavior shrinks. It works — under conditions that are rarely met — and it comes with a bill. Here is what the research shows, and what to do instead. > **Definition** > > **Positive punishment** is the process in which a behavior is followed by the *addition* of a stimulus, and as a result the behavior becomes less frequent or less likely in the future. The added stimulus is called a **punisher**. > > "Positive" means something is added, not that the procedure is good. And a stimulus is a punisher only if it actually reduces the behavior it follows; something that feels unpleasant but leaves the behavior unchanged is not punishment in the technical sense.[1] **In brief** - Positive punishment adds a stimulus after a behavior, and the behavior becomes less frequent or less likely in the future. - It can suppress behavior, but only when it is immediate, consistent, intense from the start, inescapable, and paired with a reinforced alternative. - It brings side effects — escape, avoidance, aggression, fear — and teaches nothing to do instead, so reinforcing an alternative is usually the better tool. ## How positive punishment works Positive punishment is the mirror image of [positive reinforcement](https://operantconditioning.com/positive-reinforcement/): the behavior occurs, a stimulus appears that was not there before, and the future rate of the behavior drops. - A rat presses a lever → a brief shock follows → pressing decreases. - A toddler touches the stove → it burns → stove-touching decreases (a natural punisher). - A driver runs a red light → a ticket arrives → red-light running decreases, at least at that intersection. Like reinforcement, punishment depends on *contingency* and *contiguity*. A ticket that arrives three weeks later is a weak punisher for a specific act of speeding, which is why cameras and points systems aim to make the consequence certain rather than severe. The same [A-B-C analysis](https://operantconditioning.com/abc-model/) applies: antecedent, behavior, and a consequence that changes the future probability of that behavior in that context. > **Punishers are defined by their effect, not their intent** > > A teacher who yells at a class clown intends to punish. But if the yelling is the most attention the student gets all day, clowning goes *up* — the yelling was a positive reinforcer. The only way to know whether a consequence is punishing is to measure the behavior afterward. "I punished him and it didn't work" is a contradiction in terms. ### Check any scenario in two questions Not sure whether a consequence is this quadrant or a neighbor? Answer the two questions and the checker names it. ## Examples of positive punishment | Setting | Behavior | Stimulus added | Result | | --- | --- | --- | --- | | Home | Child touches a hot pan | Burn | Touching hot pans decreases (natural punisher) | | Home | Child hits a sibling | Firm verbal reprimand | Hitting decreases — if the reprimand functions as a punisher | | Home | Teenager comes home late | Extra chores | Lateness decreases | | Classroom | Student talks during instruction | Teacher's pointed look and name said aloud | Talking decreases | | Classroom | Student writes on the desk | Must clean every desk in the room (overcorrection) | Desk-writing decreases | | Workplace | Employee skips a safety step | Written warning | Skipped steps decrease | | Driving | Driver speeds past a camera | Fine in the mail | Speeding at that spot decreases | | Driving | Driver drifts out of the lane | Rumble-strip vibration | Lane drifting decreases | | Dog training | Dog jumps on the couch | Spray of water | Couch-jumping decreases (when the owner is present) | | Dog training | Dog pulls on the leash | Leash jerk | Pulling decreases, with stress side effects | | Technology | User mistypes a password five times | Lockout and warning | Careless attempts decrease | | Health | Person skips sunscreen at the beach | Painful sunburn | Skipping sunscreen decreases (natural punisher) | | Health | Person drinks alcohol while taking disulfiram | Flushing, nausea, headache | Drinking decreases | | Sports | Player commits a foul | Yellow card and coach's rebuke | Fouling decreases | | Self-management | Person bites their nails | Bitter-tasting nail polish | Nail-biting decreases (while the taste is present) | | Nature | Young dog pounces on a bee | Sting on the nose | Pouncing on buzzing insects decreases | Notice that the natural punishers — burns, nausea — produce fast, durable learning, while many human examples carry a qualifier: "when the owner is present," "at that spot." Contrived punishment tends to suppress behavior in the punisher's presence rather than eliminate it. [Examples for every quadrant ›](https://operantconditioning.com/examples/) ## What the research says about punishment ### Thorndike and Skinner: the early doubts Thorndike's original law of effect (1898) gave punishment equal billing: responses followed by discomfort were "stamped out." By 1932 he had concluded that annoying consequences did not weaken connections nearly as directly as satisfying ones strengthened them.[2] Skinner reached the same position from the laboratory. His student W. K. Estes trained rats to press a lever for food, then briefly punished pressing with shock during extinction. Pressing dropped sharply while the shock was in effect, then recovered; by the end of extinction the punished rats had made about as many responses as unpunished ones.[3] Skinner concluded that punishment merely *suppresses* behavior temporarily, and argued against it for the rest of his career.[1] ### Azrin and Holz: when punishment does work The picture changed in the 1960s, when Nathan Azrin and colleagues ran punishment through the same parametric analysis Skinner had applied to reinforcement. Their landmark 1966 chapter concluded that punishment can produce complete and lasting suppression — but only under specific conditions.[4][5] | Condition | What the research found | Why it is hard to meet | | --- | --- | --- | | **Immediacy** | The punisher should follow within seconds; delayed punishment is far weaker. | Most human punishment is delayed by hours or weeks. | | **Intensity from the start** | A punisher introduced at full intensity suppresses behavior; one that starts mild and is gradually increased produces adaptation. | Parents and trainers naturally escalate, which trains tolerance. | | **Consistency** | Every occurrence should be punished; intermittent punishment is much weaker. | Nobody catches every instance; the behavior is reinforced on the rest. | | **No unauthorized escape** | If the organism can avoid the punisher by hiding, lying, or leaving, it learns the escape instead. | Humans are very good at finding escape routes. | | **An alternative response** | Suppression is far greater when another behavior produces the same reinforcer. | Punishment is usually applied without teaching a replacement. | | **Reduced motivation** | Punishment works better when the maintaining reinforcer is weak or unavailable. | The maintaining reinforcer is often unknown. | The same research documented the costs: escape and avoidance of the punishing agent, aggression elicited by aversive stimulation, and the odd finding that a punisher can become a signal for reinforcement, so that punishment paradoxically *increases* behavior when it predicts reward.[4] Russell Church's review reached a similar verdict: punishment is a lawful process, but its effects depend on parameters most users never control.[6] ### The applied literature In 2002, Dorothea Lerman and Christina Vorndran reviewed what applied behavior analysts actually knew about punishment: laboratory findings on intensity, immediacy, and schedule had rarely been tested with people in clinical settings; punishment sometimes remained necessary when reinforcement-based treatments failed for dangerous behavior; and the field needed better data on using it with the least intensity and fewest side effects.[7] Punishment is neither the reliable tool folk wisdom assumes nor the useless one Estes's rats suggested. ## Side effects of positive punishment Even when punishment suppresses a target behavior, it produces collateral effects that reinforcement does not.[4][7] - **Escape and avoidance.** The learner avoids the punisher — and the person, place, or task associated with it. The child punished for spilling hides spills; the employee stops reporting mistakes; the dog steals food only when no one is in the room. - **Aggression.** Aversive stimulation elicits fighting. Rats shocked together attack each other, and the same reflexive aggression appears across species.[8] Punished children and animals lash out, often at an unrelated target. - **Emotional responding.** Fear, crying, freezing, and disruption of behavior you wanted to keep. A punished student may stop talking out of turn and also stop participating. - **Modeling.** Punishment teaches that adding aversive consequences is how you change people. Children who are hit are more likely to hit. - **Suppression, not elimination.** The behavior returns when the punisher is absent or the motivation rises. Punishment does not teach what to do instead. - **The punisher becomes a conditioned aversive stimulus.** Anything reliably paired with punishment — the parent's raised voice, the classroom, the training collar — acquires aversive properties of its own, spreading escape and avoidance further. ## Corporal punishment: what the evidence shows Spanking is the most studied form of positive punishment in humans. A 2002 meta-analysis of 88 studies found corporal punishment associated with immediate compliance and with worse outcomes on essentially every other measure — aggression, antisocial behavior, mental health, the parent–child relationship.[9] A 2016 meta-analysis by Elizabeth Gershoff and Andrew Grogan-Kaylor addressed the objection that earlier work had lumped spanking together with abuse. Restricting the analysis to ordinary open-handed spanking across 75 studies and more than 160,000 children, they found spanking significantly associated with 13 of the 17 outcomes examined, all in the harmful direction; none favored spanking.[10] The evidence is largely correlational, and it is fair to ask whether difficult children simply get spanked more. But longitudinal studies that control for prior behavior still find spanking predicting later increases in problem behavior, and no comparable evidence shows benefits. In 2018 the American Academy of Pediatrics recommended that parents not spank, hit, or otherwise physically punish children, nor use verbal abuse that shames or humiliates, and instead rely on positive reinforcement, limit-setting, and brief time-out.[11] The operant analysis predicts exactly this: spanking as actually practiced — inconsistent, delayed, escalating, delivered by a person the child depends on — violates every condition on Azrin and Holz's list. ## Aversives in dog training: what the evidence shows The same question has been studied in companion dogs, with the same answer. An early owner survey found reward-based methods associated with higher obedience and punishment-based methods with more problem behaviors.[12] A survey at a veterinary behavior clinic found that confrontational techniques — hitting, alpha rolls, leash jerks — frequently provoked an aggressive response from the dog.[13] A 2017 review concluded that aversive methods carry welfare risks and no advantage in effectiveness.[14] And a 2020 study of dogs at aversive-based versus reward-based training schools found the aversively trained dogs showed more stress behaviors, higher salivary cortisol, and a more pessimistic bias in a cognitive test outside training.[15] Positive punishment trades short-term suppression for long-term stress, which is why the reward-based methods on our [dog training page](https://operantconditioning.com/dog-training/) are the standard recommendation of veterinary behavior organizations. ## When is punishment used in applied behavior analysis? Rarely, and only under conditions. In 1988 a task force of the Association for Behavior Analysis affirmed that clients have a right to *effective* treatment, which may in rare cases include restrictive procedures when less intrusive ones have failed and the behavior is dangerous.[16] The Behavior Analyst Certification Board's ethics code operationalizes that position: prioritize reinforcement-based procedures; recommend punishment only when the severity of the behavior or the failure of less intrusive procedures warrants it; always include reinforcement of alternative behavior; and monitor data and side effects continuously.[17] In practice a punishment procedure in [ABA](https://operantconditioning.com/applications/#aba) follows a functional assessment, sits inside a plan that teaches a replacement, uses the least intrusive procedure that works, requires informed consent and oversight, and is faded as soon as possible.[18] The contrast with everyday punishment — improvised, delayed, escalating, standalone — could hardly be sharper. ## Natural vs. contrived punishers A **natural punisher** is produced by the behavior itself: the burn from the stove, the fall from tipping a chair. A **contrived** punisher is delivered by another person: the reprimand, the fine, the leash jerk. Natural punishers teach efficiently because they are immediate, perfectly consistent, impossible to escape, and there is no punishing agent to become a conditioned aversive stimulus. "Let natural consequences happen" is sound advice whenever the consequence is safe: a child who forgets a coat gets cold, and cold is a better teacher than a lecture. ## Reprimands and overcorrection ### Reprimands A **reprimand** is a brief, firm verbal statement delivered immediately after the behavior. Classroom research found that *soft* reprimands, audible only to the target student, reduced disruptive behavior more than loud public ones, which can function as attention and entertainment for the class.[19] Reprimands work better delivered up close, with eye contact, than shouted across the room.[20] They lose their effect when used constantly or when the student is reinforced by the reaction. ### Overcorrection **Overcorrection** requires the person to repair the effects of the behavior beyond the original state (*restitutional* — clean the whole table, not just your spill) or to practice a correct alternative repeatedly (*positive practice* — walk calmly to the door five times after running). Foxx and Azrin developed it in the early 1970s for severe behavior in institutional settings.[21] It is effortful, can provoke aggression when the person is physically guided, and is used far less today than reinforcement-based procedures. ## Alternatives to positive punishment Because punishment does not teach a replacement, the most effective way to reduce a behavior is usually to build something else in its place. | Procedure | What you do | Example | | --- | --- | --- | | **DRA** — differential reinforcement of alternative behavior | Reinforce an appropriate behavior that serves the same function; withhold reinforcement for the problem behavior | A child who grabs toys is taught to ask, and asking gets the toy | | **DRI** — differential reinforcement of incompatible behavior | Reinforce a behavior that cannot happen at the same time | A dog that jumps on guests is reinforced for sitting when they arrive | | **DRO** — differential reinforcement of other behavior | Reinforce each interval that passes without the problem behavior | A student earns a point for each five minutes with no shouting | | **Extinction** | Identify the maintaining reinforcer and stop delivering it | Whining for candy never produces candy again | | **Antecedent changes** | Alter the setting so the behavior is less likely to be triggered or needed | Move the cookie jar; break the hard worksheet into pieces; offer a choice | | **Negative punishment** | Remove a reinforcer contingent on the behavior | Brief time-out from play after hitting | Differential reinforcement paired with [extinction](https://operantconditioning.com/extinction/) is the default treatment package in ABA and the core of evidence-based parent training. [Negative punishment](https://operantconditioning.com/negative-punishment/) — removing a reinforcer rather than adding an aversive — carries fewer side effects and is what pediatricians mean by time-out. Antecedent strategies are the most under-used option: the easiest way to reduce a behavior is often to change the situation that occasions it, as our [guide to the ABC model](https://operantconditioning.com/abc-model/) explains. ## Positive punishment vs. negative punishment vs. negative reinforcement Positive punishment *adds* an aversive to *decrease* behavior (a fine for speeding). Negative punishment *removes* a reinforcer to *decrease* behavior (losing the car keys for speeding). [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) *removes* an aversive to *increase* behavior (slowing down ends the passenger's complaints). "Positive" and "negative" only ever mean added and removed; whether it is punishment or reinforcement depends on which way the behavior moves. ## Key takeaways - "Positive" means added, not good. A stimulus is a punisher only if it reduces the behavior it follows; "I punished him and it didn't work" is a contradiction in terms. - Punishment can produce complete suppression, but only when it is immediate, at full intensity from the start, consistent, impossible to escape, and paired with a reinforced alternative. Everyday punishment is usually delayed, escalating, and inconsistent, which is why it mostly suppresses behavior in the punisher's presence. - Its side effects are escape and avoidance of the punisher, aggression, fear, modeling of aversive control, and a punisher that becomes a conditioned aversive stimulus. Punishment does not teach what to do instead. - Spanking is associated with worse outcomes on nearly every measure studied and with no long-term benefits, and pediatricians recommend positive reinforcement, limit-setting, and brief time-out instead. Aversive dog-training methods show the same pattern: more stress and no advantage in effectiveness. - The most effective way to reduce a behavior is usually to build something else in its place: differential reinforcement, extinction, antecedent changes, or negative punishment, which carries fewer side effects. ### Check yourself **A teacher yells at a class clown after every joke. Over the month, the joking increases. What was the yelling?** A positive reinforcer. Consequences are classified by their effect, not their intent: the behavior went up and something was added. If the yelling is the most attention the student gets all day, it reinforces the clowning it was meant to stop. **A dog is sprayed with water whenever it jumps on the couch, and it stops jumping — while the owner is home. What has the dog learned?** The behavior has been suppressed in the punisher's presence, not eliminated. Contrived punishment tends to work only where the punisher is, and the dog has found an escape route: jumping when no one is in the room. Nothing has taught it what to do instead. **A parent scolds mildly, then louder, then more firmly as a behavior continues. Why does this escalation tend to fail?** A punisher that starts mild and is gradually increased produces adaptation; each step trains tolerance of the next. Punishment suppresses behavior when it is introduced at full intensity from the start, which is one reason it is so hard to use well outside a laboratory. **One child forgets a coat and gets cold. Another is lectured about forgetting coats. Which consequence teaches better, and why?** The cold. It is a natural punisher: immediate, perfectly consistent, impossible to escape, and with no punishing agent to become a conditioned aversive stimulus. The lecture is a contrived punisher delivered by a person, and that person can become something the child learns to avoid. **Explain it to a friend.** Explain why punishment that seems to work in the moment usually fails in the long run, in two sentences a twelve-year-old would follow. ## Frequently asked questions **What is positive punishment in simple terms?** Adding something unpleasant after a behavior so the behavior happens less often. A dog jumps on the couch, gets sprayed with water, and jumps less. The spray is the positive punisher. "Positive" means added, not good. **What is an example of positive punishment?** A driver runs a red light and gets a ticket; red-light running decreases. Others: a child touches a hot stove and is burned, a student receives a reprimand for talking, an employee gets a written warning for skipping a safety check, a bird eats a bitter insect and vomits. **What is the difference between positive and negative punishment?** Both decrease behavior. Positive punishment adds an aversive stimulus (a reprimand, a fine, a shock). Negative punishment removes a pleasant one (a privilege, tokens, attention — as in time-out). Negative punishment has fewer side effects and is preferred when punishment is used at all. **Does positive punishment work?** It can reduce behavior when it is immediate, consistent, intense enough from the start, impossible to escape, and paired with a reinforced alternative. Those conditions are rarely met outside a laboratory. In everyday use it suppresses behavior mainly in the punisher's presence, with side effects: avoidance, aggression, fear, and imitation of aversive control. **Is spanking positive punishment?** Yes — an aversive stimulus is added after a behavior to reduce it. Meta-analyses find spanking associated with worse outcomes across the board (more aggression, more behavior problems, worse mental health) and no long-term benefits, and the American Academy of Pediatrics recommends against it. **Is positive punishment ever used in ABA therapy?** Rarely, and only within strict limits. Ethics guidelines require reinforcement-based procedures first, reserve punishment for dangerous behavior that has not responded to less intrusive treatment, require reinforcement of an alternative behavior alongside it, and require consent and monitoring of side effects. **Why is punishment less effective than reinforcement?** Punishment says what not to do but not what to do instead, so the motivation finds another outlet. It also creates escape and avoidance (including sneaking and lying), elicits aggression and fear, turns the punisher into something to avoid, and usually suppresses behavior only while the punisher is around. Reinforcement builds a replacement that persists on its own. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Thorndike, E. L. (1932). *The Fundamentals of Learning*. Teachers College, Columbia University. 3. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 4. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 5. Azrin, N. H. (1960). Effects of punishment intensity during variable-interval reinforcement. *Journal of the Experimental Analysis of Behavior, 3*(2), 123–142. 6. Church, R. M. (1963). The varied effects of punishment on behavior. *Psychological Review, 70*(5), 369–402. 7. Lerman, D. C., & Vorndran, C. M. (2002). On the status of knowledge for using punishment: Implications for treating behavior disorders. *Journal of Applied Behavior Analysis, 35*(4), 431–464. 8. Ulrich, R. E., & Azrin, N. H. (1962). Reflexive fighting in response to aversive stimulation. *Journal of the Experimental Analysis of Behavior, 5*(4), 511–520. 9. Gershoff, E. T. (2002). Corporal punishment by parents and associated child behaviors and experiences: A meta-analytic and theoretical review. *Psychological Bulletin, 128*(4), 539–579. 10. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 11. Sege, R. D., Siegel, B. S., Council on Child Abuse and Neglect, & Committee on Psychosocial Aspects of Child and Family Health. (2018). Effective discipline to raise healthy children. *Pediatrics, 142*(6), e20183112. 12. Hiby, E. F., Rooney, N. J., & Bradshaw, J. W. S. (2004). Dog training methods: Their use, effectiveness and interaction with behaviour and welfare. *Animal Welfare, 13*(1), 63–69. 13. Herron, M. E., Shofer, F. S., & Reisner, I. R. (2009). Survey of the use and outcome of confrontational and non-confrontational training methods in client-owned dogs showing undesired behaviors. *Applied Animal Behaviour Science, 117*(1–2), 47–54. 14. Ziv, G. (2017). The effects of using aversive training methods in dogs — A review. *Journal of Veterinary Behavior, 19*, 50–60. 15. Vieira de Castro, A. C., Fuchs, D., Morello, G. M., Pastur, S., de Sousa, L., & Olsson, I. A. S. (2020). Does training method matter? Evidence for the negative impact of aversive-based methods on companion dog welfare. *PLoS ONE, 15*(12), e0225023. 16. Van Houten, R., Axelrod, S., Bailey, J. S., Favell, J. E., Foxx, R. M., Iwata, B. A., & Lovaas, O. I. (1988). The right to effective behavioral treatment. *Journal of Applied Behavior Analysis, 21*(4), 381–384. 17. Behavior Analyst Certification Board. (2020). *Ethics Code for Behavior Analysts*. BACB. 18. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 19. O'Leary, K. D., Kaufman, K. F., Kass, R. E., & Drabman, R. S. (1970). The effects of loud and soft reprimands on the behavior of disruptive students. *Exceptional Children, 37*(2), 145–155. 20. Van Houten, R., Nau, P. A., MacKenzie-Keating, S. E., Sameoto, D., & Colavecchia, B. (1982). An analysis of some variables influencing the effectiveness of reprimands. *Journal of Applied Behavior Analysis, 15*(1), 65–83. 21. Foxx, R. M., & Azrin, N. H. (1973). The elimination of autistic self-stimulatory behavior by overcorrection. *Journal of Applied Behavior Analysis, 6*(1), 1–14. ## Related - [Negative punishment](https://operantconditioning.com/negative-punishment/): Time-out and response cost — removing a reinforcer to weaken behavior. - [Extinction](https://operantconditioning.com/extinction/): Reducing behavior by withholding its reinforcer, and what to expect when you do. - [Punishment: the hub](https://operantconditioning.com/punishment/): Both types, what the evidence says, and the alternatives. --- # Negative Punishment: Definition, Examples, Time-Out, and Response Cost > Negative punishment removes a reinforcer after a behavior so it decreases. Definition, examples, time-out and response cost, and how it differs from extinction. - Source: https://operantconditioning.com/negative-punishment/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Punishment · Taking something away* Take something good away after a behavior and the behavior shrinks. Time-out, fines, lost privileges, and penalty points all work this way — when they are done right, which is less often than you would think. > **Definition** > > **Negative punishment** is the process in which a behavior is followed by the *removal* of a reinforcing stimulus, and as a result the behavior becomes less frequent or less likely in the future. Its two main forms are **time-out from positive reinforcement** and **response cost**. > > "Negative" means something is subtracted; "punishment" means the behavior goes down. Whether a removal is actually punishing is decided by its effect on the behavior, not by how much the person seems to mind.[1] **In brief** - Negative punishment removes a reinforcing stimulus after a behavior, and the behavior becomes less frequent; its main forms are time-out and response cost. - Time-out only removes what time-in provides; if the situation the child leaves was not rewarding, removal is an escape, not a punishment. - It is preferred over [positive punishment](https://operantconditioning.com/positive-punishment/) because it adds no aversive stimulus, but it teaches no replacement and is weak when delayed. ## How negative punishment works Negative punishment requires that a reinforcer be present or available *before* the behavior, so that it can be taken away after. A teenager who has the car keys can lose them; a child who is playing can be removed from play; a driver with a clean license can lose points. The behavior costs the person something they had, and its future rate drops. - A child hits a sibling → is removed from the game for two minutes → hitting decreases. - A player fouls → sits out for two minutes → fouling decreases. - A driver is caught speeding → loses three points → speeding decreases. The usual rules apply. The loss must follow the behavior closely (a privilege revoked on Friday for something done on Monday teaches little), be contingent on the behavior and nothing else, and be consistent. The [A-B-C analysis](https://operantconditioning.com/abc-model/) is the same as for reinforcement, except that the consequence is a subtraction. > **The reinforcer has to be real** > > You cannot remove what the person does not value. Taking away dessert from a child who does not like dessert is not negative punishment. Sending a bored student out of a boring classroom is not either — it removes nothing the student wants, and it may remove something they wanted to escape, which makes it [negative reinforcement](https://operantconditioning.com/negative-reinforcement/) of the behavior that got them sent out. ### Check any scenario in two questions Not sure whether a consequence is this quadrant or a neighbor? Answer the two questions and the checker names it. ## Time-out from positive reinforcement ### What time-out actually is The full name matters: **time-out from positive reinforcement**. Time-out is a period during which, contingent on a behavior, the person loses access to the reinforcers that were available a moment earlier. The chair, the corner, and the bedroom are incidental; the procedure is the *removal of access to reinforcement*. It is not "go and think about what you did," and it works only when the environment the child leaves — the **time-in** — is actually rewarding. Time-out entered clinical practice in the early 1960s. In one of the first applied case studies, a three-and-a-half-year-old boy whose tantrums and self-injury had made him impossible to treat was placed briefly in his room contingent on tantrums while appropriate behavior was reinforced; the tantrums declined until he could wear his glasses, eat at the table, and eventually attend school.[2] Time-out has since become one of the most researched discipline techniques in existence and a standard component of evidence-based parent training.[3] ### Exclusionary vs. non-exclusionary time-out | Type | Procedure | Example | | --- | --- | --- | | **Non-exclusionary** (person stays in the setting) | Planned ignoring | Adult withdraws attention for a set period | | Withdrawal of a specific reinforcer | The toy or tablet is removed for two minutes | | | Contingent observation ("sit and watch") | Child sits at the edge of the activity, watches, then rejoins | | | Time-out ribbon | A ribbon signals that reinforcement is available; it is removed after misbehavior and returned after a short interval[4] | | | **Exclusionary** (person leaves the reinforcing setting) | Partition or corner | Child sits facing away from the group in the same room | | Hallway or another room | Child sits in the hallway or a quiet room without toys | | | Seclusion | An isolated time-out room — a restrictive procedure requiring oversight, reserved for dangerous behavior | | The least restrictive version that works is the right one. For most children in most homes, that is a brief non-exclusionary or corner time-out, delivered calmly. ### Why time-out fails when time-in isn't reinforcing The most common reason time-out does not work is that the child was not losing anything. A classic pair of experiments made the point. In one, time-out reduced a teenager's tantrums when the time-in environment was enriched with activities and attention, but not when it was impoverished. In the other, time-out actually *increased* a young child's problem behavior, because the time-out period gave her uninterrupted opportunity for self-stimulatory behavior she found reinforcing.[5] Time-out is only punishing relative to what it interrupts; for a child in a dull or demanding situation, being removed is an escape or an opportunity — hence the child happy to be sent to a bedroom full of toys, and the student who acts out to be sent from a hard lesson to a comfortable office. Before using time-out, ask what time-in offers. If the answer is "not much," fix that first. ### What the evidence says Decades of research, mostly with young children and their parents, support time-out for reducing aggression, noncompliance, and tantrums when it is brief, consistent, and embedded in a home rich in positive attention.[3][6] It is a component of every major evidence-based parent-management program, and the American Academy of Pediatrics recommends it as an alternative to physical punishment.[7] A 2019 review examined the claim that time-out harms attachment and concluded that, implemented as designed, the evidence does not support that concern — and that discouraging time-out risks steering parents away from one of the few discipline strategies with a strong evidence base.[8] One caution: the time-out parameters recommended in popular books and websites often diverge from what the research supports, so the details below matter.[9] ## Response cost **Response cost** is the removal of a specified amount of a reinforcer, contingent on a behavior — a fine, in other words. The term comes from laboratory work in the early 1960s in which people working for points lost some of them for responding under certain conditions, which reduced responding.[10] In applied settings it usually means losing tokens, points, money, minutes of a privilege, or the privilege itself.[11] - A student in a token economy loses two tokens for leaving her seat without permission. - A driver loses points from his license — and eventually the license — for moving violations. - A teenager loses thirty minutes of screen time for each unfinished chore. - A gym member forfeits a deposit for each missed class (a commitment contract). Response cost fits naturally inside a token economy, where reinforcers are earned for desired behavior and lost for undesired behavior. The practical rules: the person must have something to lose (a child at zero tokens has no reason to behave); the cost should matter without wiping out the day's earnings; and earnings should outpace losses. A system in which people lose more than they earn becomes an aversive environment everyone tries to escape. ## Examples of negative punishment | Setting | Behavior | Reinforcer removed | Result | | --- | --- | --- | --- | | Home | Child throws a toy at a sibling | Two-minute time-out from play | Toy-throwing decreases | | Home | Teenager misses curfew | Car keys for the weekend | Curfew violations decrease | | Home | Child interrupts repeatedly at dinner | Parent stops talking to them for a minute | Interrupting decreases | | School | Student shouts out answers | One token from the day's total | Shouting out decreases | | School | Student pushes in line | Place in line (sent to the back) | Pushing decreases | | School | Student misuses lab equipment | Lab privileges for a week | Misuse decreases | | Sports | Hockey player trips an opponent | Two minutes of play (penalty box) | Tripping decreases | | Sports | Soccer player keeps arguing after a warning | Sent off: loses the rest of the match | Arguing decreases | | Driving | Driver runs a red light | Points on the license | Red-light running decreases | | Driving | Driver parks illegally | Money (fine) | Illegal parking decreases | | Workplace | Employee is repeatedly late | Flexible-hours privilege | Lateness decreases | | Workplace | Salesperson skips compliance training | Eligibility for the quarter's bonus | Skipped trainings decrease | | Games and apps | Player attacks a teammate online | Ranking points; temporary ban | Team-killing decreases | | Games and apps | User skips a day in a streak-based app | The streak (reset to zero) | Skipped days decrease, for those who value the streak | | Self-management | You check social media during a focus block | A dollar into a jar you don't get back | Checking decreases | | Dog training | Puppy nips during play | Play; the person leaves for 30 seconds | Nipping decreases | The dog-training row is negative punishment done well: the reinforcer removed (play) is exactly the one maintaining the nipping, the removal is immediate, and play resumes as soon as the dog is calm, so the contrast is unmistakable. It teaches without adding anything aversive, which is why it is the recommended way to handle puppy biting. [Reward-based dog training ›](https://operantconditioning.com/dog-training/) · [Examples for every quadrant ›](https://operantconditioning.com/examples/) ## Negative punishment vs. extinction This is the distinction students most often get wrong, because both procedures involve a reinforcer failing to arrive and both reduce behavior. The difference is *which* reinforcer. | Aspect | Extinction | Negative punishment | | --- | --- | --- | | **What happens** | The reinforcer that has been *maintaining* the behavior is no longer delivered for it | A reinforcer the person already has is removed *because* the behavior occurred | | **Is anything taken away?** | No — a reinforcer is withheld | Yes — the person loses something | | **Which reinforcer** | The one maintaining the behavior | Usually a *different* one | | **Typical course** | Gradual decline, often after an initial burst | Usually a faster decrease | | **Child whines for candy** | Whining no longer ever produces candy | Each whine costs a token from the sticker chart | | **Dog jumps up for attention** | Jumping is never followed by attention | Jumping ends the walk for one minute | The test question: *was the removed reinforcer the one paying for the behavior?* If whining simply stops working, that is extinction — you declined to give something. If the whine costs a token, that is response cost — the token was never the reason for whining, and you removed it. The two behave differently: [extinction](https://operantconditioning.com/extinction/) produces an extinction burst and requires you to control the maintaining reinforcer, while response cost can be applied even when you cannot control what maintains the behavior (peer laughter, for instance). ## Negative punishment vs. negative reinforcement Both begin with "negative," so both involve removing something. That is where the similarity ends. | Aspect | Negative reinforcement | Negative punishment | | --- | --- | --- | | **What is removed** | An aversive stimulus | A reinforcing stimulus | | **Effect on behavior** | Increases | Decreases | | **Example** | Buckling up silences the seat-belt chime | Speeding costs points on a license | | **What the learner feels** | Relief | Loss | A single event can be both, for different people. When a parent carries a screaming toddler out of a restaurant, the toddler may be negatively punished (loses the crayons and the attention) while the parent is negatively reinforced (the embarrassment ends). And if a "punishment" removes something the person wanted to escape anyway, it has flipped into negative reinforcement. ## Effectiveness and side effects Negative punishment is the preferred form of punishment when punishment is used at all. It adds no aversive stimulus, so it produces less of the fear, reflexive aggression, and conditioned aversiveness that follow [positive punishment](https://operantconditioning.com/positive-punishment/); it is easier to apply consistently without escalation; and it fits naturally with reinforcement-based systems as the debit side of a ledger.[12][13] But it is still punishment, and it shares punishment's limits: - **It does not teach a replacement.** A time-out says what not to do. The child still needs a reinforced way to get what hitting was getting. - **Emotional responding and escape.** Children resist going to time-out, argue about fines, and sometimes escalate; the struggle to enforce the procedure can become more aversive than the procedure. - **Response cost can provoke aggression and collapse.** Losing a great deal at once elicits anger, and once a person is at zero there is nothing left to lose. - **Overuse turns a rich environment into a poor one.** A classroom where tokens are constantly taken away is one children want to leave, which brings escape-maintained behavior with it. - **Delay kills it.** "You're grounded next weekend" is a weak consequence for a Tuesday-night behavior. ## How to do time-out correctly Most failed time-outs fail on the details. The research-supported procedure looks like this:[3][9] 1. **Make time-in rich first.** Time-out only removes what time-in provides. Catch the child being good many times a day. 2. **Decide in advance which behaviors earn it.** A short list — hitting, throwing, deliberate destruction. Minor behavior is better handled by planned ignoring and reinforcing the alternative. 3. **Deliver it immediately and calmly.** One brief statement ("No hitting. Time-out."), no lecture, no negotiation. Anger and explanation are attention, and attention is a reinforcer. 4. **Keep it brief.** A few minutes is enough; about one minute per year of age is the common rule of thumb for young children. Longer is not more effective and is harder to enforce. 5. **End it on calm, not the clock alone.** Release when the interval has passed *and* the child has been quiet for a few seconds, so release does not reinforce protesting. 6. **Return to time-in without a debrief.** Back to the activity, and reinforce the first good behavior you see. The point is the contrast: good behavior gets warmth and access; hitting gets a brief, boring pause. 7. **Be consistent.** Every listed behavior, every time, from every caregiver. Intermittent time-out teaches that the behavior sometimes pays. 8. **Track it.** If the behavior is not decreasing within a couple of weeks, the procedure is not functioning as punishment. Check whether time-in is rewarding and whether time-out offers an escape. The same principles transfer to response cost: small, immediate, consistent, and embedded in a system where the person earns far more than they lose. ## Common mistakes - **Time-out from nothing.** Removing a child from a situation they wanted to leave is negative reinforcement, and it will increase the behavior. - **Time-out with a lecture.** Explaining, scolding, and checking in during the interval deliver attention and turn the procedure into a conversation. - **Making it long.** A thirty-minute time-out is less effective than a two-minute one and more likely to end in a fight. - **Fining people into the red.** Once tokens hit zero the system has no leverage. - **Confusing it with extinction.** Withholding the maintaining reinforcer is extinction; removing a different reinforcer is negative punishment. They need different planning. - **Using it instead of teaching.** Punishment of any kind supplements reinforcement of the behavior you want; it never substitutes for it. [Positive parenting strategies ›](https://operantconditioning.com/applications/#parenting) ## Key takeaways - Negative punishment needs a reinforcer that is present or available before the behavior, so that it can be taken away after. You cannot remove what the person does not value, and whether a removal is punishing is decided by its effect on the behavior. - Time-out means time-out from positive reinforcement. The chair and the corner are incidental; the procedure works only when time-in is rewarding, and for a child in a dull or demanding situation, being removed is an escape or an opportunity. - Response cost is a fine: a set amount of a reinforcer is removed each time the behavior occurs. It works only when the person has something to lose and earns more than they lose. - Extinction withholds the reinforcer that has been maintaining the behavior; negative punishment removes a reinforcer the person already has, usually a different one. Ask whether the removed reinforcer was the one paying for the behavior. - Negative punishment is the preferred form of punishment because it adds no aversive stimulus, but it still teaches no replacement, weakens with delay, and turns a rich environment poor when overused. Punishment supplements reinforcement of the behavior you want; it never substitutes for it. ### Check yourself **A bored student acts out during a hard lesson and is sent to sit in the office. Acting out increases. Was this negative punishment that failed?** No. It was negative reinforcement. The student lost nothing they wanted and escaped something they wanted to leave, so the removal strengthened the behavior. A consequence is classified by its effect, not by its label. **A child whines for candy. Parent A never gives candy for whining again. Parent B takes a token off the sticker chart for each whine. Which one is negative punishment?** Parent B. The token was never the reason for whining, and it was removed because the behavior occurred: response cost. Parent A is using extinction, in which the maintaining reinforcer is withheld and nothing is taken away, and should expect an extinction burst. **A parent puts a child in a two-minute time-out for hitting, explains during the interval why hitting is wrong, and checks in every thirty seconds. Hitting does not decrease. What went wrong?** The time-out was delivering attention. Explaining, scolding, and checking in are attention, and attention is a reinforcer, so the procedure became a conversation rather than a removal of reinforcement. One brief statement, no lecture, then back to time-in and reinforce the first good behavior. **A classroom token system fines students for every minor infraction, and several students are at zero tokens by mid-morning. Why does the system stop working?** Once a person is at zero there is nothing left to lose, so response cost has no leverage. A room where tokens are constantly taken away also becomes an aversive environment children want to leave, which brings escape-maintained behavior with it. Earnings must outpace losses. **Explain it to a friend.** Explain why a time-out only works when the situation it interrupts is rewarding, using an example from a sport or a game rather than from parenting. ## Frequently asked questions **What is negative punishment in simple terms?** Taking something good away after a behavior so the behavior happens less often. A child hits, loses two minutes of playtime, and hits less. The removed playtime is the negative punisher. "Negative" means removed, not bad. **What is an example of negative punishment?** A teenager comes home after curfew and loses the car for the weekend; curfew violations decrease. Others: time-out for hitting, losing license points for speeding, a penalty-box stint in hockey, losing tokens in a classroom system, a parking fine. **Is time-out negative punishment?** Yes, when done properly. Time-out from positive reinforcement removes access to reinforcers (play, attention, activities) for a brief period contingent on a behavior, and the behavior decreases. If the child was not enjoying the situation they were removed from, time-out is not functioning as punishment and may act as an escape instead. **What is the difference between negative punishment and negative reinforcement?** Both remove something. Negative punishment removes a reinforcer and decreases behavior (losing screen time for hitting). Negative reinforcement removes an aversive stimulus and increases behavior (buckling up to stop the seat-belt chime). Ask whether the behavior went up or down. **What is the difference between negative punishment and extinction?** Extinction withholds the specific reinforcer that has been maintaining the behavior — whining no longer produces candy. Negative punishment removes a reinforcer the person already has, usually a different one — each whine costs a token. Extinction produces an extinction burst; response cost usually works faster but does not address why the behavior occurs. **What is response cost?** A form of negative punishment in which a set amount of a reinforcer — tokens, points, money, minutes of a privilege — is removed each time a behavior occurs. Fines, penalty points, and losing tokens in a classroom are all response cost. It works best inside a system where the person earns more than they lose. **How long should a time-out be?** Brief. A few minutes is enough for young children; one minute per year of age is a common guideline. Longer time-outs are not more effective and are harder to enforce. End it when the interval has passed and the child has been calm for a few seconds. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Wolf, M. M., Risley, T., & Mees, H. (1964). Application of operant conditioning procedures to the behaviour problems of an autistic child. *Behaviour Research and Therapy, 1*(2–4), 305–312. 3. Everett, G. E., Hupp, S. D. A., & Olmi, D. J. (2010). Time-out with parents: A descriptive analysis of 30 years of research. *Education and Treatment of Children, 33*(2), 235–259. 4. Foxx, R. M., & Shapiro, S. T. (1978). The timeout ribbon: A nonexclusionary timeout procedure. *Journal of Applied Behavior Analysis, 11*(1), 125–136. 5. Solnick, J. V., Rincover, A., & Peterson, C. R. (1977). Some determinants of the reinforcing and punishing effects of timeout. *Journal of Applied Behavior Analysis, 10*(3), 415–424. 6. Kazdin, A. E. (2005). *Parent Management Training: Treatment for Oppositional, Aggressive, and Antisocial Behavior in Children and Adolescents*. Oxford University Press. 7. Sege, R. D., Siegel, B. S., Council on Child Abuse and Neglect, & Committee on Psychosocial Aspects of Child and Family Health. (2018). Effective discipline to raise healthy children. *Pediatrics, 142*(6), e20183112. 8. Dadds, M. R., & Tully, L. A. (2019). What is it to discipline a child: What should it be? A reanalysis of time-out from the perspective of child mental health, attachment, and trauma. *American Psychologist, 74*(7), 794–808. 9. Corralejo, S. M., Jensen, S. A., Greathouse, A. D., & Ward, L. E. (2018). Parameters of time-out: Research update and comparison to parenting programs, books, and online recommendations. *Behavior Therapy, 49*(1), 99–112. 10. Weiner, H. (1962). Some effects of response cost upon human operant behavior. *Journal of the Experimental Analysis of Behavior, 5*(2), 201–208. 11. Kazdin, A. E. (1972). Response cost: The removal of conditioned reinforcers for therapeutic change. *Behavior Therapy, 3*(4), 533–546. 12. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 13. Lerman, D. C., & Vorndran, C. M. (2002). On the status of knowledge for using punishment: Implications for treating behavior disorders. *Journal of Applied Behavior Analysis, 35*(4), 431–464. ## Related - [Positive punishment](https://operantconditioning.com/positive-punishment/): Adding an aversive stimulus — and why the evidence favors alternatives. - [Extinction](https://operantconditioning.com/extinction/): Withholding the maintaining reinforcer: bursts, recovery, and how to do it. - [Punishment: the hub](https://operantconditioning.com/punishment/): Both types, what the evidence says, and the alternatives. --- # Schedules of Reinforcement: Fixed, Variable, Ratio, and Interval — With an Interactive Simulator > Schedules of reinforcement decide how often a behavior pays: continuous vs. intermittent, fixed and variable ratio and interval, with a live simulator. - Source: https://operantconditioning.com/schedules-of-reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Timing and rate* A behavior's strength depends not just on *whether* it is reinforced but on *how often* and *on what rule*. Those rules are schedules, and they explain everything from slot machines to why you keep checking your phone. > **Definition** > > A **schedule of reinforcement** is the rule that determines which occurrences of a behavior are followed by a reinforcer. A **continuous** schedule reinforces every response; an **intermittent** (partial) schedule reinforces only some, based on the number of responses (*ratio* schedules) or the passage of time (*interval* schedules), on a fixed or variable basis. > > Each schedule produces a characteristic pattern of responding and a characteristic resistance to [extinction](https://operantconditioning.com/extinction/). The systematic study of schedules was the work of Charles Ferster and B. F. Skinner, whose 1957 book *Schedules of Reinforcement* reported hundreds of cumulative records from pigeons and rats.[1] **In brief** - A schedule of reinforcement is the rule deciding which responses earn a reinforcer: every response (continuous) or only some (intermittent). - Intermittent schedules are based on a count of responses (ratio) or on time elapsed (interval), with a fixed or variable requirement. - Continuous reinforcement teaches a behavior fastest; variable ratio maintains the highest, steadiest rate and is the hardest to [extinguish](https://operantconditioning.com/extinction/). ## Try it: the schedule simulator Below is a cumulative record — the same kind of graph Skinner's recorder drew on a rolling strip of paper. Time runs left to right; every response steps the line up by one; a small green tick marks each reinforcer. Pick a schedule and let the simulated subject run to watch the classic patterns appear — or switch the mode to "I'll respond" and press **Respond** yourself. The simulated subject is a stylized model of the response patterns Ferster and Skinner documented — post-reinforcement pauses on fixed ratio, steady high rates on variable ratio, the fixed-interval "scallop," and a moderate steady rate on variable interval — not a replay of real data. Try FR 20 to see a longer post-reinforcement pause, or FI 15 to see the scallop. ## Continuous versus intermittent reinforcement On a **continuous reinforcement** schedule (CRF), every response is reinforced. It is the fastest way to establish a new behavior: the relationship between behavior and consequence is unmistakable. It is also the fastest schedule to extinguish, because the first unreinforced response is immediately informative — something has changed. On an **intermittent** schedule, only some responses are reinforced. Learning is slower, but the behavior becomes far more durable. This is the **partial reinforcement extinction effect**: behavior maintained on an intermittent schedule persists much longer once reinforcement stops than behavior maintained on CRF.[2] A vending machine that fails once loses you immediately; a slot machine that fails a hundred times in a row is doing exactly what it always does. > **The practical rule** > > Reinforce continuously while a behavior is being learned. Once it is reliable, thin gradually to an intermittent schedule — ideally a variable-ratio schedule — to make it resistant to extinction. Thin too fast and you get **ratio strain**: pausing, erratic responding, and eventually extinction. ## The four basic intermittent schedules Two questions define the basic schedules. Is reinforcement based on the *number of responses* (ratio) or on *time elapsed* (interval)? And is the requirement *fixed* or *variable*?  *Idealized cumulative records. Each upward step is a response; green ticks mark reinforcers. Patterns after Ferster & Skinner (1957).* ### Fixed ratio (FR) Reinforcement follows every *n*th response. FR 1 is continuous reinforcement; FR 10 pays after every tenth response. Fixed-ratio schedules produce a high rate of responding with a distinctive **post-reinforcement pause** after each reinforcer — the larger the ratio, the longer the pause — followed by a rapid, steady run to the next one. Piece-rate pay, "buy ten coffees, get one free," and a set of ten push-ups before a rest are fixed ratios. ### Variable ratio (VR) Reinforcement follows an unpredictable number of responses that averages *n*. VR 10 might pay after 3, then 15, then 8, then 14 responses. Because the very next response might be the one that pays, there is little pausing; the result is the highest, steadiest response rate of any basic schedule and the greatest resistance to extinction. Slot machines are the textbook example. So are fishing casts, sales calls, refreshing a social feed, and opening loot boxes.[1] ### Fixed interval (FI) The first response after a fixed time has elapsed is reinforced; responses during the interval earn nothing. FI 60 s pays the first response after a minute. Animals on FI schedules produce the famous **scallop**: almost no responding right after a reinforcer, then an accelerating rate as the end of the interval approaches. Checking the oven as the timer runs down, studying more as the exam nears, and glancing at the mailbox around delivery time follow this pattern. (A salary is *not* a good example — it is paid on a time basis regardless of the number of responses, which makes it closer to a fixed-time schedule than a fixed-interval one.) ### Variable interval (VI) The first response after an unpredictable interval that averages *t* is reinforced. VI schedules produce a moderate, remarkably steady rate — the organism keeps checking, because the reinforcer could become available at any moment, but doesn't race, because responding faster doesn't make it come sooner. Checking email, waiting for a reply to a text, and a manager's random walk-throughs of the shop floor all maintain behavior on variable intervals. Because VI produces such stable baselines, it is the workhorse schedule in laboratory research on choice and drugs. | Schedule | Reinforcer delivered | Response pattern | Resistance to extinction | Everyday examples | | --- | --- | --- | --- | --- | | **Continuous (CRF)** | After every response | Steady while satiation is far off | Low | Light switch, vending machine, a new skill being taught | | **Fixed ratio (FR)** | After a fixed number of responses | High rate; post-reinforcement pause | Moderate–high | Piece-rate pay, loyalty punch cards, sets of reps | | **Variable ratio (VR)** | After a variable number of responses | Very high, steady rate; little pausing | Highest | Slot machines, sales calls, social-media feeds, fishing | | **Fixed interval (FI)** | First response after a fixed time | Scallop: pause then acceleration | Moderate | Watching the oven timer, cramming before a scheduled exam | | **Variable interval (VI)** | First response after a variable time | Moderate, steady rate | High | Checking email or texts, surprise inspections | ## Which schedule is it? A two-question rule Students confuse the four schedules almost as often as they confuse the quadrants, and the same trick works: two questions, in order. **(1) Does the reinforcer depend on how many responses were made, or on how much time has passed?** Count means *ratio*; time means *interval*. **(2) Is the requirement always the same, or does it vary around an average?** Same means *fixed*; varies means *variable*. (If every single response is reinforced, it is continuous reinforcement, and if none are, it is extinction.) A detail that catches people: on an interval schedule, responding faster does not bring the reinforcer sooner — only the *first* response after the interval counts. | Basis | Requirement fixed | Requirement varies | | --- | --- | --- | | Depends on a count of responses | Fixed ratio (FR) | Variable ratio (VR) | | Depends on time elapsed | Fixed interval (FI) | Variable interval (VI) | ### Practice: name the schedule Decide, then open the answer. **A factory pays a worker $2 for every 50 shirts sewn.** **Fixed ratio (FR 50).** The reinforcer depends on a count, and the count is always the same. Expect a brief pause after each payment, then a run. **A slot machine pays out after an unpredictable number of pulls, averaging about one in twenty.** **Variable ratio (VR 20).** Count-based, requirement varies around an average. High, steady responding; the hardest schedule to extinguish. **Your paycheck arrives every other Friday, provided you have worked that fortnight.** **Fixed interval (FI two weeks) — approximately.** Time-based and fixed. Real paychecks are a poor example of pure FI because the response requirement is not a single response after the interval, but the classic textbook classification is FI, and the "scallop" of end-of-period effort is real. **You check your phone for messages that arrive at unpredictable times; the first check after a message arrives finds it.** **Variable interval.** A message becomes available after a varying time, and only a check after that point is reinforced. Checking faster does not make messages arrive sooner — which is why VI produces a moderate, steady rate rather than a frantic one. **A dog gets a treat for every single sit during the first week of training.** **Continuous reinforcement (CRF, or FR 1).** Every response pays. Fast learning, fast extinction; the right schedule for building a behavior, the wrong one for keeping it. **A quality inspector checks a machine that jams at random intervals; she can only fix a jam once it has happened.** **Variable interval.** The opportunity to be reinforced (finding and fixing a jam) becomes available after a varying time; checks before a jam cannot be reinforced. Time-based, variable. ## Why variable ratio is so powerful — and so dangerous Three features make VR the schedule of choice for anyone who wants behavior to persist: no signal ever tells the organism that the next response is pointless; the average payout can be made arbitrarily lean once the behavior is established; and the occasional early payout after a long dry run is itself a powerful reinforcer of persistence. Casinos, mobile games, and social platforms have converged on variable-ratio designs for exactly these reasons.[3] The same mechanism that makes a behavior hard to quit makes it hard to *build* deliberately — you cannot start on VR. The order is always CRF, then a gradual stretch. ## Beyond the basic four | Schedule | Rule | What it's used for | | --- | --- | --- | | **Fixed / variable time (FT, VT)** | Reinforcer delivered after time passes, regardless of behavior | Noncontingent reinforcement; the procedure behind Skinner's "superstition" study[4] and a treatment for attention-maintained problem behavior | | **Differential reinforcement of low rate (DRL)** | Reinforced only if at least *t* seconds have passed since the last response | Slowing behavior that is fine at a low rate (talking in class, eating speed) | | **Differential reinforcement of high rate (DRH)** | Reinforced only if *n* responses occur within *t* seconds | Speeding up fluent skills (math facts, typing) | | **Differential reinforcement of other behavior (DRO)** | Reinforced if the target behavior does *not* occur for an interval | Reducing problem behavior without punishment; see [alternatives to punishment](https://operantconditioning.com/positive-punishment/) | | **Progressive ratio (PR)** | Ratio increases after each reinforcer until the organism stops (the "breakpoint") | Measuring how hard an organism will work for a reinforcer — a standard index of reinforcer value and drug abuse liability[5] | | **Concurrent schedules** | Two or more schedules available at once on different responses | Studying choice; produced Herrnstein's [matching law](https://operantconditioning.com/matching-law/): relative response rate matches relative reinforcement rate[6] | | **Chained schedules** | Completing one schedule produces the stimulus for the next; only the last delivers the primary reinforcer | Analyzing sequences of behavior; conditioned reinforcement | | **Multiple and mixed schedules** | Schedules alternate, with (multiple) or without (mixed) a signal | Studying stimulus control and behavioral contrast | ## Do humans follow these schedules? Mostly, with an important caveat. Human performance on simple schedules is often shaped by verbal rules and self-instruction as much as by the schedule itself. Adults told (or who figure out) that reinforcement is time-based may respond at a very low rate or a very high steady rate on FI rather than producing the clean scallop seen in pigeons, and instructions can override contingencies for a surprisingly long time.[7] Infants and young children, who have less verbal behavior to lean on, look more like the animal data. The lesson for anyone applying schedules to people: the *description* of the contingency is itself a variable, and rules that don't match the real schedule eventually lose to it. ## Using schedules deliberately 1. **Start continuous.** While a behavior is new — a child's first attempts at a chore, a dog's first sits, your first week of a habit — reinforce every occurrence. 2. **Stretch the ratio slowly.** Move to reinforcing two out of three, then every other, then unpredictably around one in three. Watch for ratio strain and back off if the behavior falters. 3. **Prefer variable over fixed.** Variable schedules produce steadier behavior with fewer pauses and greater persistence. 4. **Prefer ratio over interval when you want rate.** If you want more of a behavior, tie reinforcement to responses. If you want steady monitoring, an interval schedule is fine. 5. **Let natural reinforcers take over.** The goal of any contrived schedule is to hand the behavior off to the consequences the world already provides — the clean kitchen, the finished chapter, the dog that comes when called. [How to do this for your own habits ›](https://operantconditioning.com/habits/) ## Schedules and problem behavior The same principles explain why unwanted behavior is so persistent. A parent who gives in to a tantrum "only occasionally" has put tantrums on a lean variable-ratio schedule — the most extinction-resistant schedule there is. A manager who answers after-hours emails "only when urgent" has done the same for after-hours emailing. When you are trying to reduce a behavior, the first question is always: *what schedule is currently maintaining it, and who is delivering the reinforcer?* [See extinction and the extinction burst ›](https://operantconditioning.com/extinction/) ## Key takeaways - Continuous reinforcement, in which every response is reinforced, establishes a new behavior fastest and extinguishes fastest. Intermittent reinforcement teaches more slowly but makes behavior far more durable: the partial reinforcement extinction effect. - Two questions identify any basic schedule. Does the reinforcer depend on a count of responses (ratio) or on time elapsed (interval), and is the requirement always the same (fixed) or does it vary around an average (variable)? - Each schedule has a signature pattern. Fixed ratio produces a post-reinforcement pause followed by a run; variable ratio a high, steady rate; fixed interval the scallop; variable interval a moderate, steady rate. - Variable ratio is the most resistant to extinction because no signal ever tells the organism that the next response is pointless. Slot machines, social feeds, and tantrums that are given in to "only occasionally" all run on it. - To use schedules deliberately, reinforce continuously while a behavior is new, then thin gradually to an intermittent, ideally variable-ratio, schedule. Thin too fast and you get ratio strain: pausing, erratic responding, and eventually extinction. **Explain it to a friend.** Explain why a slot machine keeps people playing long after a broken vending machine would have lost them, without using the words schedule, ratio, or interval. ## Frequently asked questions **What are the four schedules of reinforcement?** Fixed ratio (reinforcement after a set number of responses), variable ratio (after an unpredictable number of responses), fixed interval (first response after a set time), and variable interval (first response after an unpredictable time). Continuous reinforcement — every response reinforced — is the fifth, simplest case. **Which schedule of reinforcement is most effective?** It depends on the goal. Continuous reinforcement is best for teaching a new behavior. Variable ratio produces the highest, steadiest rate and the greatest resistance to extinction, so it is best for maintaining an established behavior. Variable interval produces the most stable moderate rate. **Which schedule is most resistant to extinction?** Variable ratio. Because reinforcement has always been unpredictable, a long run without it is not a signal that anything has changed, so responding persists. This is why gambling and social-media checking are so hard to stop. **Is a salary a fixed interval schedule?** Not really. Fixed interval requires a response after the interval; a salary is paid on a time basis regardless of the amount of behavior, which makes it more like a fixed-time (noncontingent) schedule. Piece-rate pay and commissions are ratio schedules. This is one reason salaries do little to reinforce specific daily behaviors. **What is the fixed-interval scallop?** The characteristic pattern on a cumulative record under a fixed-interval schedule: little or no responding right after a reinforcer, then an accelerating rate as the end of the interval approaches, producing a curve that looks like a series of scallops. **What is a post-reinforcement pause?** The pause in responding that follows each reinforcer on a fixed-ratio (and fixed-interval) schedule. On fixed ratio, the pause grows with the size of the ratio — the "break" in break-and-run responding. **What is the partial reinforcement extinction effect?** The finding that behavior reinforced only some of the time persists longer during extinction than behavior that was reinforced every time. Intermittent reinforcement makes the transition to no reinforcement harder to detect, so the behavior keeps going. **What is ratio strain?** The breakdown in responding — long pauses, erratic bursts, eventually extinction — that occurs when a ratio schedule is thinned too quickly or set too high for the value of the reinforcer. The cure is to drop back to a richer schedule and stretch more gradually. ## References 1. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 2. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158. See also Mowrer, O. H., & Jones, H. (1945). Habit strength as a function of the pattern of reinforcement. *Journal of Experimental Psychology, 35*(4), 293–311. 3. Schüll, N. D. (2012). *Addiction by Design: Machine Gambling in Las Vegas*. Princeton University Press. 4. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 5. Hodos, W. (1961). Progressive ratio as a measure of reward strength. *Science, 134*(3483), 943–944. 6. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 7. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): What a reinforcer is and what makes it work. - [Extinction](https://operantconditioning.com/extinction/): What happens when reinforcement stops — bursts, recovery, resurgence. - [B. F. Skinner](https://operantconditioning.com/bf-skinner/): The man who built the box and the cumulative recorder. --- # Extinction in Operant Conditioning: Definition, the Extinction Burst, Spontaneous Recovery, and Examples > Extinction in psychology is the fading of a behavior when its reinforcer is withheld. The extinction burst, spontaneous recovery, resurgence, and how to use it. - Source: https://operantconditioning.com/extinction/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Core process* Stop paying for a behavior and it fades — but not quietly, not immediately, and not for good. What extinction in psychology really means, why behavior gets worse before it gets better, and how to use it well. > **Definition** > > **Extinction** in operant conditioning is the process in which a behavior that was previously reinforced is no longer followed by its reinforcer, and as a result the behavior decreases and eventually stops. The procedure is withholding the reinforcer; the effect is the decline in behavior. > > Extinction is not one of the four quadrants — nothing is added or removed; the consequence that used to arrive simply stops. Nor is it forgetting: extinguished behavior can return through spontaneous recovery, renewal, and resurgence, so extinction is best understood as new learning laid over the old rather than erasure.[1][2] **In brief** - Extinction is what happens when a previously reinforced behavior stops producing its reinforcer: the behavior declines and eventually stops. - Behavior can get worse before it fades — the extinction burst — and giving in at that point reinforces a more intense version, intermittently. - Extinction is not ignoring: it withholds whichever reinforcer maintains the behavior, and behavior reinforced [intermittently](https://operantconditioning.com/schedules-of-reinforcement/) resists extinction far longer. ## How extinction works Every reinforced behavior exists because of a contingency: do this, get that. Extinction breaks the contingency. The rat that has learned to press a lever for food presses, and nothing happens. It presses again. Over the next hour its rate climbs, wobbles, and slides toward zero, tracing the **extinction curve** Skinner recorded in 1938.[1] - A rat presses a lever → no pellet, ever again → pressing declines over the session. - A child whines for a cookie → whining is never again followed by a cookie → whining declines over days. - You text a friend who has stopped replying → no replies → you text less and eventually stop.  *The shape of extinction. When reinforcement stops, responding briefly rises (the burst), then declines; after a rest it returns at a lower level (spontaneous recovery) and fades again.* Two things distinguish extinction from other ways of reducing behavior. It works only on the reinforcer that has been *maintaining* the behavior; withholding something else is not extinction. And it is slow and bumpy, because a history of the behavior paying off does not vanish when the payoff stops — it is tested, protested, and eventually overwritten. How long that takes depends on the reinforcement history, as the partial reinforcement extinction effect explains. ### Watch it happen The simulator below starts in extinction: a subject with a history of reinforcement, and no more reinforcers. Run it and watch the burst, the decline, and the flattening record. Switch the schedule to compare how quickly behavior built on continuous versus intermittent reinforcement gives up. ## Operant vs. respondent extinction "Extinction" is used in both branches of conditioning. In **respondent (classical) extinction**, a conditioned stimulus is presented repeatedly without the unconditioned stimulus — the bell rings, no food follows — and the conditioned response fades. In **operant extinction**, a behavior occurs repeatedly without its reinforcer and the behavior fades. Both show extinction curves, spontaneous recovery, and renewal, and both are now understood as new inhibitory learning rather than unlearning.[2] The difference is what is disconnected: a stimulus from a stimulus, or a behavior from its consequence. [Operant vs. classical conditioning in full ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## What is an extinction burst? An **extinction burst** is a temporary increase in the frequency, intensity, or variability of a behavior immediately after its reinforcer is withheld, before the behavior declines. Press the elevator button; nothing happens; you press again, harder, three times, then hold it, then try the other button. More, harder, different — that is the burst. The organism's history says the behavior *should* work, and vigorous, varied responding is the evolved answer to a missing payoff. The burst is also the engine of [shaping](https://operantconditioning.com/shaping/): when a trainer stops reinforcing the current approximation, the burst of variability produces the slightly better response to reinforce next. But in everyday life the burst is where almost everyone gives up. The child whose whining has stopped producing candy whines louder, then cries, then throws himself on the floor — and a parent who gives in at that point has reinforced the floor-throwing. How common is it? A review of 113 published data sets in which extinction was applied to problem behavior found a burst in about a quarter of cases, more often when extinction was used alone than when combined with procedures such as differential reinforcement.[3] A later analysis of 41 cases of self-injury treated with extinction found bursts in 39% and extinction-induced aggression in 22%, again less often when extinction was combined with other treatments.[4] The burst is worth planning for but not inevitable, and the way to make it less likely is also the way to make extinction work better: reinforce something else at the same time. > **The burst is the moment that decides everything** > > If you give in during the extinction burst, you have not returned to where you started. You have reinforced a more intense version of the behavior, on an intermittent schedule — which makes it *more* resistant to extinction than before. Do not start extinction unless you can outlast the burst. If you cannot, use a different procedure. ## Extinction-induced aggression and emotional responding Withholding an expected reinforcer also elicits emotional behavior: agitation, frustration, and aggression, often aimed at whatever is nearby. In a landmark experiment, pigeons pecking a key for grain were placed on extinction, and when the grain stopped they attacked a restrained pigeon at the other end of the chamber — a target that had nothing to do with the reinforcer.[5] The human version is familiar: the vending machine gets kicked, the customer-service line gets yelled at, a child on extinction hits a sibling. Emotional responding is the second main side effect of extinction, after the burst, and a central reason it is rarely used alone for serious behavior.[4][6] ## Spontaneous recovery, resurgence, and renewal Extinguished behavior comes back in at least three distinct ways, each with its own trigger. Knowing which is which tells you what to do about it. ### Spontaneous recovery **Spontaneous recovery** is the reappearance of an extinguished behavior after time away from the extinction setting. The rat that stopped pressing by the end of Monday's session starts again on Tuesday, at a lower rate, and stops sooner. First described by Pavlov for conditioned reflexes, it occurs just as reliably for operant behavior, and each recovery is smaller than the last if reinforcement stays withheld.[7] A behavior that seems gone on Friday may be back on Monday; that is not a sign extinction failed. ### Resurgence **Resurgence** is the reappearance of a *previously* reinforced behavior when a *more recently* reinforced behavior is placed on extinction. Teach a pigeon to peck key A for food, extinguish it, teach it to peck key B, then extinguish B — and the pigeon starts pecking A again.[8][9] This explains a common form of relapse: a child learns to ask politely instead of screaming, then in a busy classroom polite requests go unanswered, and the screaming returns.[10] The remedy is to keep the replacement behavior reliably reinforced, especially where the old behavior once worked. ### Renewal **Renewal** is the return of an extinguished behavior when the context changes. Mark Bouton's research showed that extinction learning is unusually tied to the place and cues in which it happened, whereas the original learning transfers broadly — which is why a fear extinguished in a therapist's office returns in the parking garage, and a behavior extinguished at the clinic returns at home.[2][11] The remedy is to conduct extinction in every context that matters. | Phenomenon | Trigger | Example | What to do | | --- | --- | --- | --- | | **Spontaneous recovery** | Time since the last extinction session | Bedtime crying that stopped last week returns after a weekend away | Keep withholding; each recovery is weaker | | **Resurgence** | A newer behavior stops being reinforced | Polite requests go unanswered and tantrums return | Keep the replacement reliably reinforced | | **Renewal** | A change of context | Behavior extinguished at school returns at home | Extinguish in every relevant setting | ## The partial reinforcement extinction effect The most important variable in how long extinction takes is the [schedule](https://operantconditioning.com/schedules-of-reinforcement/) on which the behavior was reinforced. Behavior reinforced **intermittently** is far more resistant to extinction than behavior reinforced every time. This is the **partial reinforcement extinction effect**, first demonstrated systematically by Lloyd Humphreys in 1939: responses reinforced only half the time persisted much longer without reinforcement than responses reinforced every time.[12] The effect is paradoxical — *less* reinforcement produces *more* persistence — and two theories survived decades of testing. Abram Amsel's **frustration theory** holds that organisms on intermittent schedules learn to keep responding *through* the frustration of nonreward, so nonreward during extinction is nothing new.[13] E. J. Capaldi's **sequential theory** holds that the memory of unreinforced trials becomes a cue for responding.[14] In plain language: an organism on a continuous schedule can tell instantly that the rules have changed; one on a variable schedule cannot, because unrewarded stretches were always part of the deal. The consequences are everywhere. Slot-machine players keep pulling through long droughts. A parent who gives in to whining "only sometimes" has built whining on a variable schedule and faces a long extinction. And a habit you want to *keep* should be moved from continuous to intermittent reinforcement once established, so it survives the days the reinforcer does not show up. ## Resistance to extinction and behavioral momentum Resistance to extinction is one case of what John Nevin called **behavioral momentum**. In physics, momentum depends on velocity and mass; in Nevin's analogy, a behavior's velocity is its rate and its mass is its resistance to disruption — by extinction, satiation, distraction, or free food. His central finding is that mass depends on the *rate of reinforcement obtained in the presence of a stimulus*, largely independent of response rate: behavior from a richly reinforced context is harder to disrupt than behavior from a lean one, even at the same rate.[15][16] This has a counterintuitive implication for treatment. Reinforcing an alternative behavior makes the alternative more likely, but it also enriches the context, which can make the *problem* behavior more resistant to extinction when its reinforcement is later withheld.[16] Practitioners deliver the alternative reinforcement in clearly distinct situations and expect the persistence. For the rest of us, momentum is why a well-established habit, good or bad, shrugs off disruption. ## Extinction is not ignoring "Just ignore it" is the folk version of extinction, and it is wrong often enough to be dangerous. Ignoring is extinction only when the reinforcer maintaining the behavior is *your attention*. Extinction means withholding whichever reinforcer is doing the maintaining; a functional analysis identifies it, and the procedure must match it.[17][18] | What maintains the behavior | What extinction requires | Example | What "ignoring" would do | | --- | --- | --- | --- | | **Attention** (social positive reinforcement) | *Planned ignoring*: no eye contact, comment, or reaction | A child interrupts for attention; the parent no longer responds | Works — this is the case ignoring was made for | | **Escape from demands** (social negative reinforcement) | *Escape extinction*: the demand stays in place; the behavior no longer ends it | A student shoves the worksheet away; the teacher calmly keeps it there and prompts the next step | Makes it worse — ignoring lets the student escape the task | | **Tangible items** | The item is never delivered after the behavior | Grabbing for the phone never produces the phone | Irrelevant unless the adult was handing over the item | | **Automatic (sensory) reinforcement** | *Sensory extinction*: the sensory consequence is blocked | A child spins objects for the visual effect; the effect is removed or altered | Does nothing — the behavior reinforces itself | The escape row is the one that catches parents and teachers. A child who tantrums to get out of putting on shoes is not seeking attention, and ignoring the tantrum while the shoes stay off is not extinction; it is the reinforcer, delivered. Escape extinction means the shoes still go on — planned carefully, with calm, safety, and often an easier demand to begin with. [More on escape-maintained behavior ›](https://operantconditioning.com/negative-reinforcement/) ## Examples of extinction | Situation | Reinforcer withheld | Burst | Outcome | | --- | --- | --- | --- | | Elevator button | Elevator arriving | Repeated, harder presses; holding the button | You stop pressing and take the stairs | | Vending machine that eats a coin | Snack dropping | Pressing every button; shaking; a kick | You walk away (and avoid that machine) | | Crying at bedtime | Parent returning to the room | Louder, longer crying on the first nights | Crying declines within about ten nights if parents hold firm | | Dog begging at the table | Scraps | Whining, pawing, intense staring | Begging fades — unless one family member keeps slipping food | | Checking a phone with notifications off | New messages and likes | More frequent checks for a few days | Checking declines | | Slot machine on a losing streak | Payout | Faster play, bigger bets | Very slow decline: the variable-ratio history resists extinction | | App that has shut down | Content, updates, responses | A few repeated opens and refreshes | You stop opening it within days | | Light switch during a power cut | Light | Flipping it several times | You stop within a day — and flip it again entering a new room (renewal) | The bedtime example is a real case, published in 1959. A 21-month-old boy screamed when his parents left the room, keeping one of them there for up to two hours a night. The parents put him to bed, left, and did not return. He cried for about 45 minutes the first night, less on subsequent nights, and by the tenth night not at all. A week later an aunt put him to bed, returned when he cried, and the tantrums came back at full strength — requiring a second extinction, which again succeeded.[19] The case contains the whole story: burst, curve, accidental intermittent reinforcement, recovery, and the persistence needed to see it through. ## How to use extinction 1. **Identify the maintaining reinforcer.** Attention? Escape? An item? A sensation? If you cannot tell, the procedure cannot be designed. 2. **Make sure you control that reinforcer.** If a classroom behavior is maintained by peers' laughter, a teacher's ignoring changes nothing. 3. **Decide whether you can withstand the burst.** If the behavior is dangerous at its worst — self-injury, aggression, bolting — extinction alone is not appropriate, and a professional should design the plan. 4. **Reinforce a replacement at the same time.** Extinction says what no longer works; differential reinforcement says what does. The combination reduces bursts and aggression. 5. **Get everyone on the same page.** One family member or one weekend with the grandparents can reinforce the behavior intermittently and undo weeks of progress. 6. **Expect recovery, resurgence, and renewal.** Returns after a break, under stress, and in new settings are normal, and each is weaker if the plan holds. 7. **Track it.** Extinction curves are bumpy; without data a bad week looks like failure. ### When not to use extinction - When the behavior is dangerous and the burst could cause harm. - When you cannot control the maintaining reinforcer — peer attention, sensory consequences you cannot block, reinforcers delivered by others. - When you cannot be consistent; intermittent extinction is intermittent reinforcement. - When an antecedent change would remove the need — moving the cookie jar is easier than extinguishing cookie-seeking. [Antecedent strategies in the ABC model ›](https://operantconditioning.com/abc-model/) ## Combining extinction with differential reinforcement Extinction is rarely used alone in modern practice: by itself it produces the burst and the emotional side effects and leaves a hole where the behavior was. The standard package pairs it with reinforcement of something else.[6][20] - **DRA (alternative behavior):** reinforce an appropriate behavior that gets the same result. The child who screamed for help is taught to tap the adult's arm; tapping is answered every time, screaming never. - **DRI (incompatible behavior):** reinforce a behavior that cannot happen at the same time. Sitting is reinforced; jumping up cannot coexist with it. - **DRO (other behavior):** reinforce the passage of time without the behavior. Every five minutes without shouting earns a point. - **Noncontingent reinforcement:** deliver the maintaining reinforcer freely on a schedule, so the behavior has nothing to earn. With a replacement in place, extinction becomes a redirection rather than a standoff: the reinforcer the learner wants is still available, through the door you chose. The same applies when the learner is you — the reliable way to extinguish a habit you do not want is to make its payoff available from one you do. [Building habits with operant conditioning ›](https://operantconditioning.com/habits/) · [Positive reinforcement in depth ›](https://operantconditioning.com/positive-reinforcement/) ## Key takeaways - Extinction breaks the contingency: the behavior occurs and its reinforcer no longer arrives. Nothing is added or removed, so it is not one of the four quadrants, and it is not forgetting but new learning laid over the old. - The extinction burst is a temporary rise in the frequency, intensity, or variability of the behavior. Giving in during the burst reinforces a more intense version on an intermittent schedule, so do not start extinction unless you can outlast it. - Extinguished behavior returns in three ways: spontaneous recovery after time away, resurgence when a newer behavior stops being reinforced, and renewal when the context changes. Each is weaker if the plan holds, and none means extinction failed. - Behavior reinforced intermittently is far more resistant to extinction than behavior reinforced every time: the partial reinforcement extinction effect. Giving in "only sometimes" builds whining on a variable schedule and guarantees a long extinction. - Ignoring is extinction only when attention is the maintaining reinforcer; escape-maintained behavior requires escape extinction, in which the demand stays in place. In practice, pair extinction with reinforcement of a replacement, which reduces bursts and aggression. ### Check yourself **A child tantrums when told to put on shoes. The parent ignores the tantrum, and the shoes stay off. Is this extinction?** No. The tantrum is maintained by escape from the demand, not by attention, so ignoring it while the shoes stay off delivers the reinforcer. Escape extinction means the shoes still go on; the behavior no longer ends the demand. **A child whines for cookies. Parent A never gives a cookie for whining again. Parent B takes away screen time for each whine. Which one is extinction?** Parent A. Extinction stops delivering the reinforcer the behavior used to produce; nothing is added or removed. Parent B removes a reinforcer the child already has, which is negative punishment, not extinction. **Bedtime crying that stopped last week returns after a weekend at the grandparents' house. Has extinction failed?** Not necessarily. A return after time away from the extinction setting is spontaneous recovery, which is normal and weaker each time if reinforcement stays withheld. The real risk is that someone responded to the crying over the weekend: one intermittent reinforcement can bring the behavior back at full strength. **Two children whined for candy. One was always given candy for whining; the other got it only sometimes. Both families now stop entirely. Whose whining fades faster, and why?** The child who was always given candy. Behavior reinforced every time is far easier to extinguish than behavior reinforced intermittently, which is the partial reinforcement extinction effect. The child on the continuous schedule can tell instantly that the rules have changed; the other cannot, because unrewarded stretches were always part of the deal. **Explain it to a friend.** Explain why a behavior can get worse before it fades once its payoff stops, using an example that has nothing to do with children or pets. ## Frequently asked questions **What is extinction in psychology?** The weakening and eventual disappearance of a learned behavior when the consequence that maintained it stops occurring. In operant conditioning, a behavior is no longer followed by its reinforcer and declines. In classical conditioning, a conditioned stimulus is presented without the unconditioned stimulus and the conditioned response fades. **What is an extinction burst?** A temporary increase in the frequency, intensity, or variability of a behavior right after its reinforcer is withheld — pressing the elevator button harder and faster before giving up. Bursts occur in a substantial minority of cases when extinction is used alone and are less likely when it is combined with reinforcement of another behavior. **What is an example of extinction in operant conditioning?** A child who used to get a cookie by whining stops getting cookies for whining; after an initial increase, the whining fades over days. Others: a dog stops begging when scraps stop, you stop pressing a broken elevator button, a toddler stops crying at bedtime once parents reliably stop returning. **Is extinction the same as ignoring?** Only when attention is the reinforcer maintaining the behavior. Extinction means withholding the specific reinforcer that keeps the behavior going. If a behavior is maintained by escape from a task, ignoring it lets the escape happen and makes things worse; escape extinction means keeping the task in place. **Is extinction a form of punishment?** No. Punishment adds an aversive stimulus or removes a reinforcer the person already has. Extinction simply stops delivering the reinforcer the behavior used to produce. Both reduce behavior, but extinction is slower, produces a burst, and is not one of the four quadrants. **What is spontaneous recovery?** The return of an extinguished behavior after time away from the extinction setting — the rat that stopped pressing on Monday presses again briefly on Tuesday. Each recovery is weaker than the last if reinforcement continues to be withheld. **Why is behavior on a variable schedule harder to extinguish?** Because unrewarded stretches were always part of the experience, so the organism cannot tell that reinforcement has stopped. This is the partial reinforcement extinction effect. It is why slot machines and social media are hard to quit and why "giving in sometimes" produces the most persistent whining. **How long does extinction take?** It depends on the reinforcement history. Behavior reinforced every time may fade within a session or a few days; behavior reinforced intermittently can persist for hundreds of unreinforced responses. In the classic bedtime-crying case, tantrums disappeared within about ten nights. Consistency matters more than anything else. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Bouton, M. E. (2004). Context and behavioral processes in extinction. *Learning & Memory, 11*(5), 485–494. 3. Lerman, D. C., & Iwata, B. A. (1995). Prevalence of the extinction burst and its attenuation during treatment. *Journal of Applied Behavior Analysis, 28*(1), 93–94. 4. Lerman, D. C., Iwata, B. A., & Wallace, M. D. (1999). Side effects of extinction: Prevalence of bursting and aggression during the treatment of self-injurious behavior. *Journal of Applied Behavior Analysis, 32*(1), 1–8. 5. Azrin, N. H., Hutchinson, R. R., & Hake, D. F. (1966). Extinction-induced aggression. *Journal of the Experimental Analysis of Behavior, 9*(3), 191–204. 6. Lerman, D. C., & Iwata, B. A. (1996). Developing a technology for the use of operant extinction in clinical settings: An examination of basic and applied research. *Journal of Applied Behavior Analysis, 29*(3), 345–382. 7. Rescorla, R. A. (2004). Spontaneous recovery. *Learning & Memory, 11*(5), 501–509. See also Pavlov, I. P. (1927). *Conditioned Reflexes* (G. V. Anrep, Trans.). Oxford University Press. 8. Epstein, R. (1983). Resurgence of previously reinforced behavior during extinction. *Behaviour Analysis Letters, 3*, 391–397. 9. Epstein, R. (1985). Extinction-induced resurgence: Preliminary investigations and possible applications. *The Psychological Record, 35*(2), 143–153. 10. Lattal, K. A., & St. Peter Pipkin, C. (2009). Resurgence of previously reinforced responding: Research and application. *The Behavior Analyst Today, 10*(2), 254–266. 11. Bouton, M. E. (2002). Context, ambiguity, and unlearning: Sources of relapse after behavioral extinction. *Biological Psychiatry, 52*(10), 976–986. 12. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158. 13. Amsel, A. (1958). The role of frustrative nonreward in noncontinuous reward situations. *Psychological Bulletin, 55*(2), 102–119. 14. Capaldi, E. J. (1966). Partial reinforcement: A hypothesis of sequential effects. *Psychological Review, 73*(5), 459–477. 15. Nevin, J. A. (1974). Response strength in multiple schedules. *Journal of the Experimental Analysis of Behavior, 21*(3), 389–408. 16. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 17. Iwata, B. A., Pace, G. M., Cowdery, G. E., & Miltenberger, R. G. (1994). What makes extinction work: An analysis of procedural form and function. *Journal of Applied Behavior Analysis, 27*(1), 131–144. 18. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 19. Williams, C. D. (1959). The elimination of tantrum behavior by extinction procedures. *Journal of Abnormal and Social Psychology, 59*(2), 269. 20. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Why intermittent schedules resist extinction — with a live simulator. - [Negative punishment](https://operantconditioning.com/negative-punishment/): Time-out and response cost, and how they differ from extinction. - [Shaping](https://operantconditioning.com/shaping/): How the variability of the extinction burst builds new behavior. --- # Shaping in Operant Conditioning: Successive Approximations, Step by Step > Shaping is differential reinforcement of successive approximations to a target behavior. How it works, a step-by-step protocol, examples, shaping vs. chaining. - Source: https://operantconditioning.com/shaping/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Core process* You cannot reinforce a behavior that never happens. Shaping is how operant conditioning builds behavior that does not yet exist — one small approximation at a time, in pigeons, children, patients, and yourself. > **Definition** > > **Shaping** is the **differential reinforcement of successive approximations** to a target behavior. Responses that come closer to the goal are reinforced; earlier, cruder forms are no longer reinforced; and the criterion for reinforcement is moved step by step until the target behavior appears. > > The term is B. F. Skinner's, and so is the analogy: "Operant conditioning shapes behavior as a sculptor shapes a lump of clay." Shaping is the standard method for teaching a new response in [applied behavior analysis](https://operantconditioning.com/applications/#aba), animal training, rehabilitation, and skill learning of every kind.[1] **In brief** - Shaping is differential reinforcement of successive approximations: reinforce responses closer to the goal, extinguish cruder ones, and raise the criterion step by step. - Reinforcement acts only on behavior that occurs; shaping builds behavior that does not yet exist by selecting from the learner's natural variation. - Mark each approximation immediately, keep the steps small enough that reinforcement stays frequent, and [thin the schedule](https://operantconditioning.com/schedules-of-reinforcement/) once the terminal behavior is reliable. ## Why shaping is needed [Reinforcement](https://operantconditioning.com/positive-reinforcement/) can only act on behavior that occurs. A rat in an operant chamber will not press the lever by chance for a very long time; a child who has never said "water" cannot be praised for saying it; a stroke patient who cannot lift an arm cannot be rewarded for lifting it. If you wait for the finished behavior, you wait forever. Shaping solves the problem by reinforcing whatever the organism *can* do that resembles the goal, then demanding a little more. It relies on a fact that is easy to miss: no response is ever repeated exactly. Every lever press has a slightly different force, every attempt at a word a slightly different sound. Shaping selects from that natural variation, the way breeding selects from variation in a population.[1] ## How shaping works: the mechanism Shaping is two procedures running together — reinforcement of the current approximation and [extinction](https://operantconditioning.com/extinction/) of everything else — and it cycles through four phases:  *Shaping as a moving criterion. Reinforcing only the upper tail of what the animal currently does shifts the whole distribution; the criterion is then raised again. Variability (the width of each curve) is what makes the next step possible.* 1. **Reinforce a starting behavior.** The rat turns toward the lever; a pellet drops. Turning toward the lever increases in frequency. 2. **Variability appears.** As the behavior is repeated it varies: some turns are closer, some include a step forward, some a raised paw. Extinction increases variability further — when a response stops paying off, organisms do it in new ways.[2][3] 3. **Raise the criterion.** Once turns are reliable, only turns that include a step forward are reinforced. Plain turns go on extinction; steps forward increase. 4. **Repeat.** Steps toward, touching, pawing, pressing. Each stage is built on the reinforced behavior of the previous one and the extinction-driven variability that behavior produces. The variability step is the one people forget. Allen Neuringer's research showed that variability is not noise around a "true" response but a dimension of behavior that reinforcement controls: reinforce variety and organisms become more variable; reinforce sameness and they become stereotyped.[2] Good shaping keeps the current form reinforced often enough to persist but not so often that it hardens into a fixed habit. > **Differential reinforcement is the engine** > > "Differential" means some forms of the response are reinforced and others are not. Without the extinction half you are not shaping — you are reinforcing whatever happens, and the behavior settles at the easiest form that still pays. Without the reinforcement half, the behavior disappears. The skill is holding both at once and moving the line between them. ## Skinner's discovery: the pigeon that learned to bowl Skinner dated his understanding of shaping to a single day in 1943. He, Keller Breland, and Norman Guttman were working on [Project Pigeon](https://operantconditioning.com/bf-skinner/) on the top floor of a flour mill in Minneapolis and decided, for amusement, to teach a pigeon to swipe a small wooden ball down a miniature alley with its beak. Waiting for a full swipe went nowhere. So they reinforced any response that resembled it — a look at the ball, a move toward it, a touch — and within minutes the bird was bowling.[4][5] Skinner later wrote that the day changed how he thought about behavior: complex acts could be built rapidly from nothing by hand-delivered reinforcement. He showed the public how in a 1951 *Scientific American* article, "How to Teach Animals," which walks a reader through establishing a conditioned reinforcer (a sound paired with food) and using it to shape a dog or pigeon to a new behavior in a single session — the recipe clicker trainers still follow.[6] His pigeons went on to play a version of ping-pong, pecking a ball back and forth across a table.[7] ## How to shape a behavior: step-by-step 1. **Define the terminal behavior.** Be exact about the finished form: "sits on the mat with all four paws for 10 seconds," "says 'water' clearly," "raises the affected arm to shoulder height." A vague goal cannot be approximated. 2. **Find the starting behavior.** Something the learner already does, at least occasionally, that is on the road to the goal. Glancing at the mat. Saying "wa." Moving the arm two inches. If it never occurs, you need an easier start. 3. **Plan the approximations.** Write down the steps you expect, but hold the plan loosely — the learner's variability will suggest better ones. 4. **Reinforce immediately and every time.** Use a conditioned reinforcer (a click, a "yes," a checkmark) to mark the exact instant the criterion is met, then deliver the real reinforcer. A second's delay reinforces whatever happened in that second instead. 5. **Raise the criterion in small steps.** Move on when the current approximation is reliable — Karen Pryor's rule of thumb is when the learner succeeds most of the time — and keep each step small enough that reinforcement stays frequent.[8] 6. **Don't stay too long; don't move too fast.** Staying reinforces a plateau until it resists change. Moving too fast puts the learner on extinction: variability spikes, then the behavior collapses. If that happens, drop back a step. 7. **Thin the schedule at the end.** Once the terminal behavior is reliable, shift to [intermittent reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) and bring the behavior under the cue you want it to answer to. ## Examples of shaping across settings | Setting | Terminal behavior | Successive approximations | | --- | --- | --- | | Laboratory | Rat presses a lever | Faces lever → approaches → touches → rests paw on it → presses | | Speech / early intervention | Child says "water" | Any vocalization → "wa" → "wa-wa" → "wa-ter" → "water" clearly; the same method Lovaas used to build first words[9] | | Parenting | Toilet training | Sits on potty clothed → sits unclothed → sits at scheduled times → urinates on potty → initiates independently[10] | | Classroom | Shy student answers in class | Nods → one-word answer to a direct question → a sentence → volunteers an answer → asks a question | | Dog training | "Go to mat" | Looks at mat → steps toward → one paw on → all four paws → lies down → stays as handler moves away | | Rehabilitation | Stroke patient uses affected arm | Small movements reinforced, then larger range, then functional tasks — the core of constraint-induced movement therapy[11] | | Psychiatric care | Mute patient speaks | Eye movement toward gum → lip movement → any sound → a word → answering questions, in a classic 1960 case[12] | | Self-management | Writing 500 words a day | Open the document → one sentence → one paragraph → 100 words → 250 → 500 | | Business (analogy) | A finished product | Minimum viable version → customer feedback → iterate; a loose analogy, since the "reinforcer" is market data rather than an immediate consequence for a specific response | ## Shaping vs. chaining Shaping builds a *new form* of a single response. **Chaining** links *existing* responses into a sequence, where each step produces the cue for the next and the whole chain ends in a reinforcer. Making a bed, brushing teeth, and a dog's retrieve are chains. The stimulus each step produces does two jobs at once: it is the discriminative stimulus for the next response and a conditioned reinforcer for the one just completed, which is why chains hold together — and why a link that stops paying early in a chain lets everything after it fall apart.[16] The first step in chaining is a **task analysis**: breaking the sequence into teachable components.[13] | Aspect | Shaping | Forward chaining | Backward chaining | Total-task chaining | | --- | --- | --- | --- | --- | | **What is taught** | A new response form | A sequence, first step first | A sequence, last step first | Whole sequence every trial | | **How** | Reinforce closer approximations; extinguish earlier ones | Teach step 1 to mastery, then 1+2, and so on | Trainer does all but the last step; learner completes it and is reinforced; then last two… | Learner attempts all steps with prompts where needed | | **Reinforcer** | After each approximation | After the last mastered step | Always at the natural end of the chain | At the end, plus prompts faded | | **Best for** | Behavior that does not yet occur in any form | Learners who can already do the early steps | Learners who benefit from finishing every trial with success | Learners who can do most steps already | | **Example** | Teaching a first word | Putting on a coat | Zipping a jacket (start with the last inch) | Making a sandwich with a picture schedule | In practice the two combine: a step within a chain that the learner cannot yet perform is shaped. Backward chaining has a special advantage — the learner completes the chain and contacts the terminal reinforcer on every trial, and each newly added step is reinforced by the chance to perform the steps already mastered.[13] ## Shaping vs. prompting and fading A **prompt** is help added before or during the response to make it occur: a verbal instruction, a gesture, a model to imitate, or physical guidance. Prompting produces the behavior now; shaping waits for it to emerge. Both end in the same place — the behavior under the control of its natural cue — but prompting requires **fading**: removing the help gradually so the learner does not become dependent on it.[13] - **Most-to-least prompting** starts with full physical guidance and fades to lighter prompts; useful for learners who make many errors. - **Least-to-most prompting** gives the learner a chance to respond alone and adds help only as needed; it lets independent responding happen sooner. - **Time delay** keeps the prompt but waits progressively longer before giving it, so the learner has room to beat the prompt. The rule of thumb: prompt when the behavior is physically possible but the learner does not know what is wanted; shape when the behavior does not yet exist in the repertoire. Many programs do both — prompt a rough form, fade the prompt, then shape toward fluency. [More on cues, prompts, and stimulus control ›](https://operantconditioning.com/abc-model/) ## Shaping vs. luring in animal training Trainers distinguish **free shaping** — waiting for the animal to offer approximations and marking them — from **luring**, in which food is used to steer the animal into position (a treat moved over a dog's head produces a sit). Luring is faster for simple positions but has a cost: the dog learns to follow the food, and the lure must itself be faded or the behavior never happens without it. Shaped behavior tends to be more durable and produces animals that actively experiment; luring is a prompt and should be treated like one. [Reinforcement-based dog training ›](https://operantconditioning.com/dog-training/) ## Common shaping mistakes - **Steps that are too big.** The learner rarely meets the new criterion, reinforcement drops, and the behavior extinguishes. The fix is always to split the step. - **Staying too long on one step.** A heavily reinforced approximation becomes stereotyped and the learner stops varying. Move on while variability is still present. - **Inconsistent criteria.** Reinforcing a below-criterion response "because they tried" teaches that the criterion is negotiable. - **Late reinforcement.** A marker delivered a second late reinforces whatever followed the target — the head-turn after the sit. This is how [superstitious behaviors](https://operantconditioning.com/glossary/#superstitious-behavior) get built into a shaped response. - **Lumping instead of splitting.** Karen Pryor's term for demanding two improvements at once — a longer stay *and* a straighter sit. Raise one criterion at a time, and relax the others temporarily when you introduce a new one.[8] - **Not planning the endgame.** Shaping ends with the terminal behavior on continuous reinforcement, which is fragile. Thin the schedule and add the cue, or the behavior will not survive real life. ## Shaping yourself: technology and self-management Most successful [habit-building](https://operantconditioning.com/habits/) is shaping in disguise. The fitness watch that suggests a slightly higher daily goal after a week of hitting the old one, the running program that adds a minute each week, the language app that lengthens sessions as you improve — all are raising criteria on a behavior they reinforce immediately. The ones that fail usually fail by lumping: they ask for the terminal behavior on day one. To shape your own behavior, start with an approximation you will certainly emit — one push-up, one sentence, shoes on by the door — and reinforce it on the spot, even if only by marking it done. When the small behavior is happening on most days, raise the criterion a little. Real-world habit research suggests automaticity takes weeks to months and varies widely between people and behaviors, so raise the bar on the evidence of your own record, not a calendar.[14] A missed step is data: the criterion was raised too far, and the correct response is to split it, not to try harder. ### The percentile schedule: shaping as a formula Because human shapers drift — too generous on good days, too strict on bad ones — researchers formalized the procedure. In a **percentile schedule** the criterion is set from the learner's own recent performance: a response is reinforced if it beats a fixed percentage of the last several responses (say, half of the last ten). The criterion rises exactly as fast as the learner improves, falls if performance drops, and keeps the rate of reinforcement roughly constant. Gregory Galbicka's 1994 review proposed bringing the method into applied settings; it remains the clearest statement of what a good shaper does intuitively.[15] ## Key takeaways - Shaping is two procedures running together: reinforcement of the current approximation and extinction of everything else. Without the extinction half you are reinforcing whatever happens; without the reinforcement half the behavior disappears. - Variability is what makes each step possible. No response is ever repeated exactly, extinction increases variability further, and reinforcement itself can make behavior more variable or more stereotyped. - Define the terminal behavior exactly, start with something the learner already does, mark the instant the criterion is met with a conditioned reinforcer, and raise the criterion in small steps. Steps that are too big put the learner on extinction; staying too long hardens a plateau. - Shaping builds a new form of a single response; chaining links existing responses into a sequence; prompting adds help that must then be faded. Shape when the behavior does not yet exist, and prompt when it is possible but the learner does not know what is wanted. - Shaping ends with the terminal behavior on continuous reinforcement, which is fragile. Thin the schedule and bring the behavior under its cue, or it will not survive real life. ### Check yourself **You are shaping yourself toward writing 500 words a day. After a week at 100 words you jump to 400 and miss three days in a row. What went wrong, and what is the fix?** The criterion was raised too far, so reinforcement dropped and the behavior went on extinction. A missed step is data, not a character flaw: the fix is always to split the step, not to try harder. **A trainer shaping a dog to lie on its mat sometimes clicks when only two paws are on the mat, "because it tried." Why is this a problem?** Reinforcing a below-criterion response teaches that the criterion is negotiable. Differential reinforcement is the engine of shaping, with some forms reinforced and others not; without the extinction half, the behavior settles at the easiest form that still pays. **A trainer moves a treat over a dog's head so that it sits, and repeats this for a week. The dog now sits only when a treat is in the hand. Was the sit shaped?** No. That is luring, in which food steers the animal into position; it is a prompt, and it must be faded or the behavior never happens without it. Shaping waits for the animal to offer approximations and marks them, which produces more durable behavior. **A child can put each arm into a coat sleeve, pull the coat up, and zip it, but never does them in order without help. Shaping or chaining?** Chaining. The responses already exist; what is missing is the sequence, in which each step produces the cue for the next and the whole chain ends in a reinforcer. Shaping is for a behavior that does not yet occur in any form. **Explain it to a friend.** Explain how shaping builds a behavior that has never happened yet, without using the words "step," "approximation," or "criterion." ## Frequently asked questions **What is shaping in psychology?** Shaping is a procedure in operant conditioning for teaching a behavior that does not yet occur. The trainer reinforces successive approximations — responses that come progressively closer to the target — while withholding reinforcement from earlier, cruder forms, until the target behavior is performed. **What are successive approximations?** The series of intermediate behaviors between the learner's starting point and the goal, each a little closer to the goal than the last. Teaching a rat to press a lever, the approximations might be facing the lever, approaching it, touching it, and pressing it. **What is an example of shaping?** A parent teaching a toddler to say "water" praises "wa," then only "wa-wa," then only the full word. A dog trainer teaching "go to your mat" clicks and treats for looking at the mat, then a step toward it, then a paw on it, then all four paws, then lying down. **What is the difference between shaping and chaining?** Shaping creates a new form of a single behavior by reinforcing closer approximations. Chaining links behaviors the learner can already do into a sequence, such as the steps of brushing teeth. Chaining can be done forward, backward, or as a total task, and a step the learner cannot perform is often shaped separately. **Who invented shaping?** B. F. Skinner, with Keller Breland and Norman Guttman, discovered shaping in 1943 while teaching a pigeon to "bowl" during Project Pigeon. Skinner described the method in *Science and Human Behavior* (1953) and for a general audience in "How to Teach Animals" (*Scientific American*, 1951). **What is differential reinforcement of successive approximations?** It is the technical definition of shaping. "Differential reinforcement" means some responses are reinforced and others are placed on extinction; "successive approximations" means the reinforced responses are, in sequence, progressively closer to the target behavior. **How is shaping used in ABA therapy?** Behavior analysts use shaping to build first words and speech sounds, motor skills, feeding and self-care behaviors, social responses such as eye contact, and tolerance of medical or dental procedures. It is usually combined with prompting, fading, and chaining within a task-analyzed program. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Neuringer, A. (2002). Operant variability: Evidence, functions, and theory. *Psychonomic Bulletin & Review, 9*(4), 672–705. See also Page, S., & Neuringer, A. (1985). Variability is an operant. *Journal of Experimental Psychology: Animal Behavior Processes, 11*(3), 429–452. 3. Antonitis, J. J. (1951). Response variability in the white rat during conditioning, extinction, and reconditioning. *Journal of Experimental Psychology, 42*(4), 273–281. 4. Peterson, G. B. (2004). A day of great illumination: B. F. Skinner's discovery of shaping. *Journal of the Experimental Analysis of Behavior, 82*(3), 317–328. 5. Skinner, B. F. (1958). Reinforcement today. *American Psychologist, 13*(3), 94–99. 6. Skinner, B. F. (1951). How to teach animals. *Scientific American, 185*(6), 26–29. 7. Skinner, B. F. (1962). Two "synthetic social relations." *Journal of the Experimental Analysis of Behavior, 5*(4), 531–533. 8. Pryor, K. (1984). *Don't Shoot the Dog! The New Art of Teaching and Training*. Simon & Schuster. 9. Lovaas, O. I., Berberich, J. P., Perloff, B. F., & Schaeffer, B. (1966). Acquisition of imitative speech by schizophrenic children. *Science, 151*(3711), 705–707. 10. Azrin, N. H., & Foxx, R. M. (1971). A rapid method of toilet training the institutionalized retarded. *Journal of Applied Behavior Analysis, 4*(2), 89–99. 11. Taub, E., Crago, J. E., Burgio, L. D., Fleming, W. C., Nepomuceno, C. S., Connell, J. S., & Miller, N. E. (1994). An operant approach to rehabilitation medicine: Overcoming learned nonuse by shaping. *Journal of the Experimental Analysis of Behavior, 61*(2), 281–293. 12. Isaacs, W., Thomas, J., & Goldiamond, I. (1960). Application of operant conditioning to reinstate verbal behavior in psychotics. *Journal of Speech and Hearing Disorders, 25*(1), 8–12. 13. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 14. Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. *European Journal of Social Psychology, 40*(6), 998–1009. 15. Galbicka, G. (1994). Shaping in the 21st century: Moving percentile schedules into applied settings. *Journal of Applied Behavior Analysis, 27*(4), 739–760. 16. Kelleher, R. T., & Gollub, L. R. (1962). A review of positive conditioned reinforcement. *Journal of the Experimental Analysis of Behavior, 5*(4), 543–597. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The consequence that does the work in every approximation. - [Extinction](https://operantconditioning.com/extinction/): The other half of differential reinforcement — bursts, variability, and recovery. - [Habits](https://operantconditioning.com/habits/): Shaping your own behavior with antecedents, tiny steps, and immediate consequences. --- # Stimulus Control: Discrimination, Generalization, Peak Shift, and Errorless Learning > Stimulus control means a behavior's probability depends on an antecedent. SD vs. S-delta, discrimination, generalization, peak shift, and errorless learning. - Source: https://operantconditioning.com/stimulus-control/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Antecedents · Core concept* Consequences decide whether a behavior is learned; antecedents decide when it shows up. Here is how a stimulus comes to govern behavior without forcing it, how far that control spreads, why it can shift in strange directions, and how it is used on everything from pigeons to insomnia. > **Definition** > > A behavior is under **stimulus control** when its probability — how often it occurs, how quickly, or in what form — depends on whether a particular antecedent stimulus is present. The stimulus does not force the behavior; it changes the odds, because in the past the behavior has been reinforced in its presence and not in its absence. > > Skinner introduced the analysis, and the symbols S^D and S^Δ, in *The Behavior of Organisms* (1938); the classic experimental review is Terrace (1966).[1][2] **In brief** - A behavior is under stimulus control when its probability depends on an antecedent stimulus, because it has been reinforced in that stimulus's presence before. - A discriminative stimulus (S^D) signals that a response will be reinforced; an S-delta signals that it will not, and discrimination training alternates the two. - Generalization spreads responding to similar stimuli along a gradient; discrimination training steepens that gradient and can shift its peak away from the S-delta. ## What is stimulus control? Put a rat in a chamber where lever presses produce food only while a light is on. At first it presses at the same rate light or dark. After a few sessions it presses briskly the moment the light comes on and hardly at all when it goes off. Nothing about the light compels the press — a rat that has just eaten will ignore it — but the light has become the condition under which pressing is worth doing. Pressing is under stimulus control, and the gap between the two rates measures how tightly.[1] Human life is dense with the same relation. A ringing phone, a green light, a colleague's raised eyebrow, an "Open" sign, the first bars of a song you know: each raises the probability of a specific behavior because that behavior has paid off in its presence before. This is the antecedent term of the [three-term contingency](https://operantconditioning.com/abc-model/) — in the presence of S^D, response R produces reinforcer S^R. Reinforcement builds a behavior; stimulus control decides where and when it appears.[3] > **Control is not compulsion** > > "Control" is a technical word here, closer to the way a thermostat controls a furnace than the way a puppeteer controls a puppet. A stimulus controls an operant by changing its probability, and only because of a history of consequences; change the consequences and the same stimulus stops working. A stimulus that *forces* a response — the puff of air that makes you blink — is eliciting a reflex, a different process. [Operant vs. classical conditioning ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## S^D and S-delta: how discrimination training works A **discriminative stimulus**, S^D (pronounced "ess-dee"), is a stimulus in whose presence a response has been reinforced. An **S-delta**, S^Δ, is a stimulus in whose presence the same response has gone unreinforced. In the generalization literature the pair is often written S+ and S−. **Discrimination training** is simply reinforcement in one and extinction in the other, alternated until responding diverges. The organism is said to *discriminate* when it responds differently to the two.[1] | S^D (respond) | S^Δ (don't) | Behavior | What the S^D signals | | --- | --- | --- | --- | | Light on | Light off | Rat presses the lever | Presses will produce food | | "Sit" | Any other word | Dog sits | A treat is available for sitting now | | Teacher looking at the class | Teacher writing on the board | Student raises a hand | Hand-raising will get called on | | Friends at a bar | Grandmother at dinner | Telling a crude joke | Laughter, not a frown | | Green "Walk" signal | Red hand | Crossing the street | Crossing will be safe and unfined | | Bed, dark room, 11 p.m. | Desk, daylight | Falling asleep | Sleep will come — unless the bed has also become a cue for scrolling | Stimulus control is not automatic. A stimulus that is always present, with no difference in consequences attached, may acquire no control at all. Jenkins and Harrison trained pigeons to peck a key with a 1000-hertz tone playing throughout; tested later with tones of other pitches, the birds responded equally at every pitch — a flat gradient. The tone had been there all along, but they had never had a reason to attend to it. Birds trained with the tone on during reinforcement and off during extinction produced a sharply peaked gradient centered on 1000 hertz.[4] A cue has to *predict a difference* to gain control, which is why "sit" said to a dog that will be fed anyway teaches nothing. ## Generalization and the generalization gradient **Stimulus generalization** is the mirror image of discrimination: responding to stimuli that resemble the S^D without having been trained on them. The more similar the stimulus, the more responding it evokes — a relation that, plotted, is a **generalization gradient**. The textbook demonstration is Guttman and Kalish's. They reinforced pigeons for pecking a key lit with a single wavelength — separate groups at 530, 550, 580, and 600 nanometers — and then, in extinction so the test itself taught nothing, presented a range of wavelengths on either side. Responding peaked at the training wavelength and fell away smoothly on both sides.[5] The gradient's shape is not fixed: discrimination training steepens it, and what the organism learns is less the stimulus itself than how it differs from what surrounds it.[6] | Aspect | Discrimination | Generalization | | --- | --- | --- | | **What it is** | Responding differently to different stimuli | Responding similarly to similar stimuli | | **Laboratory sign** | A steep gradient; near-zero responding in S^Δ | A flat, wide gradient | | **Everyday example** | Answering your own ringtone, not a stranger's | Braking for a stop sign you have never seen before | | **When it fails you** | A skill that works only in the room it was taught in | A fear of one dog that spreads to every dog | | **How to get more of it** | Differential reinforcement: pay in S^D, never in S^Δ | Train with many examples in many settings | Both are adaptive and both can misfire. Generalization lets a toddler call the neighbor's terrier a dog; discrimination stops her calling the cat one. In applied work generalization is usually the scarce commodity: a behavior taught in one place with one person by one method tends to stay there unless generalization is deliberately programmed.[7] ## Peak shift: when discrimination training moves the peak Discrimination training does something stranger than sharpen the gradient. In 1937 Kenneth Spence proposed that an S+ builds up a gradient of excitation and an S− a gradient of inhibition, and that the two add algebraically. If the S− sits close to the S+ on the same dimension, the sum should peak not at the S+ but a little beyond it, on the side away from the S−.[8] The prediction sat on paper for twenty years. H. M. Hanson tested it in 1959. He reinforced pigeons for pecking at 550 nm, then extinguished pecking at a slightly yellower 555 nm (other groups had S− at 560, 570, or 590 nm; a control group had no S− at all). In the generalization test, the control birds peaked at 550, as Guttman and Kalish's had. The birds trained against 555 peaked at 540 — a wavelength they had never been reinforced for, displaced away from the S− — and pecked far more overall than the controls. The nearer the S− had been to the S+, the larger the shift.[9] This is **peak shift**: evidence that the S^Δ does not merely switch responding off but exerts an inhibitory gradient of its own, which subtracts most on the side nearest it.  *Schematic of Hanson's result. After discrimination training against a nearby S−, the peak of the gradient moves away from the S− (here from 550 to about 540 nm) and rises above the control gradient.* ## Errorless discrimination learning In ordinary discrimination training the learner meets the S^Δ at full strength, responds to it, and is extinguished: it learns by making errors. Herbert Terrace showed in 1963 that the errors are optional. He trained pigeons to peck a red key and not a green one, but introduced the green key from the first session so dimly and briefly that the birds never pecked it, then raised its brightness and duration in small steps. Pigeons trained this way acquired the discrimination with few or no errors, where birds trained conventionally made thousands.[10] In a companion study he transferred the discrimination from colors to vertical and horizontal lines by superimposing the lines on the colors and fading the colors out.[11] The errorless birds differed in more than their error count. They showed none of the agitation that conventionally trained birds displayed during the S−, and in a wavelength test they showed no peak shift: the S− had acquired no inhibitory gradient, because it had never been responded to and extinguished.[10][12] Whether a stimulus becomes aversive depends on how it was learned. The legacy is everywhere in teaching. **Prompting** — a gesture, a model, a highlighted answer, physical guidance — gets the right response before an error can occur, and **fading** transfers control from the prompt to the natural S^D in graded steps.[13] In neuropsychology, Baddeley and Wilson found that people with amnesia learned word lists better when prevented from guessing than by trial and error: without explicit memory they could not correct their mistakes, so each error was simply practiced.[14] Errorless learning has costs — a learner who never meets the S^Δ may be thrown when it finally appears, and a prompt that is never faded produces a learner who waits for it — but for fragile or fearful learners it is usually the right default. [Prompting and fading in depth ›](https://operantconditioning.com/shaping/) ## Contextual control and renewal Stimulus control belongs not only to discrete cues but to the background: the room, the time of day, the people present. Contexts acquire control more weakly than a lit key, but reliably, and their influence is sharpest at the moment of extinction. Mark Bouton's research showed that when a behavior is learned in one context and extinguished in another, returning to the first context brings it back — **renewal**. The original learning transfers across contexts; the extinction learning largely does not. Extinction does not erase what was learned but adds a second, inhibitory lesson whose retrieval depends on the context in which it was learned, so that the stimulus becomes ambiguous and the setting decides which meaning wins.[15] The practical consequences are large. A fear reduced in a therapist's office can return in the parking garage; a tantrum extinguished at the clinic can return at home; a habit dropped on holiday returns on the first morning back at work. The remedies follow directly: extinguish in every context that matters, and teach the replacement behavior where the old one used to pay. [Renewal, resurgence, and spontaneous recovery ›](https://operantconditioning.com/extinction/#renewal) ## Stimulus control in practice ### Stimulus control therapy for insomnia The most direct clinical use of the concept treats a bed that has stopped working. For a good sleeper, bed, darkness, and bedtime are S^Ds for falling asleep. For a chronic insomniac they have become cues for lying awake, worrying, checking the time, and watching television, because that is what has repeatedly happened there. Richard Bootzin's **stimulus control treatment**, introduced in 1972, re-establishes the bed as a cue for sleep and nothing else.[16] 1. **Go to bed only when sleepy**, not merely tired or because it is late. 2. **Use the bed only for sleep.** No reading, eating, screens, or worrying in bed. (Sex is the conventional exception.) 3. **If you cannot fall asleep within about 10 to 20 minutes, get up**, go to another room, and return only when sleepy. Repeat as needed. 4. **Get up at the same time every morning**, however little you slept. 5. **Do not nap** during the day. The rules are uncomfortable for a week or two and then, for most people, they work. Stimulus control is a core component of cognitive behavioral therapy for insomnia (CBT-I), which the American College of Physicians recommends as the first-line treatment for chronic insomnia in adults, ahead of medication.[17] ### "Train it everywhere": dog training A dog that sits perfectly in the kitchen has learned "sit, in the kitchen, facing my owner, treat pouch on." At the park none of those stimuli are present, and neither is the sit. The dog is not stubborn; it is discriminating exactly as its training taught it to. The fix is to program generalization: train with many examples — rooms, people, distances, distractions — until the word alone controls the behavior. Stokes and Baer's "train sufficient exemplars" is the same advice in the language of applied behavior analysis.[7] [Cues and generalization in dog training ›](https://operantconditioning.com/dog-training/) ### Habit design: choose the cue Habits are behaviors under tight stimulus control: the context evokes the response with little deliberation, and the response survives on the context rather than on intention.[18] That gives you two levers. To build a habit, attach the behavior to a specific, stable cue that already occurs — after the coffee is poured, when the front door closes — and reinforce it there until the cue does the work. To break one, remove or change the cue; a phone charging in the kitchen is not an S^D for scrolling in bed. Skinner listed "changing the stimulus" among the basic techniques of self-control for exactly this reason.[3] [The operant protocol for building habits ›](https://operantconditioning.com/habits/) ### Study in one place The classic study-skills prescription is Bootzin's rule applied to a desk. Work in one place used for nothing else; when you stop working, leave it. Over a few weeks the desk becomes an S^D for working and stops being one for daydreaming, snacking, and messaging, because those behaviors are never reinforced there. "I'll study on the couch" so rarely produces studying because the couch already controls something else. ## Which is it: discrimination or generalization? Five scenarios. Decide which process each shows before opening the answer. **1. A toddler who has learned "doggie" for the family beagle says it to a neighbor's Labrador, then to a goat.** **Generalization.** The response spreads to stimuli that resemble the training stimulus. The Labrador is a useful generalization; the goat is an overgeneralization that discrimination training — "no, that's a goat" — will correct. **2. A rat presses the lever rapidly when the light is on and almost never when it is off.** **Discrimination.** Responding differs sharply between S^D and S^Δ; the behavior is under tight stimulus control. **3. A dog that sits reliably in the kitchen ignores "sit" at the park.** **Discrimination — more than the trainer wanted.** The behavior is controlled by the kitchen's stimuli, not by the word. The trainer's job is to build generalization to the cue alone by training across settings. **4. You reach for your phone whenever you hear a notification chime, including other people's.** **Generalization.** The response evoked by your own chime spreads to similar chimes. If you later stop reaching for others' phones because only yours has ever paid off, that is discrimination developing. **5. A student swears freely with friends and never at the dinner table.** **Discrimination.** The same behavior occurs in one social setting (where it has been reinforced with laughter) and not in another (where it has met disapproval). Friends are the S^D; family dinner is the S^Δ. ## Common confusions ### Discriminative stimulus vs. conditioned stimulus | Aspect | Discriminative stimulus (S^D) | Conditioned stimulus (CS) | | --- | --- | --- | | **What it does** | *Evokes* an operant: raises its probability | *Elicits* a respondent: triggers a reflex | | **How it got its power** | The behavior was reinforced in its presence | It was paired with an unconditioned stimulus, whatever the organism did | | **Does the behavior matter?** | Yes: the consequence depends on responding | No: the US arrives regardless | | **Example** | A lit key: pecking now produces grain | A tone that precedes food: salivation follows the tone | The two often ride on the same event. The click of the food magazine is a CS (it elicits approach and salivation), a conditioned reinforcer (it strengthens the press it follows), and an S^D for going to the tray. Asking which function a stimulus is serving, rather than what it "is," is the behavior analyst's habit. ### Stimulus control vs. elicitation An S^D changes probabilities; it never guarantees a response, and a motivating operation can cancel it entirely (the lit key evokes nothing in a stuffed pigeon). Elicitation is closer to a guarantee: the reflex follows the stimulus whether or not the organism is hungry, tired, or busy. When a stimulus seems to "make" someone do something — a craving at the sight of a bar, a flinch at a raised hand — the elicited, Pavlovian component is often doing more of the work than the operant one. ### Three more - **An S^D is not a motivating operation.** The S^D signals that a reinforcer is *available*; a motivating operation changes how much it is *worth*. [Motivating operations explained ›](https://operantconditioning.com/abc-model/) - **An S^Δ is not a punisher.** It signals that responding will go unreinforced, not that it will be punished — though, as Terrace's pigeons showed, a stimulus learned through extinction can become mildly aversive in its own right. - **"Discrimination" here has no social meaning.** It is the technical word for responding differently to different stimuli, as in a discriminating palate. ## Key takeaways - Stimulus control means an antecedent changes the probability of a behavior; it does not force it. An S^D evokes an operant only because of a history of consequences, and a stimulus that forces a response is eliciting a reflex, a different process. - Discrimination training is reinforcement in the presence of the S^D and extinction in the presence of the S^Δ, alternated until responding diverges. A cue has to predict a difference in consequences to gain any control at all. - Generalization is responding to untrained stimuli that resemble the S^D, and the generalization gradient shows how responding falls off with similarity. Discrimination training steepens the gradient and, when the S− is close to the S+, shifts its peak away from the S−: peak shift. - Errors are optional. Introducing the S^Δ so faintly that it is never responded to, then fading it in, produces a discrimination with few or no errors, no agitation, and no peak shift; prompting and fading apply the same idea in teaching. - Contexts control behavior too, and extinction learning is more context-bound than the original learning, so an extinguished behavior renews when the setting changes. Stimulus control therapy for insomnia, training a dog in many settings, and choosing a cue for a habit all apply the same principle. **Explain it to a friend.** Explain why a dog that sits perfectly in the kitchen ignores "sit" at the park, in two sentences a twelve-year-old would follow. ## Frequently asked questions **What is stimulus control in psychology?** Stimulus control is the condition in which a behavior occurs more often, faster, or more reliably in the presence of a particular stimulus than in its absence, because the behavior has been reinforced in that stimulus's presence in the past. The stimulus is called a discriminative stimulus. Stimulus control is the antecedent side of operant conditioning: consequences build a behavior, and antecedents decide when it appears. **What is an example of stimulus control?** A rat that presses a lever only when a light is on; a driver who brakes at a red light and not a green one; a dog that sits when it hears "sit"; a student who swears with friends but not at the family table. In each case the behavior is possible at any time but is reliably evoked by one stimulus and not by others. **What is the difference between SD and S-delta?** An S^D (discriminative stimulus) is a stimulus in whose presence a behavior has been reinforced, so the behavior becomes more likely when it appears. An S^Δ (S-delta) is a stimulus in whose presence the same behavior has not been reinforced, so the behavior becomes less likely. Discrimination training alternates the two — reinforcement in S^D, extinction in S^Δ — until responding diverges. **What is the difference between discrimination and generalization?** Discrimination is responding differently to different stimuli: pressing when the light is on but not when it is off. Generalization is responding similarly to similar stimuli: pecking a slightly different color that was never trained. They are two ends of one continuum, and the generalization gradient — how responding falls off as a test stimulus becomes less like the training stimulus — measures where an organism sits on it. **What is peak shift?** After discrimination training between an S+ and a nearby S− on the same dimension, the peak of the generalization gradient moves away from the S−. Hanson's pigeons, reinforced at 550 nm and extinguished at 555 nm, responded most to 540 nm, a color they had never been reinforced for. Spence predicted the effect in 1937 from the idea that excitatory and inhibitory gradients add. **What is stimulus control therapy for insomnia?** A behavioral treatment, developed by Richard Bootzin in 1972, that makes the bed a cue for sleep and nothing else: go to bed only when sleepy, use the bed only for sleep, get up if you cannot sleep and return only when sleepy, rise at the same time every day, and do not nap. It is a core component of cognitive behavioral therapy for insomnia, the recommended first-line treatment for chronic insomnia. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Terrace, H. S. (1966). Stimulus control. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 271–344). Appleton-Century-Crofts. 3. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 4. Jenkins, H. M., & Harrison, R. H. (1960). Effect of discrimination training on auditory generalization. *Journal of Experimental Psychology, 59*(4), 246–253. 5. Guttman, N., & Kalish, H. I. (1956). Discriminability and stimulus generalization. *Journal of Experimental Psychology, 51*(1), 79–88. 6. Honig, W. K., & Urcuioli, P. J. (1981). The legacy of Guttman and Kalish (1956): Twenty-five years of research on stimulus generalization. *Journal of the Experimental Analysis of Behavior, 36*(3), 405–445. 7. Stokes, T. F., & Baer, D. M. (1977). An implicit technology of generalization. *Journal of Applied Behavior Analysis, 10*(2), 349–367. 8. Spence, K. W. (1937). The differential response in animals to stimuli varying within a single dimension. *Psychological Review, 44*(5), 430–444. 9. Hanson, H. M. (1959). Effects of discrimination training on stimulus generalization. *Journal of Experimental Psychology, 58*(5), 321–334. 10. Terrace, H. S. (1963). Discrimination learning with and without "errors." *Journal of the Experimental Analysis of Behavior, 6*(1), 1–27. 11. Terrace, H. S. (1963). Errorless transfer of a discrimination across two continua. *Journal of the Experimental Analysis of Behavior, 6*(2), 223–232. 12. Terrace, H. S. (1964). Wavelength generalization after discrimination learning with and without "errors." *Science, 144*(3615), 78–80. 13. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 14. Baddeley, A., & Wilson, B. A. (1994). When implicit learning fails: Amnesia and the problem of error elimination. *Neuropsychologia, 32*(1), 53–68. 15. Bouton, M. E. (2004). Context and behavioral processes in extinction. *Learning & Memory, 11*(5), 485–494. 16. Bootzin, R. R. (1972). Stimulus control treatment for insomnia. *Proceedings of the 80th Annual Convention of the American Psychological Association, 7*, 395–396. 17. Qaseem, A., Kansagara, D., Forciea, M. A., Cooke, M., & Denberg, T. D. (2016). Management of chronic insomnia disorder in adults: A clinical practice guideline from the American College of Physicians. *Annals of Internal Medicine, 165*(2), 125–133. 18. Wood, W., & Rünger, D. (2016). Psychology of habit. *Annual Review of Psychology, 67*, 289–314. ## Related - [The ABC model](https://operantconditioning.com/abc-model/): Antecedent, behavior, consequence — the three-term contingency that stimulus control is the front half of. - [Extinction](https://operantconditioning.com/extinction/): Bursts, spontaneous recovery, resurgence, and the renewal that context brings. - [Dog training](https://operantconditioning.com/dog-training/): Adding the cue, proofing it everywhere, and why a kitchen "sit" vanishes at the park. --- # The ABC Model of Behavior: Antecedent, Behavior, and Consequence Explained > The ABC model of behavior (antecedent, behavior, consequence) explained: the three-term contingency, discriminative stimuli, motivating operations, ABC data. - Source: https://operantconditioning.com/abc-model/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Core concept* Every operant behavior sits between what came before it and what came after it. Learn to read that three-part pattern and you can explain almost any behavior — and change it at three points instead of one. > **Definition** > > The **ABC model of behavior** — antecedent, behavior, consequence — is the **three-term contingency**, the basic unit of analysis in operant conditioning. An *antecedent* sets the occasion for a *behavior*; a *consequence* follows the behavior and changes how likely it is to occur again under similar antecedents. > > B. F. Skinner made the three-term contingency the centerpiece of *Science and Human Behavior* (1953), building on his laboratory work on discriminated operants. Each term names an observable event, and the relations between them can be measured.[1][2] **In brief** - The ABC model, or three-term contingency, is the basic unit of operant analysis: an antecedent sets the occasion, a behavior occurs, a consequence follows. - An antecedent evokes a behavior rather than causing it, and only because of what has followed that behavior there before. - A consequence is defined by its effect on future behavior, not by its appearance, and the model gives you three points to intervene. ## How the ABC model of behavior works The word doing the work is *contingency*. The consequence happens because the behavior happened, and the behavior is emitted in the presence of the antecedent because that is where the consequence has followed before. A rat whose lever-presses produce food only while a light is on comes to press when the light is on and to ignore the lever when it is off.[2] Nothing about the light forces the press; it matters only because of the history attached to it. The antecedent is the front of the loop, but the consequence is what loads it. ## What is an antecedent? An antecedent is anything in place before the behavior: the immediate stimuli, the wider context, and the state of the organism. Behavior analysts distinguish several kinds, because each is changed in a different way. ### The discriminative stimulus (S^D) and S-delta A **discriminative stimulus**, written S^D, is a stimulus in whose presence a behavior has been reinforced. Its counterpart, **S-delta** (S^Δ), is one in whose presence the same behavior has gone unreinforced. The lit "Open" sign is an S^D for pulling the door; the dark sign is an S^Δ. A friend's grin is an S^D for a crude joke; your grandmother's presence is an S^Δ for the same joke. The S^D does not *elicit* behavior the way a puff of air elicits a blink. It *evokes* it — raises its probability — and only because of what has followed the behavior there before. Pavlov's stimulus produces the response; Skinner's merely signals that a response will pay. [Operant vs. classical conditioning ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ### Stimulus control, discrimination, and generalization When a behavior's probability depends on whether a stimulus is present, the behavior is under **stimulus control**. It gets there through **discrimination training** — reinforcement in the presence of the S^D, extinction in the presence of the S^Δ — and its mirror image is **generalization**, responding to stimuli that resemble the S^D: pigeons reinforced for pecking a key lit with 550 nm light pecked less and less as the color moved away from it, tracing a smooth generalization gradient.[3] Discrimination training also reshapes that gradient (peak shift), can be arranged so the learner never makes an error (the basis of prompting and fading), and extends to the context itself, which is why an extinguished behavior can [renew](https://operantconditioning.com/extinction/) when the setting changes back. [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) ### Motivating operations Jack Michael pointed out in 1982 that two different antecedent functions were being lumped together. A stimulus can signal that a reinforcer is *available* (the discriminative function) or change how much the reinforcer is *worth* (the motivational function).[4] He called the second an **establishing operation**; the field later adopted **motivating operation** (MO) as the umbrella term, with an *establishing operation* increasing a reinforcer's value and an *abolishing operation* decreasing it.[5][6] Every MO has two effects at once: it alters the reinforcer's value *and* it alters behavior, evoking responses that have produced that reinforcer or abating them. Eight hours without food makes food a stronger reinforcer and makes you open the fridge; a large lunch does the opposite. A child who slept badly finds escape from demands more valuable, so escape behavior climbs. The vending machine says chips are available; hunger says they are worth having. MOs are the closest thing to "motivation" in the model, and unlike the folk concept they are manipulable: train the dog before dinner, not after. ### Setting events **Setting events** are more distant conditions — illness, poor sleep, an argument at breakfast — that change the strength of the immediate A-B-C relation without being part of it.[7] Most can be re-described as motivating operations, but the term survives in schools because it reminds observers to look further back than the last thirty seconds. ### Prompts A **prompt** is a supplementary antecedent — an instruction, a gesture, a demonstration, physical guidance — added when the natural S^D does not yet evoke the behavior. "What do you say?" prompts "thank you" until the gift itself does the job. Prompts must be faded so control transfers to the natural stimulus; prompting without fading produces a learner who waits to be told.[9] ## What counts as a behavior? ### Operational definitions and the dead-man test A behavior is something the organism *does*: observable, measurable, and defined so that two observers would agree whether it occurred — objective, clear, and complete, in the textbook formula.[9] The quickest screen is Ogden Lindsley's **dead-man test** from 1965: if a dead man can do it, it isn't behavior; if a dead man can't, it is.[8] "Sit still," "don't interrupt," and "stop snacking" all fail; a corpse manages every one. They name the absence of behavior, and an absence cannot be reinforced. Rewrite them as what you want to *see*: "completes the worksheet while seated," "raises a hand and waits," "eats an apple at 3 p.m." ### Response classes and the dimensions of behavior An operant is a **response class**: every response that produces the same consequence, whatever its form. Pressing the lever with the left paw, the right paw, or the nose is one operant; asking, pointing, and grabbing are one operant if they all get the cookie. Behavior is defined by function, not by which muscles moved. It also has measurable dimensions — **frequency**, **duration**, **latency** (how long after the antecedent it begins), and **intensity** — and the problem decides which you track. A tantrum is a duration problem; a morning run is a latency problem, the minutes between alarm and front door. ### Why "be more productive" is not a behavior "Be more productive," "eat healthier," and "be a better listener" are labels for outcomes. They cannot be observed at a moment in time or counted, so nothing can be made contingent on them. The translation is always the same: find a specific, countable act that would, repeated, produce the outcome. "Open the document and write one sentence before checking email" has a start and an end, either happened or didn't, and can be followed by a consequence within seconds. Start with the smallest version, because that is the one that actually gets emitted — the logic of [shaping](https://operantconditioning.com/shaping/) and of every effective [habit protocol](https://operantconditioning.com/habits/). ## What is a consequence? A consequence is a stimulus change that follows a behavior. Its effect on the future frequency of the behavior — not its appearance or the intent of whoever delivered it — determines what it is. Consequences sort into four quadrants plus one non-event: | Consequence | What happens after the behavior | Effect | Example | | --- | --- | --- | --- | | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | A stimulus is added | Increases | Dog sits → treat | | [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) | A stimulus is removed | Increases | You buckle up → chime stops | | [Positive punishment](https://operantconditioning.com/positive-punishment/) | A stimulus is added | Decreases | You touch the pan → burn | | [Negative punishment](https://operantconditioning.com/negative-punishment/) | A stimulus is removed | Decreases | Teen breaks curfew → loses car keys | | [Extinction](https://operantconditioning.com/extinction/) | The usual reinforcer no longer follows | Decreases, after a burst | Elevator button → nothing → you stop pressing | ### Immediacy and contingency Two properties decide how much of the loop closes. **Immediacy**: a reinforcer's effect falls steeply as the delay grows from seconds to minutes, which is why trainers use a click and why a monthly paycheck reinforces almost nothing in particular.[10] **Contingency**: the consequence must depend on the behavior. Pigeons fed on a timer regardless of what they did developed "superstitious" rituals from accidental pairings, and adding free reinforcers that arrive whether or not the animal responds weakens the behavior.[11][12] ### The four-term contingency Because a reinforcer only reinforces when its value is high, a complete account adds the motivating operation: MO → S^D → behavior → consequence, the **four-term contingency**.[9] It is 6 p.m. and you last ate at noon (MO); you walk into the kitchen (S^D); you open the fridge (B); you eat (C). Remove the MO with a late lunch and the same kitchen evokes nothing. Practitioners reach for all four terms because the fourth is often the easiest to change. ## Changing the antecedent is the underrated lever When most people say "behavior change" they mean changing the consequence. That works, but it acts at the hardest moment — after the old habit has already run. The antecedent acts *before* the moment of weakness. Skinner listed "changing the stimulus" among the basic techniques of self-control: remove the cue for an unwanted behavior, or put the cue for a wanted one where you cannot miss it.[1] Charge the phone in the kitchen. Put the guitar on a stand, not in a case. Two research programs outside behavior analysis reach the same conclusion. Peter Gollwitzer's **implementation intentions** are plans of the form "when situation X arises, I will do Y." Forming one links a concrete cue to a concrete response in advance, so the cue evokes the behavior with little deliberation; a meta-analysis of 94 studies found a medium-to-large effect on goal attainment.[13] An implementation intention is a discriminative stimulus installed verbally — antecedent control's cognitive cousin. Wendy Wood's habit research shows the other side. Habits are cued by context and survive on context rather than intention: students who transferred universities kept their exercise, reading, and TV habits only when the new setting resembled the old one, and habitual cinema popcorn-eaters ate stale popcorn as readily as fresh in a cinema but not in a meeting room.[14][15][16] To keep a behavior, keep its cue constant; to lose one, change the cue. > **Rule of thumb** > > If you keep failing at the moment of choice, you are trying to solve an antecedent problem with a consequence. Move the intervention earlier: remove the cue, add a cue, or change the motivating operation. ## Why notifications are weak antecedents Held against the three-term contingency, most attempts to build a habit are all B: a behavior with a checkbox, no antecedent, and a consequence that may or may not do anything. Where there is an antecedent it is usually a push notification, and notifications are weak antecedents for three reasons. Habituation: a stimulus repeated without any differential consequence loses its power to evoke a response, which is why the fortieth buzz of the day is background noise.[17] Many people switch them off. And, least appreciated, a notification is a discriminative stimulus for *picking up the phone* — the behavior that reinforcement has actually followed — not for flossing. The fix is a cue that already occurs reliably in the world: after the coffee, when you sit in the car, when the meeting ends. [Building habits with all three terms ›](https://operantconditioning.com/habits/) ## Three worked ABC examples The model's value is that it hands you three places to intervene. Each example is analyzed once, then attacked at A, at B, and at C. ### A child's tantrum in the supermarket | Aspect | Antecedent | Behavior | Consequence | | --- | --- | --- | --- | | **Analysis** | Checkout aisle lined with candy; parent busy paying; child skipped a nap (MO) | Screaming, lying on the floor | Parent hands over candy. The tantrum is positively reinforced; the giving-in is negatively reinforced by the silence. | | **Intervene** | Shop after the nap; use self-checkout; state the rule before entering ("no candy today, you choose the cereal") | Teach and prompt a replacement: "Can I have a snack at home?" — reinforced every time at first | Candy never follows screaming (expect an [extinction burst](https://operantconditioning.com/extinction/)); attention and a small privilege follow calm behavior | ### An employee who keeps missing deadlines | Aspect | Antecedent | Behavior | Consequence | | --- | --- | --- | --- | | **Analysis** | Tasks assigned verbally, without written dates; competing requests from three managers | Submits work days late | Late work is accepted with a mild remark; on-time work draws no comment. Delay is negatively reinforced (the aversive task is postponed) and never costs anything. | | **Intervene** | Every task in writing with a date and a check-in three days before; one prioritized queue | Define it as "send the draft by Thursday noon," not "be more reliable" | Same-day acknowledgment of on-time delivery; a late submission triggers an immediate re-planning conversation, not a note in the [annual review](https://operantconditioning.com/applications/#workplace) | ### Your own doom-scrolling | Aspect | Antecedent | Behavior | Consequence | | --- | --- | --- | --- | | **Analysis** | Phone on the nightstand; in bed, lights off; tired (MO: fatigue makes escape from effort more valuable) | Unlock, open the feed, scroll | Unpredictable novelty on a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/); escape from boredom and unwelcome thoughts | | **Intervene** | Charge the phone in the kitchen; delete the app; leave a book on the pillow; plan "when I get into bed, I open the book" | Replace with a behavior that serves the same function — one page of a novel, a podcast | Log out after every session so the app opens to a login screen (response cost); mark the page read and let that be the consequence you chose | ## ABC data collection in practice ### The ABC recording chart **ABC recording** means writing down, for each occurrence of a target behavior, what came immediately before and after. Formalized for field research by Bijou, Peterson, and Ault in 1968, it is now the standard first step in schools and clinics.[18] | Date / time | Antecedent | Behavior | Consequence | Possible function | | --- | --- | --- | --- | --- | | Mon 9:05 | Teacher hands out a math worksheet | Jamal tears the worksheet and shouts | Sent to the hallway for 10 minutes | Escape from the task | | Mon 10:30 | Teacher helping another student | Jamal shouts | Teacher comes over and talks with him | Attention | | Tue 9:10 | Math worksheet, harder problems | Jamal pushes the worksheet off the desk | Aide removes the worksheet | Escape from the task | Ten to twenty entries usually reveal a pattern — here, shouting reliably follows academic demands and is reliably followed by their removal. Write what you saw rather than what you inferred, and fill in the "possible function" column after the observations, not before. ### Functional behavior assessment A **functional behavior assessment** (FBA) combines indirect methods (interviews, rating scales), descriptive methods (ABC recording), and sometimes experimental methods to reach a hypothesis about what a problem behavior gets or avoids; U.S. special-education law has required one in certain disciplinary situations since 1997. The output is a behavior intervention plan that changes the antecedents, teaches a replacement behavior, and rearranges the consequences — all three terms, deliberately. ### Functional analysis and the four functions of behavior ABC recording shows correlation; a **functional analysis** tests causation. In the landmark 1982 study, Brian Iwata and colleagues exposed nine children who injured themselves to systematically arranged conditions — adult attention contingent on self-injury, escape from demands contingent on it, being alone, and a play control — and measured self-injury in each.[19] The same behavior was maintained by attention in one child, escape in another, and its own sensory consequences in a third. A later summary of 152 such analyses found escape the most common function, followed by social-positive reinforcement (attention or tangibles) and automatic reinforcement.[20] Practitioners therefore speak of four common functions — **attention**, **escape**, **access to tangibles**, and **automatic** (sensory) reinforcement — and the function dictates the treatment. Escape-maintained shouting is treated by teaching the child to ask for a break and honoring the request, not by sending him to the hallway, which is the very consequence keeping the behavior alive.[21] [How applied behavior analysis uses this ›](https://operantconditioning.com/applications/#aba) ## Common misconceptions about the ABC model - **It is not the ABC model of cognitive-behavioral therapy.** Albert Ellis's ABC, from rational emotive behavior therapy, runs Activating event → Belief → emotional Consequence: the B is a thought and the C is a feeling.[22] In the behavioral ABC, the B is an observable act and the C is an environmental event that follows it. Both date from the 1950s; they describe different things. - **The antecedent does not cause the behavior.** It evokes it, and only because of a consequence history. Change the consequences and the same antecedent stops working. - **ABC data does not tell you the function.** Descriptive records suggest hypotheses; only a functional analysis tests them. - **"Consequence" does not mean punishment.** It includes reinforcement, punishment, and the absence of either. - **A → B is not classical conditioning.** A conditioned stimulus elicits a reflex regardless of what the organism does; a discriminative stimulus sets the occasion for a behavior whose fate is decided by what follows. ## Key takeaways - The three-term contingency is antecedent, behavior, consequence. The consequence happens because the behavior happened, and the behavior is emitted in the presence of the antecedent because that is where the consequence has followed before. - Antecedents come in kinds that are changed in different ways: a discriminative stimulus signals that a reinforcer is available, a motivating operation changes how much it is worth, and a prompt is a supplementary cue that must be faded. Adding the motivating operation gives the four-term contingency. - A behavior is something observable and countable that a dead man could not do, defined by its function rather than its form. "Be more productive" is an outcome label; "write one sentence before checking email" is a behavior. - A consequence is sorted by its effect on future frequency, not by its appearance or the deliverer's intent: reinforcement increases behavior, punishment decreases it, and extinction is the usual reinforcer no longer arriving. Immediacy and contingency decide how much of the loop closes. - Changing the antecedent is the underrated lever, because it acts before the moment of weakness. ABC recording reveals patterns, but only a functional analysis tests what a behavior gets or avoids, and the function dictates the treatment. ### Check yourself **A teacher sends Jamal to the hallway every time he shouts during math. Over the month, shouting increases. Is the hallway a punishment?** No. A consequence is defined by its effect on future behavior, not by how it looks, and shouting went up, so the hallway is reinforcing it: shouting produces escape from the math demand. The fix is to teach Jamal to ask for a break and honor the request, not to keep delivering the very consequence that maintains the behavior. **Your phone buzzes and you pick it up. Does the buzz cause the behavior the way a puff of air causes a blink?** No. The buzz is a discriminative stimulus: it evokes picking up the phone only because that behavior has been reinforced after buzzes before, whereas the air puff elicits a reflex regardless of any history. Change the consequences and the same buzz stops working. **It is 6 p.m., you last ate at noon, you walk into the kitchen, and you open the fridge. Which part is the discriminative stimulus and which is the motivating operation?** The kitchen is the discriminative stimulus, because it signals that food is available; the six hours without food are the motivating operation, because they make food worth more and evoke the behavior that has produced it. Remove the motivating operation with a late lunch and the same kitchen evokes nothing. **A parent's goal for their child is "stop interrupting." Is that a behavior you can reinforce?** No. A dead man can manage not interrupting, so it fails the dead-man test: it names the absence of behavior, and an absence cannot be reinforced. Rewrite it as something you want to see, such as "raises a hand and waits," and reinforce that. **Explain it to a friend.** Explain why the antecedent does not cause the behavior, using something you did today as the example and without using the words reinforcement or punishment. ## Frequently asked questions **What is the ABC model of behavior?** A framework for analyzing any behavior in three parts: the antecedent (what came before and set the occasion), the behavior itself, and the consequence (what followed and made the behavior more or less likely in future). It is the three-term contingency that B. F. Skinner placed at the center of operant conditioning. **What is the three-term contingency?** The relation between a discriminative stimulus, a response, and a reinforcing or punishing consequence — the same thing as the ABC model. "Contingency" means the consequence depends on the behavior, and the behavior comes to depend on the antecedent because of that history. **What is a discriminative stimulus?** A stimulus in whose presence a behavior has been reinforced, so that the behavior becomes more likely when the stimulus is present. A ringing phone is a discriminative stimulus for answering; a lit "Open" sign is one for entering. Its opposite, the S-delta, is a stimulus in whose presence the behavior has not paid off. **What is the difference between a discriminative stimulus and a motivating operation?** A discriminative stimulus signals that a reinforcer is *available*; a motivating operation changes how much the reinforcer is *worth* and evokes the behavior that has produced it. A vending machine is a discriminative stimulus for inserting coins; hunger is a motivating operation that makes the snack worth buying. **What is ABC data collection?** Recording, for each instance of a target behavior, what happened immediately before (antecedent) and immediately after (consequence), usually in a table with the date and time. After ten to twenty entries a pattern typically emerges that suggests what the behavior is getting or avoiding. It is the descriptive core of a functional behavior assessment. **What is the four-term contingency?** The three-term contingency with the motivating operation added in front: MO → discriminative stimulus → behavior → consequence. It recognizes that a reinforcer only reinforces when the organism's state makes it valuable, so a full analysis includes deprivation, satiation, and similar conditions. **Is the ABC model the same as the ABC model in CBT?** No. In Albert Ellis's cognitive model, A is an activating event, B is a belief about it, and C is the emotional consequence of that belief. In the behavioral ABC model, B is an observable behavior and C is an environmental event that follows it and changes its future frequency. Same letters, different variables. **How do I use the ABC model to change my own behavior?** Write out the antecedent, behavior, and consequence for the behavior you want to change, then intervene at all three points: change the cue, define the behavior you want in small and countable terms, and arrange an immediate consequence that actually matters to you. [The full self-management protocol ›](https://operantconditioning.com/habits/) ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Guttman, N., & Kalish, H. I. (1956). Discriminability and stimulus generalization. *Journal of Experimental Psychology, 51*(1), 79–88. 4. Michael, J. (1982). Distinguishing between discriminative and motivational functions of stimuli. *Journal of the Experimental Analysis of Behavior, 37*(1), 149–155. 5. Michael, J. (1993). Establishing operations. *The Behavior Analyst, 16*(2), 191–206. 6. Laraway, S., Snycerski, S., Michael, J., & Poling, A. (2003). Motivating operations and terms to describe them: Some further refinements. *Journal of Applied Behavior Analysis, 36*(3), 407–414. 7. Wahler, R. G., & Fox, J. J. (1981). Setting events in applied behavior analysis: Toward a conceptual and methodological expansion. *Journal of Applied Behavior Analysis, 14*(3), 327–338. 8. Lindsley, O. R. (1991). From technical jargon to plain English for application. *Journal of Applied Behavior Analysis, 24*(3), 449–458. 9. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 10. Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139. 11. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 12. Hammond, L. J. (1980). The effect of contingency upon the appetitive conditioning of free-operant behavior. *Journal of the Experimental Analysis of Behavior, 34*(3), 297–304. 13. Gollwitzer, P. M. (1999). Implementation intentions: Strong effects of simple plans. *American Psychologist, 54*(7), 493–503. See also Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. *Advances in Experimental Social Psychology, 38*, 69–119. 14. Wood, W., Tam, L., & Witt, M. G. (2005). Changing circumstances, disrupting habits. *Journal of Personality and Social Psychology, 88*(6), 918–933. 15. Neal, D. T., Wood, W., Wu, M., & Kurlander, D. (2011). The pull of the past: When do habits persist despite conflict with motives? *Personality and Social Psychology Bulletin, 37*(11), 1428–1437. 16. Wood, W., & Rünger, D. (2016). Psychology of habit. *Annual Review of Psychology, 67*, 289–314. 17. Rankin, C. H., Abrams, T., Barry, R. J., et al. (2009). Habituation revisited: An updated and revised description of the behavioral characteristics of habituation. *Neurobiology of Learning and Memory, 92*(2), 135–138. 18. Bijou, S. W., Peterson, R. F., & Ault, M. H. (1968). A method to integrate descriptive and experimental field studies at the level of data and empirical concepts. *Journal of Applied Behavior Analysis, 1*(2), 175–191. 19. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1982/1994). Toward a functional analysis of self-injury. *Analysis and Intervention in Developmental Disabilities, 2*(1), 3–20. Reprinted in *Journal of Applied Behavior Analysis, 27*(2), 197–209. 20. Iwata, B. A., Pace, G. M., Dorsey, M. F., et al. (1994). The functions of self-injurious behavior: An experimental-epidemiological analysis. *Journal of Applied Behavior Analysis, 27*(2), 215–240. 21. Carr, E. G., & Durand, V. M. (1985). Reducing behavior problems through functional communication training. *Journal of Applied Behavior Analysis, 18*(2), 111–126. 22. Ellis, A. (1962). *Reason and Emotion in Psychotherapy*. Lyle Stuart. ## Related - [Build habits with operant conditioning](https://operantconditioning.com/habits/): The A-B-C protocol applied to yourself, step by step. - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The consequence that does most of the work, and how to deliver it well. - [Operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/): Why an antecedent that evokes is not a stimulus that elicits. --- # Operant vs. Classical Conditioning: The Difference, With Examples > Classical conditioning pairs stimuli to transfer a reflex; operant conditioning changes behavior through consequences. Comparison table, examples, checklist. - Source: https://operantconditioning.com/operant-vs-classical-conditioning/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Comparison* Pavlov's dogs and Skinner's rats learned two different things. Here is the difference between classical and operant conditioning, the science behind each, how they work together in real life, and a checklist for telling them apart in the wild. > **Definitions** > > **Classical conditioning** (also Pavlovian or respondent conditioning) is learning in which a neutral stimulus comes to elicit a reflexive response because it reliably predicts a stimulus that already elicits that response. The organism learns that one stimulus signals another; the key event comes *before* the response.[1] > > **Operant conditioning** (also instrumental conditioning) is learning in which a voluntary behavior becomes more or less frequent because of the consequences that follow it. The organism learns that its own behavior produces an outcome; the key event comes *after* the response.[2] **In brief** - In classical conditioning a stimulus that predicts another comes to elicit a reflex; the key event comes before the response and the organism is passive. - In operant conditioning a voluntary behavior changes because of the consequence that follows it and depends on it. - Real scenarios usually contain both: a classically conditioned emotion and an operant action, so name each part rather than forcing one label. ## Operant vs. classical conditioning: the short answer In classical conditioning, two stimuli are paired and a reflex transfers from one to the other. Pavlov's dogs already salivated at food; after a metronome repeatedly preceded food, they salivated at the metronome. Nothing the dog *did* changed what happened — food came regardless. In operant conditioning, a behavior is followed by a consequence and the behavior changes as a result. Skinner's rats pressed a lever, food arrived *because* they pressed, and pressing increased. The dog learned "metronome means food." The rat learned "pressing gets food." That difference — signal versus consequence, elicited reflex versus emitted action — is the whole distinction, and everything else in this article is a consequence of it.[3]  *The two kinds of learning as sequences. In classical conditioning a stimulus comes to elicit a reflex; in operant conditioning a consequence changes how often a voluntary behavior recurs.* ## The difference between classical and operant conditioning, side by side | Aspect | Classical (Pavlovian) conditioning | Operant (instrumental) conditioning | | --- | --- | --- | | **What is learned** | A relation between two stimuli: the CS predicts the US | A relation between a behavior and its consequence, in a context | | **Type of behavior** | Reflexive, involuntary: salivation, blinking, nausea, fear, arousal, heart rate | Voluntary, "emitted": pressing, walking, speaking, studying, scrolling | | **Role of the organism** | Passive — the stimuli arrive whatever it does | Active — the consequence depends on what it does | | **Timing of the key event** | The conditioned stimulus comes *before* the response, usually by less than a few seconds (hours, in taste aversion) | The consequence comes *after* the response, ideally within seconds | | **Is the outcome contingent on behavior?** | No. Food follows the signal regardless of salivation | Yes. Food follows the press only if the press occurs | | **Key figures** | Ivan Pavlov (1890s–1927); John B. Watson (1920); Robert Rescorla (1960s–1980s) | Edward Thorndike (1898); B. F. Skinner (1937–1990) | | **Key terms** | Unconditioned stimulus and response (US, UR); conditioned stimulus and response (CS, CR); neutral stimulus | Reinforcement, punishment (positive and negative); discriminative stimulus (S^D); schedules; shaping | | **Typical laboratory preparation** | Dog in a harness with a salivary fistula; rabbit eyeblink conditioning; conditioned suppression in rats | Thorndike's puzzle box; Skinner's operant chamber with a lever or key and a cumulative recorder | | **Acquisition** | CR strengthens over repeated CS–US pairings that carry predictive information | Response rate rises as responses are reinforced; complex behavior is built by shaping | | **Extinction** | Present the CS without the US; the CR declines | Stop delivering the reinforcer after the response; the response declines, often after a burst | | **Spontaneous recovery** | Yes — the CR returns partially after a rest | Yes — the response returns partially after a rest | | **Generalization and discrimination** | CR spreads to similar stimuli; discrimination training narrows it | Behavior spreads to similar contexts; discrimination training brings it under stimulus control | | **Everyday examples** | Flinching at the dentist's drill; nausea in the chemo clinic waiting room; excitement at the sound of a treat bag | Studying for grades; buckling up to stop the chime; a dog sitting for a treat; checking a phone for likes | ## How to tell which is which: a decision checklist 1. **Name the response.** Is it something the organism does with its skeletal muscles — walks, presses, says, buys — or something that happens to it — salivates, flinches, feels sick, feels afraid, heart races? Voluntary points to operant; reflexive or emotional points to classical. 2. **Find the key event and check its timing.** Does the important stimulus come *before* the response as a signal, or *after* it as a result? Before is classical; after is operant. 3. **Test contingency.** Would the outcome have happened anyway? If food comes whether or not the dog salivates, it is classical. If food comes only if the dog sits, it is operant. 4. **State what was learned in one sentence.** "X predicts Y" is classical. "Doing X produces Y" is operant. 5. **Look for both.** Most real scenarios contain a classically conditioned emotion and an operant action. Name each part separately rather than forcing one label on the whole scene. ## Worked examples: classical or operant? | Scenario | Answer | Why | | --- | --- | --- | | You tense up when you hear a dentist's drill. | **Classical** | The drill sound (CS) preceded pain (US) in the past; tension (CR) is elicited, not chosen, and comes before anything you do. | | A teenager cleans his room and gets the car keys for the evening. | **Operant** (positive reinforcement) | A voluntary behavior is followed by an added consequence that depends on it; cleaning increases. | | A cat comes running when it hears the can opener. | **Both** | The sound is a CS that elicits excitement and salivation (classical). Running to the kitchen is an operant reinforced by food (operant). | | A chemotherapy patient feels nauseated in the clinic waiting room. | **Classical** | Clinic cues (CS) preceded the drug (US) that caused nausea (UR); now the cues elicit anticipatory nausea (CR). Nothing the patient does changes the outcome. | | A student stops raising her hand after the teacher never calls on her. | **Operant** (extinction) | A behavior that was once reinforced by being called on no longer is, and it declines. | | Your smoke alarm shrieks every time you make toast. Now you flinch when you push the lever down — and you open a window before you start. | **Both** | Flinching at the lever is a CR to a CS (classical). Opening the window is an operant, negatively reinforced by preventing the alarm (avoidance). | | You got a stomach bug hours after eating clams and now can't stand the smell of them. | **Classical** (taste aversion) | One pairing, a long delay, and a response — disgust — you cannot decide not to have. Biological preparedness at work. | | A puppy wags and drools when it sees the treat pouch, then sits when asked and gets a treat. | **Both** | The wagging and drooling are CRs to the pouch (classical). The sit is an operant reinforced by the treat (operant). The pouch is also becoming an S^D for sitting. | More practice: the [examples page](https://operantconditioning.com/examples/) has more than fifty operant scenarios sorted by quadrant, and the [quiz](https://operantconditioning.com/quiz/) mixes classical and operant items with instant explanations. ## Where people mix them up - **Deciding by whether the stimulus is pleasant.** Both kinds of conditioning use pleasant and unpleasant stimuli. Decide by timing and contingency, not by valence. - **Calling the bell the unconditioned stimulus.** The US is the stimulus that works *without* learning — the food. The bell (or metronome) is neutral, then conditioned. - **Assuming the CR is a copy of the UR.** Often it is weaker; sometimes it is the opposite (Siegel's compensatory responses), and the CR to a shock-predicting tone is freezing, not the jump the shock itself produces. - **Filing negative reinforcement under classical conditioning** because it "involves something unpleasant." Negative reinforcement is operant: a behavior removes an aversive stimulus and increases. Classical conditioning does not have reinforcement or punishment in Skinner's sense at all. - **Confusing the CS with the discriminative stimulus.** A CS elicits a reflex regardless of behavior; an S^D signals that a behavior will be reinforced. - **Treating classical conditioning as mere pairing.** Since Rescorla, the CS must *predict* the US. Pairings without predictive value produce little learning. - **Treating extinction as forgetting.** In both kinds of conditioning, extinction is new learning that inhibits the old, which is why spontaneous recovery and renewal occur. - **Forcing one label on a scene that has both.** "The dog gets excited at the leash and then sits for a treat" contains a classical part and an operant part. Full credit requires naming each. For the people, dates, and disputes behind both traditions, see the [history of operant conditioning](https://operantconditioning.com/history/). ## Going further The sections below go past the short answer: how classical conditioning really works, where the two kinds of learning meet, and the cases that blur the line. ## Classical conditioning, explained properly ### Pavlov's discovery Ivan Pavlov was a physiologist studying digestion — work that earned him the 1904 Nobel Prize — when he noticed that his dogs began salivating before food arrived: at the sight of the food dish, at the footsteps of the attendant. He called these "psychic secretions" and spent the rest of his career studying them with the rigor of a physiologist, using a surgically implanted tube to measure drops of saliva. In many of his experiments the signal was a metronome, a buzzer, a light, or a touch rather than the bell of legend. The results were published in English in 1927 as *Conditioned Reflexes*.[1] (Pavlov's own word was "conditional" — the reflex was conditional on the pairing — and "conditioned" is an early translation that stuck.) ### The four terms - **Unconditioned stimulus (US or UCS):** a stimulus that elicits a response without any learning. Food in the mouth. - **Unconditioned response (UR or UCR):** the unlearned response to it. Salivation to food. - **Conditioned stimulus (CS):** a previously neutral stimulus that, after predicting the US, elicits a response on its own. The metronome. - **Conditioned response (CR):** the learned response to the CS. Salivation to the metronome — usually similar to the UR, but often weaker and sometimes different in form. ### What Pavlov found **Acquisition** is gradual: the CR grows over pairings. **Extinction** follows when the CS is presented repeatedly without the US — but Pavlov noticed that an extinguished response reappears after a rest (**spontaneous recovery**), which told him extinction was new learning laid over the old, not erasure. Modern work confirms this: extinguished responses also return when the context changes (renewal) or when the US is encountered again (reinstatement).[4] A CR trained to one tone appears, more weakly, to similar tones (**generalization**); pairing one tone with food and another with nothing narrows the response to the first (**discrimination**). When Pavlov's laboratory made a circle-versus-ellipse discrimination progressively harder, a previously calm dog became agitated and uncooperative — the first "experimental neurosis."[1] Finally, an established CS can itself condition a new stimulus (**higher-order conditioning**), which is how a word like "dinner" ends up doing what the metronome did. ### Little Albert: the famous study and its problems In 1920 John B. Watson and Rosalie Rayner reported conditioning fear in an infant, "Albert B.," about eleven months old. Albert initially reached for a white rat without fear. Watson then struck a steel bar with a hammer behind Albert's head whenever the rat appeared. After a handful of pairings, Albert cried and turned away at the sight of the rat alone, and his distress generalized to a rabbit, a dog, a fur coat, and a Santa Claus mask.[5] The study is in every textbook, and it should be read with its problems attached. Ethically, it deliberately induced fear in an infant who could not consent, and Watson and Rayner made no attempt to remove the fear before Albert left the hospital, though they knew in advance when he would leave. Methodologically, it was a single case with no control condition; fear was rated subjectively; some of Albert's reactions were mild or inconsistent; and the responses were "freshened up" with additional pairings between tests. A 1979 review found that textbooks had for decades embellished the results, reporting deconditioning that never happened and generalization that was never tested.[6] Albert's real identity has been the subject of competing claims by historians. What survives is the modest, real finding: a fear response can be conditioned to a neutral stimulus in a human infant, and it generalizes. ### It's not what you think it is: contingency, not pairing The textbook story — "pair two stimuli enough times and the reflex transfers" — turns out to be wrong in an important way. In 1968 Robert Rescorla gave rats a tone followed by shock, but for some groups he added shocks during the silent periods too, so that the tone no longer *predicted* any change in the likelihood of shock. Those rats received exactly as many tone–shock pairings as the others, and they learned almost nothing.[7] Leon Kamin's blocking experiment made the same point from another direction: if a light already predicts shock, adding a tone alongside it teaches the animal nothing about the tone, because the shock is no longer surprising.[8] Rescorla and Allan Wagner turned these findings into a mathematical model in which learning is driven by prediction error — how much the outcome differs from what was expected.[9] Rescorla summarized the modern view in a 1988 paper whose title says it all: "Pavlovian conditioning: It's not what you think it is." Classical conditioning is not the mechanical transfer of a reflex by contiguity; it is the organism learning the predictive structure of its environment — which events signal which others — and adjusting a whole set of responses accordingly.[10] ### Taste aversion and biological preparedness The other crack in the simple story came from John Garcia. In 1966 Garcia and Robert Koelling let rats drink "bright, noisy" water — sweetened, and accompanied by a light and a click with every lick. Rats that were then made ill (with X-rays or lithium chloride) later avoided the sweet taste but drank the noisy, bright water happily. Rats that were shocked instead avoided the light and click but not the taste.[11] The animals were not equally ready to associate any stimulus with any outcome: tastes go with illness, sights and sounds go with pain. Taste aversion also broke the timing rule — it formed after a single trial with delays of an hour or more between taste and illness. Martin Seligman called this **preparedness**: evolution has made some associations easy to learn and others nearly impossible.[12] ## Operant conditioning, briefly Edward Thorndike's cats, escaping from puzzle boxes, showed that responses followed by satisfying outcomes are "stamped in" — the law of effect.[13] [B. F. Skinner](https://operantconditioning.com/bf-skinner/) made this a laboratory science: an animal in a chamber, a lever or key, a consequence delivered by the apparatus, and a cumulative record of responses over time. In a 1937 paper he formally separated the two kinds of learning, calling Pavlov's "Type S" (stimulus-elicited, respondent) and his own "Type R" (response-emitted, operant).[3] Every consequence falls into one of four quadrants — [positive reinforcement](https://operantconditioning.com/positive-reinforcement/), [negative reinforcement](https://operantconditioning.com/negative-reinforcement/), positive punishment, negative punishment — and how often it arrives is governed by [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/). Behavior that no longer pays off undergoes [extinction](https://operantconditioning.com/extinction/), complex behavior is built by [shaping](https://operantconditioning.com/shaping/), and the [antecedent](https://operantconditioning.com/abc-model/) that signals when a behavior will be reinforced brings it under stimulus control. [The complete guide to operant conditioning ›](https://operantconditioning.com/) > **The CS and the S^D are not the same thing** > > Both are "cues," which is why students confuse them. A **conditioned stimulus** *elicits* a response: the metronome makes the dog salivate whether or not it does anything. A **[discriminative stimulus](https://operantconditioning.com/glossary/#discriminative-stimulus)** *sets the occasion* for a response: the green light tells the rat that pressing will now be reinforced, but the food still depends on the press. The test is contingency. If the outcome arrives no matter what the organism does, the cue is a CS. If the outcome depends on the behavior, the cue is an S^D. ## How classical and operant conditioning work together Outside the laboratory the two processes almost never run alone, and the most useful accounts of real behavior combine them. ### Two-factor theory of avoidance Why does a rat keep jumping a barrier to avoid a shock that never comes any more? O. H. Mowrer's answer was two processes in sequence: first, a warning signal is classically conditioned to elicit fear; second, the avoidance response is operantly reinforced by escape from that fear.[14] The same logic explains why **phobias** persist: the fear may be acquired classically (a dog bite, or even a frightening story), but it is maintained operantly, because avoiding dogs is negatively reinforced by relief and the fear never gets a chance to extinguish. Exposure therapy is the deliberate blocking of the operant half so the classical half can extinguish. ### Conditioned reinforcers are classically conditioned A clicker is silent to a dog that has never heard one paired with food. Pair click and treat a few dozen times and the click acquires two properties at once. It is a **CS**: it elicits the anticipatory excitement that food elicits. And it is a **conditioned reinforcer**: delivered right after a behavior, it strengthens that behavior, bridging the seconds until the treat arrives.[15] Money, praise, grades, and the green checkmark work the same way — classically conditioned value, operantly deployed. [How clicker training uses both ›](https://operantconditioning.com/dog-training/) ### Conditioned suppression One of the cleanest laboratory demonstrations of the two processes interacting is also one of the oldest. In 1941 Estes and Skinner had rats pressing a lever for food, then sounded a tone that ended in shock. As the tone acquired fear (classical), the rats' ongoing lever pressing slowed or stopped during it (operant behavior suppressed).[16] The "conditioned emotional response" became a standard way to measure Pavlovian fear — by its effect on operant behavior. ### Addiction cues Drug taking is operant: the drug's effects reinforce the behavior that produces them. But the paraphernalia, the place, and the people become classically conditioned cues. Shepard Siegel showed that rats given morphine in a familiar environment developed conditioned *compensatory* responses — the body preparing to counteract the drug — so that tolerance was partly a learned response to the setting, and the same dose in a new place was more dangerous.[17] Craving triggered by cues, and relapse in old environments, is classical conditioning driving the person back toward the operant. ### Pavlovian-instrumental transfer A classically conditioned cue can also change the *vigor* of operant behavior without ever having been part of it. Train a rat to press a lever for food; separately, in another session, pair a tone with free food. Then play the tone while the rat is pressing: pressing speeds up, even though the tone was never a signal for pressing and pressing has never paid off during it. William Estes reported the effect in 1948, and it is now called **Pavlovian-instrumental transfer** (PIT).[18][19] It is the laboratory version of a familiar experience: the smell of the bakery does not teach you to walk in, but it makes you walk in faster. In addiction research PIT is one of the main models of how drug cues energize drug seeking. ## Behavior without reinforcement? Autoshaping and contrafreeloading Some observations look, at first, like operant behavior that no reinforcement produced, and they mark the edge of the law of effect. ### Autoshaping and sign-tracking In 1968 Brown and Jenkins lit a pigeon's response key for a few seconds and then delivered grain — whether or not the bird did anything. After a few dozen pairings the pigeons began pecking the lit key. Nobody had shaped the peck; the bird had "auto-shaped" it.[20] The following year Williams and Williams arranged that pecking the key *cancelled* the grain, so that the only way to be fed was not to peck. The pigeons kept pecking — less, but persistently — and lost food for it.[21] This **omission** result rules out reinforcement as the cause: the pecks were being punished by food loss and continued anyway. Jenkins and Moore then showed that the form of the peck matched the reinforcer — birds autoshaped with grain pecked the key as if eating it, birds autoshaped with water pecked as if drinking.[22] The behavior is directed at the signal as if it were the reward, which is why Hearst and Jenkins named it **sign-tracking**.[23] The accepted interpretation is that autoshaping is classical conditioning: the key light is a CS, the grain a US, and approaching and pecking a food-predicting stimulus is the conditioned response, as inevitable in a pigeon as salivation in a dog. The autoshaping procedure has, in fact, become one of the standard ways to *measure* Pavlovian conditioning. It also uncovered stable individual differences: some rats become sign-trackers who approach the cue, others goal-trackers who go straight to the food cup, and dopamine appears to be required for the first kind of learning but not the second.[24] The lesson for the operant–classical distinction is not that the law of effect is wrong but that a response can look operant and be Pavlovian, and the experimenter's job is to find out which contingency is actually controlling it. ### Contrafreeloading Give a rat free food in a dish and a lever that delivers the same food, and it will press the lever for a substantial share of its meals. Jensen reported the effect in 1963, and it has been found in most species tested, with cats the notable exception.[25][26] Working for food that is freely available is not what a naive reading of reinforcement predicts. It is compatible with a fuller one: the opportunity to explore, manipulate, and gather information is itself reinforcing, and a well-designed environment for a captive animal — or a person — is one that lets it work. ## Key takeaways - Classical conditioning pairs two stimuli so that a reflex transfers from one to the other; operant conditioning follows a behavior with a consequence that changes its future rate. Signal versus consequence, elicited reflex versus emitted action, is the whole distinction. - To tell them apart, name the response (reflexive or voluntary), check whether the key event comes before or after it, and test contingency. If the outcome would have happened anyway, it is classical; if it depends on the behavior, it is operant. - A conditioned stimulus elicits a response regardless of behavior; a discriminative stimulus sets the occasion for a behavior whose outcome still depends on it. Negative reinforcement is operant, not classical, because a behavior removes the aversive stimulus and increases. - Classical conditioning is prediction, not mere pairing: a stimulus that does not predict the outcome teaches almost nothing. Taste aversion shows that some associations are biologically prepared, forming in one trial across delays of an hour or more. - The two processes usually work together. A phobia is acquired classically and maintained operantly by avoidance, a clicker is a conditioned stimulus and a conditioned reinforcer at once, and autoshaping shows that a response can look operant and be Pavlovian. ### Check yourself **You buckle your seat belt to stop the chime. Because the chime is unpleasant, a classmate files this under classical conditioning. Is that right?** No. This is operant conditioning, specifically negative reinforcement: a voluntary behavior removes an aversive stimulus and becomes more frequent. Whether a stimulus is pleasant or unpleasant does not decide the type of learning; timing and contingency do, and classical conditioning has no reinforcement in Skinner's sense at all. **A pigeon pecks a lit key that is followed by grain no matter what the bird does. Since the peck looks like a lever press, is it operant behavior?** No. This is autoshaping, and the accepted interpretation is classical conditioning: the key light is a conditioned stimulus for grain, and pecking a food-predicting signal is the conditioned response. Pigeons keep pecking even when a peck cancels the grain, which rules out reinforcement as the cause. **A cat comes running when it hears the can opener. Classical or operant?** Both. The sound is a conditioned stimulus that elicits excitement and salivation, which is classical; running to the kitchen is an operant reinforced by food. Full credit means naming each part rather than forcing one label on the scene. **In Pavlov's experiment, which is the unconditioned stimulus: the metronome or the food?** The food. The unconditioned stimulus is the one that elicits the response without any learning, and food in the mouth produces salivation from the start. The metronome begins as a neutral stimulus and becomes a conditioned stimulus only after it reliably predicts the food. **Explain it to a friend.** Explain the difference between the two kinds of conditioning using a single everyday example that contains both, and say which part is which without using the words voluntary or involuntary. ## Frequently asked questions **What is the main difference between classical and operant conditioning?** Classical conditioning pairs two stimuli so that an involuntary response (salivation, fear, nausea) transfers from one to the other; the signal comes before the response and the outcome does not depend on what the organism does. Operant conditioning changes a voluntary behavior through the consequence that follows it; the consequence comes after the behavior and depends on it. **Is Pavlov's dog classical or operant conditioning?** Classical. The metronome or bell predicted food, and the dog's salivation transferred to the signal. The dog did not have to do anything for the food to arrive. If the dog had been required to press a lever to get food, that would be operant conditioning. **Can classical and operant conditioning happen at the same time?** Yes, and in real life they usually do. A clicker is a classically conditioned stimulus and an operant reinforcer at once. A phobia is typically acquired classically and maintained operantly by avoidance. Mowrer's two-factor theory of avoidance is built on exactly this combination. **Is negative reinforcement classical or operant conditioning?** Operant. Negative reinforcement means a behavior removes or prevents an aversive stimulus and becomes more frequent — buckling a seat belt to stop the chime. Classical conditioning does not involve reinforcement or punishment of behavior at all; it involves one stimulus coming to predict another. **What is an example of classical conditioning in everyday life?** Feeling your mouth water when you smell bread baking; flinching at the sound of a dentist's drill; feeling anxious when you hear the ringtone assigned to your boss; a dog getting excited at the jingle of the leash; feeling queasy at the sight of a food that once made you ill. **Which is stronger, classical or operant conditioning?** Neither — they do different jobs. Classical conditioning is the fastest way to attach an emotional or physiological response to a cue, sometimes in one trial. Operant conditioning is the only way to build a new voluntary skill or change how often someone does something. Most effective behavior change uses both: make the cue mean something, and make the behavior pay off. **What is the difference between a conditioned stimulus and a discriminative stimulus?** A conditioned stimulus elicits a reflexive response by itself, because it predicts an unconditioned stimulus; the dog salivates at the tone whatever it does. A discriminative stimulus signals that a voluntary behavior will now be reinforced; the light tells the rat that pressing will produce food, but it still has to press. **Who discovered classical and operant conditioning?** Ivan Pavlov described classical conditioning in dogs beginning in the 1890s and published *Conditioned Reflexes* in 1927. Edward Thorndike described the law of effect in 1898; B. F. Skinner named operant conditioning in 1937 and developed the experimental science of it from *The Behavior of Organisms* (1938) onward. ## References 1. Pavlov, I. P. (1927). *Conditioned Reflexes: An Investigation of the Physiological Activity of the Cerebral Cortex* (G. V. Anrep, Trans.). Oxford University Press. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279. 4. Bouton, M. E. (2004). Context and behavioral processes in extinction. *Learning & Memory, 11*(5), 485–494. 5. Watson, J. B., & Rayner, R. (1920). Conditioned emotional reactions. *Journal of Experimental Psychology, 3*(1), 1–14. 6. Harris, B. (1979). Whatever happened to Little Albert? *American Psychologist, 34*(2), 151–160. 7. Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. *Journal of Comparative and Physiological Psychology, 66*(1), 1–5. 8. Kamin, L. J. (1969). Predictability, surprise, attention, and conditioning. In B. A. Campbell & R. M. Church (Eds.), *Punishment and Aversive Behavior* (pp. 279–296). Appleton-Century-Crofts. 9. Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), *Classical Conditioning II: Current Research and Theory* (pp. 64–99). Appleton-Century-Crofts. 10. Rescorla, R. A. (1988). Pavlovian conditioning: It's not what you think it is. *American Psychologist, 43*(3), 151–160. 11. Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. *Psychonomic Science, 4*(1), 123–124. 12. Seligman, M. E. P. (1970). On the generality of the laws of learning. *Psychological Review, 77*(5), 406–418. 13. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 14. Mowrer, O. H. (1947). On the dual nature of learning — a re-interpretation of "conditioning" and "problem-solving." *Harvard Educational Review, 17*, 102–148. 15. Williams, B. A. (1994). Conditioned reinforcement: Experimental and theoretical issues. *The Behavior Analyst, 17*(2), 261–285. 16. Estes, W. K., & Skinner, B. F. (1941). Some quantitative properties of anxiety. *Journal of Experimental Psychology, 29*(5), 390–400. 17. Siegel, S. (1975). Evidence from rats that morphine tolerance is a learned response. *Journal of Comparative and Physiological Psychology, 89*(5), 498–506. 18. Estes, W. K. (1948). Discriminative conditioning II: Effects of a Pavlovian conditioned stimulus upon a subsequently established operant response. *Journal of Experimental Psychology, 38*(2), 173–177. 19. Lovibond, P. F. (1983). Facilitation of instrumental behavior by a Pavlovian appetitive conditioned stimulus. *Journal of Experimental Psychology: Animal Behavior Processes, 9*(3), 225–247. 20. Brown, P. L., & Jenkins, H. M. (1968). Auto-shaping of the pigeon's key-peck. *Journal of the Experimental Analysis of Behavior, 11*(1), 1–8. 21. Williams, D. R., & Williams, H. (1969). Auto-maintenance in the pigeon: Sustained pecking despite contingent non-reinforcement. *Journal of the Experimental Analysis of Behavior, 12*(4), 511–520. 22. Jenkins, H. M., & Moore, B. R. (1973). The form of the auto-shaped response with food or water reinforcers. *Journal of the Experimental Analysis of Behavior, 20*(2), 163–181. 23. Hearst, E., & Jenkins, H. M. (1974). *Sign-Tracking: The Stimulus-Reinforcer Relation and Directed Action*. Psychonomic Society. 24. Flagel, S. B., Clark, J. J., Robinson, T. E., Mayo, L., Czuj, A., Willuhn, I., Akers, C. A., Clinton, S. M., Phillips, P. E. M., & Akil, H. (2011). A selective role for dopamine in stimulus–reward learning. *Nature, 469*(7328), 53–57. 25. Jensen, G. D. (1963). Preference for bar pressing over "freeloading" as a function of number of rewarded presses. *Journal of Experimental Psychology, 65*(5), 451–454. 26. Inglis, I. R., Forkman, B., & Lazarus, J. (1997). Free food or earned food? A review and fuzzy model of contrafreeloading. *Animal Behaviour, 53*(6), 1171–1191. ## Related - [Operant conditioning: the complete guide](https://operantconditioning.com/): The four quadrants, schedules, extinction, shaping, and how to use it on yourself. - [History of operant conditioning](https://operantconditioning.com/history/): Thorndike, Watson, Skinner, the cognitive revolution, and reinforcement learning. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): Escape, avoidance, and the two-factor theory in depth. --- # Avoidance Learning: Escape, Avoidance, and Why It Persists > Escape and avoidance learning: signaled and Sidman avoidance, the avoidance paradox, two-factor vs. one-factor theory, and learned helplessness. - Source: https://operantconditioning.com/avoidance-learning/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Negative reinforcement · Escape and avoidance* A dog that learned to jump a barrier after a handful of shocks kept jumping for hundreds of trials without ever being shocked again. How can the absence of something reinforce behavior? The answer took thirty years and explains phobias, procrastination, and safety rituals. > **Definition** > > **Escape learning** is operant learning in which a behavior *terminates* an aversive stimulus that is already present. **Avoidance learning** is operant learning in which a behavior *prevents or postpones* an aversive stimulus that has not yet occurred. Both are forms of [negative reinforcement](https://operantconditioning.com/negative-reinforcement/): the behavior increases because something is removed or kept away.[1] > > Shielding your eyes from the sun is escape; putting on sunglasses before you go outside is avoidance. Almost every avoidance response begins its life as an escape response that got earlier. **In brief** - Escape ends an aversive stimulus that is already present; avoidance prevents or postpones one that has not yet occurred, and both are [negative reinforcement](https://operantconditioning.com/negative-reinforcement/). - In successful avoidance the consequence is that nothing happens, yet avoidance is among the most persistent behavior known: the avoidance paradox. - Avoidance persists because a successful avoider never tests the contingency, which is why blocking the response works when simple extinction does not. ## Escape: the easy case Escape is learned quickly because the organism feels the contrast directly: the shock is on, the lever is pressed, the shock is off. Thorndike's cats escaping from puzzle boxes were the first laboratory case, and the relief of a headache after aspirin is the everyday one.[2] Nothing about escape is paradoxical. The consequence is present, immediate, and obviously reinforcing. ## Discriminated (signaled) avoidance The classic procedure adds a warning signal. A light or tone comes on; a few seconds later a shock begins; a response made during the signal cancels the shock, and a response made after the shock begins ends it. Early trials are escape trials. As learning progresses the response moves earlier, into the signal, and the animal stops receiving shocks altogether: the trials have become avoidance trials. Richard Solomon and Lyman Wynne ran the definitive version in 1953 with dogs in a two-compartment shuttle box. After a few intense shocks the dogs began jumping the barrier during the signal, their latencies shortened to a second or two, and they continued to jump on trial after trial — hundreds of them — without receiving another shock. Several dogs never received more than a handful of shocks in the whole experiment.[3] When the shock generator was later switched off entirely, the jumping persisted; ordinary extinction barely touched it. What worked best was a glass barrier that physically prevented the jump, so that the dogs had to stay in the compartment and discover that no shock came — most effectively combined with a shock for jumping — and even then several dogs never fully stopped, which led Solomon and Wynne to speak of the "partial irreversibility" of the response.[4][12] ## Free-operant (Sidman) avoidance In 1953 Murray Sidman removed the warning signal altogether. Rats received brief shocks on a timer — say, every 5 seconds — unless they pressed a lever, in which case the next shock was postponed by a fixed period. Two intervals define the procedure: the **shock–shock (S–S) interval**, the time between shocks if the animal does nothing, and the **response–shock (R–S) interval**, the shock-free period each response buys. Every press restarts the R–S clock.[5][6] With nothing to warn them, rats still learned to press at a steady rate that kept shocks rare, and response rate depended in an orderly way on both intervals. Sidman avoidance mattered because it seemed to strip the procedure to its bones: no signal, so no fear of a signal, and yet learning. What exactly was being reinforced? ## The avoidance paradox Reinforcement, by definition, is a consequence that follows a response. In successful avoidance the consequence is that nothing happens. A rat that presses every few seconds never gets shocked, and an event that never occurs cannot follow a response. Worse, the better the animal performs, the less contact it has with the contingency: a perfect avoider could not tell whether the shock generator was still connected. Yet avoidance is among the most persistent behavior known. The theories below are attempts to say what the real consequence is. ## Two-factor theory O. H. Mowrer's answer, first sketched in 1939 and developed into "two-factor" theory from 1947, is that two learning processes run in sequence.[7][8] 1. **Classical conditioning of fear.** The warning signal is paired with shock and comes to elicit a conditioned emotional response — fear — with the racing heart and freezing that go with it. Neal Miller showed in 1948 that this conditioned fear works as an acquired drive: rats would learn a brand-new response, turning a wheel, simply to get out of a compartment where they had once been shocked.[9] 2. **Operant reinforcement by fear reduction.** The avoidance response terminates the signal, and with it the fear. The response is therefore reinforced — not by the absence of shock, which is nothing, but by escape from an aversive internal state, which is something. On this account the animal is not "avoiding" in the sense of anticipating the shock; it is escaping the fear the signal produces. Leon Kamin confirmed in 1956 that both parts contribute: rats learned best when a response both turned off the signal and cancelled the shock, and each element alone supported weaker learning.[10] The theory also explains the persistence: because the animal responds early, it never stays with the signal long enough to discover that shock no longer follows, so the fear never extinguishes and neither does the response. ### Problems with two-factor theory - **Well-trained avoiders show little fear.** Solomon and Wynne's dogs, after the first few trials, looked calm and businesslike. Kamin, Brimer, and Black measured fear of the signal directly, by how much it suppressed food-reinforced lever pressing, and found that fear *declined* as avoidance became well learned — while the avoidance response stayed strong.[11] If fear reduction is the reinforcer, the reinforcer seems to fade while the behavior does not. - **Sidman avoidance has no signal.** Two-factor theorists replied that the passage of time since the last response, and the animal's own proprioceptive feedback, serve as internal signals — a reasonable move, but one that makes the theory hard to test. - **Extinction is far too slow.** With the shock off, the signal should lose its fear and the response should fade. In practice avoidance can outlast measurable fear by hundreds of trials.[4][12] ## One-factor theory: the missed shock is the reinforcer Richard Herrnstein and Philip Hineline argued in 1966 that avoidance needs only one process if "consequence" is understood at the level of rates rather than single events. They built a procedure in which a response did not postpone any particular shock; it merely switched the rat from a schedule delivering shocks at a high rate to one delivering them at a lower rate, and shocks still arrived, unpredictably, after responses. Rats learned to respond anyway. The reinforcer, they concluded, is a **reduction in the overall frequency of aversive stimulation** — something organisms are demonstrably sensitive to.[13][14] James Dinsmoor pushed the point further: stimuli that reliably accompany the avoidance response — the feel of the lever, the click of a relay, or an explicit "safety signal" — are paired with shock-free time and become conditioned reinforcers in their own right. Add a brief tone after each avoidance response and rats learn faster and respond more.[15] On this view the missed shock is not nothing at all; it is a period of safety, and safety has cues. ## Expectancy and modern hybrid accounts Martin Seligman and James Johnston proposed in 1973 that avoiders learn two expectancies — "if I respond, no shock; if I don't, shock" — and that the response persists as long as the first expectancy is confirmed, which in successful avoidance it always is.[16] Cognitive language, but the same structural insight as one-factor theory: a well-trained avoider never tests the alternative. Current accounts generally combine Pavlovian fear (which clearly drives acquisition), operant reinforcement by safety and by aversive-rate reduction (which maintains the behavior), and expectancies about the contingency (which explain the immunity to extinction).[17] ## Not every response can be an avoidance response Rats learn to run or jump to avoid shock in a handful of trials and learn to press a lever to avoid it slowly and unreliably, sometimes never. Robert Bolles explained the discrepancy in 1970 with **species-specific defense reactions**: every species has a small innate repertoire for danger — freezing, fleeing, fighting — and an avoidance response is learned readily only if it is one of these or compatible with them. A rat's natural response to a threatening chamber is to freeze or flee, not to manipulate an object; requiring a lever press pits the contingency against biology.[18] It is the aversive counterpart of the Brelands' "misbehavior of organisms," and a reminder that reinforcement selects from what an animal is built to do. ## Learned helplessness: when escape is never learned In 1967 Seligman, Steven Maier, and Bruce Overmier gave dogs a series of inescapable shocks in a harness, then placed them in a shuttle box where a jump would end the shock. Dogs that had first received *escapable* shock — or no shock — learned to jump quickly. Most of the dogs that had received inescapable shock did not: they whimpered, lay down, and took the shock, even after occasionally jumping and ending it.[19][20] The decisive comparison was the "triadic design": animals in the two shocked groups received exactly the same shocks, one group's controllable and the other's not. Only uncontrollability produced the deficit. Seligman called it **learned helplessness** and proposed that the animals had learned that responding and outcomes were independent. Fifty years on, Maier and Seligman reversed the interpretation on the strength of the neuroscience. Passivity in the face of prolonged aversive stimulation turned out to be the *default*, driven by serotonergic neurons in the dorsal raphe nucleus; what animals with escapable shock actually learn is that they have control, and that learning, mediated by the ventromedial prefrontal cortex, inhibits the default and allows escape and avoidance to be learned later.[21] Helplessness is not learned. Control is. ## Why avoidance matters outside the laboratory Avoidance is the operant engine inside most anxiety problems. Fear may be acquired classically — a panic attack in a supermarket, a dog bite, a humiliating presentation — but it is *maintained* operantly, because avoiding the supermarket, the dog, or the meeting is negatively reinforced by relief, and the avoidance prevents the fear from ever being tested. Paul Salkovskis showed that even subtle "safety behaviors" — gripping the trolley, sitting near the exit, rehearsing every sentence — work the same way: they feel protective, and they keep the catastrophe uncontradicted.[22] Compulsions in obsessive–compulsive disorder are avoidance responses to an internal threat, and the treatment that works, exposure with response prevention, is Solomon and Wynne's cure applied to people: block the response so that the feared outcome can fail to arrive.[23][4] | Everyday avoidance | Aversive event kept away | Why it persists | | --- | --- | --- | | Procrastinating on a hard task | The discomfort of starting | Every postponement is reinforced immediately; the task's real cost arrives much later | | Checking the stove three times | Imagined fire | The house never burns down, which "confirms" that checking works | | Never speaking in meetings | Possible embarrassment | Silence is safe every single time; the belief is never tested | | Leaving early to beat traffic | The jam | The jam is never experienced, so the response never meets extinction | | Ordering tests a patient does not need ("defensive medicine") | A malpractice suit | The suit never comes, which looks like proof the tests work; the cost lands on someone else | | Buying insurance | Financial loss | Rational avoidance: the rate of catastrophe really is reduced | | Streak-keeping in an app | Losing the count | A designed avoidance contingency; the aversive event is manufactured | Not all avoidance is pathological — sunglasses, seat belts, and vaccines are avoidance, and adaptive. The problem cases are the ones in which the feared event would not occur, or would be tolerable, and the avoidance itself costs more than the thing avoided. The diagnostic question is always the same: has the contingency been tested lately? > **Three confusions worth clearing up** > > **Avoidance is not punishment.** In avoidance a behavior *increases* because it prevents something aversive; in punishment a behavior *decreases* because it produces something aversive. **Escape is not avoidance.** Escape ends a stimulus that is present; avoidance prevents one that is not. **Negative does not mean bad.** Both are negative reinforcement because a stimulus is subtracted, and both strengthen behavior. ## Key takeaways - Escape terminates an aversive stimulus that is present; avoidance prevents or postpones one that has not yet occurred. Both are negative reinforcement, and almost every avoidance response begins as an escape response that got earlier. - In successful avoidance nothing happens, so the reinforcer is not obvious, and a perfect avoider never contacts the change when the shock is switched off. That is why avoidance can outlast measurable fear by hundreds of trials and why response prevention works when ordinary extinction does not. - Two-factor theory says the warning signal is classically conditioned to elicit fear and the response is reinforced by escape from that fear, but well-trained avoiders show little fear and Sidman avoidance has no signal. One-factor theory names the reinforcer as a reduction in the overall rate of aversive events plus the safety signals that accompany responding; modern accounts combine both with expectancies. - Avoidance is the operant engine inside most anxiety problems: fear is acquired classically but maintained by relief, and safety behaviors keep the catastrophe uncontradicted. Exposure with response prevention blocks the response so the feared outcome can fail to arrive. - Animals given inescapable shock later fail to escape when they can, and only uncontrollability produces the deficit. Maier and Seligman later reversed the interpretation: passivity is the default, and what is learned is control. ### Check yourself **Every time a warning tone sounds, a rat jumps a barrier and the shock never comes. Since shock is aversive, is the jumping being punished?** No. Punishment decreases a behavior by producing something aversive; here jumping increases because it prevents something aversive, which is negative reinforcement in the form of avoidance. Negative does not mean bad: a stimulus is kept away, and the behavior strengthens. **Solomon and Wynne switched off the shock generator, yet the dogs kept jumping for hundreds of trials. Did extinction fail because the dogs were still terrified?** Not mainly. Well-trained avoiders show little fear; the response persists because a dog that jumps early on every trial receives exactly what it always received, no shock, so nothing signals that the contingency has changed. What worked was a barrier that prevented the jump, so the dogs had to stay put and discover that no shock came. **You take a painkiller once a headache has started. Your roommate takes one before a long day at the screen. Which is escape and which is avoidance?** Taking the painkiller once the headache is present is escape, because the behavior terminates an aversive stimulus that is already there. Taking it beforehand is avoidance, because the behavior prevents an aversive stimulus that has not yet occurred. Both are negative reinforcement. **Two groups of dogs receive exactly the same shocks in a harness, but only one group can turn them off. Which group later fails to escape in the shuttle box, and what does the modern account say the other group learned?** The group that could not control the shock: only uncontrollability produces the deficit, which is what the triadic design showed. On the modern account, passivity is the default response to prolonged aversive stimulation, and the group with escapable shock learned that it had control, which inhibits that default and allows escape and avoidance to be learned later. **Explain it to a friend.** Explain why a safety ritual that "works every time" is the hardest kind of habit to drop, using one of your own habits as the example. ## Frequently asked questions **What is the difference between escape and avoidance learning?** In escape learning the aversive stimulus is already present and the behavior terminates it (turning off a shock, taking a painkiller). In avoidance learning the behavior occurs before the aversive stimulus and prevents or postpones it (jumping during the warning signal, leaving early to miss traffic). Both are negative reinforcement. **What is the avoidance paradox?** Reinforcement is supposed to be a consequence that follows a response, but in successful avoidance the "consequence" is that the aversive event does not happen. Something that never occurs cannot follow anything. Theories of avoidance are attempts to identify the actual reinforcer: escape from conditioned fear (two-factor theory), reduction in the overall rate of aversive events and the safety signals that accompany responding (one-factor theory), or confirmed expectancies. **What is two-factor theory?** Mowrer's proposal that avoidance involves two learning processes: first the warning signal is classically conditioned to elicit fear, then the avoidance response is operantly reinforced because it terminates the signal and the fear. It explains acquisition well and persistence partly, but well-trained animals show little fear, and avoidance can be learned without any signal. **What is Sidman avoidance?** Free-operant avoidance, introduced by Murray Sidman in 1953: shocks arrive on a timer unless the animal responds, and each response postpones the next shock for a set interval. There is no warning signal. Rats learn to respond at a steady rate that keeps shocks rare, which was hard for signal-based theories to explain. **Why is avoidance so hard to extinguish?** Because a successful avoider never experiences the change. If the shock generator is switched off, an animal that responds early on every trial receives exactly what it always received — no shock — so nothing signals that responding is now unnecessary. Extinction requires contact with the new contingency, which is why response prevention (blocking the response) works when simple extinction does not. **Is learned helplessness still an accepted theory?** The phenomenon is solid — animals and people exposed to uncontrollable aversive events later fail to escape controllable ones — but the explanation has been revised by its own authors. Maier and Seligman (2016) concluded that passivity is the brain's default response to prolonged aversive stimulation, and that what is learned is control, which inhibits the default. The practical lesson is unchanged: experiences of control are protective. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 3. Solomon, R. L., & Wynne, L. C. (1953). Traumatic avoidance learning: Acquisition in normal dogs. *Psychological Monographs, 67*(4), 1–19. 4. Solomon, R. L., Kamin, L. J., & Wynne, L. C. (1953). Traumatic avoidance learning: The outcomes of several extinction procedures with dogs. *Journal of Abnormal and Social Psychology, 48*(2), 291–302. 5. Sidman, M. (1953). Avoidance conditioning with brief shock and no exteroceptive warning signal. *Science, 118*(3058), 157–158. 6. Sidman, M. (1953). Two temporal parameters of the maintenance of avoidance behavior by the white rat. *Journal of Comparative and Physiological Psychology, 46*(4), 253–261. 7. Mowrer, O. H. (1947). On the dual nature of learning — a re-interpretation of "conditioning" and "problem-solving." *Harvard Educational Review, 17*, 102–148. 8. Mowrer, O. H. (1960). *Learning Theory and Behavior*. Wiley. 9. Miller, N. E. (1948). Studies of fear as an acquirable drive: I. Fear as motivation and fear-reduction as reinforcement in the learning of new responses. *Journal of Experimental Psychology, 38*(1), 89–101. 10. Kamin, L. J. (1956). The effects of termination of the CS and avoidance of the US on avoidance learning. *Journal of Comparative and Physiological Psychology, 49*(4), 420–424. 11. Kamin, L. J., Brimer, C. J., & Black, A. H. (1963). Conditioned suppression as a monitor of fear of the CS in the course of avoidance training. *Journal of Comparative and Physiological Psychology, 56*(3), 497–501. 12. Solomon, R. L., & Wynne, L. C. (1954). Traumatic avoidance learning: The principles of anxiety conservation and partial irreversibility. *Psychological Review, 61*(5), 353–385. 13. Herrnstein, R. J., & Hineline, P. N. (1966). Negative reinforcement as shock-frequency reduction. *Journal of the Experimental Analysis of Behavior, 9*(4), 421–430. 14. Herrnstein, R. J. (1969). Method and theory in the study of avoidance. *Psychological Review, 76*(1), 49–69. 15. Dinsmoor, J. A. (2001). Stimuli inevitably generated by behavior that avoids electric shock are inherently reinforcing. *Journal of the Experimental Analysis of Behavior, 75*(3), 311–333. 16. Seligman, M. E. P., & Johnston, J. C. (1973). A cognitive theory of avoidance learning. In F. J. McGuigan & D. B. Lumsden (Eds.), *Contemporary Approaches to Conditioning and Learning* (pp. 69–110). Winston-Wiley. 17. Krypotos, A.-M., Effting, M., Kindt, M., & Beckers, T. (2015). Avoidance learning: A review on theoretical and experimental approaches. *Frontiers in Behavioral Neuroscience, 9*, 189. 18. Bolles, R. C. (1970). Species-specific defense reactions and avoidance learning. *Psychological Review, 77*(1), 32–48. 19. Seligman, M. E. P., & Maier, S. F. (1967). Failure to escape traumatic shock. *Journal of Experimental Psychology, 74*(1), 1–9. 20. Overmier, J. B., & Seligman, M. E. P. (1967). Effects of inescapable shock upon subsequent escape and avoidance responding. *Journal of Comparative and Physiological Psychology, 63*(1), 28–33. 21. Maier, S. F., & Seligman, M. E. P. (2016). Learned helplessness at fifty: Insights from neuroscience. *Psychological Review, 123*(4), 349–367. 22. Salkovskis, P. M. (1991). The importance of behaviour in the maintenance of anxiety and panic: A cognitive account. *Behavioural Psychotherapy, 19*(1), 6–19. 23. Meyer, V. (1966). Modification of expectations in cases with obsessional rituals. *Behaviour Research and Therapy, 4*(4), 273–280. ## Related - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): The quadrant avoidance belongs to, and why it is not punishment. - [Operant vs. classical](https://operantconditioning.com/operant-vs-classical-conditioning/): Two-factor theory is where the two kinds of learning meet. - [Extinction](https://operantconditioning.com/extinction/): Why avoidance resists it, and what response prevention does. --- # The Matching Law: Herrnstein's Equation for Choice, Explained > The matching law: behavior is allocated in proportion to reinforcement. Herrnstein's experiment, the equations, sports and classroom evidence, self-control. - Source: https://operantconditioning.com/matching-law/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Choice · How time gets divided* Every behavior is a choice among alternatives, and organisms divide their behavior among alternatives in proportion to what each one pays. That single regularity, found in pigeons in 1961, now predicts basketball shot selection, classroom disruption, and why we buy things we do not need. > **Definition** > > The **matching law** states that when two or more responses are available, the *relative rate* of each response equals — matches — the *relative rate of reinforcement* it produces. For two alternatives: > > (B1)/(B1 + B2) = (R1)/(R1 + R2) > > where B is the rate of each behavior and R the rate of reinforcement it earns. An option that delivers 70% of the available reinforcement gets about 70% of the behavior.[1] > > **A worked example.** A pigeon can peck either of two keys. Key A pays off about 30 times an hour, key B about 10 times an hour, so A delivers 30 ÷ (30 + 10) = 75% of the reinforcement. The matching law predicts that the pigeon will make about 75% of its pecks on A and 25% on B — not all of them on A, even though A is plainly better. That the pigeon does not simply pick the better key every time, and that the poorer option keeps a share proportional to what it pays, is the surprise of the law: on schedules like these, behavior is spread in proportion to payoff rather than piled onto the best option. **In brief** - Organisms allocate behavior in proportion to reinforcement: an option that delivers 70% of the available reinforcement gets about 70% of the behavior. - A behavior's strength depends not only on what it earns but on what everything else earns, so enriching the alternatives reduces it without punishment. - Because a reinforcer's value falls hyperbolically with delay, preference flips toward a small, soon reward as it approaches; commitment means choosing before the flip. ## Herrnstein's experiment Richard Herrnstein put pigeons in a chamber with two response keys, each paying grain on its own [variable-interval schedule](https://operantconditioning.com/schedules-of-reinforcement/) — a **concurrent VI VI** schedule. A bird could switch between keys whenever it liked. He varied how the reinforcement was divided between the keys while holding the total constant (VI 3-minute against VI 3-minute, VI 2.25 against VI 4.5, and so on), and to stop the birds simply alternating, he added a **changeover delay**: a brief period after each switch during which no reinforcement could be collected.[1] The result was startlingly orderly. Plot the proportion of pecks on the left key against the proportion of reinforcers earned on the left key and the points fall on the diagonal. Birds did not go exclusively to the richer key, and they did not split their pecks evenly; they matched.  *Schematic of the matching relation. Real data cluster around the diagonal; systematic departures from it are captured by the generalized matching law below.* ## From matching to a new law of effect In 1970 Herrnstein drew the larger conclusion. If behavior is allocated in proportion to reinforcement, then even a single response in a Skinner box is a choice — between pressing the lever and everything else the rat could do (grooming, exploring, resting), all of which produce some reinforcement of their own. Writing the "everything else" reinforcement as Re, the matching law for one measured response becomes a hyperbola: B = (kR)/(R + Re) Response rate rises with reinforcement rate but with diminishing returns, leveling off at a maximum k. This fit decades of single-schedule data that the original law of effect had described only qualitatively, and Herrnstein proposed it as the quantitative law of effect.[2] It also carried a practical implication: a behavior's strength depends not just on what it earns but on what everything else earns. Enrich the alternatives and the behavior declines without any punishment at all. ## The generalized matching law Real organisms do not match perfectly. William Baum showed in 1974 that the departures are systematic and can be captured by two parameters. Taking logarithms of the ratio form: log ( (B1)/(B2) ) = a log ( (R1)/(R2) ) + log b The slope **a** is sensitivity. When a = 1, matching is perfect; the usual finding is **undermatching**, a slope around 0.8, meaning organisms are somewhat less extreme in their preference than the reinforcement ratio warrants. The intercept **b** is bias: a constant preference for one alternative unrelated to reinforcement — a key that is easier to reach, a side the animal favors.[3][4] The generalized form has been fitted to hundreds of data sets across species, and its parameters turn out to be sensitive to procedural details in interpretable ways: shorter changeover delays produce more undermatching, for instance, because switching itself is reinforced. Reinforcement rate is not the only thing organisms match to. Reinforcer magnitude, delay, and quality all enter the equation, which is what makes the framework a general account of choice rather than a fact about grain.[4] ## Matching beyond the pigeon ### Sports A basketball player choosing between a two-point and a three-point attempt is on a concurrent schedule. Vollmer and Bourret analyzed a season of college basketball and found that the proportion of three-point shots taken by teams and by individual players matched the proportion of points those shots produced.[5] Reed, Critchfield, and Martens applied the generalized matching law to NFL play-calling and found that the ratio of passing to rushing plays tracked the ratio of yards each type gained, with the undermatching and bias the generalized law allows for.[6] No coach was computing logarithms; the law describes what allocation looks like when consequences are doing the selecting. ### Classrooms and problem behavior Martens and Houk observed a student whose disruptive and on-task behavior each drew teacher attention at different rates, and found the two behaviors allocated in proportion to the attention each earned.[7] This is the theoretical spine of [differential reinforcement of alternative behavior](https://operantconditioning.com/glossary/#differential-reinforcement-of-alternative-behavior): to reduce a problem behavior maintained by attention, you do not need to punish it. You need the alternative to pay better — more attention, more reliably, sooner. A large applied literature has since evaluated problem behavior as choice, with concurrent-schedule arrangements that make the appropriate response the richer option.[8] ### Everyday human behavior McDowell argued in 1988 that Herrnstein's hyperbola predicts a common frustration: adding a little reinforcement for a desired behavior has a large effect in an environment that is otherwise barren and almost none in an environment that is already rich. The same praise that transforms a child's behavior in a bleak classroom does nothing in one full of competing reinforcers.[9] Humans, it must be said, match less cleanly than pigeons; people given instructions or forming their own rules about a schedule often follow the rule rather than the contingency, a theme that runs through all human operant research. ## Melioration and the mechanism of matching The matching law is a description, not a mechanism. Two candidate mechanisms competed. **Maximizing** accounts, borrowed from economics, hold that organisms distribute behavior so as to obtain the most total reinforcement, and on concurrent VI VI schedules matching happens to be nearly optimal.[10] Herrnstein and Vaughan's **melioration** holds instead that organisms shift behavior toward whichever alternative currently has the higher *local* rate of return until the local rates are equal — a myopic rule that produces matching without any computation of totals, and that predicts the systematically suboptimal choices people make when a locally better option worsens the long-run outcome.[11] Melioration is one reason the matching law connects so naturally to impulsiveness. ## Self-control: when the alternatives differ in time The most consequential extension of matching is to choices between a smaller reinforcer available sooner and a larger one available later. Rachlin and Green showed in 1972 that pigeons facing that choice directly took the small immediate grain, but that if the choice was made well in advance, the same pigeons committed themselves to the larger, later reward — the first laboratory demonstration of a commitment device.[12] George Ainslie explained why in 1975: if the value of a reinforcer falls with delay along a *hyperbola* rather than an exponential curve, the curves for a small-soon and a large-late reward cross, so preference reverses as the small reward approaches.[13] James Mazur's adjusting-delay procedure confirmed the hyperbolic shape precisely.[14] Steep [delay discounting](https://operantconditioning.com/glossary/#delay-discounting) — a fast drop in value with delay — has since been documented in people with substance-use disorders, in problem gamblers, and in smokers, and is studied as a process that cuts across many conditions.[15] The everyday translation: the environment that makes you impulsive is one in which the small reward is near and the large reward is far, and the fix is to move the choice point earlier, when the curves have not yet crossed. Organisms are not uniformly impulsive, though. Cole found that rats on a schedule in which retrieving food pellets from the tray started a one-minute period without further pellets learned to let pellets accumulate and collect them in batches — **operant hoarding**, a form of self-control the impulsivity findings would not have predicted, and a reminder that the details of the contingency matter.[16] ## Behavioral economics: demand, price, and elasticity Once behavior is allocation, the tools of economics apply. Steven Hursh proposed in 1980 that a schedule requirement is a **price** (responses per reinforcer), that consumption plotted against price gives a **demand curve**, and that the slope of that curve — **elasticity** — measures how essential a reinforcer is.[17] Food in a closed economy, where the animal earns all of its food in the chamber, is inelastic: raise the price and the animal works harder to keep consumption up. Sweetened water in an open economy is elastic: raise the price and consumption collapses. Whether reinforcers are **substitutes** (one replaces another) or **complements** (consumed together) can be measured the same way.[18] Kagel, Battalio, and Green showed, in a research program running from the 1970s onward, that rats and pigeons obey demand theory in detail, including some of its odder predictions, such as Giffen goods.[19] The approach has direct policy uses. Demand curves for cigarettes, alcohol, and drugs measured in the laboratory predict how consumption responds to taxation, and "essential value" derived from demand analysis compares the reinforcing efficacy of drugs on a common scale.[20] It also explains a stubborn feature of behavior change: a reinforcer you offer competes in a market, and if the problem behavior is a cheap, inelastic, non-substitutable good, small incentives for the alternative will not move it. ## Limits and criticisms - **Matching is descriptive.** It tells you the outcome of allocation, not how the organism gets there; melioration, maximizing, and momentary-maximizing accounts all reproduce it under various conditions and are hard to separate. - **Humans often follow rules instead.** Verbal instructions and self-generated rules can override contingencies, so human matching is weaker and more variable than animal matching unless the schedule is hard to describe.[21] - **Ratio schedules break the pattern.** On concurrent ratio schedules, exclusive preference for the better option is the optimal strategy and is what animals do, so matching in its simple form applies mainly to interval schedules, where spreading behavior across options pays. - **Parameters need estimating.** The generalized law fits almost anything with a free slope and intercept; its value lies in the parameters being stable and interpretable, which they generally are, not in the fit alone. Within those limits, the matching law is the closest thing behavior analysis has to a physical law. It made choice measurable, connected the laboratory to economics, and gave clinicians a simple instruction that holds up: to change what someone does, change what the alternatives pay. ## Key takeaways - On concurrent variable-interval schedules, organisms do not pile behavior onto the best option; they spread it in proportion to payoff. Herrnstein's pigeons matched the proportion of pecks on each key to the proportion of reinforcers it delivered. - Even a single response is a choice against everything else the organism could do. Herrnstein's hyperbola makes response rate rise with reinforcement at diminishing returns, which is why the same praise transforms behavior in a barren environment and does almost nothing in a rich one. - Real organisms deviate systematically. The generalized matching law adds sensitivity (usually undermatching, a slope around 0.8) and bias (a constant preference unrelated to reinforcement), and reinforcer magnitude, delay, and quality enter the equation alongside rate. - Matching is a description, not a mechanism; melioration, which shifts behavior toward the locally richer option, is one candidate. Humans match less cleanly because rules can override contingencies, and on concurrent ratio schedules exclusive preference for the better option is optimal and is what animals do. - The practical instruction: to change what someone does, change what the alternatives pay. Problem behavior maintained by attention is reduced by making the appropriate alternative pay better, and impulsive choices are avoided by moving the choice point earlier, before the value curves cross. ### Check yourself **A pigeon can peck key A, which pays about 30 times an hour, or key B, which pays about 10. A classmate predicts the bird will peck A almost exclusively, since A is plainly better. What does the matching law predict?** About 75% of pecks on A and 25% on B, because A delivers 30 out of every 40 reinforcers. On concurrent variable-interval schedules behavior is spread in proportion to payoff rather than piled onto the best option; the poorer key keeps a share proportional to what it pays. **A praise program that transformed a student's behavior in one classroom does nothing in another. Was the praise too weak?** Not necessarily. In Herrnstein's hyperbola a behavior's strength depends on its reinforcement relative to the reinforcement for everything else, so adding a little reinforcement has a large effect in a barren environment and almost none in one already full of competing reinforcers. The praise is the same; the alternatives are not. **A student's disruptive behavior earns teacher attention more reliably than on-task behavior does. Without punishing anything, how does the matching law say to reduce the disruption?** Make the alternative pay better: attend to on-task behavior more, more reliably, and sooner, so that it earns the larger share of attention and therefore draws the larger share of behavior. This is differential reinforcement of alternative behavior, and the matching law is its theoretical spine. **Pigeons choosing between a small immediate grain and a larger delayed one take the small one, yet when the same choice is offered well in advance they commit to the larger one. Why does preference reverse?** Because value falls with delay along a hyperbola rather than an exponential curve, the value curves for the small-soon and large-late rewards cross. Far from both rewards the larger one is worth more; as the small one becomes imminent it overtakes, so choosing early, before the curves cross, is a commitment device. **Explain it to a friend.** Explain the matching law to someone who dislikes math, using either the basketball or the classroom example and no numbers at all. ## Frequently asked questions **What is the matching law in simple terms?** Organisms spread their behavior across options in proportion to how much reinforcement each option provides. If one option delivers twice as much reinforcement as another, it gets about twice as much behavior. Richard Herrnstein discovered it in pigeons in 1961, and it holds, with some systematic deviations, across species and settings. **What is the generalized matching law?** Baum's 1974 extension, which adds two parameters: sensitivity (how strongly behavior tracks reinforcement; usually a little less than 1, called undermatching) and bias (a constant preference for one option unrelated to reinforcement). It is written as a straight line in logarithmic ratios and fits most choice data. **How does the matching law explain problem behavior?** Problem behavior and appropriate behavior are alternatives on a concurrent schedule. If misbehavior earns attention more reliably than good behavior, the matching law predicts a lot of misbehavior. The treatment is to make the appropriate alternative pay more — differential reinforcement of alternative behavior — rather than to punish the problem. **What is the difference between the matching law and the law of effect?** Thorndike's law of effect says responses followed by satisfying consequences are strengthened. Herrnstein's matching law quantifies it: response strength is relative, depending on the reinforcement for a behavior compared with the reinforcement for everything else. Herrnstein proposed the hyperbolic form of matching as the quantitative law of effect. **Does the matching law apply to humans?** Yes, though less cleanly. Sports play-calling, classroom behavior, and conversation have been shown to match. Human deviations mostly come from rules and instructions: people who can describe a schedule often follow their description rather than the contingency. **What does the matching law have to do with self-control?** Choices between a smaller, sooner reward and a larger, later one are matching choices in which delay reduces value. Because value falls hyperbolically with delay, preference flips toward the small reward as it becomes imminent. Making the choice early, before the flip, is what commitment devices do. ## References 1. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 2. Herrnstein, R. J. (1970). On the law of effect. *Journal of the Experimental Analysis of Behavior, 13*(2), 243–266. 3. Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. *Journal of the Experimental Analysis of Behavior, 22*(1), 231–242. 4. Baum, W. M. (1979). Matching, undermatching, and overmatching in studies of choice. *Journal of the Experimental Analysis of Behavior, 32*(2), 269–281. 5. Vollmer, T. R., & Bourret, J. (2000). An application of the matching law to evaluate the allocation of two- and three-point shots by college basketball players. *Journal of Applied Behavior Analysis, 33*(2), 137–150. 6. Reed, D. D., Critchfield, T. S., & Martens, B. K. (2006). The generalized matching law in elite sport competition: Football play calling as operant choice. *Journal of Applied Behavior Analysis, 39*(3), 281–297. 7. Martens, B. K., & Houk, J. L. (1989). The application of Herrnstein's law of effect to disruptive and on-task behavior of a retarded adolescent girl. *Journal of the Experimental Analysis of Behavior, 51*(1), 17–27. 8. Fisher, W. W., & Mazur, J. E. (1997). Basic and applied research on choice responding. *Journal of Applied Behavior Analysis, 30*(3), 387–410. 9. McDowell, J. J. (1988). Matching theory in natural human environments. *The Behavior Analyst, 11*(2), 95–109. 10. Rachlin, H., Green, L., Kagel, J. H., & Battalio, R. C. (1976). Economic demand theory and psychological studies of choice. In G. H. Bower (Ed.), *The Psychology of Learning and Motivation* (Vol. 10). Academic Press. 11. Herrnstein, R. J., & Vaughan, W. (1980). Melioration and behavioral allocation. In J. E. R. Staddon (Ed.), *Limits to Action: The Allocation of Individual Behavior*. Academic Press. 12. Rachlin, H., & Green, L. (1972). Commitment, choice and self-control. *Journal of the Experimental Analysis of Behavior, 17*(1), 15–22. 13. Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. *Psychological Bulletin, 82*(4), 463–496. 14. Mazur, J. E. (1987). An adjusting procedure for studying delayed reinforcement. In M. L. Commons, J. E. Mazur, J. A. Nevin, & H. Rachlin (Eds.), *Quantitative Analyses of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value* (pp. 55–73). Erlbaum. 15. Bickel, W. K., & Marsch, L. A. (2001). Toward a behavioral economic understanding of drug dependence: Delay discounting processes. *Addiction, 96*(1), 73–86. 16. Cole, M. R. (1990). Operant hoarding: A new paradigm for the study of self-control. *Journal of the Experimental Analysis of Behavior, 53*(2), 247–262. 17. Hursh, S. R. (1980). Economic concepts for the analysis of behavior. *Journal of the Experimental Analysis of Behavior, 34*(2), 219–238. 18. Hursh, S. R. (1984). Behavioral economics. *Journal of the Experimental Analysis of Behavior, 42*(3), 435–452. 19. Kagel, J. H., Battalio, R. C., & Green, L. (1995). *Economic Choice Theory: An Experimental Analysis of Animal Behavior*. Cambridge University Press. 20. Hursh, S. R., & Silberberg, A. (2008). Economic demand and essential value. *Psychological Review, 115*(1), 186–198. 21. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley. ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): The concurrent schedules matching was discovered on. - [Premack principle](https://operantconditioning.com/premack-principle/): The other great relativity of reinforcement. - [Building habits](https://operantconditioning.com/habits/): Move the choice point before the curves cross. --- # The Premack Principle: Definition, Examples, and the Response Deprivation Hypothesis > The Premack principle: a more probable behavior can reinforce a less probable one. Premack's experiments, Grandma's rule, response deprivation, and mistakes. - Source: https://operantconditioning.com/premack-principle/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Grandma’s rule* "Eat your vegetables, then you can have dessert." Every grandmother knows the rule; David Premack showed why it works, discovered that it runs backward under the right conditions, and changed what a reinforcer is. > **Definition** > > The **Premack principle** states that the opportunity to engage in a *more probable* behavior will reinforce a *less probable* behavior when access to the first is made contingent on performing the second. Reinforcers, on this view, are not stimuli but **behaviors**: it is not the dessert that reinforces vegetable-eating but the eating of the dessert.[1] > > Informally it is called **Grandma's rule**, or in classrooms the **first–then** rule: first the work, then the play. **In brief** - The opportunity to perform a more probable behavior reinforces a less probable one when access to the first is made contingent on the second. - Reinforcers are behaviors, not objects, and the relation reverses with circumstances: water-deprived rats ran to drink, running-deprived rats drank to run. - The response deprivation hypothesis is the better-supported statement: a contingency reinforces whenever it restricts the contingent behavior below its free baseline. ## Premack's experiments David Premack's 1959 study gave first-grade children free access to a candy dispenser and a pinball machine and measured which each child used more. He then made the two contingent on each other. For children who preferred pinball, making the machine available only after they ate a candy increased candy eating; making candy available only after a game did nothing much. For children who preferred candy, the result reversed: candy eating reinforced pinball, and pinball did not reinforce candy eating.[1] Whether an activity worked as a reinforcer depended entirely on whether it was the more probable of the pair for that child. The decisive test came in 1962 with rats, running wheels, and water. Rats deprived of water would run in a wheel to earn access to drinking — the standard result, drinking reinforcing running. But rats given free water and *deprived of running* learned to drink in order to unlock the wheel: running reinforced drinking.[2] The same two behaviors reversed roles with the animal's circumstances. Drinking was not a reinforcer in itself; it was a reinforcer only when drinking was the more probable behavior.  *Premack's reversal. Whichever behavior is more probable at baseline can reinforce the other; deprivation changes which one that is.* > **Why this was radical** > > Thorndike and Skinner had treated reinforcers as things — food, water, a pellet — that were reinforcing more or less everywhere. Premack argued that reinforcement is a *relation between behaviors*: any behavior can reinforce any less probable one and be reinforced by any more probable one. The list of reinforcers is not a list of objects; it is a ranking of what the organism would do right now if it could.[3] ## The punishment side Premack later extended the principle to punishment. If being made to perform a *less* probable behavior is made contingent on a more probable one, the more probable one declines. Forcing a rat that would rather drink to run a wheel after each drink suppresses drinking.[4] Reinforcement and punishment become two sides of a single relation: the direction of the probability difference determines which you get. ## The response deprivation hypothesis The principle has a known failure. Eisenberger, Karpman, and Trattner found in 1967 that a *less* probable behavior can reinforce a more probable one, provided the schedule restricts the less probable behavior below the level the organism would normally choose.[5] William Timberlake and James Allison built that into a more general rule in 1974, the **response deprivation hypothesis**. Measure how much of each behavior the organism performs when both are freely available — the baseline. A contingency will make one behavior reinforce another if, under the contingency, performing the instrumental behavior at its baseline level would leave the organism with *less* of the contingent behavior than its baseline. Formally, with I the required instrumental responses, C the contingent responses earned, and O the baseline levels: (I)/(C) > (OI)/(OC) When the inequality holds, the organism is "response deprived," and it will increase the instrumental behavior to recover access to the contingent one — regardless of which was more probable to begin with.[6] Premack's principle turns out to be the special case in which the contingent behavior is the more probable one, which almost guarantees deprivation. Timberlake and Allison's version explains the reversals, and it explains why a small requirement often fails: "read one page, then play for an hour" deprives the child of nothing. Applied tests followed. Konarski and colleagues showed in classrooms that schedules meeting the response-deprivation condition increased children's academic work even when the contingent activity was the *less* preferred one, exactly as the hypothesis and not the original principle predicted.[7] A later review concluded that the response deprivation hypothesis is the better-supported statement, and that both are best understood through the lens of [motivating operations](https://operantconditioning.com/glossary/#motivating-operation): restricting a behavior below baseline is an establishing operation for it.[8] ## How to use the Premack principle 1. **Observe before you arrange.** Watch what the person (or animal, or you) actually does when free to choose. The most probable behaviors are the reinforcers available to you, whatever anyone says they enjoy. 2. **Make the probable behavior contingent on the improbable one.** First homework, then screen; first the walk, then the coffee; first three sits, then the game of tug. Access to the reinforcing activity is granted *only* after the target behavior. 3. **Keep the ratio deprivational.** The requirement has to leave the person with less of the preferred activity than they would otherwise have taken, or there is no reason to work for it. Short requirements, delivered often, are usually better than a huge requirement for a huge reward. 4. **Deliver promptly.** The preferred activity should begin within seconds of the target behavior ending, or be bridged with a conditioned reinforcer such as a token or a check-off. 5. **Watch for satiation.** Probabilities change. After an hour of screen time, screen time is no longer the most probable behavior, and the contingency stops working until it is restored. ## Examples | Setting | Less probable behavior (first) | More probable behavior (then) | | --- | --- | --- | | Parenting | Clearing the table | Playing outside | | Classroom | Ten minutes of math problems | Five minutes of free choice | | Special education | Completing a task strip | Time with a preferred toy, shown on a "first–then" board | | Dog training | Sitting at the door | Going through it for the walk | | Self-management | Writing 200 words | Checking messages | | Exercise | The workout | The podcast you only listen to at the gym | | Workplace | Filing the expense report | Starting the interesting design task | The dog example is the one most people already use without a name. A dog that wants to go outside will sit, wait, and make eye contact for the privilege, and the door opening is a more reliable reinforcer than any treat because, at that moment, going through the door is the dog's most probable behavior.[9] The classroom "first–then" board, standard in [applied behavior analysis](https://operantconditioning.com/applications/#aba), is Premack made visible. ## Common mistakes - **Reversing the order.** "You can play now if you promise to do homework after" delivers the reinforcer before the behavior. It reinforces promising. - **Guessing the probabilities.** Parents and managers routinely pick a "reward" nobody would choose. The only test is observation: what does the person do when free? - **Requirements that deprive nobody.** If the person could get as much of the preferred activity as they want anyway, the contingency has no bite. Access must actually be restricted. - **Making the target behavior aversive.** A punishing requirement — an hour of tedium for five minutes of play — teaches avoidance of the whole arrangement. Small, frequent contingencies work better. - **Forgetting that the reinforcer is an activity.** Ending the preferred activity abruptly to start the next requirement can function as negative punishment and provoke resistance. Signal transitions in advance. ## Where it fits in the theory The Premack principle and its successor sit alongside the [matching law](https://operantconditioning.com/matching-law/) as the two great "relativity" results of operant research. Matching says a behavior's strength depends on what the alternatives pay; Premack says whether something reinforces at all depends on what the organism would otherwise be doing. Both replaced the picture of reinforcers as fixed objects with a picture of organisms distributing their time among activities, and both feed directly into [motivating operations](https://operantconditioning.com/glossary/#motivating-operation), the modern term for the deprivation and satiation that make an activity more or less probable.[8] ## Key takeaways - A more probable behavior reinforces a less probable one when access to it is granted only after the less probable one. Grandma's rule and the classroom first–then board are the same idea. - Reinforcement is a relation between behaviors, not a property of objects. Which behavior reinforces which depends on what the organism would otherwise be doing, deprivation can reverse the roles, and running the relation the other way produces punishment. - The response deprivation hypothesis is the more accurate statement: a contingency reinforces the instrumental behavior whenever it leaves the organism with less of the contingent behavior than its free baseline, regardless of which behavior was more probable to begin with. - To use it, observe what the person actually does when free, make the probable behavior contingent on the improbable one, keep the requirement deprivational but small and frequent, deliver promptly, and watch for satiation. - The common mistakes are delivering the preferred activity first, which reinforces promising; guessing the probabilities instead of observing them; requirements that deprive nobody; and requirements so large that the whole arrangement becomes aversive. ### Check yourself **A manager announces a team lunch as a reward for finishing reports on time. Reports get no faster. Why might the lunch have failed?** A reinforcer is not whatever the manager thinks people enjoy; it is what people would actually be doing if free to choose, and the only test is observation. If the lunch is not a more probable activity than what it displaces, or if people would get it anyway, the contingency has no bite. **A child spends more free time on math worksheets than on coloring. The teacher makes coloring available only after math and restricts it below the amount the child would normally do. Can coloring reinforce math, the more probable behavior?** Yes. The original Premack principle says only a more probable behavior can reinforce, but the response deprivation hypothesis says a contingency reinforces whenever it restricts the contingent behavior below its baseline, whichever behavior was more probable. Konarski's classroom studies found exactly this, which is why the response deprivation hypothesis is the better-supported statement. **A parent says, "You can play video games now if you promise to do your homework afterward." Is this the Premack principle?** No. The preferred activity is delivered before the target behavior, so the arrangement reinforces promising, not homework. Access to the probable behavior has to come only after the less probable one. **"Read one page, then you can play for an hour" produces almost no reading. What went wrong?** The requirement deprives the child of nothing: one page for an hour of play leaves the child with as much play as ever, so there is no reason to work for it. The ratio has to be deprivational, and short requirements delivered often work better than a huge requirement for a huge reward. **Explain it to a friend.** Explain why "first the work, then the play" works and when it stops working, in three sentences that use two activities from your own routine. ## Frequently asked questions **What is the Premack principle in simple terms?** A behavior you are likely to do can be used to reinforce a behavior you are unlikely to do, if the likely one is only allowed after the unlikely one. Vegetables before dessert; homework before video games; the walk before the coffee. It is often called Grandma's rule. **What is an example of the Premack principle?** A child who would rather play than tidy up is told that play begins once the toys are put away. Tidying increases. In the laboratory, water-deprived rats ran in a wheel to earn drinking, while rats deprived of running drank to earn access to the wheel — the same two behaviors, reversed by circumstances. **What is the response deprivation hypothesis?** Timberlake and Allison's 1974 refinement: a contingency reinforces the instrumental behavior whenever it restricts the contingent behavior below the amount the organism would freely perform. It explains cases where a less-preferred activity reinforces a more-preferred one, which the original Premack principle cannot, and it is generally considered the more accurate statement. **Is the Premack principle positive reinforcement?** Yes. Access to the preferred activity is added after the target behavior, and the target behavior increases. The novelty is in what counts as the reinforcer — an opportunity to behave, rather than a stimulus — not in the quadrant. **Does the Premack principle work on yourself?** Yes, with the caveat that you are both the person setting the rule and the person tempted to break it. It works best when the preferred activity is physically gated — the podcast that only exists at the gym, the café you only visit after the writing — so that the contingency does not depend on willpower. **Who was David Premack?** An American psychologist (1925–2015) who, besides the principle that bears his name, was a pioneer of primate cognition research: he taught the chimpanzee Sarah to communicate with plastic symbols and, with Guy Woodruff, coined the term "theory of mind" in 1978. ## References 1. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 2. Premack, D. (1962). Reversibility of the reinforcement relation. *Science, 136*(3512), 255–257. 3. Premack, D. (1965). Reinforcement theory. In D. Levine (Ed.), *Nebraska Symposium on Motivation* (Vol. 13, pp. 123–180). University of Nebraska Press. 4. Premack, D. (1971). Catching up with common sense or two sides of a generalization: Reinforcement and punishment. In R. Glaser (Ed.), *The Nature of Reinforcement* (pp. 121–150). Academic Press. 5. Eisenberger, R., Karpman, M., & Trattner, J. (1967). What is the necessary and sufficient condition for reinforcement in the contingency situation? *Journal of Experimental Psychology, 74*(3), 342–350. 6. Timberlake, W., & Allison, J. (1974). Response deprivation: An empirical approach to instrumental performance. *Psychological Review, 81*(2), 146–164. 7. Konarski, E. A., Johnson, M. R., Crowell, C. R., & Whitman, T. L. (1980). Response deprivation and reinforcement in applied settings: A preliminary analysis. *Journal of Applied Behavior Analysis, 13*(4), 595–609. 8. Klatt, K. P., & Morris, E. K. (2001). The Premack principle, response deprivation, and establishing operations. *The Behavior Analyst, 24*(2), 173–180. 9. Pryor, K. (1984). *Don't Shoot the Dog! The New Art of Teaching and Training*. Simon & Schuster. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): Types of reinforcers, including activity reinforcers. - [The matching law](https://operantconditioning.com/matching-law/): The other relativity: strength depends on the alternatives. - [Building habits](https://operantconditioning.com/habits/): Put the principle to work on yourself. --- # The Neuroscience of Operant Conditioning: Dopamine, Reward Prediction Error, and Habit > What happens in the brain during operant conditioning: Olds and Milner, dopamine as a prediction-error signal, wanting vs. liking, how actions become habits. - Source: https://operantconditioning.com/neuroscience/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *The brain · Dopamine and prediction error* Skinner deliberately treated the organism as a black box. Seventy years of neuroscience have opened it, and what is inside looks remarkably like the law of effect: a broadcast teaching signal, a window of a few seconds, and two separate systems for wanting and liking. > **In one paragraph** > > Reinforcement has a physical address. When a consequence is better than the brain predicted, midbrain **dopamine** neurons fire a brief burst that is broadcast to the striatum and frontal cortex, strengthening whichever synapses were active in the preceding seconds. When the consequence is exactly as predicted, they stay quiet; when an expected reinforcer fails to arrive, they dip below baseline. This **reward prediction error** is the teaching signal of operant conditioning, and it is why immediacy, contingency, and unpredictability matter so much.[4][5] **In brief** - Midbrain dopamine neurons signal reward prediction error: a burst when a consequence is better than predicted, silence when expected, a dip when worse. - Dopamine drives wanting, not liking: rats without dopamine still enjoy sugar but stop seeking it, and addiction is sensitized wanting. - Dopamine strengthens only synapses active in the preceding seconds, which is why immediacy matters, and extended training turns goal-directed actions into [habits](https://operantconditioning.com/habits/). ## Reinforcement has an address: Olds and Milner In 1953 James Olds and Peter Milner, working at McGill, implanted an electrode into the septal area of a rat's brain and arranged for a lever press to deliver a brief electrical pulse. The rat pressed, and kept pressing; they published the finding the following year. Olds went on to map the sites that supported **intracranial self-stimulation**: rats would respond thousands of times an hour and cross electrified grids to reach the lever. Routtenberg and Lindy later gave hungry rats a daily hour with both a food lever and a stimulation lever; some spent the hour on stimulation and lost weight.[1][2][3] For the first time, a reinforcer had been produced by acting directly on the nervous system, and the anatomy of the effective sites — the medial forebrain bundle and the pathways it carries — pointed toward a particular chemical system. ## Dopamine is a prediction-error signal, not a pleasure signal The pathways Olds had stimulated carry axons of dopamine neurons from the midbrain (the ventral tegmental area and substantia nigra) to the striatum and prefrontal cortex. Through the 1980s the working assumption was that dopamine *was* pleasure. Wolfram Schultz's recordings from monkeys overturned that. Schultz trained monkeys on a simple task in which a light or sound was followed, a second or two later, by a squirt of juice, while recording individual dopamine neurons. Early in training the neurons fired when the juice arrived. Once the cue reliably predicted juice, the burst moved to the *cue*, and the fully predicted juice produced no response at all. And when the cue was followed by no juice, the neurons paused — activity dropped below baseline at the moment the juice should have come.[4][5] That is exactly the profile of a **prediction error**: positive for better-than-expected, zero for as-expected, negative for worse-than-expected. In 1996 Montague, Dayan, and Sejnowski had shown that this is the quantity a temporal-difference learning algorithm needs, and the 1997 paper by Schultz, Dayan, and Montague joined the two: the brain appeared to be running the same computation that computer scientists had derived from the law of effect.[6][4][7] Correlation became causation with optogenetics. Activating dopamine neurons with light at the moment a reward is delivered is enough to make rats learn about a cue that would otherwise be blocked, and phasic stimulation of these neurons is sufficient to produce a conditioned preference for a place.[8][9] The dopamine burst does not merely accompany learning; it drives it. > **Why this explains schedules of reinforcement** > > A reinforcer that is fully predictable generates no prediction error and, eventually, no dopamine response: a continuous schedule teaches fast, then goes flat. A reinforcer that arrives after an unpredictable number of responses cannot be fully predicted, so every one produces a burst. Schultz's group later found that dopamine neurons also show a slow ramp of activity that is largest when reward probability is 50% — maximal uncertainty.[10] Whether that ramp is what makes gambling compelling is still argued, but the [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/) has, at minimum, a plausible neural signature. ## Wanting is not liking If dopamine is not pleasure, what is it for? Kent Berridge and Terry Robinson answered with a dissociation that has held up for three decades. Rats whose dopamine neurons were destroyed with the toxin 6-hydroxydopamine stopped eating — they would starve unless tube-fed — yet when sugar was placed on their tongues they showed exactly the lip-licking, paw-licking facial reactions that intact rats show. They still *liked* sugar. What they had lost was *wanting*: the motivation to work for it, approach it, and treat cues for it as attractive.[11][12] Berridge and Robinson called this attribution of attractiveness to reinforcers and their cues **incentive salience**. Liking, meanwhile, turned out to depend on tiny opioid "hedonic hotspots" in the nucleus accumbens and elsewhere, anatomically separate from the dopamine system.[13] Reinforcement, at the level of neurons, is therefore mostly about wanting. Dopamine also sets how much effort an animal will spend: with dopamine reduced, rats still choose food, but they shift from a lever that pays well and requires many presses to freely available, less preferred chow.[14] The everyday consequence is familiar to anyone who has kept doing something they no longer enjoy. ### Addiction as sensitized wanting The dissociation explains the central puzzle of addiction: people keep wanting drugs that have long since stopped delivering much pleasure. Robinson and Berridge's incentive-sensitization theory proposes that repeated drug exposure sensitizes the dopamine system's response to the drug and to its cues, so that wanting grows even as liking shrinks. The syringe, the bar, the friend, the time of day become cues with enormous incentive salience, which is why relapse is so often triggered by the environment rather than by withdrawal.[15][16] In operant terms: drug taking is positively reinforced by the drug and negatively reinforced by relief from withdrawal, and the cues that precede it become discriminative stimuli and conditioned reinforcers with a sensitized grip. ## Learning from the stick: punishment in the brain Worse-than-expected outcomes produce a dopamine dip, and the brain appears to learn from dips through a different route than from bursts. Michael Frank and colleagues gave a probabilistic learning task to people with Parkinson's disease, a condition of dopamine loss. Off medication, patients were better at learning from negative feedback — which choices to avoid — than from positive feedback. On dopamine-replacing medication, the pattern reversed: they learned from positive feedback and became worse at learning from negative.[17] Frank's model attributes the two to separate striatal pathways, one facilitating action when dopamine is high and one suppressing it when dopamine is low. Separately, neurons in the lateral habenula fire when a reward is omitted or a punishment predicted, and inhibit dopamine neurons — a candidate source of the dip.[18][36] Reinforcement and punishment are not mirror images at the neural level any more than they are at the behavioral one. ## The window of a few seconds: why immediacy matters Every practical guide to reinforcement says the consequence must be immediate. The neural reason is now visible. A dopamine burst cannot strengthen every synapse in the striatum; it strengthens the ones that were recently active, which carry a short-lived molecular "eligibility trace." Yagishita and colleagues, using glutamate uncaging on single dendritic spines, found that dopamine enlarged a spine only if it arrived within roughly 0.3 to 2 seconds after the spine had been stimulated. Earlier or later, nothing happened.[19] That window is the cellular basis of [contiguity](https://operantconditioning.com/glossary/#contiguity): a reinforcer that comes thirty seconds after a behavior finds the trace gone and strengthens whatever happened in the last two seconds instead. Dopamine is not the only teaching signal. Neurons of the nucleus basalis release acetylcholine across the cortex and respond to reinforcers and to the stimuli that predict them; pairing a tone with electrical stimulation of the nucleus basalis is enough to expand the tone's representation in the auditory cortex, without any behavior at all.[20][21] Reinforcement, in other words, reshapes perception as well as action. ## From action to habit: two systems in the striatum Press a lever a few hundred times for food and then make the food unappealing — pair it with a mild poison, or feed the animal to satiety. A rat with moderate training stops pressing; it "knows" what the lever produces and no longer wants it. A rat with extensive training keeps pressing anyway.[22][23] Anthony Dickinson used this reinforcer-devaluation test to distinguish **goal-directed actions**, which are sensitive to the current value of their outcome, from **habits**, which are triggered by antecedents and run off regardless. The two have different homes. Lesions of the dorsomedial striatum leave animals unable to act on outcome value, while lesions of the dorsolateral striatum prevent habits from forming, so that over-trained animals stay sensitive to devaluation.[24][25][26] Training on interval schedules produces habits faster than training on ratio schedules, presumably because on an interval schedule the connection between how much you respond and how much you get is loose.[27] Human imaging finds the same division of labor: prediction errors in the ventral striatum during learning, with the dorsal striatum engaged when the learning must guide action.[28] Everitt and Robbins argued that addiction is this transition run to its end — from action to habit to compulsion, with control migrating from ventral to dorsal striatum as the behavior becomes cue-driven and insensitive to consequences.[29] This is the neuroscience behind a piece of practical advice on [the habits page](https://operantconditioning.com/habits/): a well-formed habit survives the loss of motivation, for good and ill. The cue keeps producing the behavior after the outcome has lost its appeal, which is why habits are hard to break by deciding to and easier to break by changing the antecedent. ## Operant conditioning in a single neuron The principle scales down remarkably far. In 1969 Eberhard Fetz reinforced monkeys with food pellets whenever a single recorded neuron in motor cortex fired faster; within minutes the monkeys raised that neuron's firing rate, with no instruction about what they were doing.[30] In the sea slug *Aplysia*, Brembs and colleagues reinforced a feeding movement by stimulating a dopaminergic nerve immediately after it, and then reproduced the learning in a single identified neuron in a dish: contingent dopamine applied to neuron B51 changed its excitability the way training changed the whole animal's behavior.[31] Operant conditioning is not a trick of large brains; it is a property of neurons. ## The brain as a reinforcement learner Reinforcement learning, the branch of artificial intelligence in which an agent learns from reward signals, was built on Thorndike and Skinner and on temporal-difference learning, and the dopamine findings turned it into a theory of the brain.[7][32] The current picture has two learners running in parallel: a "model-free" system that caches the value of actions from prediction errors — the habit system — and a "model-based" system that plans using a map of how actions lead to outcomes — the goal-directed system — with control shifting between them according to which is more reliable.[33] The same framework connects to classical conditioning through the Rescorla–Wagner model of 1972, which was itself a prediction-error rule.[34] ## What this means in practice - **Immediacy is not a rule of thumb; it is a molecular window.** If the real reinforcer must be delayed, deliver a conditioned reinforcer — a click, a word, a checkmark — inside the window and let it bridge the gap. - **Predictable reinforcers stop teaching.** Once a behavior is learned, thinning to an intermittent schedule keeps prediction errors, and dopamine, alive. The same mechanism is what makes slot machines and feeds hard to leave. - **Wanting and liking come apart.** A behavior can be maintained by cues long after its outcome stopped being enjoyable; treat the cues, not just the outcome. - **Habits outlive motivation.** An over-trained behavior is insensitive to devaluation. To change it, change the antecedent or make the response impossible rather than relying on wanting it less. - **Reinforcement and punishment use different circuitry.** Which one a person learns from best can depend on the state of their dopamine system — one reason blanket claims that "punishment doesn't work" or "rewards don't work" are both too simple. ## What is still unsettled The prediction-error account is the best-supported theory of dopamine, not the whole story. Dopamine also ramps up as animals approach rewards, participates in movement and vigor, and is released in patterns that a single scalar error signal does not obviously explain; some researchers argue it broadcasts several different messages on different timescales.[35][14] Most of the causal work is in rodents, and human evidence rests on imaging and on patient groups. And none of it changes the functional definitions: a reinforcer is still whatever increases the behavior it follows. The neuroscience explains why the law of effect holds; it does not replace it. ## Key takeaways - Dopamine is a prediction-error signal, not a pleasure signal: midbrain dopamine neurons burst when a consequence is better than predicted, stay quiet when it is as predicted, and dip when it is worse. Because a fully predictable reinforcer produces no error, continuous reinforcement teaches fast and then goes flat, while intermittent schedules keep prediction errors alive. - Reinforcement and punishment are not mirror images in the brain. Learning from dips runs through a different striatal pathway than learning from bursts, and which one a person learns from best can depend on the state of their dopamine system. - Wanting and liking are separate systems: dopamine drives wanting (incentive salience) and effort, while liking depends on opioid hotspots. Addiction is sensitized wanting, which is why cues trigger relapse long after the drug stopped being enjoyable. - A dopamine burst strengthens only synapses that were active within roughly 0.3 to 2 seconds before it, the cellular basis of contiguity. A delayed reinforcer strengthens whatever happened just before it arrived, so bridge any delay with a conditioned reinforcer. - With extended training, control shifts from a goal-directed system in the dorsomedial striatum, sensitive to outcome value, to a habit system in the dorsolateral striatum, triggered by cues regardless of value. Habits outlive motivation, so they are easier to break by changing the antecedent than by wanting the outcome less. ### Check yourself **A monkey has learned that a light predicts juice. When the juice arrives exactly as predicted, do its dopamine neurons fire?** No. Once the cue reliably predicts juice, the burst moves to the cue and the fully predicted juice produces no response, because dopamine signals prediction error rather than pleasure. If the juice is then omitted, the neurons dip below baseline at the moment it should have come. **A rat whose dopamine neurons have been destroyed stops eating and would starve unless tube-fed. Has it lost the ability to enjoy food?** No. When sugar is placed on its tongue it shows the same lip-licking and paw-licking reactions as an intact rat, so liking is intact. What it has lost is wanting: the motivation to seek food, work for it, and treat its cues as attractive, which depends on dopamine while liking depends on separate opioid hotspots. **A trainer gives a treat about thirty seconds after a good sit, once she has found the treat bag. What does the dopamine burst strengthen?** Whatever the dog did in the last couple of seconds before the treat arrived, not the sit. Dopamine enlarges a synapse only if it arrives within roughly 0.3 to 2 seconds of the synapse's activity, and the sit's eligibility trace is long gone. The fix is a conditioned reinforcer, such as a click, delivered inside the window to bridge the gap. **Two rats were trained to press a lever for food, one moderately and one extensively. The food is then made unappealing. Which rat keeps pressing, and why?** The extensively trained rat. Its pressing has become a habit, supported by the dorsolateral striatum and triggered by antecedents regardless of the outcome's current value. The moderately trained rat's pressing is still goal-directed, supported by the dorsomedial striatum, so it stops once the outcome is devalued. **Explain it to a friend.** Explain why a reinforcer that arrives every single time eventually stops teaching anything, using the word "surprise" and without using the word "dopamine." ## Frequently asked questions **Is dopamine the "pleasure chemical"?** No. Dopamine neurons signal reward prediction error — how much better or worse an outcome was than expected — and drive wanting (motivation and cue attraction). Pleasure, or "liking," depends on separate opioid systems. Animals with dopamine removed still show every sign of enjoying sugar; they simply stop seeking it. **What is a reward prediction error?** The difference between the reward that arrives and the reward that was predicted. A positive error (better than expected) produces a burst of dopamine and strengthens the preceding behavior; zero error (as expected) produces nothing; a negative error (worse than expected, including an omitted reward) produces a dip. It is the neural version of the law of effect and the core of reinforcement-learning algorithms. **Which part of the brain is responsible for operant conditioning?** No single part. Dopamine neurons in the midbrain provide the teaching signal; the ventral striatum learns predictions; the dorsomedial striatum supports goal-directed action; the dorsolateral striatum supports habits; the prefrontal cortex supports planning and rule use; and the amygdala and lateral habenula handle aversive outcomes. Operant learning also occurs in invertebrates with far simpler nervous systems, and even in single neurons. **Why must reinforcement be immediate?** Because the synapses that were active during a behavior stay "eligible" for strengthening only briefly. In mouse striatal neurons, dopamine strengthened a synapse only if it arrived within about 0.3–2 seconds of the synapse's activity. A delayed reinforcer strengthens whatever happened just before it arrived, not the behavior you intended. **Does the brain treat punishment as the opposite of reinforcement?** Not exactly. Omitted rewards and predicted punishments produce dopamine dips, partly driven by the lateral habenula, and learning from them appears to run through a different striatal pathway than learning from rewards. In people with Parkinson's disease, dopamine medication improves learning from positive feedback and worsens learning from negative feedback. **How does this relate to habits?** With extended practice, control of a behavior shifts from a goal-directed system (dorsomedial striatum, sensitive to whether the outcome is still valuable) to a habit system (dorsolateral striatum, triggered by cues regardless of outcome value). That is why long-standing habits continue after their rewards have lost appeal, and why changing cues works better than willpower. ## References 1. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427. 2. Olds, J. (1958). Self-stimulation of the brain. *Science, 127*(3294), 315–324. 3. Routtenberg, A., & Lindy, J. (1965). Effects of the availability of rewarding septal and hypothalamic stimulation on bar pressing for food under conditions of deprivation. *Journal of Comparative and Physiological Psychology, 60*(2), 158–161. 4. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 5. Schultz, W. (1998). Predictive reward signal of dopamine neurons. *Journal of Neurophysiology, 80*(1), 1–27. 6. Montague, P. R., Dayan, P., & Sejnowski, T. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. *Journal of Neuroscience, 16*(5), 1936–1947. 7. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. 8. Steinberg, E. E., Keiflin, R., Boivin, J. R., Witten, I. B., Deisseroth, K., & Janak, P. H. (2013). A causal link between prediction errors, dopamine neurons and learning. *Nature Neuroscience, 16*(7), 966–973. 9. Tsai, H.-C., Zhang, F., Adamantidis, A., Stuber, G. D., Bonci, A., de Lecea, L., & Deisseroth, K. (2009). Phasic firing in dopaminergic neurons is sufficient for behavioral conditioning. *Science, 324*(5930), 1080–1084. 10. Fiorillo, C. D., Tobler, P. N., & Schultz, W. (2003). Discrete coding of reward probability and uncertainty by dopamine neurons. *Science, 299*(5614), 1898–1902. 11. Berridge, K. C., Venier, I. L., & Robinson, T. E. (1989). Taste reactivity analysis of 6-hydroxydopamine-induced aphagia: Implications for arousal and anhedonia hypotheses of dopamine function. *Behavioral Neuroscience, 103*(1), 36–45. 12. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? *Brain Research Reviews, 28*(3), 309–369. 13. Peciña, S., & Berridge, K. C. (2005). Hedonic hot spot in nucleus accumbens shell: Where do μ-opioids cause increased hedonic impact of sweetness? *Journal of Neuroscience, 25*(50), 11777–11786. 14. Salamone, J. D., & Correa, M. (2012). The mysterious motivational functions of mesolimbic dopamine. *Neuron, 76*(3), 470–485. 15. Robinson, T. E., & Berridge, K. C. (1993). The neural basis of drug craving: An incentive-sensitization theory of addiction. *Brain Research Reviews, 18*(3), 247–291. 16. Volkow, N. D., Koob, G. F., & McLellan, A. T. (2016). Neurobiologic advances from the brain disease model of addiction. *New England Journal of Medicine, 374*(4), 363–371. 17. Frank, M. J., Seeberger, L. C., & O'Reilly, R. C. (2004). By carrot or by stick: Cognitive reinforcement learning in parkinsonism. *Science, 306*(5703), 1940–1943. 18. Matsumoto, M., & Hikosaka, O. (2007). Lateral habenula as a source of negative reward signals in dopamine neurons. *Nature, 447*(7148), 1111–1115. 19. Yagishita, S., Hayashi-Takagi, A., Ellis-Davies, G. C. R., Urakubo, H., Ishii, S., & Kasai, H. (2014). A critical time window for dopamine actions on the structural plasticity of dendritic spines. *Science, 345*(6204), 1616–1620. 20. Richardson, R. T., & DeLong, M. R. (1990). Context-dependent responses of primate nucleus basalis neurons in a go/no-go task. *Journal of Neuroscience, 10*(8), 2528–2540. 21. Kilgard, M. P., & Merzenich, M. M. (1998). Cortical map reorganization enabled by nucleus basalis activity. *Science, 279*(5357), 1714–1718. 22. Adams, C. D., & Dickinson, A. (1981). Instrumental responding following reinforcer devaluation. *Quarterly Journal of Experimental Psychology B, 33*(2), 109–121. 23. Dickinson, A. (1985). Actions and habits: The development of behavioural autonomy. *Philosophical Transactions of the Royal Society B, 308*(1135), 67–78. 24. Yin, H. H., Knowlton, B. J., & Balleine, B. W. (2004). Lesions of dorsolateral striatum preserve outcome expectancy but disrupt habit formation in instrumental learning. *European Journal of Neuroscience, 19*(1), 181–189. 25. Yin, H. H., Ostlund, S. B., Knowlton, B. J., & Balleine, B. W. (2005). The role of the dorsomedial striatum in instrumental conditioning. *European Journal of Neuroscience, 22*(2), 513–523. 26. Yin, H. H., & Knowlton, B. J. (2006). The role of the basal ganglia in habit formation. *Nature Reviews Neuroscience, 7*(6), 464–476. 27. Dickinson, A., Nicholas, D. J., & Adams, C. D. (1983). The effect of the instrumental training contingency on susceptibility to reinforcer devaluation. *Quarterly Journal of Experimental Psychology B, 35*(1), 35–51. 28. O'Doherty, J., Dayan, P., Schultz, J., Deichmann, R., Friston, K., & Dolan, R. J. (2004). Dissociable roles of ventral and dorsal striatum in instrumental conditioning. *Science, 304*(5669), 452–454. 29. Everitt, B. J., & Robbins, T. W. (2005). Neural systems of reinforcement for drug addiction: From actions to habits to compulsion. *Nature Neuroscience, 8*(11), 1481–1489. 30. Fetz, E. E. (1969). Operant conditioning of cortical unit activity. *Science, 163*(3870), 955–958. 31. Brembs, B., Lorenzetti, F. D., Reyes, F. D., Baxter, D. A., & Byrne, J. H. (2002). Operant reward learning in *Aplysia*: Neuronal correlates and mechanisms. *Science, 296*(5573), 1706–1709. 32. Dayan, P., & Niv, Y. (2008). Reinforcement learning: The good, the bad and the ugly. *Current Opinion in Neurobiology, 18*(2), 185–196. 33. Daw, N. D., Niv, Y., & Dayan, P. (2005). Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control. *Nature Neuroscience, 8*(12), 1704–1711. 34. Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), *Classical Conditioning II: Current Research and Theory* (pp. 64–99). Appleton-Century-Crofts. 35. Berke, J. D. (2018). What does dopamine mean? *Nature Neuroscience, 21*(6), 787–793. 36. Matsumoto, M., & Hikosaka, O. (2009). Representation of negative motivational value in the primate lateral habenula. *Nature Neuroscience, 12*(1), 77–84. ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Why unpredictable reinforcers hold behavior — now with a neural reason. - [Building habits](https://operantconditioning.com/habits/): Use the action-to-habit transition on purpose. - [History](https://operantconditioning.com/history/): From Thorndike to reinforcement learning. --- # The Skinner Box (Operant Conditioning Chamber): What It Is, How It Works, and Why It Mattered > A Skinner box (operant conditioning chamber) is the apparatus Skinner built to measure how consequences change behavior: parts, a typical session, and myths. - Source: https://operantconditioning.com/skinner-box/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Apparatus · Methods* A small chamber, a lever, a food tray, and a pen that stepped up once per press. Here is what the operant conditioning chamber is, part by part, what it measured that nothing before it could, how to read the records it drew — and what the "baby in a box" story gets wrong. > **Definition** > > A **Skinner box** — formally an **operant conditioning chamber** — is an enclosed apparatus in which an animal can make a simple, repeatable response, such as pressing a lever or pecking a lit disk, that automatic equipment records and can follow with a programmed consequence, usually food. B. F. Skinner built the first versions at Harvard in the early 1930s to measure the *rate* of a freely emitted behavior over time. > > Its purpose is isolation, not confinement: one response, one consequence, and nothing else changing, so that the relation between them can be measured for hours without an experimenter in the room.[1] **In brief** - A Skinner box, formally an operant conditioning chamber, lets an animal make one simple response that equipment records and can follow with a consequence. - Unlike puzzle boxes and mazes, it records the rate of a free operant continuously; on the [cumulative record](https://operantconditioning.com/glossary/#cumulative-record), slope is rate. - Skinner never called it a Skinner box, and the "baby in a box" was a climate-controlled crib, not an experiment. ## What is a Skinner box? Most students meet the Skinner box as a picture: a rat in a metal cage pressing a bar. The picture is accurate but misses the point. The chamber was never the discovery; it was the instrument that made a kind of discovery possible. It let the animal respond whenever and as often as it liked and recorded every response automatically. That made *how often* an animal did something a measurable quantity, and almost everything in [operant conditioning](https://operantconditioning.com/) — reinforcement, extinction, [schedules](https://operantconditioning.com/schedules-of-reinforcement/), [shaping](https://operantconditioning.com/shaping/), [stimulus control](https://operantconditioning.com/stimulus-control/) — was first seen as a change in that quantity. By Skinner's account the box evolved by simplification: a runway along which a rat ran for food, whose timing proved surprisingly orderly, was shortened step by step until the rat needed only to press a horizontal bar to work a food dispenser.[2] *The Behavior of Organisms* (1938), which laid out the resulting science, is built almost entirely on records from it.[1] ### Why Skinner never called it a Skinner box The name was not his. It came out of Clark Hull's laboratory at Yale, which adopted a version of the apparatus in the 1930s; Skinner never used it, and the field settled on *operant conditioning chamber*.[3] The eponym misleads in two ways. It makes a method sound like a gadget, when the box was only a way of getting at the method — measuring the rate of a free operant. And within a few years the name attached itself to something else entirely: the enclosed crib Skinner built for his daughter, the subject of the most durable myth about him (below). ## The parts of a Skinner box  *An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a pellet on any schedule; the cue light signals when pressing will pay off; the recorder draws responses against time.* | Part | Rat version | Pigeon version | What it is for | | --- | --- | --- | --- | | **Operandum** | A small lever that closes a switch when pressed | A translucent disk, the "key," at head height, lit from behind; a peck closes the switch | Defines the response that counts. Anything that closes the switch — left paw, right paw, nose — is the same operant. | | **Food magazine** | A dispenser drops a 45-milligram pellet into a tray | A hopper of grain is raised into an opening, lit, for a few seconds | Delivers the reinforcer. Its click precedes every pellet and becomes a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer) that bridges press and food. | | **Stimulus lights** | A house light and cue lights above the lever | Colors or patterns projected onto the key | Signal when responding will pay: the discriminative stimulus and its opposite, the S-delta. | | **Speaker** | Tones, clicks, or white noise; the cubicle's fan masks outside sounds | Auditory stimuli, and a chamber that sounds the same every session. | | | **Grid floor** | Parallel metal rods over a droppings tray | Easy to clean; in punishment and [avoidance](https://operantconditioning.com/avoidance-learning/) experiments the rods can carry a brief shock. | | | **Programming and recording** | Originally relays and timers; now a computer | Runs the [schedule](https://operantconditioning.com/schedules-of-reinforcement/) and timestamps every response, so the experimenter can leave the room. | | ## How a Skinner box experiment runs A chamber with an untrained rat in it produces nothing: the rat sniffs the corners, grooms, and occasionally leans on the lever by accident. A working session is built in stages, each depending on the one before. 1. **Deprivation.** The animal is kept mildly hungry — typically around 80 to 85 percent of its free-feeding weight — so that food functions as a reinforcer. Without this [motivating operation](https://operantconditioning.com/abc-model/), nothing else works. 2. **Magazine training.** Pellets arrive on a timer, no response required. Within a few dozen deliveries the rat goes to the tray the instant it hears the click, now a conditioned reinforcer that can be delivered from across the chamber. 3. **Shaping.** The experimenter delivers a pellet by hand switch whenever the rat does something closer to a lever press: facing it, approaching, rearing, touching. Each approximation is reinforced until frequent, then dropped for the next. [How shaping works ›](https://operantconditioning.com/shaping/) 4. **Continuous reinforcement, then schedules.** Every press pays at first. Then the requirement is thinned — every fifth press, an unpredictable number, the first press after a fixed time — and the recorder shows each schedule's signature. 5. **Discrimination training.** Presses pay only while a light is on. Pressing comes under stimulus control — high in the light, near zero in the dark — the laboratory version of the antecedent in the [A-B-C model](https://operantconditioning.com/abc-model/). [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) 6. **Extinction.** Food stops. The record shows a burst of rapid, variable pressing, then a decline to almost nothing, with some recovery at the next session. [What extinction looks like ›](https://operantconditioning.com/extinction/) Later experiments are built on this base: punishment (a brief shock for a press that is still reinforced), choice (two keys on different schedules, the procedure behind the [matching law](https://operantconditioning.com/matching-law/)), drugs (a dose before the session, with the change in rate as the measure). You can run the magazine-training, shaping, schedule, and extinction stages in the virtual chamber below. ## Run a Skinner box yourself This lab puts an untrained, mildly hungry rat in a chamber and hands you the pellet button. Magazine-train it, shape a lever press one approximation at a time, put the press on a schedule, then stop the food and watch extinction. The cumulative record under the chamber draws every press. The rat is a stylized model — its tendencies shift with what you reinforce, drift back when you don't, and follow the schedule patterns Ferster and Skinner documented — not a replay of real data. Real shaping takes longer and real rats are more surprising. ## What a Skinner box measures, and why that mattered Edward Thorndike's cats, in 1898, were placed one at a time in a latched "puzzle box" with food outside; Thorndike timed each escape and plotted the times across trials.[4] The mazes and runways that dominated animal psychology for the next forty years worked the same way. These are **discrete-trial** methods: the experimenter decides when the animal may behave, each trial yields one number, and learning appears as a curve across trials. In the operant chamber the lever is always there. The animal decides when to respond and how often — a **[free operant](https://operantconditioning.com/glossary/#free-operant)** — and the natural measure becomes the **rate of response**, presses per minute, moment by moment. Skinner argued that rate was the datum psychology had been missing: it changes continuously rather than in trial-sized lumps, it can be read at a glance from a cumulative record, and it proved sensitive to almost every variable an experimenter could manipulate.[1][5] | Aspect | Thorndike's puzzle box | Maze or runway | Operant chamber | | --- | --- | --- | --- | | **Who starts a trial** | Experimenter places the cat | Experimenter places the rat at the start | The animal, whenever it likes | | **What is measured** | Seconds to escape | Time to the goal; wrong turns | Responses per unit time, continuously | | **What a session yields** | One number per trial | One or two numbers per trial | A record of every response and reinforcer | | **Best suited to** | Showing that consequences "stamp in" behavior | Spatial learning and motivation | Schedules, extinction, stimulus control, choice, drug effects | Some of the field's landmarks were accidents of the apparatus, and Skinner said so. The cumulative record began as a lucky by-product of the way his early food magazine was built. The first extinction curve he ever saw appeared when the pellet dispenser jammed and the rat went on pressing. And the first intermittent schedule was born when, running short of the pellets he made by hand, he reinforced only one press a minute to save them — and found the record perfectly orderly.[2] In the same kind of chamber, feeding pigeons every fifteen seconds regardless of what they did produced "superstitious" rituals in six of eight birds, the classic demonstration that accidental contingencies shape behavior.[6] The papers describing the method and these discoveries are collected in *Cumulative Record*.[7] ## Pigeons vs. rats *The Behavior of Organisms* is a book about rats. Skinner switched to pigeons during the Second World War, when Project Pigeon trained them to guide a missile by pecking at a target image, and largely stayed with them.[8] The bird pecks quickly and can sustain thousands of responses an hour, has sharp color vision (ideal for work on stimulus control), is cheap and hardy, and lives for a decade or more, so one subject can be studied for years. *Schedules of Reinforcement*, the 1957 catalog of schedule effects, is built mostly on pigeon records.[9] Rats never went away. They are nocturnal and nearly color-blind by comparison, but their physiology, genetics, and brain are far better mapped, which is why the rat (and now the mouse) chamber is standard in pharmacology and neuroscience. The logic is the same for any species — a defined response, an automated consequence, a continuous record of rate — and chambers have been built for monkeys, fish, and humans. ## How to read a cumulative record Skinner's second invention made the data visible. In the **cumulative recorder**, a strip of paper moves under a pen at constant speed; each response steps the pen a small fixed distance upward, and each reinforcer is marked with a short diagonal tick, or "pip." When the pen reaches the top of the paper it drops back to the bottom and continues. The [cumulative record](https://operantconditioning.com/glossary/#cumulative-record) is the field's native graph, and every schedule leaves a signature on it.[9]  *Reading a cumulative record: slope is rate, ticks are reinforcers, flat stretches are pauses, and the curved "scallop" is the fixed-interval pattern. The line only ever goes up; a behavior that has stopped draws a horizontal line.* - **Slope is rate.** Steep means fast responding; flat means the animal has stopped. A change in slope is a change in behavior, visible the moment it happens. - **Pauses.** A flat stretch right after a tick is the [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause), which lengthens as ratios and intervals grow. The fixed-ratio record is "break and run": pause, then a burst at full speed. - **Scallops.** On a fixed-interval schedule the line curves — nothing just after a reinforcer, then accelerating responding as the interval runs out. Variable schedules erase both pause and curve and draw a steady, nearly straight line. - **Extinction.** When reinforcement stops, the record rises steeply for a while (the extinction burst) and then bends toward horizontal. Ferster and Skinner's book contains hundreds of such records; the reader is expected to see the effects, not compute them. [Draw your own records in the schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) ## The "baby in a box": the Skinner box myth In 1944, before his second daughter Deborah was born, Skinner built what he called a **baby tender** and later an **air crib**: an enclosed crib with a large safety-glass front and a stretched canvas floor, supplied with filtered, warmed air, so that a baby could sleep and play in a diaper alone — no blankets, layers of clothing, or bars — at a constant temperature. It was a bed with climate control. Deborah was taken out to be fed, changed, held, and played with like any other infant. *Ladies' Home Journal* published Skinner's article about it in October 1945 under a headline he did not write, "Baby in a Box," and a number of families later used commercial versions.[10] > **What is false** > > The headline, and the fact that "Skinner box" was already a phrase, produced a rumor that circulated for decades: that Skinner raised his daughter in a Skinner box as an experiment, and that she became psychotic, sued him, or died by suicide. None of it is true. No experiment was ever run in the air crib; it had no lever, no dispenser, and no contingencies. Deborah Skinner Buzan grew up to be an artist living in Britain, has described a happy childhood with an affectionate father, and answered the rumor in 2004 in an article titled "I was not a lab rat," after a popular book repeated it.[11] The myth survives because it fits a picture of Skinner as a cold manipulator. The facts of the box cut against it: it was a measuring instrument, its reinforcers were overwhelmingly food rather than shock, and Skinner spent fifty years arguing against punishment. [B. F. Skinner: life, work, critics, and myths ›](https://operantconditioning.com/bf-skinner/) ## Modern descendants of the Skinner box The chamber remains standard laboratory equipment, though it rarely looks like Skinner's. - **Touchscreen and home-cage chambers.** The lever becomes a screen a rat or mouse touches with its nose, so tasks used to test memory and attention in human patients can be run in animals. In home-cage systems the operandum lives inside the animals' enclosure; they identify themselves electronically and work when they choose, around the clock, without handling stress. - **Drug self-administration.** Since James Weeks fitted rats with an intravenous line in 1962 so that a lever press delivered a dose of morphine, the operant chamber has been the principal animal model of addiction: the drugs people abuse are the ones animals will work for.[12] - **Brain stimulation and optogenetics.** In 1954 James Olds and Peter Milner showed that a rat would press a lever for a brief pulse of current to a region of its own brain — the first direct evidence that reinforcement has a physical address.[13] In the current version, the lever fires a laser that switches on a genetically targeted set of neurons, so the reinforcing effect of specific dopamine cells can be tested press by press. [The neuroscience of reinforcement ›](https://operantconditioning.com/neuroscience/) ### The human "Skinner box" The phrase now describes any product engineered to keep people responding: slot machines, infinite feeds, loot boxes. The analogy has substance. Skinner himself pointed to the gambling device as a variable-ratio schedule, the arrangement that produces the highest and most persistent rates in the laboratory, and the same schedule runs the pull-to-refresh feed.[14] > **A caution about the metaphor** > > A chamber isolates one response and one consequence under an experimenter's complete control. An app competes with every other reinforcer in your life, and you can walk away in a way a rat cannot. The metaphor is useful for naming the schedule and misleading when it suggests people are helpless: the same analysis says what to do — change the cue, the schedule, or the cost of responding. [Operant conditioning in technology and product design ›](https://operantconditioning.com/applications/#technology) ## Criticisms and limitations of the Skinner box **Artificiality.** The chamber studies a hungry animal, a bare environment, and a single arbitrary response, and critics have argued from the start that this tells us about rats in boxes rather than organisms in the world. Two findings from inside the operant tradition gave the objection teeth. Keller and Marian Breland, who left Skinner's laboratory to train animals commercially, reported that trained behavior drifted toward species-typical patterns — raccoons "washing" the coins they were supposed to deposit — whatever the contingencies said.[15] And the pigeon's key peck, the chamber's signature response, turned out to be partly Pavlovian: a pigeon will begin pecking a lit key that merely precedes food, even though pecking is not required and changes nothing.[16] The box does not create behavior from nothing; it selects from what the species brings. **Generalization to humans.** Noam Chomsky's charge was that terms precise in the pigeon laboratory become loose metaphors when stretched over human language and thought.[17] Skinner's later work conceded part of the point: humans follow rules and instructions, and rule-governed behavior can look quite different from behavior shaped directly by contingencies.[18] Human volunteers on laboratory schedules, who arrive with instructions and hypotheses of their own, often fail to show the animal patterns for that reason. **The reply.** The chamber is a method for isolating variables, as a physicist's frictionless plane is, and its findings have to be tested outside it. They were: reinforcement, extinction, schedule effects, shaping, and stimulus control replicated across species and then across classrooms, clinics, zoos, and workplaces, where [applied behavior analysis](https://operantconditioning.com/applications/) measures them against real behavior. Whether that vindicates the philosophy Skinner built on the box is a different, and still open, question. [The history of operant conditioning, including the cognitive critique ›](https://operantconditioning.com/history/) ## Key takeaways - The chamber was never the discovery; it was the instrument. It isolates one response and one consequence, with nothing else changing, so their relation can be measured for hours without an experimenter in the room. - Its measure is the rate of a free operant. Unlike Thorndike's puzzle box or a maze, where the experimenter starts each trial and gets one number, the animal decides when and how often to respond, and every response is recorded. - On a cumulative record, slope is rate, ticks are reinforcers, and flat stretches are pauses. Each schedule leaves a signature: the fixed-ratio break and run, the fixed-interval scallop, the steady line of variable schedules, and the burst-then-flatten of extinction. - A session is built in stages that each depend on the one before: deprivation, magazine training, shaping, continuous reinforcement then schedules, discrimination training, and extinction. Punishment, choice, and drug experiments are built on that base. - The "baby in a box" was an air crib, a climate-controlled bed with no lever, dispenser, or contingencies; no experiment was ever run in it. The chamber's real limitation is artificiality: it selects from what the species brings rather than creating behavior from nothing. ### Check yourself **On a cumulative record, the line goes flat for a stretch right after every reinforcer tick, then shoots up steeply. A student concludes the pellet must be punishing the rat, since pressing stops each time one arrives. What is wrong with that reading?** The flat stretch is the post-reinforcement pause, the first half of the fixed-ratio "break and run" pattern, and the steep run that follows shows pressing is as strong as ever. A behavior that has actually stopped draws a line that stays horizontal; a pause followed by a burst at full speed is the schedule's signature, not evidence of punishment. **Thorndike's cat and Skinner's rat both learn to work a mechanism to get food. Why did Skinner treat his chamber as a different kind of measurement?** The puzzle box is a discrete-trial method: the experimenter starts each trial and gets one number, the time to escape. In the chamber the lever is always there, so the animal decides when and how often to respond, and the measure is the rate of a free operant, recorded continuously and read from the slope of the record. **A pigeon begins pecking a lit key that comes on just before free food, even though pecking is not required and changes nothing. Does this show operant conditioning at work?** Not on its own. Because pecking has no consequence, the light merely preceding food is enough to produce it, which is why the key peck is described as partly Pavlovian. It is one of the findings behind the criticism that the box does not create behavior from nothing but selects from what the species brings. **An untrained rat is placed in a chamber with the lever already wired to the pellet dispenser. Why does nothing happen, and which two stages have to come first?** The rat has no reason to press and presses only by accident. It must first be mildly hungry, so that food functions as a reinforcer, and then magazine-trained, so that the dispenser's click becomes a conditioned reinforcer that can be delivered from across the chamber. Only then can shaping reinforce closer and closer approximations to a press. **Explain it to a friend.** Explain what a Skinner box measures that Thorndike's puzzle box could not, without using the words "rate" or "trial." ## Frequently asked questions **What is a Skinner box?** An enclosed apparatus, formally called an operant conditioning chamber, in which an animal can make a simple response — pressing a lever or pecking a lit key — that equipment records automatically and can follow with a consequence such as food. B. F. Skinner built it in the early 1930s to measure the rate of freely emitted behavior and study how its consequences change it. **What did Skinner's box experiment show?** There was no single experiment but a research program. Its central findings: behavior followed by a reinforcer becomes more frequent and fades when the reinforcer stops (extinction); how often reinforcement is delivered — the schedule — controls the rate, pattern, and persistence of behavior; and behavior comes under the control of stimuli that signal when reinforcement is available. The 1948 "superstition" study showed that accidental reinforcement is enough to build a ritual. **Is the Skinner box cruel?** In the typical experiment the animal is kept mildly hungry, works for food in a daily session of an hour or so, and is otherwise housed and fed normally; the great majority of operant research uses positive reinforcement. Some experiments on punishment and avoidance use brief shock through the grid floor, and those are the ones most people object to; today all such work is reviewed by institutional animal-care committees. Given free food in the chamber, many animals still work the lever — a phenomenon called [contrafreeloading](https://operantconditioning.com/operant-vs-classical-conditioning/#contrafreeloading). **Are Skinner boxes still used?** Yes. Modern chambers with touchscreens, home-cage versions, and computer control are standard in neuroscience, pharmacology, and behavioral genetics, and drug self-administration in an operant chamber is the main animal model of addiction. The principles first found in the box also underlie applied behavior analysis and reinforcement-based animal training. **Did Skinner put his daughter in a Skinner box?** No. He built a climate-controlled crib, the air crib, in which his daughter Deborah slept as an infant. It was a bed, not an experiment, and had no lever, dispenser, or contingencies. The rumor that she was harmed by it, sued him, or died by suicide is false; she said so herself in a 2004 newspaper article. **What is the difference between a Skinner box and a puzzle box?** Thorndike's puzzle box (1898) was a discrete-trial method: the cat was put in, escaped by working a latch, and was put back for the next trial, with escape time as the measure. A Skinner box is a free-operant method: the animal stays in, responds whenever it likes, and the rate of responding is recorded continuously. The puzzle box showed that consequences strengthen behavior; the Skinner box made it possible to measure exactly how. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Skinner, B. F. (1956). A case history in scientific method. *American Psychologist, 11*(5), 221–233. 3. Skinner, B. F. (1979). *The Shaping of a Behaviorist*. Knopf. 4. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 5. Skinner, B. F. (1950). Are theories of learning necessary? *Psychological Review, 57*(4), 193–216. 6. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 7. Skinner, B. F. (1959). *Cumulative Record*. Appleton-Century-Crofts. 8. Skinner, B. F. (1960). Pigeons in a pelican. *American Psychologist, 15*(1), 28–37. 9. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 10. Skinner, B. F. (1945, October). Baby in a box. *Ladies' Home Journal*. 11. Buzan, D. S. (2004, March 12). I was not a lab rat. *The Guardian*. 12. Weeks, J. R. (1962). Experimental morphine addiction: Method for automatic intravenous injections in unrestrained rats. *Science, 138*(3537), 143–144. 13. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427. 14. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 15. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 16. Brown, P. L., & Jenkins, H. M. (1968). Auto-shaping of the pigeon's key-peck. *Journal of the Experimental Analysis of Behavior, 11*(1), 1–8. 17. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58. 18. Skinner, B. F. (1969). *Contingencies of Reinforcement: A Theoretical Analysis*. Appleton-Century-Crofts. ## Related - [B. F. Skinner](https://operantconditioning.com/bf-skinner/): The life, the work, the critics, and the myths of the man who built the box. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): The patterns the cumulative recorder revealed — with a live simulator. - [Shaping](https://operantconditioning.com/shaping/): How a rat that has never seen a lever learns to press it, one approximation at a time. --- # B. F. Skinner: Biography, the Skinner Box, and Operant Conditioning > B. F. Skinner (1904–1990) founded radical behaviorism and the science of operant conditioning. His life, the Skinner box, major works, critics, and the myths. - Source: https://operantconditioning.com/bf-skinner/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *People* The psychologist who turned the law of effect into a laboratory science, built the box that carries his name, and spent fifty years arguing that behavior is selected by its consequences. His life, his work, his critics — and the myths that still follow him. > **Who he was** > > **B. F. Skinner** (Burrhus Frederic Skinner, March 20, 1904 – August 18, 1990) was an American psychologist at Harvard University who founded **radical behaviorism** and the **experimental analysis of behavior** — the science of [operant conditioning](https://operantconditioning.com/). He coined the term "operant," invented the operant conditioning chamber (the "Skinner box") and the cumulative recorder, and argued that behavior is selected by its consequences in the way species are selected by their environments. > > Born in Susquehanna, Pennsylvania, he died of leukemia in Cambridge, Massachusetts, eight days after receiving the American Psychological Association's first citation for an outstanding lifetime contribution to psychology. A 2002 survey ranked him the most eminent psychologist of the twentieth century.[1] **In brief** - B. F. Skinner (1904–1990) founded radical behaviorism and the experimental analysis of behavior, arguing that behavior is selected by its consequences. - He invented the [operant conditioning chamber](https://operantconditioning.com/skinner-box/) and the cumulative recorder, which made the rate of freely emitted behavior a measurable quantity. - Radical behaviorism does not deny thoughts and feelings; it treats them as private behavior to be explained rather than as ultimate causes. ## Who was B. F. Skinner? Skinner grew up in a small railroad town in northeastern Pennsylvania, the son of a lawyer, and spent his boyhood building things — wagons, rafts, a perpetual-motion machine that did not work. He went to Hamilton College intending to become a writer, and after graduating with a degree in English in 1926 he tried. Robert Frost had read three of his short stories and sent an encouraging letter. A year in his parents' house produced almost nothing, and Skinner later concluded that he had failed as a writer because he had nothing important to say. He decided that literature described behavior and that science might explain it.[2] What pushed him toward psychology was reading: John B. Watson's *Behaviorism*, Bertrand Russell's essays on Watson, and Ivan Pavlov's *Conditioned Reflexes*, newly translated into English in 1927. He entered Harvard's graduate program in 1928 with no formal training in the subject, worked largely in the physiology laboratory of William Crozier, and set himself a schedule of study so strict that he later described it with some embarrassment. By 1931 he had a PhD; by 1938 he had a book that defined a new field.[3] The thread running through everything he did afterward is a single ambition: to make psychology a natural science of behavior, with the same standing as physics or biology, by finding orderly relationships between what an organism does and the conditions under which it does it — without appealing to a mind inside to do the explaining. ## B. F. Skinner's life and career: a timeline - **1904** — Born in Susquehanna, Pennsylvania March 20. His father, William, was a lawyer; his mother, Grace, gave him her maiden name, Burrhus. Friends and family called him Fred. - **1926** — Hamilton College, then the "dark year" Graduates with a BA in English. Spends a year at home in Scranton trying to write fiction, then a few months in Greenwich Village. Reads Watson, Pavlov, and Russell, and decides to become a psychologist.[2] - **1928–1931** — Harvard graduate school Builds his own apparatus — a series of increasingly automated boxes for studying rats' eating and lever-pressing — and, almost by accident, the cumulative recorder. Completes his PhD in 1931 with a thesis on the concept of the reflex.[3] - **1933–1936** — Harvard Society of Fellows One of the first Junior Fellows of the newly founded Society, which gives him three years to do research with no teaching duties. In 1935 he publishes the paper distinguishing two types of conditioned reflex that will lead to the word "operant."[4] - **1936** — University of Minnesota Takes his first faculty post. Marries Yvonne (Eve) Blue the same year. Daughter Julie is born in 1938; Deborah in 1944. - **1937** — "Operant" and "respondent" In a reply to the Polish physiologists Konorski and Miller, Skinner introduces the terms that separate behavior *emitted* and controlled by its consequences from reflexive behavior *elicited* by a prior stimulus.[5] - **1938** — *The Behavior of Organisms* His first book lays out the experimental analysis of behavior: rate of response as the basic datum, reinforcement, extinction, discrimination, and the first schedule effects, all from rats in his apparatus.[6] - **1940–1944** — Project Pigeon Trains pigeons to guide a missile by pecking at a target image projected inside its nose cone. The birds worked; the military, committed to radar and electronic guidance, canceled the project in 1944. Along the way, in 1943, Skinner and his collaborators discover [shaping](https://operantconditioning.com/shaping/) while teaching a pigeon to "bowl."[7][8] - **1945** — Indiana, the air crib, and radical behaviorism Becomes chair of psychology at Indiana University. Publishes "Baby in a Box" in *Ladies' Home Journal* describing the climate-controlled crib he built for Deborah, and "The Operational Analysis of Psychological Terms," the founding statement of radical behaviorism.[9][10] - **1948** — Back to Harvard; *Walden Two*; superstition Returns to Harvard, where he stays for the rest of his career. Publishes the utopian novel *Walden Two*, written in seven weeks in 1945, and the "superstition in the pigeon" experiment.[11] - **1953** — *Science and Human Behavior* and the teaching machine Extends operant analysis to thinking, self-control, government, religion, and education. In November, a visit to his daughter Deborah's fourth-grade arithmetic class convinces him that classrooms violate everything known about reinforcement, and he builds his first teaching machine within days.[12][13] - **1957** — Two big books *Schedules of Reinforcement*, with Charles Ferster, catalogs thousands of hours of cumulative records. *Verbal Behavior*, developed from his 1947 William James Lectures, treats language as operant behavior.[14][15] - **1958–1959** — A journal, a chair, and a famous review The *Journal of the Experimental Analysis of Behavior* is founded. Skinner is named Edgar Pierce Professor of Psychology and receives the APA's Distinguished Scientific Contribution Award. In 1959 Noam Chomsky publishes his review of *Verbal Behavior*.[16] - **1968–1971** — National Medal of Science; *Beyond Freedom and Dignity* Receives the National Medal of Science in 1968. In 1971 *Beyond Freedom and Dignity* becomes a bestseller and puts him on the cover of *Time*, arguing that "autonomous man" is a fiction and that a culture must design its own contingencies.[17] - **1974** — *About Behaviorism*; retirement Retires from Harvard as professor emeritus and publishes his clearest statement of what radical behaviorism does and does not claim, written to correct twenty common misreadings.[18] - **1976–1983** — The autobiography Publishes a three-volume life — *Particulars of My Life*, *The Shaping of a Behaviorist*, and *A Matter of Consequences* — written in the same plain, external style he used for everything else. - **1990** — Final address and death On August 10, at the APA convention in Boston, he accepts the first Citation for Outstanding Lifetime Contribution to Psychology and delivers a last talk, "Can Psychology Be a Science of Mind?" He dies of leukemia at home in Cambridge, Massachusetts, on August 18.[19] ## What is a Skinner box? A **Skinner box** — Skinner's own term was **operant conditioning chamber** — is an enclosed, sound-attenuated space in which an animal can make a simple, repeatable response that the apparatus records and, under programmed conditions, rewards. Its purpose is not to trap the animal but to isolate one behavior and one consequence so that the relationship between them can be measured precisely, for hours at a time, without an experimenter in the room. [A full tour of the box, part by part ›](https://operantconditioning.com/skinner-box/) ### How the operant conditioning chamber works - **The operandum.** For rats, a small lever that closes a switch when pressed. For pigeons, a translucent plastic disk — the "key" — mounted at head height and lit from behind; a light peck closes a switch and registers a response. - **The food magazine.** A dispenser that drops a 45-milligram pellet into a tray for rats, or raises a hopper of grain for a few seconds for pigeons. The sound of the magazine operating becomes a conditioned reinforcer, bridging the gap between the response and the food. - **Stimuli.** Lights, tones, and the color of the key light let the experimenter signal when responding will pay off — the basis of [stimulus control](https://operantconditioning.com/abc-model/) and the discriminative stimulus. - **Programming.** Originally electromechanical relays and timers, later computers, deliver reinforcement on whatever [schedule](https://operantconditioning.com/schedules-of-reinforcement/) the experiment requires and record every response with its time. The crucial feature is that the animal is free to respond at any time and at any rate. This **free-operant** method, as opposed to discrete trials in mazes and runways, is what made *rate of response* available as a measure — and rate turned out to be exquisitely sensitive to the conditions of reinforcement.[6] ### The cumulative recorder Skinner's second invention was the instrument that made the data visible. A strip of paper moves at a constant speed under a pen; each response steps the pen a small fixed distance upward, and a diagonal tick marks each reinforcement. The result is a **cumulative record** whose slope is the rate of responding: steep means fast, flat means the animal has stopped, and the shape of the curve between reinforcers reveals the pattern a schedule produces. The famous "scallop" of the fixed-interval schedule and the "break-and-run" of the fixed-ratio schedule were first seen as pen strokes on this paper.[14] ### Why Skinner disliked the name The nickname "Skinner box" came from Clark Hull's laboratory at Yale, not from Skinner, who never used it and objected to it — partly because it suggested a gadget rather than a method, and partly because the term became attached in the popular mind to something quite different, which is the next story.[3] ### The "baby in a box": the air crib and a persistent rumor In 1944, before his second daughter Deborah was born, Skinner built what he called a **baby tender** and later an **air crib**: an enclosed crib with a large safety-glass front, a stretched canvas floor, and filtered, warmed air, so that the baby could sleep and play in a diaper alone, without blankets or layers of clothing, at a controlled temperature. It was a crib, used for sleeping and unattended play; Deborah was taken out to be fed, held, and played with like any other infant. *Ladies' Home Journal* published his article about it in October 1945 under a headline he did not choose, "Baby in a Box," and a number of families later used commercial versions.[9] > **What is false** > > A rumor circulated for decades that Skinner raised his daughter in a "Skinner box" as an experiment, and that she became psychotic, sued her father, or died by suicide. None of it is true. Deborah Skinner Buzan is an artist living in Britain, has said she was a happy child with a loving father, and publicly answered the rumor in 2004 in an article titled "I was not a lab rat" after a popular book repeated it.[20] The air crib was a bed, not an operant chamber, and no experiments were run in it. ## Skinner's famous experiments Four experiments come up again and again, partly because they are vivid and partly because each one makes a point that still matters. ### "Superstition" in the pigeon (1948) Skinner put hungry pigeons in a chamber in which the food hopper appeared for a few seconds at regular intervals — every fifteen seconds — no matter what the bird was doing. Six of eight birds developed a distinctive ritual: one turned counterclockwise around the cage between feedings, another thrust its head repeatedly into an upper corner, a third made a "tossing" motion as if lifting an invisible bar with its head, two swung their heads and bodies like a pendulum, and one made brushing movements toward the floor. Skinner's explanation was **adventitious reinforcement**: whatever the bird happened to be doing when food arrived was strengthened, which made it more likely to be under way at the next delivery, which strengthened it again. Contingency in the world is not required; contiguity is enough.[11] The experiment is also a lesson in how science corrects itself. When Staddon and Simmelhag repeated it in 1971 and recorded behavior throughout each interval, they found that the "terminal" behaviors just before food were much the same across birds — mostly pecking near the hopper — and looked less like accidentally reinforced rituals than like responses induced by the periodic arrival of food itself. The rituals were real; the explanation was more complicated than Skinner's.[30] The glossary entry on [superstitious behavior](https://operantconditioning.com/glossary/#superstitious-behavior) gives the everyday version. ### Project Pigeon (1940–1944) During the Second World War Skinner trained pigeons to guide a missile. A lens in the nose cone projected the image of the target onto a screen; the pigeon, harnessed inside, pecked at the image, and the position of its pecks steered the missile. Three pigeons voted for reliability. With support first from General Mills and then, in 1943, from the government's Office of Scientific Research and Development, the birds learned to track a target through distortion, noise, and cold, and kept pecking steadily under conditions that unsettled human observers. The project was cancelled in 1944 — the committee could not take a pigeon-guided bomb seriously against electronic guidance — and was briefly revived by the Navy after the war as Project ORCON.[7] Its lasting product was accidental: in 1943, while teaching a pigeon to "bowl" a ball down a miniature alley, Skinner and his colleagues discovered how quickly behavior could be built by reinforcing successive approximations by hand — [shaping](https://operantconditioning.com/shaping/).[8] ### The ping-pong pigeons In a demonstration still shown in classrooms, two pigeons stand at either end of a small table and bat a ping-pong ball back and forth with their beaks. Each bird was first shaped to peck the ball, then to peck it toward the other end; once both could rally, food was delivered to the bird that got the ball past its opponent. Skinner presented it as a "synthetic social relation": competition assembled from individual contingencies, with no need to assume the birds understood the game.[31] The companion demonstration in the same paper built cooperation the same way — two pigeons reinforced only when they pecked matching keys at nearly the same moment. ### Teaching machines (1954–1958) On a visit to his daughter Deborah's fourth-grade arithmetic class in November 1953, Skinner watched some children sit idle after finishing a problem sheet while others struggled, all of them waiting a day or more to learn whether their answers were right. By his own account he built a prototype teaching machine within days. The machine presented material in small steps, required the student to compose an answer rather than pick one, showed the correct answer immediately, and let each student move at his or her own pace. He described the principles at a 1954 conference and set out the case for "teaching machines" in *Science* in 1958, crediting Sidney Pressey's self-scoring devices of the 1920s as a precursor.[13][22] The machines themselves went out of fashion within a decade; the principles — small steps, active responding, immediate feedback, self-pacing — reappear in programmed instruction and in most learning software written since. ## Skinner's key contributions to psychology | Year | Contribution | Why it matters | | --- | --- | --- | | 1930–1938 | Operant chamber, cumulative recorder, rate as the datum | Made moment-to-moment behavior measurable and reproducible; the method the entire field still uses.[6] | | 1935–1937 | Operant vs. respondent behavior | Separated behavior controlled by consequences from reflexes elicited by stimuli; "operant conditioning" gets its name.[4][5] | | 1938 | *The Behavior of Organisms* | First systematic account of reinforcement, extinction, discrimination, and schedule effects in the free-operant method.[6] | | 1940–1944 | Project Pigeon | Demonstrated fine stimulus control in pigeons; produced the discovery of shaping; a famous example of practical operant technology ahead of its time.[7] | | 1943; 1951 | Shaping by successive approximation | Showed how new behavior can be built by reinforcing closer and closer approximations; explained to the public in "How to Teach Animals."[8][21] | | 1945 | Radical behaviorism | Brought private events — thinking, feeling — inside the analysis as behavior, rather than excluding them as Watson had.[10] | | 1948 | "Superstition" in the pigeon | Showed that accidental reinforcement produces and maintains behavior; contingency, not the animal's understanding, does the work.[11] | | 1948 | *Walden Two* | A novel imagining a community designed on behavioral principles; inspired real intentional communities and decades of argument. | | 1953 | *Science and Human Behavior* | The textbook that applied operant analysis to self-control, thinking, social behavior, and institutions.[12] | | 1954; 1958 | Teaching machines and programmed instruction | Small steps, active responding, immediate feedback, self-pacing — principles that survive in modern learning software.[13][22] | | 1957 | *Schedules of Reinforcement* (with Ferster) | Established that how often behavior is reinforced controls its rate, pattern, and persistence.[14] | | 1957 | *Verbal Behavior* | A functional analysis of language (mands, tacts, intraverbals) that underlies much of modern language intervention in applied behavior analysis.[15] | | 1971 | *Beyond Freedom and Dignity* | Argued that behavior is always controlled and that cultures should design their contingencies deliberately; his most controversial book.[17] | | 1974 | *About Behaviorism* | His definitive answer to critics and misreadings of the philosophy behind the science.[18] | | 1981 | "Selection by Consequences" | Placed operant conditioning alongside natural selection and cultural evolution as a third kind of selection.[23] | ## What is radical behaviorism? **Radical behaviorism** is Skinner's philosophy of the science of behavior. "Radical" means thoroughgoing, not extreme: where Watson's *methodological* behaviorism ruled private experience out of psychology because it could not be observed by a second person, Skinner ruled it *in*. Thinking, feeling, imagining, and seeing with your eyes closed are, in his account, behavior — covert, but subject to the same variables as any other behavior, and open to study through the verbal reports a community teaches each of us to make about them.[10][18] What radical behaviorism rejects is not the *existence* of mental life but its use as an *explanation*. To say a student studies "because she is motivated" or a rat presses "because it expects food" is, for Skinner, to stop the analysis one step too soon: the motivation and the expectation must themselves be explained, and when they are, the explanation turns out to be a history of consequences and a present set of conditions. He called inner causes that merely restate the behavior they are supposed to explain "explanatory fictions."[12] Three commitments follow: - **Behavior is a subject matter in its own right**, not a symptom of something happening elsewhere (the mind, the brain). Neuroscience is welcome; it fills in mechanism, but it does not replace the functional relations. - **The causes of behavior lie in the organism's genetic endowment, its history of reinforcement, and its current environment.** Feelings are real but are collateral products of those same contingencies — we feel the effects of the causes, not the causes themselves. - **Selection by consequences.** Just as natural selection explains the appearance of design in organisms without a designer, reinforcement explains the appearance of purpose in behavior without purpose as a prior cause.[23] > **The most common misreading** > > Skinner did not deny that people think and feel, and he did not say the organism is "empty." The first chapter of *About Behaviorism* opens with a list of twenty things commonly said about behaviorism — that it ignores consciousness, that it treats people as robots, that it cannot account for creativity — and calls every one of them wrong.[18] The dispute is about where the causes are, not about what exists. ## Controversies and critiques of Skinner ### Chomsky's review of *Verbal Behavior* The most consequential criticism Skinner ever received was Noam Chomsky's 1959 review of *Verbal Behavior*. Chomsky argued that terms like "stimulus," "response," and "reinforcement," precise in the pigeon laboratory, became vacuous metaphors when stretched over human language, and that children acquire grammar far too fast, from far too little input, for reinforcement to be the mechanism.[16] The review is often credited with helping launch the cognitive revolution. Skinner never published a reply, and later said he had read only part of it. The detailed answer came from Kenneth MacCorquodale in 1970, who argued that Chomsky had reviewed Hull-style stimulus–response psychology rather than Skinner's book, had misdescribed reinforcement as a theory of drive reduction, and had criticized the book for failing to be a theory of grammar when it set out to be a functional analysis of why speakers say what they say.[24] Both essays remain worth reading; on the specific question of syntax, most linguists sided with Chomsky, while *Verbal Behavior*'s functional categories went on to become the basis of a large applied literature on teaching language to children with autism. ### Free will and *Beyond Freedom and Dignity* Skinner's 1971 book argued that the idea of an autonomous inner agent who freely chooses is a pre-scientific holdover, and that the real question is not whether behavior will be controlled — it always is — but by what and for whose benefit.[17] The reaction was intense. Chomsky reviewed it, too, under the title "The Case Against B. F. Skinner," and philosophers objected that a science of behavior cannot settle a metaphysical question by fiat. Skinner's answer was that the traditional view had not produced a technology for solving problems like overpopulation, pollution, or war, and that dignity, properly understood, is the credit we give people when we cannot see what controls them. > Give me the specifications, and I'll give you the man! > > *— Frazier, the fictional founder of the community in Skinner's *Walden Two*, 1948* ### The ethics of behavioral control Lines like Frazier's are why critics heard totalitarianism in *Walden Two*. Skinner's position was that a science of behavior is a tool, that the alternative to designed contingencies is accidental ones, and that the safeguard against abuse is **countercontrol**: the controlled must be able to control the controller. Whether that is enough is a live question, and it is why the ethics codes of applied behavior analysis put consent, least-restrictive procedures, and the client's own goals at the center. [How ABA is practiced today ›](https://operantconditioning.com/applications/#aba) ### Skinner on punishment Skinner was, throughout his career, an opponent of punishment as a method of control. His view rested on early evidence — his own 1938 experiments and W. K. Estes's 1944 dissertation — that punishment only temporarily suppressed responding while the reinforced behavior remained intact underneath, together with its side effects: fear, aggression, escape, and avoidance.[6][25] Later work by Azrin and Holz showed that sufficiently immediate and intense punishment *can* produce lasting suppression, so Skinner's empirical claim was too strong.[26] His practical conclusion — build behavior with [reinforcement](https://operantconditioning.com/positive-reinforcement/), and treat [punishment](https://operantconditioning.com/positive-punishment/) as a last resort — is nonetheless where the applied field ended up. ## Skinner's legacy Few psychologists have left as many working descendants: - **Applied behavior analysis (ABA).** A licensed profession in most U.S. states, built directly on Skinner's principles and his students' work, used in autism intervention, developmental disabilities, brain-injury rehabilitation, and behavioral medicine. - **Education.** Programmed instruction, mastery learning, precision teaching, and school-wide Positive Behavioral Interventions and Supports (PBIS) all trace to his 1950s work on teaching. [Operant conditioning in education ›](https://operantconditioning.com/applications/#education) - **Animal training.** Shaping, conditioned reinforcers, and the clicker were carried from his laboratory into the training world by Keller and Marian Breland and later Karen Pryor. [Reinforcement-based dog training ›](https://operantconditioning.com/dog-training/) - **Behavioral economics.** The matching law, delay discounting, and the quantitative analysis of choice grew out of the operant laboratory, largely through his student Richard Herrnstein. - **Reinforcement learning in AI.** The branch of machine learning behind game-playing programs, robotics, and the fine-tuning of large language models is a mathematical formalization of learning from consequences, and its founders cite the operant tradition explicitly.[27] - **Self-management.** Skinner's own late work on managing one's own behavior in old age, and the whole tradition of [habit formation through antecedents and consequences](https://operantconditioning.com/habits/), applies the three-term contingency to oneself. His works are kept in print by the B. F. Skinner Foundation, led by his daughter Julie S. Vargas. [The full history of operant conditioning, from Thorndike to reinforcement learning ›](https://operantconditioning.com/history/) ## Myths about B. F. Skinner | Myth | What is actually true | | --- | --- | | He raised his daughter in a Skinner box. | He built a climate-controlled crib. She slept and played in it as an infant, was raised normally, and is alive; she has publicly refuted the story.[20] | | He denied that thoughts and feelings exist. | Radical behaviorism treats thoughts and feelings as private behavior, real and worth studying. What he denied is that they are the ultimate *causes* of what we do.[18] | | He believed people are blank slates shaped entirely by environment. | He wrote explicitly that behavior is the joint product of genetic endowment, individual history, and current setting, and devoted a 1966 paper to the role of evolution.[28] | | He said "give me a dozen healthy infants" and he could make them into anything. | That is John B. Watson, writing in 1924, decades before Skinner's career — and Watson himself added that he was going beyond the evidence.[29] | | His main method was shocking animals. | His research was overwhelmingly about food reinforcement in mildly hungry rats and pigeons. He argued against punishment for fifty years; the definitive punishment research was done by others.[26] | | Chomsky "destroyed" behaviorism in 1959. | The review was influential on the study of language, but the experimental analysis of behavior grew steadily afterward, founding its applied journal in 1968 and a profession in the decades since. | ## Selected bibliography - *The Behavior of Organisms: An Experimental Analysis* (1938) - *Walden Two* (1948) - *Science and Human Behavior* (1953) - *Schedules of Reinforcement*, with C. B. Ferster (1957) - *Verbal Behavior* (1957) - *Cumulative Record* (1959; expanded editions 1961, 1972) — collected papers - *The Technology of Teaching* (1968) - *Contingencies of Reinforcement: A Theoretical Analysis* (1969) - *Beyond Freedom and Dignity* (1971) - *About Behaviorism* (1974) - *Particulars of My Life* (1976), *The Shaping of a Behaviorist* (1979), *A Matter of Consequences* (1983) — autobiography - *Enjoy Old Age: A Program of Self-Management*, with M. E. Vaughan (1983) None of Skinner's books is in the public domain, so none can be republished here. The books he built on — Thorndike, Morgan, James, Yerkes, Pfungst, Darwin — are, and they are in the [library](https://operantconditioning.com/library/) in full text. ## Key takeaways - Skinner's single ambition was a natural science of behavior: orderly relations between what an organism does and the conditions under which it does it, without appealing to a mind inside to do the explaining. - The operant chamber and the cumulative recorder made rate of response the basic datum. Because the animal is free to respond at any time, rate became available as a measure, and it proved sensitive to the conditions of reinforcement. - Radical behaviorism rules private experience in, not out: thinking and feeling are covert behavior, subject to the same variables as any other behavior. What it rejects is using mental states as explanations, which Skinner called explanatory fictions. - Chomsky's 1959 review charged that laboratory terms become metaphors when stretched over language, and most linguists sided with him on syntax; Azrin and Holz showed that immediate, intense punishment can produce lasting suppression, so Skinner's empirical claim against punishment was too strong. His practical conclusion, build behavior with reinforcement and treat punishment as a last resort, is where the applied field ended up. - The "baby in a box" was a climate-controlled crib, not an experiment. Skinner did not deny that thoughts and feelings exist, did not claim people are blank slates, and the "dozen healthy infants" line belongs to Watson. ### Check yourself **A classmate says Skinner, like Watson, held that psychology should ignore thoughts and feelings because a second person cannot observe them. Is that right?** No. That is Watson's methodological behaviorism, which ruled private experience out. Skinner's radical behaviorism ruled it in: thinking and feeling are covert behavior, real and worth studying; what he rejected was treating them as the ultimate causes of what a person does. **In the 1948 experiment, pigeons were fed every fifteen seconds no matter what they did, yet six of eight developed rituals. A student concludes the birds had figured out that the ritual caused the food. What was Skinner's explanation, and what did later work add?** Skinner's explanation was adventitious reinforcement: whatever the bird happened to be doing when food arrived was strengthened, making it more likely to be under way at the next delivery. No understanding is required; contiguity is enough. Staddon and Simmelhag later found that the behaviors just before food were much the same across birds and looked induced by the periodic arrival of food itself, so the rituals were real but the explanation was more complicated. **Skinner argued for fifty years that punishment only temporarily suppresses behavior. Did the evidence support that claim?** Only partly. His view rested on his own 1938 experiments and Estes's 1944 dissertation, but Azrin and Holz later showed that sufficiently immediate and intense punishment can produce lasting suppression, so the empirical claim was too strong. His practical conclusion still stands in applied work: build behavior with reinforcement and treat punishment as a last resort. **Why did Skinner treat the free-operant method as a breakthrough rather than just a new gadget?** In mazes and runways the experimenter runs discrete trials. In the chamber the animal is free to respond at any time and at any rate, which made rate of response available as a measure, and rate turned out to be exquisitely sensitive to the conditions of reinforcement. The cumulative recorder then made that rate visible as the slope of a line. **Explain it to a friend.** Explain the difference between Watson's behaviorism and Skinner's radical behaviorism in two sentences a twelve-year-old would follow. ## Frequently asked questions **What is B. F. Skinner best known for?** Operant conditioning — the science of how consequences change behavior. He named it, invented the apparatus (the operant conditioning chamber, or Skinner box) and the cumulative recorder used to study it, discovered shaping and schedules of reinforcement, and founded radical behaviorism, the philosophy behind the science. **What does B. F. stand for?** Burrhus Frederic. Burrhus was his mother's maiden name. Family and colleagues called him Fred. **What is a Skinner box and what was it used for?** A Skinner box, or operant conditioning chamber, is an enclosed apparatus in which an animal can make a simple response — pressing a lever, pecking a lit key — that the equipment records and can follow with a consequence such as food. It was used to study how reinforcement, extinction, schedules, and discriminative stimuli control the rate of behavior, with the results drawn on a cumulative recorder. **Did B. F. Skinner really raise his daughter in a box?** No. He built an enclosed, temperature-controlled crib called the air crib in which his daughter Deborah slept as an infant. It was a bed, not an experiment. The rumor that she was damaged by it, sued him, or died by suicide is false; she has said so publicly and is alive and well. **What is Skinner's theory of operant conditioning?** That behavior is selected by its consequences. Responses followed by reinforcement become more frequent; those followed by punishment or by nothing at all become less frequent. The unit of analysis is the three-term contingency — antecedent, behavior, consequence — and the basic measure is the rate of responding. [Read the complete guide to operant conditioning](https://operantconditioning.com/). **What is the difference between Skinner and Pavlov?** Pavlov studied respondent (classical) conditioning, in which a reflex such as salivation comes to be elicited by a stimulus that precedes it. Skinner studied operant conditioning, in which voluntary behavior is strengthened or weakened by the consequences that follow it. Skinner coined "respondent" and "operant" in 1937 to mark the distinction. [Full comparison](https://operantconditioning.com/operant-vs-classical-conditioning/). **What is radical behaviorism in simple terms?** The view that thoughts and feelings are real but are themselves behavior to be explained, not the ultimate explanation of what we do. The causes of behavior lie in genetics, learning history, and the present environment. It differs from Watson's behaviorism, which excluded private experience from science altogether. **When and how did B. F. Skinner die?** He died of leukemia on August 18, 1990, at his home in Cambridge, Massachusetts, at age 86. Eight days earlier he had received the American Psychological Association's first Citation for Outstanding Lifetime Contribution to Psychology and given his final public talk. ## References 1. Haggbloom, S. J., Warnick, R., Warnick, J. E., Jones, V. K., Yarbrough, G. L., Russell, T. M., Borecky, C. M., McGahhey, R., Powell, J. L., Beavers, J., & Monte, E. (2002). The 100 most eminent psychologists of the 20th century. *Review of General Psychology, 6*(2), 139–152. 2. Skinner, B. F. (1976). *Particulars of My Life*. Knopf. 3. Skinner, B. F. (1979). *The Shaping of a Behaviorist*. Knopf. See also Skinner, B. F. (1956). A case history in scientific method. *American Psychologist, 11*(5), 221–233. 4. Skinner, B. F. (1935). Two types of conditioned reflex and a pseudo type. *Journal of General Psychology, 12*, 66–77. 5. Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279. 6. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 7. Skinner, B. F. (1960). Pigeons in a pelican. *American Psychologist, 15*(1), 28–37. 8. Peterson, G. B. (2004). A day of great illumination: B. F. Skinner's discovery of shaping. *Journal of the Experimental Analysis of Behavior, 82*(3), 317–328. 9. Skinner, B. F. (1945, October). Baby in a box. *Ladies' Home Journal*. 10. Skinner, B. F. (1945). The operational analysis of psychological terms. *Psychological Review, 52*, 270–277. 11. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 12. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 13. Skinner, B. F. (1954). The science of learning and the art of teaching. *Harvard Educational Review, 24*(2), 86–97. 14. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 15. Skinner, B. F. (1957). *Verbal Behavior*. Appleton-Century-Crofts. 16. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58. 17. Skinner, B. F. (1971). *Beyond Freedom and Dignity*. Knopf. 18. Skinner, B. F. (1974). *About Behaviorism*. Knopf. 19. Skinner, B. F. (1990). Can psychology be a science of mind? *American Psychologist, 45*(11), 1206–1210. 20. Buzan, D. S. (2004). I was not a lab rat. *The Guardian*. 21. Skinner, B. F. (1951). How to teach animals. *Scientific American, 185*(6), 26–29. 22. Skinner, B. F. (1958). Teaching machines. *Science, 128*(3330), 969–977. 23. Skinner, B. F. (1981). Selection by consequences. *Science, 213*(4507), 501–504. 24. MacCorquodale, K. (1970). On Chomsky's review of Skinner's *Verbal Behavior*. *Journal of the Experimental Analysis of Behavior, 13*(1), 83–99. 25. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 26. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 27. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. 28. Skinner, B. F. (1966). The phylogeny and ontogeny of behavior. *Science, 153*, 1205–1213. 29. Watson, J. B. (1930). *Behaviorism* (rev. ed.). Norton. (First edition 1924.) 30. Staddon, J. E. R., & Simmelhag, V. L. (1971). The "superstition" experiment: A reexamination of its implications for the principles of adaptive behavior. *Psychological Review, 78*(1), 3–43. 31. Skinner, B. F. (1962). Two "synthetic social relations." *Journal of the Experimental Analysis of Behavior, 5*(4), 531–533. ## Related - [History of operant conditioning](https://operantconditioning.com/history/): From Thorndike's cats to dopamine neurons and reinforcement learning. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): The cumulative records Skinner and Ferster drew — in a live simulator. - [Shaping](https://operantconditioning.com/shaping/): The technique Skinner discovered while teaching a pigeon to bowl. --- # The History of Operant Conditioning: From Thorndike's Cats to Reinforcement Learning > Who discovered operant conditioning? Thorndike found the law of effect in 1898; Skinner named and systematized it in 1937–38. The timeline, Aristotle to AI. - Source: https://operantconditioning.com/history/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *History* Who discovered operant conditioning, where behaviorism came from, what the cognitive revolution did and did not overturn, and how a law about cats in boxes became a profession, a finding about dopamine, and the algorithm behind modern AI. > **Who discovered operant conditioning?** > > The principle was discovered by **Edward L. Thorndike**, whose puzzle-box experiments with cats (1898) produced the **law of effect**: responses followed by a satisfying consequence are strengthened, and those followed by discomfort are weakened. **B. F. Skinner** named it *operant conditioning* in 1937, invented the methods for studying it, and built the science around it beginning with *The Behavior of Organisms* (1938).[4][11] > > So the honest answer is two names: Thorndike found the law; Skinner turned it into a field. Behind both stand a century of earlier thinking about how animals learn from the results of their actions. **In brief** - Thorndike found the law of effect with his 1898 puzzle-box cats; [Skinner](https://operantconditioning.com/bf-skinner/) named operant conditioning in 1937 and built the science around it. - The cognitive revolution bounded the law of effect rather than overturning it: reinforcement still changes behavior, within limits set by evolution. - The tradition continues as applied behavior analysis, in the dopamine prediction-error signal, and in reinforcement learning, a formalization of learning from consequences. ## Who discovered operant conditioning: Thorndike or Skinner? Thorndike was the first to show experimentally, with learning curves rather than anecdotes, that the consequences of a response change its future probability. Skinner was the first to treat that fact as the foundation of a science: he defined the **operant** as a class of behavior controlled by its consequences, distinguished it from Pavlov's reflexes, built the free-operant chamber and cumulative recorder, and discovered the phenomena — [schedules](https://operantconditioning.com/schedules-of-reinforcement/), [shaping](https://operantconditioning.com/shaping/), [extinction](https://operantconditioning.com/extinction/), stimulus control — that make up the subject today. Thorndike's own term, *instrumental learning*, survives as "instrumental conditioning," a synonym still used in the discrete-trial tradition. ## Timeline: the history of operant conditioning - **c. 350 BCE** — Aristotle on habit In the *Nicomachean Ethics* Aristotle argues that character is built by repetition — “we become just by doing just acts” — and that the pleasure or pain attached to an act is what decides whether it is repeated. It is not the law of effect: there is no experiment, no measurement of response frequency, and no functional definition of a reinforcer. But the intuition that consequences shape conduct is about as old as written philosophy, and Thorndike's contribution was to make it testable rather than to think of it.[48] - **1855–1882** — Bain and Romanes The Scottish philosopher Alexander Bain describes how spontaneous movements that happen to produce pleasure are retained and those producing pain dropped — a law of effect before the name. George Romanes collects anecdotes of animal cleverness in *Animal Intelligence* (1882), a method soon discredited.[1][2] - **1894** — Morgan's canon C. Lloyd Morgan rules that no animal behavior should be explained by a higher mental faculty if a lower process will do. Comparative psychology adopts the rule, and the search for simple learning processes begins.[3] - **1898** — Thorndike's puzzle boxes For his Columbia dissertation, Edward Thorndike puts hungry cats in latched boxes with food outside and times their escapes. Escape times fall gradually, with no sudden insight; he proposes that satisfying consequences "stamp in" the connection between situation and response.[4] - **1911–1932** — The law of effect, stated and then truncated *Animal Intelligence* (1911) names the law of effect. Two decades later, on the basis of human learning experiments, Thorndike concludes that "annoyers" do not weaken connections the way "satisfiers" strengthen them, and drops the negative half of the law — the first evidence that punishment and reinforcement are not mirror images.[5][6] - **1913** — Watson's behaviorist manifesto John B. Watson declares that psychology's subject matter is behavior, not consciousness, and its goal prediction and control. Behaviorism becomes a movement — though its early learning theory rests on Pavlov's reflex rather than Thorndike's law.[7] - **1927** — Pavlov in English Ivan Pavlov's *Conditioned Reflexes* is translated, giving American psychology the terms "reinforcement," "extinction," "generalization," and "discrimination" — all of which Skinner will borrow for a different kind of learning.[8] - **1928** — Konorski and Miller's "type II" reflex Two Polish physiologists, Jerzy Konorski and Stefan Miller, passively flex a dog's leg and then feed it; the dog begins flexing on its own. They call this a second type of conditioned reflex, distinct from Pavlov's, and later argue the point with Skinner in print.[9][10] - **1930–1938** — Skinner: the chamber, the operant, the book At Harvard, [B. F. Skinner](https://operantconditioning.com/bf-skinner/) builds the [operant chamber](https://operantconditioning.com/skinner-box/) and cumulative recorder and defines behavior by its function rather than its form (1935); at Minnesota, where he takes his first faculty post in 1936, he coins "operant" and "respondent" (1937) and publishes *The Behavior of Organisms* (1938).[10][11] - **1943–1944** — Hull's system; Estes on punishment Clark Hull's *Principles of Behavior* offers a rival, mathematical behaviorism in which reinforcement works by reducing a biological drive. Skinner's student W. K. Estes shows that punishment suppresses lever-pressing only temporarily, evidence Skinner will cite against punishment for the rest of his life.[12][13] - **1948** — Cognitive maps and superstitious pigeons Edward Tolman argues from latent-learning and maze studies that rats form "cognitive maps," not chains of responses. The same year Skinner shows that pigeons fed on a timer develop idiosyncratic rituals — accidental contingencies shape behavior with no understanding required.[14][15] - **1950** — A textbook and a manifesto Fred Keller and William Schoenfeld's *Principles of Psychology* teaches introductory psychology entirely from reinforcement principles. Skinner's "Are Theories of Learning Necessary?" argues for describing functional relations rather than inventing internal mechanisms.[16][17] - **1957–1958** — Schedules; a journal Ferster and Skinner publish *Schedules of Reinforcement*. The *Journal of the Experimental Analysis of Behavior* is founded in 1958, giving the field its own outlet and its own methods: single organisms, steady states, within-subject replication.[18][19] - **1959–1961** — Premack, Herrnstein, the Brelands, Chomsky David Premack shows that a more probable behavior can reinforce a less probable one. Richard Herrnstein finds that pigeons match response ratios to reinforcement ratios — the [matching law](https://operantconditioning.com/matching-law/). The Brelands report trained animals drifting toward instinct — "instinctive drift." Noam Chomsky's review of *Verbal Behavior* opens the cognitive revolution.[20][21][22][23] - **1968** — Applied behavior analysis is born Teodoro Ayllon and Nathan Azrin publish *The Token Economy*, from their work on a psychiatric ward. The *Journal of Applied Behavior Analysis* launches, and Baer, Wolf, and Risley's paper in its first issue defines what makes an analysis "applied."[24][25] - **1972** — The Rescorla–Wagner model A model of classical conditioning in which learning is driven by *surprise* — the gap between what was predicted and what occurred. Though about Pavlovian learning, it reshapes all of learning theory and, decades later, turns out to describe what dopamine neurons compute.[26] - **1982–1987** — Functional analysis, momentum, and Lovaas Brian Iwata and colleagues show that self-injury can be experimentally traced to its reinforcer, making assessment of *function* the first step in treatment. John Nevin's behavioral momentum research shows that persistence depends on the rate of reinforcement in a context. Ivar Lovaas reports that intensive early intervention brought nine of nineteen autistic children to normal educational functioning.[27][28][29] - **1997–1998** — Dopamine, reinforcement learning, and a credential Schultz, Dayan, and Montague show that midbrain dopamine neurons signal a reward prediction error. Sutton and Barto publish *Reinforcement Learning: An Introduction*. The Behavior Analyst Certification Board is founded, and behavior analysis becomes a credentialed profession.[30][31] - **2015–2025** — Deep reinforcement learning Reinforcement learning combined with deep neural networks masters Atari games from pixels and then Go; the same family of methods is used to fine-tune large language models from human feedback. In 2025 Sutton and Barto are awarded the ACM Turing Award (for 2024) for founding the field.[32][33] > **Read the sources.** The books this timeline begins with are in the public domain and republished in full in the [library](https://operantconditioning.com/library/): Thorndike's [*Animal Intelligence*](https://operantconditioning.com/library/thorndike-animal-intelligence/) (with the 1898 monograph), Morgan's [*Animal Behaviour*](https://operantconditioning.com/library/morgan-animal-behaviour/), James's [*Principles of Psychology*](https://operantconditioning.com/library/james-principles-of-psychology-vol-1/), Yerkes's [*Dancing Mouse*](https://operantconditioning.com/library/yerkes-the-dancing-mouse/), Pfungst's [*Clever Hans*](https://operantconditioning.com/library/pfungst-clever-hans/), Romanes and Darwin. Skinner's own books are still under copyright. ## From the law of effect to the operant: what Skinner changed It is tempting to read Skinner as Thorndike with better equipment. The differences are deeper, and they explain why the field is called the *experimental analysis of behavior* rather than the study of trial-and-error learning. | Aspect | Thorndike (1898–1932) | Skinner (1938 onward) | | --- | --- | --- | | **Basic measure** | Time to escape on each trial | **Rate of response** — continuous, sensitive, and the measure that made schedule effects visible[11] | | **Method** | Discrete trials; experimenter resets the box | **Free operant**; the animal responds whenever it likes and the cumulative recorder draws rate as a slope | | **Unit of behavior** | A specific movement connected to a situation | The **operant**: a class of responses defined by their common consequence, whatever their form (left paw, right paw, nose)[34] | | **Why reinforcement works** | "Satisfiers" stamp in connections (Hull later: drive reduction) | No theory required: a reinforcer is whatever strengthens the behavior it follows; deprivation is an operation you perform and measure[17] | | **Research design** | Learning curves across trials, later group comparisons | **Single organisms** studied to steady state, with effects shown by reversal in the same animal — codified by Sidman in 1960[35] | | **Scope** | Animal learning and education | All behavior, including thinking, language, and culture — via the three-term contingency of [antecedent, behavior, consequence](https://operantconditioning.com/abc-model/) | Add the distinction from Pavlovian conditioning and you have the framework every page on this site uses. [Operant vs. classical conditioning ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## The cognitive revolution: what it overturned and what it didn't Between roughly 1956 and 1970 psychology's center of gravity moved from behavior to mind. George Miller's paper on the limits of short-term memory (1956), Chomsky's review of *Verbal Behavior* (1959), and Ulric Neisser's *Cognitive Psychology* (1967) are the usual markers.[23][36] Within learning research itself, a series of findings showed that consequences were not the whole story: - **Learning without performance.** Tolman's rats explored mazes without reward and then, once food appeared, ran them almost immediately — they had learned the layout all along.[14] - **Learning by observation.** Albert Bandura's children imitated an adult's aggression toward an inflatable doll without ever being reinforced for it.[37] - **Biological constraints.** The Brelands' raccoons "washed" the coins they were supposed to deposit; John Garcia's rats associated taste with nausea over hours but could not associate a light with nausea at all. Organisms come prepared to learn some things and not others.[22][38] - **Contingency, not contiguity.** Robert Rescorla showed that pairing alone does not produce conditioning; what matters is whether one event *predicts* another. The Rescorla–Wagner model made this quantitative.[26] > **What survived** > > None of these findings overturned the law of effect; they bounded it. Reinforcement still changes behavior, schedules still produce their signature patterns, extinction bursts still occur, and shaping and stimulus control still work — in rats, pigeons, dolphins, children, and adults. What changed is that operant conditioning is now understood as one powerful learning process among several, operating within limits set by evolution, and describable in the language of prediction error. Behavior analysts kept publishing in their own journals throughout; the "death of behaviorism" was mostly a change in who wrote the introductory textbook. ## Behavior analysis today The experimental tradition became a profession in the decades after 1968. **Applied behavior analysis** now has a certifying body (the BACB, founded in 1998), a credential (the Board Certified Behavior Analyst), licensure in most U.S. states, insurance coverage for autism services in every state, and a research literature spanning developmental disabilities, education, organizational management, addiction treatment, brain-injury rehabilitation, and animal welfare.[39] Its methods — [reinforcement](https://operantconditioning.com/positive-reinforcement/)-based teaching, functional assessment before treatment, single-subject designs with continuous measurement — descend directly from the laboratory. [How ABA is practiced ›](https://operantconditioning.com/applications/#aba) ### The neurodiversity critique An honest history has to note that ABA is also contested. Many autistic adults and advocates object to its early history (Lovaas's original 1960s program used aversives, including electric shock, a practice long since abandoned), to goals that aimed at making autistic children "indistinguishable" from peers, to suppressing harmless self-stimulatory behavior, and to intensive programs delivered without the child's assent.[40] Practitioners have responded that modern ABA is reinforcement-based, individualized, and increasingly assent-driven, and the field's own journals now publish on these concerns.[40] The debate is about goals, consent, and history, not about whether reinforcement changes behavior — which everyone in it agrees it does. ### Beyond the clinic Outside ABA, the operant tradition underlies reinforcement-based [animal training](https://operantconditioning.com/dog-training/), contingency management in addiction medicine, school-wide positive behavior supports, organizational behavior management, and the analysis of choice that became behavioral economics. Herrnstein's matching law and George Ainslie's hyperbolic discounting — the finding that we overvalue immediate consequences along a predictable curve — gave economists a behavioral account of impulsiveness years before it was fashionable.[21][41] ## The neuroscience of reinforcement The last chapter of the history is being written by neuroscience. In 1953 James Olds and Peter Milner found that a rat would press a lever for hours to deliver a pulse of current to its own brain — the first demonstration that reinforcement had an identifiable neural substrate.[42] In 1997 Wolfram Schultz, Peter Dayan, and Read Montague showed that midbrain **dopamine** neurons report a *prediction error* — firing to unexpected reward, falling silent to predicted reward, dipping when a predicted reward fails — precisely the quantity in the Rescorla–Wagner model and in temporal-difference learning.[30][43] Kent Berridge and Terry Robinson then separated "wanting" from "liking": animals without dopamine still enjoy sugar but no longer work for it.[44] The operant laboratory's strangest findings — the grip of the variable-ratio schedule, the fading of a fully predictable reinforcer — suddenly had a mechanism. [What happens in the brain, in depth ›](https://operantconditioning.com/neuroscience/) ## Operant conditioning in artificial intelligence **Reinforcement learning** is the branch of machine learning in which an agent learns by acting in an environment and receiving a numerical reward. Its founders were explicit about the lineage: Richard Sutton and Andrew Barto's textbook opens with Thorndike's law of effect, and their early work with Charles Anderson on "neuronlike adaptive elements" was an attempt to build a learning rule that behaved like an animal under reinforcement.[31][45] Sutton's 1988 **temporal-difference learning** algorithm updates predictions from the difference between successive predictions — the computational cousin of Rescorla–Wagner — and it was TD error that Montague, Dayan, and Sejnowski proposed in 1996 as the thing dopamine neurons encode, a year before Schultz's data confirmed it.[43][46] The vocabulary crossed over intact. AI researchers speak of *reward*, *policy*, *exploration and exploitation*, and *reward shaping* — the practice of adding intermediate rewards to guide an agent toward a hard-to-reach goal, named for Skinner's procedure.[47] Deep reinforcement learning, which pairs these algorithms with neural networks, learned Atari games from raw pixels in 2015 and beat the world's best Go players in 2016; *reinforcement learning from human feedback* is now a standard stage in training large language models, with human preferences standing in for the food hopper.[32][33] The differences matter too. An RL agent's reward is written by an engineer, not discovered functionally; the agent can explore millions of episodes an animal never could; and the algorithms include explicit value estimates that Skinner would have called explanatory fictions. But the central claim — that a system with no built-in knowledge of a task can acquire complex, purposeful behavior purely from the consequences of its actions — is Thorndike's and Skinner's, vindicated at a scale neither imagined. ## Key takeaways - Two names answer the question of discovery. Thorndike showed experimentally, with learning curves, that the consequences of a response change its future probability; Skinner treated that fact as the foundation of a science, defining the operant and inventing the free-operant methods that revealed schedules, shaping, extinction, and stimulus control. - Skinner was not Thorndike with better equipment. He replaced time to escape with rate of response, discrete trials with the free operant, a specific movement with an operant defined by its consequence, theories of why reinforcement works with a functional definition, and group comparisons with single organisms studied to steady state. - Thorndike himself dropped the negative half of the law of effect in the 1930s, the first evidence that punishment and reinforcement are not mirror images. Estes's 1944 finding that punishment suppresses responding only temporarily became Skinner's lifelong argument against it. - The cognitive revolution bounded the law of effect rather than overturning it. Latent learning, observational learning, biological constraints, and contingency over contiguity showed that reinforcement is one powerful process among several; its findings still replicate, and behavior analysts kept publishing throughout. - The tradition became a credentialed profession, applied behavior analysis, whose current debate is about goals, consent, and history rather than whether reinforcement works. It found a mechanism in the dopamine prediction-error signal and was formalized as reinforcement learning, whose founders cite the law of effect directly. ### Check yourself **Thorndike or Skinner: who discovered operant conditioning? Make the case for each.** Thorndike discovered the principle: his 1898 puzzle-box experiments showed, with learning curves rather than anecdotes, that consequences change the future probability of a response, and he named it the law of effect. Skinner named the process in 1937, defined the operant as a class of behavior controlled by its consequences, built the chamber and cumulative recorder, and discovered schedules, shaping, extinction, and stimulus control. Thorndike found the law; Skinner turned it into a field. **Behaviorism began with Watson's 1913 manifesto, so it seems safe to say early behaviorist learning theory was built on Thorndike's law of effect. Is that right?** No. Watson's early learning theory rested on Pavlov's reflex, not on Thorndike's law. It was Skinner, in the 1930s, who borrowed Pavlov's vocabulary of reinforcement, extinction, generalization, and discrimination and applied it to a different kind of learning, behavior controlled by its consequences rather than elicited by a prior stimulus. **A textbook states that Chomsky's 1959 review and the cognitive revolution killed behaviorism. What does the historical record show?** Mainstream psychology's attention did shift to mental processes, and findings such as latent learning, observational learning, and biological constraints showed that consequences are not the whole story. But those findings bounded the law of effect rather than overturning it: reinforcement, schedule effects, extinction bursts, shaping, and stimulus control still replicate, and behavior analysts kept publishing in their own journals, founding an applied journal in 1968 and a credentialing body in 1998. The "death of behaviorism" was mostly a change in who wrote the introductory textbook. **Sutton and Barto open their reinforcement learning textbook with Thorndike. What does a reinforcement learning agent share with a cat in a puzzle box, and where does the comparison break down?** Both acquire complex, purposeful behavior with no built-in knowledge of the task, purely from the consequences of their actions, and the vocabulary of reward, exploration, and reward shaping crossed over intact. The differences: an agent's reward is written by an engineer rather than discovered functionally, the agent can explore millions of episodes an animal never could, and its algorithms carry explicit value estimates that Skinner would have called explanatory fictions. **Explain it to a friend.** Explain why the cognitive revolution did not overturn the law of effect, without using the words "behaviorism" or "cognitive." ## Frequently asked questions **Who discovered operant conditioning?** Edward Thorndike discovered the underlying principle, the law of effect, in his 1898 puzzle-box experiments with cats. B. F. Skinner named the process "operant conditioning" in 1937, invented the experimental methods to study it, and founded the science with *The Behavior of Organisms* (1938). Both are correctly credited. **When was operant conditioning discovered?** The law of effect dates to Thorndike's 1898 monograph and was named in his 1911 book. The term "operant conditioning" dates to 1937, and the systematic science to 1938. **Who is the father of operant conditioning?** B. F. Skinner is usually called the father of operant conditioning because he named it, built the apparatus and methods, and developed the concepts — reinforcement schedules, shaping, stimulus control, the operant itself. Thorndike is the grandfather: he found the law of effect on which everything rests. **What is the difference between Thorndike and Skinner?** Thorndike studied discrete trials, measured time to respond, and explained learning as connections "stamped in" by satisfaction. Skinner let animals respond freely, measured rate of response, defined reinforcement purely by its effect, and rejected explanations in terms of satisfaction, drive, or connections. **What are the origins of behaviorism?** Behaviorism as a movement began with John B. Watson's 1913 paper "Psychology as the Behaviorist Views It," which argued that psychology should study observable behavior rather than consciousness. Its roots include Pavlov's conditioned reflexes, Thorndike's animal learning, and Morgan's canon of parsimony. Skinner's radical behaviorism, from 1945, is a later and different philosophy. **Did the cognitive revolution disprove operant conditioning?** No. It showed that reinforcement is not the only way organisms learn — observation, latent learning, and biological preparedness matter too — and it shifted mainstream psychology's attention to mental processes. The findings of operant conditioning still replicate, and the field continued as behavior analysis, with its own journals, profession, and applications. **How is operant conditioning related to artificial intelligence?** Reinforcement learning, a major branch of machine learning, is a mathematical formalization of learning from consequences. Its founders, Sutton and Barto, cite Thorndike's law of effect directly; temporal-difference learning parallels the Rescorla–Wagner model; and "reward shaping" takes its name from Skinner's procedure. The same prediction-error signal appears in the firing of dopamine neurons. **Is behaviorism still used today?** Yes, as applied behavior analysis (a licensed profession), reinforcement-based animal training, contingency management for addiction, positive behavior supports in schools, organizational behavior management, and the self-management methods behind habit-building apps. Its concepts also run through behavioral economics, neuroscience, and AI. ## References 1. Bain, A. (1855). *The Senses and the Intellect*. John W. Parker. See also Bain, A. (1859). *The Emotions and the Will*. John W. Parker. 2. Romanes, G. J. (1882). *Animal Intelligence*. Kegan Paul, Trench. 3. Morgan, C. L. (1894). *An Introduction to Comparative Psychology*. Walter Scott. 4. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 5. Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. 6. Thorndike, E. L. (1932). *The Fundamentals of Learning*. Teachers College, Columbia University. 7. Watson, J. B. (1913). Psychology as the behaviorist views it. *Psychological Review, 20*(2), 158–177. 8. Pavlov, I. P. (1927). *Conditioned Reflexes* (G. V. Anrep, Trans.). Oxford University Press. 9. Miller, S., & Konorski, J. (1928). Sur une forme particulière des réflexes conditionnels. *Comptes Rendus des Séances de la Société de Biologie, 99*, 1155–1157. English translation: Miller, S., & Konorski, J. (1969). On a particular form of conditioned reflex. *Journal of the Experimental Analysis of Behavior, 12*(1), 187–189. 10. Konorski, J., & Miller, S. (1937). On two types of conditioned reflex. *Journal of General Psychology, 16*, 264–272; Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279. 11. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 12. Hull, C. L. (1943). *Principles of Behavior: An Introduction to Behavior Theory*. Appleton-Century. 13. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 14. Tolman, E. C. (1948). Cognitive maps in rats and men. *Psychological Review, 55*(4), 189–208. See also Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. *University of California Publications in Psychology, 4*, 257–275. 15. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 16. Keller, F. S., & Schoenfeld, W. N. (1950). *Principles of Psychology: A Systematic Text in the Science of Behavior*. Appleton-Century-Crofts. 17. Skinner, B. F. (1950). Are theories of learning necessary? *Psychological Review, 57*(4), 193–216. 18. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 19. Skinner, B. F. (1956). A case history in scientific method. *American Psychologist, 11*(5), 221–233. 20. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 21. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 22. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 23. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58. 24. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 25. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97. 26. Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), *Classical Conditioning II: Current Research and Theory* (pp. 64–99). Appleton-Century-Crofts. See also Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. *Journal of Comparative and Physiological Psychology, 66*(1), 1–5. 27. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 28. Nevin, J. A., Mandell, C., & Atak, J. R. (1983). The analysis of behavioral momentum. *Journal of the Experimental Analysis of Behavior, 39*(1), 49–59. See also Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 29. Lovaas, O. I. (1987). Behavioral treatment and normal educational and intellectual functioning in young autistic children. *Journal of Consulting and Clinical Psychology, 55*(1), 3–9. 30. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 31. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. (First edition 1998.) 32. Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2015). Human-level control through deep reinforcement learning. *Nature, 518*(7540), 529–533. 33. Silver, D., Huang, A., Maddison, C. J., et al. (2016). Mastering the game of Go with deep neural networks and tree search. *Nature, 529*(7587), 484–489. See also Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. *Advances in Neural Information Processing Systems, 30*. 34. Skinner, B. F. (1935). The generic nature of the concepts of stimulus and response. *Journal of General Psychology, 12*, 40–65. 35. Sidman, M. (1960). *Tactics of Scientific Research: Evaluating Experimental Data in Psychology*. Basic Books. 36. Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. *Psychological Review, 63*(2), 81–97; Neisser, U. (1967). *Cognitive Psychology*. Appleton-Century-Crofts. 37. Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. *Journal of Abnormal and Social Psychology, 63*(3), 575–582. 38. Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. *Psychonomic Science, 4*(1), 123–124. See also Seligman, M. E. P. (1970). On the generality of the laws of learning. *Psychological Review, 77*(5), 406–418. 39. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 40. Leaf, J. B., Cihon, J. H., Leaf, R., McEachin, J., Liu, N., Russell, N., Unumb, L., Shapiro, S., & Khosrowshahi, D. (2022). Concerns about ABA-based intervention: An evaluation and recommendations. *Journal of Autism and Developmental Disorders, 52*(6), 2838–2853. On the early use of aversives, see Lovaas, O. I., Schaeffer, B., & Simmons, J. Q. (1965). Building social behavior in autistic children by use of electric shock. *Journal of Experimental Research in Personality, 1*, 99–109. 41. Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. *Psychological Bulletin, 82*(4), 463–496. 42. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427. 43. Montague, P. R., Dayan, P., & Sejnowski, T. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. *Journal of Neuroscience, 16*(5), 1936–1947. 44. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? *Brain Research Reviews, 28*(3), 309–369. 45. Barto, A. G., Sutton, R. S., & Anderson, C. W. (1983). Neuronlike adaptive elements that can solve difficult learning control problems. *IEEE Transactions on Systems, Man, and Cybernetics, 13*(5), 834–846. 46. Sutton, R. S. (1988). Learning to predict by the methods of temporal differences. *Machine Learning, 3*(1), 9–44. 47. Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. In *Proceedings of the Sixteenth International Conference on Machine Learning* (pp. 278–287). Morgan Kaufmann. 48. Aristotle. (c. 350 BCE). *Nicomachean Ethics*, Book II (W. D. Ross, Trans.). See especially 1103a–1103b on habituation and 1104b on pleasure and pain as signs of character. ## Related - [B. F. Skinner](https://operantconditioning.com/bf-skinner/): Biography, the Skinner box, radical behaviorism, and the myths. - [Operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/): Pavlov's reflexes and Skinner's operants, side by side. - [Applications](https://operantconditioning.com/applications/): Where the science is used today: ABA, education, parenting, animals, work, health, technology. --- # Applications of Operant Conditioning: Therapy, Education, Parenting, Work, Health, Technology, and More > Operant conditioning in the classroom, parenting, the workplace, therapy, animal training, technology, and AI — and how strong the evidence is for each. - Source: https://operantconditioning.com/applications/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications* The same three-term contingency runs a therapy session, a classroom, a family dinner, a factory floor, a dolphin show, and a slot machine. Here is where operant conditioning is used, how, and how strong the evidence really is in each field. > **Definition** > > The **applications of operant conditioning** are the deliberate arrangements of antecedents and consequences to change behavior that matters — in clinics, schools, homes, workplaces, zoos, software, and in the design of one's own life. > > The discipline built to do this systematically is **applied behavior analysis**, founded as a field in 1968, but the principles are used far beyond it, in some places rigorously and in others loosely.[1] The table below rates each field's evidence as candidly as the literature allows. **In brief** - Applications of operant conditioning are deliberate arrangements of antecedents and consequences to change behavior in clinics, schools, homes, workplaces, zoos, and software. - Applied behavior analysis, founded in 1968, is the discipline built to do this systematically; its core tool is the functional analysis, and treatment follows function. - The evidence varies by field: strong for classroom programs, parent training, contingency management, and behavioral activation; weaker for early intensive autism intervention and gamification. | Field | Core techniques | Evidence strength | Learn more | | --- | --- | --- | --- | | **Applied behavior analysis** | Functional analysis, reinforcement, differential reinforcement, shaping, prompting and fading | Strong for targeted behavior change in single-case research; low-certainty and contested for early intensive intervention in autism | ABA › | | **Education** | Token economies, Good Behavior Game, PBIS, precision teaching, direct instruction, praise ratios | Strong: randomized trials with follow-ups into adulthood; large meta-analyses for direct instruction | Classroom › | | **Parenting** | Specific praise, planned ignoring, time-out, response cost, consistency | Strong for behavioral parent-training programs; strong evidence *against* corporal punishment | Parenting › | | **Self-management** | Antecedent control, tiny behaviors, self-monitoring, immediate self-administered consequences | Moderate: self-monitoring and implementation intentions well supported; "self-reinforcement" mechanism debated | Self-management › · [Habits guide ›](https://operantconditioning.com/habits/) | | **Health and clinical** | Contingency management, behavioral activation, exposure with response prevention, habit reversal | Strong: contingency management and behavioral activation are among the best-supported psychosocial treatments in their areas | Health › | | **Animal training** | Marker signals, shaping, reinforcement-based husbandry training | Strong for reinforcement-based methods; aversive methods linked to welfare costs | Animals › · [Dog training ›](https://operantconditioning.com/dog-training/) | | **Workplace** | Pinpointing, performance feedback, behavior-based safety, incentives | Moderate to strong within-organization studies; feedback reliably works; gamification mixed | Workplace › | | **Technology and product design** | Variable-ratio feeds, notifications, streaks, loot boxes | Strong that the mechanisms drive engagement; contested whether population-level harms are large | Technology › | | **Behavioral economics and AI** | Matching law, delay discounting, reinforcement learning | Strong: quantitative laws replicated across species; reinforcement learning is foundational to modern AI | Economics & AI › | > **How to read the evidence column** > > "Strong" means randomized trials or many replicated single-case experiments with durable effects. "Moderate" means consistent results from weaker designs, or strong results for some components and not others. Where a field's reputation outruns its data, the table says so. ## Applied behavior analysis (ABA) **Applied behavior analysis** is the science in which procedures derived from the principles of behavior are applied to improve socially significant behavior, and experimentation is used to show that the procedures were responsible for the change.[2] Baer, Wolf, and Risley's founding paper set out seven dimensions that still define the field: work should be applied, behavioral, analytic, technological, conceptually systematic, effective, and general.[1] Its core tool is the **functional analysis**: experimentally testing what consequence maintains a problem behavior before trying to change it, a method established by Iwata and colleagues' 1982 study of self-injury.[3] Treatment then follows function. If a child hits to get attention, the analyst teaches a replacement — tapping a shoulder and saying "excuse me" — and reinforces it while the hitting no longer produces attention. That is **differential reinforcement of alternative behavior** (DRA), the most-used procedure in the field; its cousins reinforce the *absence* of the behavior for a set interval (DRO) or a behavior physically *incompatible* with it (DRI).[4] [How ABC data and functional analysis work ›](https://operantconditioning.com/abc-model/) The evidence deserves a careful statement. For targeted behavior change — reducing self-injury, teaching communication, building daily-living skills — hundreds of single-case experiments show large, reliable effects. The claim that early intensive behavioral intervention (EIBI) transforms outcomes in autism rests on a smaller base. Lovaas reported in 1987 that 47% of children receiving roughly 40 hours a week of one-to-one therapy reached normal-range IQ and unsupported first-grade placement, versus 2% of a comparison group.[5] Later reviews were more modest: a Cochrane review found only low-certainty evidence, from a handful of mostly non-randomized studies, that EIBI improves adaptive behavior and IQ, and a 2020 meta-analysis found support for behavioral approaches weakened substantially when limited to randomized trials with outcomes not reported by caregivers.[6][7] ABA also faces a principled critique from autistic adults and the neurodiversity movement: that its historical goal of making autistic children "indistinguishable" from peers, its early use of aversives (Lovaas's 1960s studies used contingent electric shock), and its emphasis on compliance were harmful whatever the outcome measures said.[8] Contemporary practice has moved toward assent, client-chosen goals, and reinforcement-only procedures. Practitioners are credentialed by the Behavior Analyst Certification Board (the BCBA credential dates from 1998), and the profession is licensed in most U.S. states. ## Operant conditioning in the classroom Schools were the first institutions to use operant technology at scale, and they still produce some of its best evidence. The token economy — points earned for target behaviors and exchanged later for backup reinforcers — moved from a psychiatric ward into classrooms almost as soon as Ayllon and Azrin described it, and it works because tokens are generalized conditioned reinforcers that can be delivered the instant a behavior occurs.[9] The Good Behavior Game, a team-based contingency introduced in a fourth-grade classroom in 1969, is among the most thoroughly studied classroom procedures of all: in randomized trials that followed first-graders into their twenties, the boys who had played it showed lower rates of substance-use and antisocial-personality disorders than controls.[10][11] School-wide positive behavioral supports scale the same logic across a building, and Skinner's own contribution — programmed instruction, with its small steps, active responding, and immediate feedback — survives in adaptive software and Direct Instruction.[12][13][14] Praise, the cheapest reinforcer in the room, works only when it is contingent, specific, and credible.[56] [The classroom in depth: praise, token economies, the Good Behavior Game, PBIS, and what the evidence says ›](https://operantconditioning.com/classroom/) ## Operant conditioning in parenting A parent's most powerful reinforcer is attention, and the most common parenting error is spending it on misbehavior. "Catching them being good" reverses the allocation: notice the behavior you want and name it specifically and immediately. Gerald Patterson's observations of families showed how the opposite pattern escalates — the child's tantrum is negatively reinforced when the parent gives in, and the parent's giving in is negatively reinforced when the tantrum stops, a *coercive family process* that trains both sides.[15] Parent management training, the best-supported treatment for childhood conduct problems, teaches the reverse contingencies: specific praise, planned ignoring, effective commands, brief time-out from reinforcement, and point systems.[16][17] On the punishment side, the largest meta-analysis of spanking — 75 studies and more than 160,000 children — found it associated with worse outcomes and no benefit.[18] [Parenting in depth: tantrums, time-out, sticker charts, bedtime, and the evidence ›](https://operantconditioning.com/parenting/) ## Self-management The three-term contingency turned inward. Skinner argued that a person controls their own behavior with the same tools used on anyone else's — changing the stimulus, restraining themselves physically, arranging deprivation and satiation, and delivering their own consequences.[19] The best-supported components are antecedent control (implementation intentions have a medium-to-large meta-analytic effect on goal attainment) and self-monitoring (recording your own behavior reliably changes it).[20][21] Whether a self-delivered reward is "really" reinforcement is debated; what is not debated is that a cue you cannot miss, a behavior small enough to emit, and a consequence that arrives immediately outperform resolve. [The complete habit-building protocol ›](https://operantconditioning.com/habits/) ## Health and clinical applications **Contingency management** for substance use is the clearest case of operant principles applied to a hard medical problem. In 1991 Stephen Higgins and colleagues offered cocaine-dependent outpatients vouchers exchangeable for retail goods, contingent on cocaine-negative urine samples, with the value escalating for consecutive clean samples and resetting after a positive one; retention and abstinence far exceeded standard counseling.[22] Two 2006 meta-analyses confirmed moderate, reliable effects across drugs, strongest for stimulants and opioids, and a 2008 meta-analytic review found contingency management produced the largest effects of any psychosocial treatment examined.[23][24][25] Its limits are what the theory predicts: effects fade after incentives stop unless natural reinforcers have taken over, and uptake has been slowed more by regulation and squeamishness about "paying people to stay clean" than by the data. Chronic pain was one of the first medical problems given an operant analysis. Wilbert Fordyce observed that "pain behaviors" — guarding, limping, grimacing, resting, taking medication, talking about pain — are behaviors, and that they are reinforced: by attention and sympathy, by relief from unwanted duties, and by medication given in response to complaints. His inpatient program at the University of Washington, reported in 1973, reversed the contingencies. Staff attended to activity rather than complaints, exercise quotas were raised gradually with rest as the reinforcer for meeting them rather than for pain, and medication was given on a time schedule instead of on demand; patients' activity rose and their medication use and reported pain fell.[57][58] Operant-behavioral treatment remains a component of modern pain programs; in a randomized trial for fibromyalgia, operant-behavioral and cognitive-behavioral treatments each outperformed an attention-placebo condition, with the operant program's largest gains in physical functioning.[59] Nobody in this literature claims the pain is imaginary; the claim is that what a person does about pain is learned, and can be re-learned. **Behavioral activation** treats depression as a collapse in response-contingent positive reinforcement, an analysis Charles Ferster offered in 1973.[26] The treatment schedules activities, grades difficult tasks, and dismantles avoidance so the person contacts reinforcement again. A 1996 dismantling study found the activation component alone matched the full cognitive therapy package; a 2006 randomized trial found it comparable to antidepressant medication and superior to cognitive therapy among more severely depressed patients; and a 2016 trial found it non-inferior to CBT when delivered by junior mental-health workers at lower cost.[27][28][29] **Exposure therapy** has an operant engine inside its classical shell. In Mowrer's two-factor account, fear is classically conditioned, but *avoidance* is negatively reinforced by the relief it brings, and avoidance is what prevents the fear from extinguishing.[30] Exposure with response prevention blocks the escape so extinction can proceed. **Habit reversal training** — awareness training, a competing response, and social support — was introduced by Azrin and Nunn in 1973 for tics and nervous habits and is now the core of the leading behavioral treatment for Tourette syndrome.[31][32] **Biofeedback** is operant conditioning of physiological responses; Neal Miller's 1969 claims of visceral learning in rats proved hard to replicate, but clinical biofeedback has a real evidence base for specific problems such as incontinence and tension headache.[33] And financial incentives for **medication adherence** consistently improve adherence while they are in place.[34] ## Animal training Modern animal training is operant conditioning with the lecture removed. A **marker signal** — a clicker or a word — is a conditioned reinforcer that bridges the gap between the behavior and the treat; [shaping](https://operantconditioning.com/shaping/) builds complex behavior from approximations; and reinforcement-based methods have displaced the dominance-and-correction approaches of the mid-twentieth century. [Dog training with operant conditioning ›](https://operantconditioning.com/dog-training/) The commercial lineage runs through Keller and Marian Breland, Skinner's former students, who founded Animal Behavior Enterprises in 1943 and trained thousands of animals for advertising, exhibitions, and their "IQ Zoo." Their 1961 paper "The misbehavior of organisms," reporting raccoons that "washed" coins instead of depositing them, remains the classic demonstration that reinforcement works with, not against, an animal's evolved tendencies.[35] Marine-mammal training grew from the same roots; in a famous 1969 study, dolphins reinforced only for behaviors they had never shown before began producing novel actions.[36] The quietest revolution is in zoos and laboratories. **Husbandry training** teaches animals to present a limb for a blood draw, open their mouths for dental checks, and enter a crate voluntarily, replacing restraint and anesthesia with cooperation and reducing measurable stress.[37] Reviews of aversive dog-training methods, meanwhile, link them to stress and problem behaviors without any advantage in effectiveness.[38] ## Operant conditioning in the workplace **Organizational behavior management** (OBM) applies the same analysis to employees: pinpoint a behavior, measure it, give feedback, and reinforce it. The field has had its own journal since 1977. The signature demonstration is **behavior-based safety**: in a 1978 study at a food-manufacturing plant, researchers defined specific safe behaviors, posted graphed feedback on how often they occurred, and watched safe performance rise sharply — then fall when feedback was withdrawn and rise again when it returned.[39] A review of performance-feedback studies from 1985 to 1998 found feedback effective in most applications, most consistently when it was graphic, frequent, and combined with goals and reinforcement; a meta-analysis of behavior-modification programs across organizations reported an average performance improvement of about 17%.[40][41] Aubrey Daniels, the field's best-known popularizer, summarizes the timing problem with a three-letter test: consequences that are positive, immediate, and certain control behavior; consequences that are negative, delayed, or uncertain barely register.[42] The annual performance review fails on every count. It arrives months after the behavior it addresses, it is delivered once, and for most people it is aversive, so it mainly evokes escape — the polished self-assessment, the defensive meeting — rather than changing what anyone does on Tuesday. Daily feedback from a supervisor who knows what to look for costs less and does more. Two cautions. **Gamification** — points, badges, leaderboards — reliably produces short-term engagement and unreliably produces lasting change; a review of the empirical studies found positive but context-dependent effects that often faded with novelty.[43] And reinforcing a proxy reinforces the proxy. When sales targets and incentives at Wells Fargo rewarded accounts opened, employees opened millions of accounts customers had not asked for. The contingency worked as designed; the design was the problem. ## Technology and product design Consumer software is the largest deployment of operant conditioning in history, and much of it is aimed at the user rather than for them. A social feed is a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/): pull to refresh, and the reinforcer — a like, a message, something novel — arrives after an unpredictable number of pulls, the schedule that produces the highest response rates and the most persistence when reinforcement stops. Notifications are discriminative stimuli for opening the app. Streaks convert a habit into an avoidance contingency — the reinforcer becomes not losing the count — and borrow loss aversion to make a missed day feel like a fine. Gambling is the original engineered variable-ratio schedule, and slot machines add a refinement the laboratory did not anticipate: the **near miss**. Two cherries and a lemon is a loss, but it looks like almost winning, and Reid argued in 1986 that near misses encourage continued play as if they were partial reinforcement.[60] Brain-imaging work later found that near misses recruit some of the same reward circuitry as wins and increase the urge to keep playing, more so in people who gamble more heavily.[61] Game designers borrowed the whole toolkit early; a 2001 article by the psychologist John Hopson laid out how to apply schedules of reinforcement to game design so that players keep playing; the industry's later name for the resulting anticipation–activity–reward cycle is the **compulsion loop**, and most games since have one.[62] **Loot boxes** in video games are the starkest case. Researchers who examined popular games found that a large share of loot-box systems met the psychological criteria for gambling, and spending on them correlates with problem-gambling severity; Belgium's gaming regulator ruled paid loot boxes illegal gambling in 2018.[44][45] BJ Fogg's **behavior model** — behavior occurs when motivation, ability, and a prompt converge — is a design-oriented cousin of the four-term contingency: the prompt is the antecedent, ability stands in for response effort, and motivation for the motivating operation.[46] When those tools trick users into choices they would not otherwise make, the designs are called **dark patterns**, and they are increasingly the target of regulation. Honesty requires a caveat: whether the resulting engagement causes large population-level harm is contested. Large-sample analyses of adolescent well-being find the association with digital technology use to be tiny.[47] That the mechanisms work is not in doubt; how much damage they do is. The same science runs the other way: environment design against the feeds — the phone in the kitchen, the app logged out — and tools that make the user set the antecedent and consequence for behavior they have chosen. [Using the loop for your own goals ›](https://operantconditioning.com/habits/) ## Behavioral economics and artificial intelligence Operant research turned quantitative in 1961, when Richard Herrnstein showed that pigeons offered two keys, each paying on its own variable-interval schedule, distributed their pecks in proportion to the reinforcement each key delivered — the **[matching law](https://operantconditioning.com/matching-law/)**.[48] Matching turned choice into something that could be modeled with the tools of economics, and by the 1990s laboratory animals had been shown to obey demand curves, substitute between goods, and respond to price much as human consumers do.[49] The most consequential offshoot is **delay discounting**. Using an adjusting procedure, James Mazur showed that the value of a delayed reinforcer falls along a hyperbola rather than the exponential curve standard economics assumed.[50] George Ainslie had already worked out the implication: hyperbolic curves cross, so a person who prefers the larger, later reward from a distance will flip to the smaller, sooner one as it approaches — a mathematical account of impulsiveness and of why commitment devices are needed.[51] Steeper discounting has since been documented in smokers, in people with substance-use disorders, and in problem gamblers, and it is now studied as a process cutting across many conditions.[52] Public policy has borrowed from both halves of the contingency. **Nudges** — Thaler and Sunstein's term for changes in "choice architecture" that steer behavior without changing incentives, such as making retirement saving the default — are antecedent interventions: they alter the stimulus conditions under which a choice is made, not its consequences, and they are cheap precisely because no reinforcer has to be delivered.[63] Incentive programs — conditional cash transfers, deposit contracts, sin taxes — work on the consequence side. The behavioral analysis predicts what the evaluations find: nudges produce modest, durable effects when the behavior is a one-off choice, and incentives produce larger effects that fade when the incentive stops unless a natural reinforcer takes over. The law of effect also became an algorithm. **Reinforcement learning**, the branch of machine learning in which an agent learns by acting and receiving reward signals, descends directly from Thorndike and Skinner by way of temporal-difference learning, and its standard textbook says so in its opening pages.[53] In 1997 Schultz, Dayan, and Montague showed that midbrain dopamine neurons fire in the pattern of a temporal-difference prediction error — the brain running the same computation.[54] [The neuroscience of operant conditioning ›](https://operantconditioning.com/neuroscience/) Reinforcement learning drove the systems that mastered Go, and reinforcement learning from human feedback is one of the methods used to train large language models to behave as people prefer.[55] [The history from Thorndike to reinforcement learning ›](https://operantconditioning.com/history/) ## Military training The most frequently cited military application is also the most disputed. After the Second World War, the U.S. Army historian S. L. A. Marshall reported, on the basis of group interviews, that only about 15–20% of riflemen had fired their weapons at the enemy in combat.[64] Training changed in response: bullseye targets were replaced by man-shaped silhouettes that pop up briefly and fall when hit — immediate feedback on a realistic discriminative stimulus — and the conditioned response was practiced hundreds of times. Dave Grossman argued in *On Killing* that this was operant conditioning in all but name, and credited it with raising reported firing rates to about 55% in Korea and over 90% in Vietnam.[65] The account should be read with care: Marshall's figures have been challenged by historians who found no record of the systematic interviews he described, and the later rates are not measured the same way.[66] What is not in dispute is the design of the training itself, which is a textbook contingency, or Skinner's own wartime contribution: Project Pigeon, in which he shaped pigeons to peck at an image of a target so that their pecks could steer a glide bomb — a working system that the military declined to deploy.[67] ## Coercion: the dark side of the contingency The same principles that build a token economy can hold a person in a harmful relationship or a toxic workplace, and behavior analysts have said so plainly. Murray Sidman's *Coercion and Its Fallout* catalogued what aversive control does to the people subjected to it: it produces escape and avoidance, countercontrol, aggression, and a generalized suppression of behavior that reaches far beyond the punished response.[68] **Traumatic bonding** is the clearest case. Donald Dutton and Susan Painter proposed that two features of abusive relationships — a power imbalance and the *intermittency* of abuse, with cycles of cruelty and reconciliation — produce unusually strong attachment to the abuser, on the same logic by which [intermittent reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) produces the most persistent behavior in the laboratory. Their follow-up study of women who had left abusive partners found that intermittency of abuse and dominance predicted continued attachment months later.[69][70] Popular accounts of psychological manipulation list the same operations — intermittent reward, negative reinforcement by the removal of hostility, punishment, and one-trial traumatic learning — and they are recognizable to anyone who has read this far. Organizations run the same contingencies at scale. Blake Ashforth's analysis of "petty tyranny" in management described supervisors who rule through arbitrary punishment, belittling, and the withholding of consideration, and traced the consequences the laboratory predicts: helplessness, low initiative, and the disappearance of any behavior not strictly required.[71] A culture of fear is a culture of avoidance, and avoidance, as [its own page](https://operantconditioning.com/avoidance-learning/) explains, is the behavior least sensitive to whether the threat is real. The ethical guidance of behavior analysis — reinforcement first, the least restrictive procedure, consent, the client's own goals — exists because the field knows exactly how effective the alternative is. ## Key takeaways - The same three-term contingency runs every application. Applied behavior analysis is the discipline built to apply it systematically, and its core tool is the functional analysis: find what consequence maintains a behavior, then teach and reinforce an alternative while the problem behavior no longer pays. - Evidence strength varies by field, and a field's reputation can outrun its data. The Good Behavior Game, behavioral parent training, contingency management, and behavioral activation have randomized trials with durable effects; early intensive behavioral intervention in autism rests on low-certainty evidence and faces a principled critique from autistic adults. - Consequences control behavior when they are positive, immediate, and certain; delayed, infrequent, and aversive consequences such as the annual review mainly evoke escape. Reinforcing a proxy reinforces the proxy. - Consumer software is the largest deployment of operant conditioning in history: feeds run variable-ratio schedules, notifications are discriminative stimuli, streaks are avoidance contingencies, and loot boxes reproduce the structure of gambling. That the mechanisms work is not in doubt; how much population-level harm they do is. - The principles cut both ways. Intermittent abuse produces traumatic bonding on the same logic that makes intermittent reinforcement persistent, and coercive management produces avoidance and helplessness. The field's ethics, reinforcement first, the least restrictive procedure, and consent, exist because it knows how effective the alternative is. ### Check yourself **A parent gives in to a tantrum and the tantrum stops. A classmate says the child has been rewarded and the parent has been punished by losing the standoff. What is actually happening, in operant terms?** Both behaviors are being strengthened, not one. The child's tantrum is negatively reinforced when the parent gives in, and the parent's giving in is negatively reinforced when the tantrum stops. This is Patterson's coercive family process: each side trains the other, which is why parent management training teaches the reverse contingencies. **A clinic offers vouchers for cocaine-negative urine samples. A critic objects that paying people to stay clean cannot work because motivation has to come from within. What does the evidence say, and what is the program's real limitation?** The evidence says it works: meta-analyses find moderate, reliable effects across drugs, strongest for stimulants and opioids, and a review of psychosocial treatments found contingency management had the largest effects of any approach examined. Its real limitation is what the theory predicts: effects fade after the incentives stop unless natural reinforcers have taken over. **A bank pays incentives on the number of accounts opened, and the number of accounts opened soars. Did the contingency fail?** No. It worked exactly as designed, and the design was the problem. Reinforcing a proxy reinforces the proxy, so employees opened millions of accounts customers had not asked for. Organizational behavior management starts by pinpointing the behavior that matters, not a stand-in for it. **One agency wants more people to enroll in a retirement plan. Another wants people to keep exercising for a year. Which tool suits each, a nudge or an incentive, and why?** Enrollment is a one-off choice, so a nudge fits: making saving the default changes the stimulus conditions under which the choice is made, delivers no reinforcer, and produces a modest, durable effect. Exercising for a year is ongoing behavior, so an incentive works on the consequence side and produces a larger effect, but that effect will fade when the incentive stops unless a natural reinforcer has taken over. **Explain it to a friend.** Explain what a classroom token economy and a slot machine have in common, and what separates them, without using the word "reinforcement." ## Frequently asked questions **What are the main applications of operant conditioning?** Applied behavior analysis (including autism and developmental-disability services), classroom management and instruction, parenting and parent training, animal training, organizational behavior management and workplace safety, clinical treatments such as contingency management and behavioral activation, product and game design, self-management and habit formation, and — in its mathematical form — behavioral economics and reinforcement learning in AI. **How is operant conditioning used in the classroom?** Through token economies, the Good Behavior Game, school-wide PBIS, high praise-to-reprimand ratios, and instructional methods built on immediate feedback such as programmed instruction, precision teaching, and Direct Instruction. The Good Behavior Game has randomized trials with follow-ups showing benefits into early adulthood. **How is operant conditioning used in parenting?** By reinforcing wanted behavior with specific, immediate attention; using planned ignoring for attention-maintained misbehavior; using brief, calm time-out and response cost for serious misbehavior; and being consistent rather than severe. Behavioral parent-training programs such as PMT and PCIT package these skills and are among the best-supported treatments for childhood behavior problems. **How is operant conditioning used in the workplace?** Organizational behavior management pinpoints specific behaviors, measures them, and delivers frequent feedback and reinforcement. Behavior-based safety programs and graphic performance feedback have decades of supporting studies. Annual reviews fail as consequences because they are delayed, infrequent, and aversive. **Is ABA therapy the same as operant conditioning?** No. Operant conditioning is the basic learning process; applied behavior analysis is the professional discipline that applies its principles (and others) to socially important behavior, with its own methods, credentials, and ethics code. ABA is one application of operant conditioning among many. **Is contingency management effective for addiction?** Yes. Meta-analyses find it produces reliable reductions in drug use, with the strongest effects for stimulants and opioids, and a broad review of psychosocial treatments found it had the largest effects of any approach. Its main limitation is that gains can fade after incentives end unless natural reinforcers have taken over. **How do apps and games use operant conditioning?** Feeds deliver unpredictable rewards on a variable-ratio schedule, notifications act as cues to open the app, streaks turn use into loss-avoidance, and loot boxes reproduce the structure of gambling. The same principles can be used deliberately for your own goals by controlling the cues in your environment and attaching immediate consequences to behaviors you choose. **Does operant conditioning apply to artificial intelligence?** Yes. Reinforcement learning — the AI method in which an agent learns from reward signals — is a direct mathematical descendant of the law of effect, and dopamine neurons in the brain have been shown to compute the same prediction-error signal its algorithms use. Reinforcement learning underlies game-playing systems and is used in training large language models. ## References 1. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97. 2. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 3. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1982/1994). Toward a functional analysis of self-injury. *Analysis and Intervention in Developmental Disabilities, 2*(1), 3–20. Reprinted in *Journal of Applied Behavior Analysis, 27*(2), 197–209. 4. Carr, E. G., & Durand, V. M. (1985). Reducing behavior problems through functional communication training. *Journal of Applied Behavior Analysis, 18*(2), 111–126. 5. Lovaas, O. I. (1987). Behavioral treatment and normal educational and intellectual functioning in young autistic children. *Journal of Consulting and Clinical Psychology, 55*(1), 3–9. 6. Reichow, B., Hume, K., Barton, E. E., & Boyd, B. A. (2018). Early intensive behavioral intervention (EIBI) for young children with autism spectrum disorders (ASD). *Cochrane Database of Systematic Reviews*, Issue 5, CD009260. 7. Sandbank, M., Bottema-Beutel, K., Crowley, S., et al. (2020). Project AIM: Autism intervention meta-analysis for studies of young children. *Psychological Bulletin, 146*(1), 1–29. 8. Lovaas, O. I., Schaeffer, B., & Simmons, J. Q. (1965). Building social behavior in autistic children by use of electric shock. *Journal of Experimental Research in Personality, 1*, 99–109. 9. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 10. Barrish, H. H., Saunders, M., & Wolf, M. M. (1969). Good behavior game: Effects of individual contingencies for group consequences on disruptive behavior in a classroom. *Journal of Applied Behavior Analysis, 2*(2), 119–124. 11. Kellam, S. G., Brown, C. H., Poduska, J. M., et al. (2008). Effects of a universal classroom behavior management program in first and second grades on young adult behavioral, psychiatric, and social outcomes. *Drug and Alcohol Dependence, 95*(Suppl. 1), S5–S28. 12. Bradshaw, C. P., Mitchell, M. M., & Leaf, P. J. (2010). Examining the effects of schoolwide positive behavioral interventions and supports on student outcomes: Results from a randomized controlled effectiveness trial in elementary schools. *Journal of Positive Behavior Interventions, 12*(3), 133–148. 13. Skinner, B. F. (1954). The science of learning and the art of teaching. *Harvard Educational Review, 24*(2), 86–97. See also Skinner, B. F. (1958). Teaching machines. *Science, 128*(3330), 969–977. 14. Stockard, J., Wood, T. W., Coughlin, C., & Rasplica Khoury, C. (2018). The effectiveness of Direct Instruction curricula: A meta-analysis of a half century of research. *Review of Educational Research, 88*(4), 479–507. 15. Wolf, M. M., Risley, T. R., & Mees, H. (1964). Application of operant conditioning procedures to the behaviour problems of an autistic child. *Behaviour Research and Therapy, 1*, 305–312. 16. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 17. Patterson, G. R. (1982). *Coercive Family Process*. Castalia. 18. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 19. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 20. Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. *Advances in Experimental Social Psychology, 38*, 69–119. 21. Harkin, B., Webb, T. L., Chang, B. P. I., et al. (2016). Does monitoring goal progress promote goal attainment? A meta-analysis of the experimental evidence. *Psychological Bulletin, 142*(2), 198–229. 22. Higgins, S. T., Delaney, D. D., Budney, A. J., Bickel, W. K., Hughes, J. R., Foerg, F., & Fenwick, J. W. (1991). A behavioral approach to achieving initial cocaine abstinence. *American Journal of Psychiatry, 148*(9), 1218–1224. 23. Prendergast, M., Podus, D., Finney, J., Greenwell, L., & Roll, J. (2006). Contingency management for treatment of substance use disorders: A meta-analysis. *Addiction, 101*(11), 1546–1560. 24. Lussier, J. P., Heil, S. H., Mongeon, J. A., Badger, G. J., & Higgins, S. T. (2006). A meta-analysis of voucher-based reinforcement therapy for substance use disorders. *Addiction, 101*(2), 192–203. 25. Dutra, L., Stathopoulou, G., Basden, S. L., Leyro, T. M., Powers, M. B., & Otto, M. W. (2008). A meta-analytic review of psychosocial interventions for substance use disorders. *American Journal of Psychiatry, 165*(2), 179–187. 26. Ferster, C. B. (1973). A functional analysis of depression. *American Psychologist, 28*(10), 857–870. 27. Jacobson, N. S., Dobson, K. S., Truax, P. A., et al. (1996). A component analysis of cognitive-behavioral treatment for depression. *Journal of Consulting and Clinical Psychology, 64*(2), 295–304. 28. Dimidjian, S., Hollon, S. D., Dobson, K. S., et al. (2006). Randomized trial of behavioral activation, cognitive therapy, and antidepressant medication in the acute treatment of adults with major depression. *Journal of Consulting and Clinical Psychology, 74*(4), 658–670. 29. Richards, D. A., Ekers, D., McMillan, D., et al. (2016). Cost and outcome of behavioural activation versus cognitive behavioural therapy for depression (COBRA): A randomised, controlled, non-inferiority trial. *The Lancet, 388*(10047), 871–880. 30. Mowrer, O. H. (1960). *Learning Theory and Behavior*. Wiley. 31. Azrin, N. H., & Nunn, R. G. (1973). Habit-reversal: A method of eliminating nervous habits and tics. *Behaviour Research and Therapy, 11*(4), 619–628. 32. Piacentini, J., Woods, D. W., Scahill, L., et al. (2010). Behavior therapy for children with Tourette disorder: A randomized controlled trial. *JAMA, 303*(19), 1929–1937. 33. Miller, N. E. (1969). Learning of visceral and glandular responses. *Science, 163*(3866), 434–445. 34. DeFulio, A., & Silverman, K. (2012). The use of incentives to reinforce medication adherence. *Preventive Medicine, 55*(Suppl.), S86–S94. 35. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 36. Pryor, K. W., Haag, R., & O'Reilly, J. (1969). The creative porpoise: Training for novel behavior. *Journal of the Experimental Analysis of Behavior, 12*(4), 653–661. 37. Laule, G. E., Bloomsmith, M. A., & Schapiro, S. J. (2003). The use of positive reinforcement training techniques to enhance the care, management, and welfare of primates in the laboratory. *Journal of Applied Animal Welfare Science, 6*(3), 163–173. 38. Ziv, G. (2017). The effects of using aversive training methods in dogs — A review. *Journal of Veterinary Behavior, 19*, 50–60. 39. Komaki, J., Barwick, K. D., & Scott, L. R. (1978). A behavioral approach to occupational safety: Pinpointing and reinforcing safe performance in a food manufacturing plant. *Journal of Applied Psychology, 63*(4), 434–445. 40. Alvero, A. M., Bucklin, B. R., & Austin, J. (2001). An objective review of the effectiveness and essential characteristics of performance feedback in organizational settings (1985–1998). *Journal of Organizational Behavior Management, 21*(1), 3–29. 41. Stajkovic, A. D., & Luthans, F. (1997). A meta-analysis of the effects of organizational behavior modification on task performance, 1975–95. *Academy of Management Journal, 40*(5), 1122–1149. 42. Daniels, A. C. (2000). *Bringing Out the Best in People: How to Apply the Astonishing Power of Positive Reinforcement* (2nd ed.). McGraw-Hill. 43. Hamari, J., Koivisto, J., & Sarsa, H. (2014). Does gamification work? — A literature review of empirical studies on gamification. *Proceedings of the 47th Hawaii International Conference on System Sciences*, 3025–3034. 44. Drummond, A., & Sauer, J. D. (2018). Video game loot boxes are psychologically akin to gambling. *Nature Human Behaviour, 2*(8), 530–532. 45. Zendle, D., & Cairns, P. (2018). Video game loot boxes are linked to problem gambling: Results of a large-scale survey. *PLoS ONE, 13*(11), e0206767. 46. Fogg, B. J. (2009). A behavior model for persuasive design. *Proceedings of the 4th International Conference on Persuasive Technology*, Article 40. 47. Orben, A., & Przybylski, A. K. (2019). The association between adolescent well-being and digital technology use. *Nature Human Behaviour, 3*(2), 173–182. 48. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 49. Kagel, J. H., Battalio, R. C., & Green, L. (1995). *Economic Choice Theory: An Experimental Analysis of Animal Behavior*. Cambridge University Press. 50. Mazur, J. E. (1987). An adjusting procedure for studying delayed reinforcement. In M. L. Commons, J. E. Mazur, J. A. Nevin, & H. Rachlin (Eds.), *Quantitative Analyses of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value* (pp. 55–73). Erlbaum. 51. Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. *Psychological Bulletin, 82*(4), 463–496. 52. Bickel, W. K., Odum, A. L., & Madden, G. J. (1999). Impulsivity and cigarette smoking: Delay discounting in current, never, and ex-smokers. *Psychopharmacology, 146*(4), 447–454. 53. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. 54. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 55. Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. *Advances in Neural Information Processing Systems, 30*. 56. Brophy, J. (1981). Teacher praise: A functional analysis. *Review of Educational Research, 51*(1), 5–32. 57. Fordyce, W. E., Fowler, R. S., Lehmann, J. F., DeLateur, B. J., Sand, P. L., & Trieschmann, R. B. (1973). Operant conditioning in the treatment of chronic pain. *Archives of Physical Medicine and Rehabilitation, 54*(9), 399–408. 58. Fordyce, W. E. (1976). *Behavioral Methods for Chronic Pain and Illness*. Mosby. 59. Thieme, K., Flor, H., & Turk, D. C. (2006). Psychological pain treatment in fibromyalgia syndrome: Efficacy of operant behavioural and cognitive behavioural treatments. *Arthritis Research & Therapy, 8*(4), R121. 60. Reid, R. L. (1986). The psychology of the near miss. *Journal of Gambling Behavior, 2*(1), 32–39. 61. Clark, L., Lawrence, A. J., Astley-Jones, F., & Gray, N. (2009). Gambling near-misses enhance motivation to gamble and recruit win-related brain circuitry. *Neuron, 61*(3), 481–490. 62. Hopson, J. (2001, April 27). Behavioral game design. *Gamasutra*. 63. Thaler, R. H., & Sunstein, C. R. (2008). *Nudge: Improving Decisions About Health, Wealth, and Happiness*. Yale University Press. 64. Marshall, S. L. A. (1947). *Men Against Fire: The Problem of Battle Command in Future War*. William Morrow. 65. Grossman, D. (1995). *On Killing: The Psychological Cost of Learning to Kill in War and Society*. Little, Brown. 66. Spiller, R. J. (1988). S.L.A. Marshall and the ratio of fire. *RUSI Journal, 133*(4), 63–71. 67. Skinner, B. F. (1960). Pigeons in a pelican. *American Psychologist, 15*(1), 28–37. 68. Sidman, M. (1989). *Coercion and Its Fallout*. Authors Cooperative. 69. Dutton, D. G., & Painter, S. L. (1981). Traumatic bonding: The development of emotional attachments in battered women and other relationships of intermittent abuse. *Victimology, 6*(1–4), 139–155. 70. Dutton, D. G., & Painter, S. (1993). Emotional attachments in abusive relationships: A test of traumatic bonding theory. *Violence and Victims, 8*(2), 105–120. 71. Ashforth, B. (1994). Petty tyranny in organizations. *Human Relations, 47*(7), 755–778. ## Related - [Build habits with operant conditioning](https://operantconditioning.com/habits/): The self-management protocol, step by step. - [Dog training](https://operantconditioning.com/dog-training/): Markers, shaping, and the evidence on aversive methods. - [50+ examples](https://operantconditioning.com/examples/): Every quadrant, every setting. --- # Operant Conditioning in the Classroom: Praise, Token Economies, the Good Behavior Game, and What the Evidence Says > Operant conditioning in the classroom: what the evidence says about praise, token economies, the Good Behavior Game, PBIS, and Skinner's teaching machines. - Source: https://operantconditioning.com/classroom/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications · Education* Schools were the first institutions to use operant conditioning at scale, and they still produce some of its best evidence. What works — praise, tokens, the Good Behavior Game, school-wide systems, feedback-driven teaching — and what quietly backfires. > **Definition** > > **Operant conditioning in the classroom** is the deliberate arrangement of antecedents (rules, cues, task design) and consequences (attention, praise, tokens, feedback, lost privileges) so that learning-related behavior increases and disruptive behavior decreases. > > Every classroom already runs on consequences; the question is whether they are arranged on purpose. The first controlled demonstration that a teacher's attention could switch a child's studying on and off appeared in the first issue of the *Journal of Applied Behavior Analysis*.[1] **In brief** - Every classroom already runs on consequences; operant conditioning in the classroom arranges antecedents and consequences on purpose so learning behavior rises and disruption falls. - Contingent, specific praise is the cheapest reinforcer in the building, and the Good Behavior Game is one of the most thoroughly studied classroom management procedures there is. - A consequence is whatever the behavior says it is: a reprimand can be attention, and removal from class can reinforce misbehavior that escapes work. ## Why classrooms adopted operant conditioning first Skinner's interest in education began on a school visit. In 1953 he watched his daughter's fourth-grade class work identical arithmetic problems at the same pace, with results returned a day later — feedback so delayed that, by everything he knew from the laboratory, it could hardly teach. His 1954 lecture "The science of learning and the art of teaching" diagnosed education's failures as failures of reinforcement: too few reinforced responses, too long a delay, no program of small steps.[2] Psychiatric wards ran the earliest behavior-change programs, but classrooms were where the technology first spread widely; the early issues of the *Journal of Applied Behavior Analysis*, launched in 1968, were full of schools.[1][3][4] The reasons are structural: one adult controlling most of the consequences, behavior that is visible and countable, a fixed schedule, twenty-five learners at once — the operant chamber with the roles reversed. And discipline was a problem everyone agreed on: measurable, urgent, and public. ## The four quadrants in a classroom | Procedure | What happens | Classroom example | Effect | | --- | --- | --- | --- | | Positive reinforcement | Something is added | A student raises her hand instead of calling out; the teacher calls on her: "Thanks for raising your hand." | Hand-raising increases | | Negative reinforcement | Something is removed | Students who get the first ten problems right are excused from the last ten. | Careful work increases | | Positive punishment | Something aversive is added | A student shouts across the room; the teacher gives a quiet, immediate reprimand. | Shouting decreases — unless, for this student, a reprimand is attention | | Negative punishment | A reinforcer is removed | A student pushes in line and loses two minutes of recess (response cost). | Pushing decreases | | Extinction | The maintaining reinforcer stops | A student blurts out answers for the teacher's reaction; the teacher stops reacting and calls only on raised hands. | Blurting rises briefly, then declines | Two rows deserve a second look. For a student who gets little adult attention, a reprimand *is* attention, and the shouting will go up. And extinction works only on the reinforcer actually maintaining the behavior; if the class is laughing, the teacher's silence changes nothing. A consequence is whatever the behavior says it is. [More examples for every quadrant ›](https://operantconditioning.com/examples/) ## Praise that works and praise that doesn't Teacher attention is the cheapest reinforcer in the building. When teachers attended to six elementary pupils only while they studied, studying rose; when attention went back to dawdling, it fell; when contingent attention returned, it rose again.[1] In a companion study, posting rules did little, adding ignoring did little, and disruption fell only when the teachers added praise for appropriate behavior — probably, the authors concluded, the key to classroom management.[3] The reverse holds too: withdraw approval from a well-behaved class and it turns disruptive; add disapproval and it gets worse.[5] Yet most classroom praise is not reinforcement. Jere Brophy's review found that teachers praise infrequently, often without regard to what the student has just done, and often as a management reflex; vague, non-contingent praise changes nothing.[6] A later review added that praise supports intrinsic motivation when it is sincere, names effort or strategy rather than ability, and sets a reachable standard — and undermines it when it compares students with one another or gushes over easy work.[7] | Praise that reinforces | Praise that doesn't | | --- | --- | | Contingent and specific: "You showed every step of the working" | Global or vague: "good class today," "nice job" | | Credible: matched to the difficulty, in a normal voice | Effusive praise for trivial work, which signals low expectations | | About effort, strategy, and improvement on the student's own past work | About ability ("you're so smart") or rank against classmates | | In a form the student can accept — privately, for many adolescents | Public praise that embarrasses, which functions as punishment | Quantity matters too. A three-year observational study of elementary classrooms found that the higher a teacher's praise-to-reprimand ratio, the more time students spent on task, with no ceiling.[8] Many programs recommend at least four praise statements per reprimand; most classrooms run far below that, because misbehavior is more salient than quiet work, so the ratio has to be engineered — a tally on the desk, a timer that prompts a scan of the room. ## Token economies: how to build one that works A [token economy](https://operantconditioning.com/glossary/#token-economy) delivers points, stars, or chips contingent on target behaviors and lets students exchange them later for backup reinforcers. Ayllon and Azrin developed it on a psychiatric ward in the early 1960s,[9] and classrooms adopted it almost immediately: in a 1967 program, a public-school class of seventeen children with serious behavior problems earned ratings exchangeable for small prizes, and disruption dropped sharply and stayed down.[4] Tokens work because they are [generalized conditioned reinforcers](https://operantconditioning.com/positive-reinforcement/#generalized-reinforcers): instant to deliver, slow to satiate, exchangeable for whatever each student values. Alan Kazdin's reviews, from 1972 to 1982, judged the evidence strong across schools, wards, prisons, and homes, and named the recurring weaknesses: gains faded when tokens were withdrawn, did not transfer to settings without tokens, and depended on staff who often stopped delivering them.[10][11] A well-built system anticipates all three. 1. **Define three to five target behaviors as things to do.** "Start work within a minute of the bell," not "don't waste time." 2. **Deliver the token in a second, with specific praise every time.** Praise inherits the token's power and still works when the tokens are gone. 3. **Build the menu from what students do when free to choose,** start cheap, and exchange the same day. Long saving periods are a schedule most children have not yet learned to work on. 4. **Keep fines rare and small.** [Response cost](https://operantconditioning.com/negative-punishment/) works only while students have tokens to lose. 5. **Plan the fade from day one.** Raise prices, space the exchanges, shift from points to praise to natural consequences — and count the target behavior weekly. ## The Good Behavior Game The **Good Behavior Game** is one of the most thoroughly studied classroom management procedures there is, and it costs nothing. In the 1969 original, a fourth-grade class was split into two teams during math and reading; a mark went against a team whenever any member left their seat or talked out; and teams finishing under a set number of marks — both could win — earned privileges such as end-of-day free time and lining up first. Out-of-seat and talking-out behavior fell dramatically and returned when the game was withdrawn.[12] It is an *interdependent group contingency*: each student's behavior affects the team's outcome, so peers prompt and reinforce one another. The follow-up is what makes it remarkable. In the mid-1980s Sheppard Kellam's group randomly assigned first-grade classrooms in Baltimore public schools to the game, to a curriculum intervention, or to standard practice, and followed the children into adulthood. At ages 19 to 21, the men who had been the most aggressive first-graders and had played the game showed lower rates of drug and alcohol disorders, regular smoking, and antisocial personality disorder than their counterparts from control classrooms.[13] > **How to run the Good Behavior Game** > > 1. **Pick two or three rules** stated as visible behaviors: in your seat, hand up to talk, hands to yourself. > 2. **Split the class into teams** of roughly equal difficulty, and announce when the game is on — one short period a day at first, in the lesson that suffers most. > 3. **Mark violations calmly and visibly,** with no lecture, and keep teaching. > 4. **Let every team under the criterion win,** immediately at first — five minutes of a preferred activity — then at the end of the day, then the week. > 5. **Expand slowly,** stop announcing the start, and give a deliberate saboteur a team of one. ## School-wide PBIS **Positive Behavioral Interventions and Supports** (PBIS) applies the same logic to a whole building: three to five expectations, taught explicitly in each setting; tickets or praise for meeting them; a predictable response to violations; monthly review of office-referral data; and tiered support, from the universal system up to individual plans built on a functional behavioral assessment. The evidence is respectable rather than spectacular: in a randomized trial of 37 Maryland elementary schools, the 21 trained in PBIS showed reductions in office discipline referrals and suspensions relative to the 16 that were not.[14] PBIS is not a product. It is the adults in a building agreeing on what they will reinforce, and then doing it. ## Skinner's teaching machines, precision teaching, and Direct Instruction Skinner's own contribution was to instruction. His **teaching machine** presented a frame of material, required a composed answer, revealed the correct one immediately, and moved on through steps so small that the student was nearly always right.[2][15] He credited Sidney Pressey's devices of the 1920s for scoring answers automatically; what was new was the program. **Programmed instruction** named the principles — small steps, active responding, immediate feedback, self-pacing — so that the reinforcement of being right arrives thousands of times a semester rather than a few dozen. The machines are gone; the principles survive in mastery learning and adaptive tutoring software. **Precision teaching**, developed by Skinner's student Ogden Lindsley, kept the laboratory's measure: rate. Students do brief timed practice and chart correct and incorrect responses per minute, and the teacher changes the teaching when the curve flattens. The target is *fluency* — accuracy plus speed — and the rule is that the child knows best: if the child is not learning, the program is wrong.[16] **Direct Instruction**, Siegfried Engelmann's scripted, fast-paced lessons with choral responding, immediate correction, and mastery before advancement, is the most evaluated instructional program of the past half-century; a 2018 meta-analysis of more than 300 studies found consistently positive effects across subjects and across five decades of research.[17] All three share what Skinner saw in 1953: many responses, immediate consequences, a program that keeps the learner succeeding. ## The negative-reinforcement trap: misbehavior that escapes work The commonest classroom mistake is not a failure of punishment but a misreading of function. A student who finds the work aversive — too hard, too long, humiliating in front of peers — discovers that acting out makes it go away: the lesson stops, the worksheet is put aside, the student is sent to the hallway or the office. Each of those is [negative reinforcement](https://operantconditioning.com/negative-reinforcement/), and the behavior that produced it will happen more. The intended punishment is functioning as a reward. In functional analyses of problem behavior, escape from demands is the single most common maintaining consequence identified — 38% of 152 cases of self-injury in the largest series,[18] and time-out is the classic casualty: in a pair of experiments it reduced problem behavior when the environment the child left was rich, and *increased* it when leaving offered an escape.[19] Since 1997, U.S. special-education law has required a functional behavioral assessment when a student with a disability is removed from school for behavior that turns out to be a manifestation of the disability, for exactly this reason. The escape-maintained student needs work pitched where success is likely, an acceptable way to ask for a break that is honored every time, and — where safe — the demand kept in place so that acting out no longer ends it. [Escape extinction vs. ignoring ›](https://operantconditioning.com/extinction/#extinction-is-not-ignoring) ## Extinction and planned ignoring: what they can and can't do **Planned ignoring** — withholding attention from a behavior that attention has been maintaining — is [extinction](https://operantconditioning.com/extinction/), and it fits exactly one class of behavior: attention-seeking aimed at the teacher. It is the wrong tool for behavior maintained by peers' laughter, for escape, and for anything dangerous. Even where it fits, it works slowly, produces an [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) first, and on its own did little in the classic experiment; combined with praise for the behavior wanted instead, it worked well.[3] The professional version is differential reinforcement: the blurt is ignored and the raised hand is called on within seconds, every time. > **Ignoring is a decision, not a default** > > Ignore a behavior only after answering three questions: is my attention what maintains it, can I withhold it every single time, and can I outlast the burst? Ignoring for four minutes and reacting on the fifth puts the behavior on an intermittent schedule and makes it stronger. ## Do rewards kill intrinsic motivation? The overjustification caveat The objection every teacher hears is that rewards turn learning into work. The evidence is real but narrow. In the classic study, preschoolers who liked drawing were promised a certificate for drawing with markers; a week or two later they drew less in free play than children given the same certificate unexpectedly, or none at all.[20] A meta-analysis of 128 experiments found that *expected, tangible* rewards for an already interesting activity reduced later free-choice engagement, while praise did not — it raised it in college students and had no reliable effect in children; a competing meta-analysis found the effect small and confined to narrow conditions.[21][22] So: use tokens and tangible rewards for behavior students do not already do, fade them as the behavior meets its natural reinforcers, use praise and progress feedback freely, and never pay students for what they already do for pleasure. [The intrinsic-motivation evidence in detail ›](https://operantconditioning.com/positive-reinforcement/#does-positive-reinforcement-undermine-intrinsic-motivation) ## Before you reach for consequences: an antecedent checklist Consequences are the last third of the [ABC model](https://operantconditioning.com/abc-model/). Most persistent classroom problems have an antecedent solution that costs less than any reward or penalty. - **Taught, not just posted?** Rules alone change little.[3] Model, practice, and reinforce each one where it applies. - **Doable?** Work that is too hard is the antecedent for escape; work that is too easy, for everything else. - **Transitions signaled?** A two-minute warning and a visible timer are discriminative stimuli for stopping. - **Enough chances to respond?** Choral responses, whiteboards, and turn-and-talk multiply the opportunities for reinforcement; a lecture offers almost none. - **Motivating operations?** A hungry, tired, or anxious student values escape more and praise less. - **A cue for the behavior you want?** A voice-level chart, a first–then card, a checklist on the desk: [stimulus control](https://operantconditioning.com/stimulus-control/) is cheaper than consequence control. ## Common mistakes with operant conditioning in the classroom - **Reprimanding the attention-starved.** For some students a scolding is the richest attention of the day. If the behavior rises, the reprimand is reinforcing it. - **Sending the escape-maintained student out.** The hallway, the office, and in-school suspension all end the lesson. Check function before removing anyone. - **Praising the class instead of the behavior.** "Good job, everyone" reinforces nothing in particular. - **Announcing consequences you will not deliver.** Unenforced threats teach that the rule is inactive. Small consequences every time beat large ones occasionally. - **Never fading.** Tokens are scaffolding. If the system is the same in June as in September, the behavior belongs to the system, not the student. The through-line is the same as everywhere else: identify the behavior, find what actually reinforces it, make the right behavior easy, reinforce it immediately and often, and let the data decide. [All applications of operant conditioning ›](https://operantconditioning.com/applications/) · [Operant conditioning in parenting ›](https://operantconditioning.com/parenting/) ## Key takeaways - Every classroom already runs on consequences; the question is whether they are arranged on purpose. Skinner diagnosed education's failures as failures of reinforcement: too few reinforced responses, too long a delay, no program of small steps. - Teacher attention is the cheapest reinforcer in the building, but most classroom praise is not reinforcement. Praise works when it is contingent, specific, credible, and about effort or strategy, and the praise-to-reprimand ratio has to be engineered because misbehavior is more salient than quiet work. - Tokens are generalized conditioned reinforcers and the evidence for token economies is strong, but gains fade when tokens are withdrawn unless the fade is planned from day one. The Good Behavior Game, an interdependent group contingency, costs nothing and has randomized trials with benefits into adulthood. - A consequence is whatever the behavior says it is. A reprimand can be attention, and sending an escape-maintained student out of the room is negative reinforcement, so function has to be checked before anything is removed or ignored. - Expected, tangible rewards for an already interesting activity can reduce later free-choice engagement; praise and informational feedback do not. Use tangible rewards for behavior students do not already do, fade them, and fix antecedents first: rules taught, work doable, transitions signaled. ### Check yourself **A teacher gives a quiet reprimand every time a student shouts across the room, and over the month the shouting increases. What is going on?** For a student who gets little adult attention, a reprimand is attention, so the intended punishment is functioning as positive reinforcement. The behavior is the evidence: if it rises, the reprimand is reinforcing it. A consequence is whatever the behavior says it is, not what the teacher meant it to be. **A student acts out during long worksheets and is sent to the office each time, yet the acting out gets worse. The teacher concludes the office is not a harsh enough punishment. What is the better analysis?** The behavior is escape-maintained. Being sent out ends work the student finds aversive, so each trip to the office is negative reinforcement, and the behavior that produced it will happen more. The fix is antecedent and functional: work pitched where success is likely, an acceptable way to ask for a break that is honored every time, and, where safe, the demand kept in place so that acting out no longer ends it. **A teacher decides to ignore a student's blurting, holds out for four minutes, then reacts on the fifth. What has happened to the blurting?** It has been put on an intermittent schedule, which makes it stronger and more persistent. Planned ignoring is extinction, and it works only when the teacher's attention is the maintaining reinforcer, only when it is withheld every single time, and only when the teacher can outlast the extinction burst, ideally while calling on raised hands within seconds instead. **Two classes run token economies. In one the system is identical in June and September; in the other, tokens have been faded toward praise and natural consequences. Which class's behavior is more likely to survive the end of the year, and why?** The faded one. Kazdin's reviews found that gains faded when tokens were withdrawn and did not transfer to settings without tokens, so a system that never fades leaves the behavior belonging to the system, not the student. Pairing every token with specific praise lets praise inherit the token's power and keep working when the tokens are gone. **Explain it to a friend.** Explain why sending a disruptive student to the hallway can make the disruption worse, without using the words "positive" or "negative." ## Frequently asked questions **Does positive reinforcement work in the classroom?** Yes; it is the best-supported classroom-management tool there is. Contingent teacher attention and specific praise reliably increase on-task behavior, and token systems and the Good Behavior Game have decades of controlled studies behind them. It fails when praise is vague or non-contingent, when the "reward" does not actually reinforce that student, or when it arrives too late. **What is a token economy in the classroom?** A system in which students earn tokens — points, stamps, chips — immediately after target behaviors and exchange them later for privileges, activities, or small prizes. Good systems start with cheap, same-day exchanges, pair every token with specific praise, keep fines rare, and fade the tokens as praise and natural consequences take over. **What is the Good Behavior Game?** A classroom procedure, first published in 1969, in which the class is split into teams, a mark is recorded against a team whenever a member breaks a posted rule, and every team finishing under a set number of marks wins a small privilege. Randomized trials in Baltimore found that first-graders who played it had lower rates of substance-use disorders and antisocial personality disorder as young adults. **Do rewards kill intrinsic motivation?** Only under specific conditions. Expected, tangible rewards for an activity students already find interesting can reduce their later free-choice engagement with it; praise, informational feedback, and rewards for behavior students would not otherwise do generally do not. Use tangible rewards to establish behavior that is not happening, fade them, and rely on praise and feedback for the rest. **What is planned ignoring?** Deliberately withholding attention from a behavior that attention has been maintaining, so that it extinguishes. It works only when the teacher's attention is the reinforcer, only when applied every time, and only alongside reinforcement of the behavior you want instead. It is the wrong tool for behavior maintained by peers, for behavior that escapes work, and for anything dangerous. **How do I use operant conditioning in my classroom?** Start with antecedents: teach the expectations, make the work doable, signal transitions. Raise your ratio of specific praise to reprimands well above four to one, ignore minor attention-seeking while reinforcing the alternative, and add the Good Behavior Game or a simple token economy for the periods that need it. Before punishing persistent misbehavior, ask what it is getting the student; if the answer is escape from work, removal will make it worse. ## References 1. Hall, R. V., Lund, D., & Jackson, D. (1968). Effects of teacher attention on study behavior. *Journal of Applied Behavior Analysis, 1*(1), 1–12. 2. Skinner, B. F. (1954). The science of learning and the art of teaching. *Harvard Educational Review, 24*(2), 86–97. 3. Madsen, C. H., Jr., Becker, W. C., & Thomas, D. R. (1968). Rules, praise, and ignoring: Elements of elementary classroom control. *Journal of Applied Behavior Analysis, 1*(2), 139–150. 4. O'Leary, K. D., & Becker, W. C. (1967). Behavior modification of an adjustment class: A token reinforcement program. *Exceptional Children, 33*(9), 637–642. 5. Thomas, D. R., Becker, W. C., & Armstrong, M. (1968). Production and elimination of disruptive classroom behavior by systematically varying teacher's behavior. *Journal of Applied Behavior Analysis, 1*(1), 35–45. 6. Brophy, J. (1981). Teacher praise: A functional analysis. *Review of Educational Research, 51*(1), 5–32. 7. Henderlong, J., & Lepper, M. R. (2002). The effects of praise on children's intrinsic motivation: A review and synthesis. *Psychological Bulletin, 128*(5), 774–795. 8. Caldarella, P., Larsen, R. A. A., Williams, L., Downs, K. R., Wills, H. P., & Wehby, J. H. (2020). Effects of teachers' praise-to-reprimand ratios on elementary students' on-task behaviour. *Educational Psychology, 40*(10), 1306–1322. 9. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 10. Kazdin, A. E., & Bootzin, R. R. (1972). The token economy: An evaluative review. *Journal of Applied Behavior Analysis, 5*(3), 343–372. See also Kazdin, A. E. (1977). *The Token Economy: A Review and Evaluation*. Plenum Press. 11. Kazdin, A. E. (1982). The token economy: A decade later. *Journal of Applied Behavior Analysis, 15*(3), 431–445. 12. Barrish, H. H., Saunders, M., & Wolf, M. M. (1969). Good behavior game: Effects of individual contingencies for group consequences on disruptive behavior in a classroom. *Journal of Applied Behavior Analysis, 2*(2), 119–124. 13. Kellam, S. G., Brown, C. H., Poduska, J. M., et al. (2008). Effects of a universal classroom behavior management program in first and second grades on young adult behavioral, psychiatric, and social outcomes. *Drug and Alcohol Dependence, 95*(Suppl. 1), S5–S28. 14. Bradshaw, C. P., Mitchell, M. M., & Leaf, P. J. (2010). Examining the effects of schoolwide positive behavioral interventions and supports on student outcomes: Results from a randomized controlled effectiveness trial in elementary schools. *Journal of Positive Behavior Interventions, 12*(3), 133–148. 15. Skinner, B. F. (1958). Teaching machines. *Science, 128*(3330), 969–977. 16. Lindsley, O. R. (1992). Precision teaching: Discoveries and effects. *Journal of Applied Behavior Analysis, 25*(1), 51–57. 17. Stockard, J., Wood, T. W., Coughlin, C., & Rasplica Khoury, C. (2018). The effectiveness of Direct Instruction curricula: A meta-analysis of a half century of research. *Review of Educational Research, 88*(4), 479–507. 18. Iwata, B. A., Pace, G. M., Dorsey, M. F., Zarcone, J. R., Vollmer, T. R., Smith, R. G., Rodgers, T. A., Lerman, D. C., Shore, B. A., Mazaleski, J. L., Goh, H.-L., Cowdery, G. E., Kalsher, M. J., McCosh, K. C., & Willis, K. D. (1994). The functions of self-injurious behavior: An experimental-epidemiological analysis. *Journal of Applied Behavior Analysis, 27*(2), 215–240. The method itself: Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 19. Solnick, J. V., Rincover, A., & Peterson, C. R. (1977). Some determinants of the reinforcing and punishing effects of timeout. *Journal of Applied Behavior Analysis, 10*(3), 415–424. 20. Lepper, M. R., Greene, D., & Nisbett, R. E. (1973). Undermining children's intrinsic interest with extrinsic reward: A test of the "overjustification" hypothesis. *Journal of Personality and Social Psychology, 28*(1), 129–137. 21. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 22. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423. ## Related - [All applications](https://operantconditioning.com/applications/): ABA, health, work, animals, technology, and more — with the evidence rated. - [Operant conditioning in parenting](https://operantconditioning.com/parenting/): Tantrums, time-out, sticker charts, and what the evidence says. - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): What a reinforcer is, what makes it work, and why so many attempts fail. --- # Operant Conditioning in Parenting: Praise, Tantrums, Time-Out, Sticker Charts, Bedtime, and What the Evidence Says > Operant conditioning in parenting: praise, the coercion trap behind tantrums, time-out, sticker charts, bedtime extinction, and the evidence on spanking. - Source: https://operantconditioning.com/parenting/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications · Parenting* Every household runs on consequences, whether or not anyone planned them. What the best-tested parenting programs teach, why tantrums train parents as surely as parents train children, and what the evidence says about time-out, sticker charts, bedtime, and spanking. > **Definition** > > **Operant conditioning in parenting** is the use of consequences — attention, praise, privileges, and their removal — and of antecedents such as routines, clear instructions, and cues to make wanted behavior more frequent and unwanted behavior less frequent. > > Parents cannot opt out; the only choice is whether the contingencies are deliberate. The deliberate version, behavioral parent training, has been refined in controlled trials for half a century and is among the best-supported treatments for childhood behavior problems.[1][2] **In brief** - Parents cannot opt out of consequences; the only choice is whether the contingencies are deliberate, and behavioral parent training is the deliberate version. - A parent's most powerful reinforcer is attention, and households usually pay it to misbehavior; catching them being good reverses the payroll. - When a parent gives in to a tantrum, both the tantrum and the giving in are [negatively reinforced](https://operantconditioning.com/negative-reinforcement/#the-negative-reinforcement-trap-how-it-maintains-problem-behavior): Patterson's coercive family process. ## The four quadrants at home | Procedure | What the parent does | Example | Where it goes wrong | | --- | --- | --- | --- | | Positive reinforcement | Adds something after the behavior | "You put your plate in the sink — thank you." Plates go in the sink more often. | Delivered late, vaguely, or for behavior that was not happening | | Negative reinforcement | Removes something after the behavior | The nagging stops the moment the coat goes on. The coat goes on faster. | Usually runs the other way: the whining stops when the parent gives in | | Positive punishment | Adds something aversive | A sharp "No" as a hand reaches for the stove. Reaching decreases. | Scolding is attention, and attention reinforces | | Negative punishment | Removes a reinforcer | Hitting a sibling ends the game: two minutes of time-out. Hitting decreases. | Time-out from a boring room is an escape, not a loss | | Extinction | Stops delivering the reinforcer | Whining for dessert never again produces dessert. Whining declines, after a burst. | Giving in on the fifth whine teaches five-whine persistence | The rows are not equal. Positive reinforcement builds behavior, and parent-training programs spend most of their time on it. Negative punishment is for the short list of behaviors that must stop. Positive punishment has the weakest evidence and the worst side effects. Extinction is the tool everyone uses accidentally, in reverse. [Examples for every quadrant ›](https://operantconditioning.com/examples/) ## Catch them being good: the attention economy of a household A parent's most powerful reinforcer is attention, and a household has a fixed budget of it. Think of the budget as a payroll. Misbehavior is paid immediately, reliably, and in full: the shout, the lecture, the parent crossing the room. Good behavior — playing quietly, sharing unasked, shoes by the door — mostly goes unpaid, because a quiet child is a chance to do something else. Over a thousand repetitions the child learns exactly which behaviors the household pays for. **Catching them being good** reverses the payroll. Notice the behavior you want and name it, specifically and at once: "You waited until I was off the phone — that was patient," not a global "good girl" an hour later. Parent-training programs treat this **labeled praise** as a skill to be modeled and rehearsed in session; Parent–Child Interaction Therapy, for instance, coaches parents until they deliver ten labeled praises in five minutes of play, because that is what it takes to overturn the existing allocation.[1][19] > **A scolding is still attention** > > If a behavior keeps happening after it has been scolded a hundred times, the scolding is not punishing it. It is probably paying for it. Watch what the behavior does, not what you meant. ## The reinforcement trap: Patterson's coercive family process Gerald Patterson's observations of families of aggressive children in their homes found a pattern that is the most important idea on this page. The parent makes a request; the child whines, argues, or screams; the parent, worn down, withdraws the request or hands over what the child wanted. Two things were just learned. The child's aversive behavior was *negatively reinforced*: it made the demand disappear. The parent's giving in was also negatively reinforced: it made the screaming stop. Each has trained the other, and neither meant to.[3] The cycle escalates because it also involves [shaping](https://operantconditioning.com/shaping/). A parent who holds out through the whine and gives in at the scream has reinforced screaming; next time the child starts closer to the scream. A parent who gives in only sometimes has put the tantrum on an intermittent schedule, the schedule that produces the most persistent behavior of all. Patterson called the result the **coercive family process**, and his group's longitudinal work traced where it leads: coercive exchanges in early childhood predict conduct problems, rejection by ordinary peers, school failure, drift toward deviant friends, and delinquency in adolescence.[4] [More on negative-reinforcement traps ›](https://operantconditioning.com/negative-reinforcement/#the-negative-reinforcement-trap-how-it-maintains-problem-behavior) > **Both of you are being trained** > > The question after any standoff is not "who won?" but "what did each of us just get?" If the child got out of the task and you got out of the noise, both behaviors will be back tomorrow, slightly stronger. ## What parent management training teaches The remedy Patterson's group developed became **Parent Management Training** (PMT), now a family of programs with a shared core.[5] Alan Kazdin's version, refined over decades of trials with children referred for oppositional and aggressive behavior, teaches parents a small set of skills through modeling, role-play, and rehearsal — not lectures — and sends them home with practice assignments.[1] Reviews applying the strictest criteria for evidence-based treatment consistently put parent training at the top of the list for childhood disruptive behavior; the Oregon model was among the first to meet the "well-established" standard.[2] The order matters: praise and attending come first and take most of the program's time, commands next, time-out and point systems last. | Skill | What it looks like | Operant mechanism | | --- | --- | --- | | **Specific, immediate praise** | "You started your homework the first time I asked." | Positive reinforcement; pairing praise with attention makes praise itself a reinforcer | | **Planned ignoring** | No eye contact, comment, or reaction to whining; full attention the moment it stops | Extinction of attention-maintained behavior, plus reinforcement of its absence | | **Effective commands** | One instruction at a time, stated as a direction ("Put the blocks in the box"), up close, followed by a short wait | A clear discriminative stimulus; vague, chained, or question-form commands evoke less compliance[6] | | **Time-out** | Brief, calm, immediate, for a short list of serious behaviors | Negative punishment: removal of access to reinforcement | | **Point charts** | Points for two or three target behaviors, exchanged daily for privileges | A token economy; conditioned reinforcers bridge the delay | | **Monitoring** (older children) | Knowing where the adolescent is and with whom | Keeps consequences contingent; unmonitored behavior meets no parental consequence | ## Time-out done correctly **Time-out** is short for time-out from positive reinforcement, and the full name is the instruction. It works only if the environment the child leaves — the **time-in** — is warm, engaged, and reinforcing; a child removed from a dull room to a bedroom full of toys has been rewarded, not punished. Done properly it is immediate, brief (a few minutes; roughly a minute per year of age for young children), calm, free of lectures, and ended when the child has been quiet for a moment rather than while she is still protesting. Time-out has been studied with parents for more than fifty years and is part of every major evidence-based parenting program.[7] The claim that it damages attachment has been examined directly: a 2019 review concluded that, implemented as designed, the evidence does not support the concern, and that steering parents away from time-out risks pushing them toward tools with far worse evidence.[8] The American Academy of Pediatrics, which advises against spanking and recommends positive reinforcement and limit-setting in its place, teaches time-out in its guidance for parents.[9] [The step-by-step time-out procedure ›](https://operantconditioning.com/negative-punishment/#how-to-do-time-out-correctly) ## Sticker charts and token systems that work A sticker chart is a token economy, and it obeys the same rules as the ones on hospital wards and in classrooms.[10] Most home charts fail on one of four details. 1. **Immediacy.** The sticker goes on within seconds of the behavior, with praise — not at bedtime for "being good today." For a young child the sticker is the consequence; the prize it buys is a bonus. 2. **Small steps.** "A clean room for a week" is a goal, not a behavior. Start with "clothes in the hamper before dinner" and shape upward once that is reliable. 3. **Cheap, quick exchanges.** Three stickers should buy something today — choosing dinner, ten extra minutes before bed, a game with a parent. 4. **A planned fade.** Once the behavior is steady, raise the price, space the exchanges, and let praise and natural consequences take over. A chart still running unchanged six months later has become the reason for the behavior. Two cautions. Do not fine stickers off the chart for misbehavior; a child at zero has nothing to work for. And do not pay for things the child already loves: tangible rewards for an activity that was already interesting can reduce interest in it once the rewards stop, whereas praise, and rewards for behavior the child would not otherwise do, generally do not.[11] Charts are for building behavior that is not happening, not for decorating behavior that is. ## Bedtime and the extinction burst Bedtime is where most parents meet extinction for the first time, usually by accident. A child who cries when the parent leaves, and whose parent returns, has been reinforced for crying; stop returning, and the crying goes through an [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) — louder and longer on the first nights — before it declines. The [classic 1959 case](https://operantconditioning.com/extinction/#examples-of-extinction) followed exactly this course, including a relapse when a relative went back in.[12] Sleep researchers have tested both the unmodified version and **graduated extinction**, in which the parent checks in briefly at lengthening intervals rather than not at all. A review of 52 treatment studies found that the large majority reported clinically significant improvements, with unmodified extinction and preventive parent education best supported and graduated extinction close behind.[13] A randomized trial that followed infants for a year found graduated extinction reduced the time to fall asleep and the number of night wakings, with no rise in stress hormones — infant cortisol was, if anything, slightly lower — and no effect on attachment or later emotional and behavioral problems;[14] a five-year follow-up of a larger trial found neither lasting harms nor lasting benefits.[15] > **What honesty about bedtime extinction sounds like** > > It works for many families, it is not the only option, and the burst is not a sign that it is failing — it is the procedure working. Decide in advance whether you can hold the line for a week, get every caregiver to agree, and expect a smaller recovery after any night the crying is answered. A gentler method you will actually follow beats a faster one you will abandon on night three, because abandoning it on night three is intermittent reinforcement of the loudest crying yet. ## Spanking: what the evidence says Spanking is [positive punishment](https://operantconditioning.com/positive-punishment/), and it has been studied more than any other parenting consequence. The most careful meta-analysis, restricted to ordinary open-handed spanking rather than abuse, pooled 75 studies covering 160,927 children and found spanking significantly associated with 13 of the 17 outcomes examined — more aggression, more antisocial behavior, more mental-health problems, worse parent–child relationships — every one in the harmful direction and none in the beneficial direction.[16] The data are mostly correlational, and difficult children are spanked more; but the associations held in the longitudinal studies the review examined, and no comparable evidence finds benefits. In 2018 the American Academy of Pediatrics advised parents not to spank, hit, or otherwise physically punish children, not to use words that shame or humiliate, and to rely instead on positive reinforcement, limit-setting, and brief time-out.[9] The operant analysis predicts the result. Spanking as practiced is delayed, inconsistent, and escalating; it delivers a burst of parental attention; it models the aggression it is meant to stop; and it makes the parent an aversive stimulus, which produces avoidance and lying. [The corporal-punishment evidence in detail ›](https://operantconditioning.com/positive-punishment/#corporal-punishment-what-the-evidence-shows) ## Consistency beats severity Parents whose consequences are not working tend to make them bigger. The laboratory says to make them more reliable instead. In Azrin and Holz's classic review, punishment delivered intermittently was far less effective than punishment delivered every time, and punishment introduced mildly and then escalated produced adaptation: the organism learned to tolerate each new level.[17] In practice consistency means three things. **Immediacy**: the consequence follows within seconds or minutes, which is why "wait until your father gets home" changes nothing except the child's feelings about the front door. **Contingency**: it follows the behavior and nothing else — not withheld because the parent is busy, not delivered because the parent is tired. **Agreement**: both parents, the grandparents, and the babysitter enforce the same short list of rules, because a rule enforced by one adult in three is a rule on a variable schedule. A rule you cannot enforce every time is better dropped than announced. ## The Premack principle at home "First homework, then screen" is the [Premack principle](https://operantconditioning.com/premack-principle/): a more probable behavior can reinforce a less probable one.[18] Grandmothers knew it as "first your vegetables, then dessert." Its power at home is that it needs no prizes: the reinforcers are the things the child already does when free — playing outside, watching a show, being read to — made contingent on the things he does not. The preferred activity comes *after*, promptly, and in an amount small enough that the child could not have had it anyway. "You can play now if you promise to do homework later" reverses the order and reinforces promising. ## Screens and the variable-ratio problem Screens complicate every contingency in the house for a specific reason: the device delivers its own reinforcement on a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/#why-variable-ratio-is-so-powerful-and-so-dangerous), the arrangement that produces the highest rates of behavior and the greatest resistance to extinction. A game or a feed pays out unpredictably, so "one more minute" is not defiance; it is the burst a pigeon shows when the key stops paying. Ending screen time is therefore a loss, the transition off a screen is the most reliable trigger of protest in the modern household, and a parent who hands the tablet back to end the protest has reinforced it. - Put screens on the "then" side of every first–then, never the "first." - Signal the end with a timer the child can see, so the timer rather than the parent is the cue for stopping. - Never hand over a screen to stop a tantrum; that is the supermarket candy with a battery. - Set the daily amount when everyone is calm; a limit renegotiated during a protest has just been placed on an intermittent schedule. [How apps and games use operant conditioning ›](https://operantconditioning.com/applications/#technology) ## What not to do - **Bribing before the behavior.** "If you stop screaming I'll buy you the toy" delivers the reinforcer for screaming. Reinforcement follows the behavior you want; a bribe precedes it and is triggered by the one you don't. - **Delayed consequences.** A privilege lost next weekend for something done on Tuesday is felt as arbitrary, and by Saturday it is punishing Saturday's behavior. - **Punishing after the fact.** The mess discovered an hour later cannot be connected to the act, for a puppy or a toddler. A consequence discovered late is better skipped than delivered. - **Rewards that satiate.** Candy after every good act stops working by mid-afternoon. Attention, activities, and choices satiate far more slowly, and they are free. - **Threats you will not carry out.** An unenforced warning teaches that warnings are noise, so the next real one is ignored too. ## A worked example: the supermarket tantrum in ABC The [ABC model page analyzes a supermarket tantrum](https://operantconditioning.com/abc-model/#a-child-s-tantrum-in-the-supermarket) and shows how to intervene at the antecedent, the behavior, and the consequence. Here is what that analysis leaves out: how the tantrum was built, and why there are two contingencies in the aisle rather than one. | Whose behavior | Antecedent | Behavior | Consequence | What was learned | | --- | --- | --- | --- | --- | | **The child's** | Checkout aisle, candy at eye level, parent busy, no snack since lunch | Asks, then whines, then screams and drops to the floor | Candy appears; so does the parent's full attention | Screaming is positively reinforced by candy and attention | | **The parent's** | Screaming child, a queue watching, a card machine waiting | Hands over the candy | The screaming stops; the stares stop | Giving in is negatively reinforced by the end of the noise and the embarrassment | Play it forward. On the first Saturday the parent gives in at the whine. On the second, resolved to be firmer, she holds out through the whine and gives in at the scream — which shapes the scream. On the third she holds out through the scream and gives in when the child hits the floor. By the fourth Saturday the child starts on the floor, because that is the response that has been paid, and the parent gives in at once, because that is the response that ends it fastest. Nobody in this story is weak or bad. Two ordinary learners have shaped each other on an intermittent schedule, and the aisle now controls both of them. The way out is the one the ABC page lays out: change the antecedent, teach and reinforce a replacement, and make sure candy never again follows a scream. The honest addition is that the next scream will be the loudest, that giving in to it will make the one after worse, and that the parent who has decided in advance what she will do — with a snack and a job for the child already in her bag — rarely reaches the floor at all. [All applications of operant conditioning ›](https://operantconditioning.com/applications/) · [Operant conditioning in the classroom ›](https://operantconditioning.com/classroom/) ## Key takeaways - Every household runs on consequences whether or not anyone planned them. Positive reinforcement builds behavior and takes most of a parent-training program's time; negative punishment is for the short list of behaviors that must stop; positive punishment has the weakest evidence and the worst side effects. - Attention is the household's most powerful reinforcer, and misbehavior is usually paid first and in full. Labeled praise, specific and immediate, reverses the allocation; a scolding is still attention, so a behavior that survives a hundred scoldings is probably being paid by them. - In the coercive family process, the child's whining is negatively reinforced when the demand disappears and the parent's giving in is negatively reinforced when the noise stops. Holding out and then giving in shapes a louder version and puts it on an intermittent schedule, the most persistent of all. - Time-out works only when time-in is warm and reinforcing, and done as designed it does not damage attachment. A sticker chart is a token economy: immediate delivery, small steps, cheap same-day exchanges, and a planned fade. - Spanking is associated with worse outcomes on 13 of 17 measures and better outcomes on none, and the operant analysis predicts why. Consistency beats severity: immediacy, contingency, and agreement among every adult, and a rule that cannot be enforced every time is better dropped than announced. ### Check yourself **A four-year-old keeps drawing on the wall even though he is scolded every time. The parent concludes the punishment is not strong enough. What does the behavior say?** A scolding is attention, and attention reinforces. If the behavior keeps happening after it has been scolded a hundred times, the scolding is not punishing it; it is probably paying for it. Watch what the behavior does, not what the consequence was meant to do, and move the attention to the behavior wanted instead. **During a dull afternoon a child hits her brother and is sent to her toy-filled bedroom for two minutes. Is this time-out?** Not in function. Time-out is short for time-out from positive reinforcement, so it works only if the time-in the child leaves is warm, engaged, and reinforcing. Leaving a boring room for a room full of toys is an escape and a reward, not a loss, and hitting is likely to go up. **A parent who used to give in at the first whine resolves to be firmer. She now holds out through the whine and gives in when the child screams. Has she made progress?** No. She has shaped screaming: the louder response is the one that was paid, so next time the child starts closer to the scream. Because she gives in only sometimes, the tantrum is now on an intermittent schedule, the schedule that produces the most persistent behavior of all. The way out is decided in advance, not during the standoff. **A parent stops returning to a crying child at bedtime. On the first two nights the crying is louder and longer than ever. Is the procedure failing?** No. That is the extinction burst, and it is the procedure working: crying that returning had reinforced gets louder before it declines. Giving in on night three would be intermittent reinforcement of the loudest crying yet, which is why the decision to hold the line, and every caregiver's agreement, has to come before night one. **Explain it to a friend.** Explain how a tantrum trains a parent, using a standoff from your own household or childhood as the example. ## Frequently asked questions **Does positive reinforcement work on children?** Yes; it is the core of every evidence-based parenting program. Specific praise and attention delivered immediately after a behavior reliably increase it. It fails when the praise is vague or late, when the "reward" is not actually reinforcing for that child, or when misbehavior is still being paid more reliably in attention. **Is time-out harmful?** Not when done as designed: brief, calm, immediate, for a short list of serious behaviors, from a home rich in positive attention. A 2019 review of the claim that time-out damages attachment found the evidence does not support it, and the American Academy of Pediatrics, which advises against physical punishment, includes time-out in its guidance for parents. It fails — and can backfire — when time-in is not rewarding or the child is glad to leave. **Do sticker charts work?** They work when they follow the rules of a token economy: the sticker arrives within seconds of a small, specific behavior; stickers buy something the same day at first; nothing is taken away for misbehavior; and the chart is faded once the behavior is steady. They fail when the goal is too large, the prize too distant, the chart abandoned after a week, or the reward attached to something the child already enjoys. **Why does my child ignore consequences?** Usually because the consequence is not functioning as one. It may be too delayed, delivered only sometimes, or not something the child values losing; it may even be reinforcing — a scolding is attention, and being sent to a room full of toys is a reward. Ask what the behavior is getting the child (attention, escape, an item) and whether your consequence removes that or supplies it. **Is spanking effective?** Not by the evidence. The largest meta-analysis of spanking alone, covering 75 studies and more than 160,000 children, found it associated with worse outcomes on 13 of 17 measures — including more aggression and antisocial behavior — and better outcomes on none. The American Academy of Pediatrics advises against it and recommends positive reinforcement, limit-setting, redirection, and clear expectations instead. **How do I stop giving in to tantrums?** Decide before the tantrum, not during it: giving in at the loudest point reinforces the loudest version and puts the tantrum on an intermittent schedule. Change the antecedents (a snack, a job, the rule stated before entering the store), teach and reinforce a replacement way of asking, ignore the tantrum if attention maintains it while keeping the child safe, and expect it to get louder for a few episodes before it fades. ## References 1. Kazdin, A. E. (2005). *Parent Management Training: Treatment for Oppositional, Aggressive, and Antisocial Behavior in Children and Adolescents*. Oxford University Press. 2. Eyberg, S. M., Nelson, M. M., & Boggs, S. R. (2008). Evidence-based psychosocial treatments for children and adolescents with disruptive behavior. *Journal of Clinical Child & Adolescent Psychology, 37*(1), 215–237. 3. Patterson, G. R. (1982). *Coercive Family Process*. Castalia. 4. Patterson, G. R., DeBaryshe, B. D., & Ramsey, E. (1989). A developmental perspective on antisocial behavior. *American Psychologist, 44*(2), 329–335. 5. Forgatch, M. S., & Patterson, G. R. (2010). Parent Management Training — Oregon Model: An intervention for antisocial behavior in children and adolescents. In J. R. Weisz & A. E. Kazdin (Eds.), *Evidence-Based Psychotherapies for Children and Adolescents* (2nd ed.). Guilford Press. 6. Forehand, R. L., & McMahon, R. J. (1981). *Helping the Noncompliant Child: A Clinician's Guide to Parent Training*. Guilford Press. 7. Everett, G. E., Hupp, S. D. A., & Olmi, D. J. (2010). Time-out with parents: A descriptive analysis of 30 years of research. *Education and Treatment of Children, 33*(2), 235–259. 8. Dadds, M. R., & Tully, L. A. (2019). What is it to discipline a child: What should it be? A reanalysis of time-out from the perspective of child mental health, attachment, and trauma. *American Psychologist, 74*(7), 794–808. 9. Sege, R. D., Siegel, B. S., Council on Child Abuse and Neglect, & Committee on Psychosocial Aspects of Child and Family Health. (2018). Effective discipline to raise healthy children. *Pediatrics, 142*(6), e20183112. 10. Kazdin, A. E. (1977). *The Token Economy: A Review and Evaluation*. Plenum Press. 11. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 12. Williams, C. D. (1959). The elimination of tantrum behavior by extinction procedures. *Journal of Abnormal and Social Psychology, 59*(2), 269. 13. Mindell, J. A., Kuhn, B., Lewin, D. S., Meltzer, L. J., & Sadeh, A. (2006). Behavioral treatment of bedtime problems and night wakings in infants and young children. *Sleep, 29*(10), 1263–1276. 14. Gradisar, M., Jackson, K., Spurrier, N. J., et al. (2016). Behavioral interventions for infant sleep problems: A randomized controlled trial. *Pediatrics, 137*(6), e20151486. 15. Price, A. M. H., Wake, M., Ukoumunne, O. C., & Hiscock, H. (2012). Five-year follow-up of harms and benefits of behavioral infant sleep intervention: Randomized trial. *Pediatrics, 130*(4), 643–651. 16. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 17. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 18. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 19. McNeil, C. B., & Hembree-Kigin, T. L. (2010). *Parent–Child Interaction Therapy* (2nd ed.). Springer. ## Related - [All applications](https://operantconditioning.com/applications/): ABA, health, work, animals, technology, and more — with the evidence rated. - [Operant conditioning in the classroom](https://operantconditioning.com/classroom/): Praise, token economies, the Good Behavior Game, and what the evidence says. - [Negative punishment](https://operantconditioning.com/negative-punishment/): Time-out and response cost, step by step. --- # Operant Conditioning in Dog Training: The Four Quadrants, the Evidence, and How to Train > Operant conditioning in dog training: the four quadrants, the evidence on reward-based vs. aversive methods, clicker training, timing, and a recall protocol. - Source: https://operantconditioning.com/dog-training/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications · Animal training* Every dog trainer uses operant conditioning, whether they know the vocabulary or not. Here are the four quadrants with dog examples, what the research actually says about reward-based and aversive methods, how markers, shaping, timing, and schedules work, and a protocol for a recall you can trust. > **Definition** > > **Operant conditioning in dog training** is the deliberate use of consequences to change what a dog does: behaviors that are followed by reinforcement (a treat, play, release to sniff, or relief from pressure) happen more often, and behaviors followed by punishment (an added aversive or a lost reward) happen less often. Trainers arrange the [antecedent](https://operantconditioning.com/abc-model/) — the cue and the setting — and the consequence, and the dog's behavior changes in between. > > A "reward" only counts as a reinforcer if the behavior actually increases, and a "correction" only counts as a punisher if the behavior actually decreases. The dog, not the trainer, decides what works.[1] **In brief** - The [four quadrants](https://operantconditioning.com/#the-four-quadrants-reinforcement-and-punishment) are not equal choices: R+ and P− need nothing unpleasant present, while R− and P+ depend on an aversive the trainer introduces. - No controlled study has found aversive methods more effective than reward-based ones, and aversive methods are associated with stress and aggressive responses. - A reward counts as a reinforcer only if the behavior actually increases; the dog, not the trainer, decides what works. ## The four quadrants in dog training Every consequence a dog experiences can be sorted by two questions: was something *added* or *removed*, and did the behavior go *up* or *down*? Trainers usually abbreviate the results as R+, R−, P+, and P−. | Quadrant | What happens | Dog training example | Effect and notes | | --- | --- | --- | --- | | R+ Positive reinforcement | Something the dog wants is added after the behavior | Dog sits → treat, or a thrown ball, or the door opens | Sitting increases. The foundation of modern, reward-based training. | | R− Negative reinforcement | Something the dog dislikes is removed after the behavior | Steady leash pressure → dog steps toward the handler → pressure released | Moving toward the handler increases. Requires an aversive to be present first; used in traditional and some "balanced" training. | | P+ Positive punishment | Something the dog dislikes is added after the behavior | Dog pulls → leash pop; dog barks → spray bottle | Pulling or barking decreases, if it works at all. Associated with stress and aggressive responses (see the evidence below). | | P− Negative punishment | Something the dog wants is removed after the behavior | Dog jumps up → person turns away; dog mouths hand → play ends for 30 seconds | Jumping or mouthing decreases. The mildest way to reduce behavior, and it pairs naturally with R+ for an alternative. | Two things follow from the table. First, "positive" and "negative" describe adding and removing, not kind and cruel: the spray bottle is *positive* punishment. Second, the quadrants are not four equal choices. R+ and P− require nothing unpleasant to be present, while R− and P+ both depend on an aversive stimulus the trainer has to introduce, which brings side effects with it. That asymmetry is why most evidence-based trainers work almost entirely in the top-left cell of the [quadrant grid](https://operantconditioning.com/#the-four-quadrants-reinforcement-and-punishment) (R+) and reach for P− when a behavior needs to shrink. ## What the evidence says about reward-based vs. aversive dog training Trainers argue about methods constantly; the research is smaller than the argument, but it points consistently in one direction. ### Owner surveys Hiby, Rooney, and Bradshaw surveyed 364 dog owners about the methods they used for common tasks and the behavior of their dogs. Owners who relied on reward-based methods reported higher obedience, and the use of punishment-based methods was associated with a higher number of problem behaviors.[2] Herron, Shofer, and Reisner surveyed 140 owners whose dogs had been referred to a veterinary behavior clinic. For several confrontational techniques — hitting or kicking the dog, growling at it, the "alpha roll," staring it down — roughly a quarter or more of owners reported that the dog had responded with aggression. Reward-based techniques rarely produced aggressive responses.[3] ### Reviews and welfare studies Ziv's 2017 review of the literature concluded that aversive methods carry welfare risks and that there is no evidence they are more effective than reward-based methods.[4] Vieira de Castro and colleagues then went beyond questionnaires: they filmed dogs from training schools that used reward-based methods and schools that used aversive methods, sampled salivary cortisol, and ran a cognitive bias test. Dogs from aversive-based schools showed more stress-related behaviors and body postures during training, higher post-training cortisol, and in the cognitive bias test they approached an ambiguous bowl more slowly — the "pessimistic" pattern seen in animals in poorer welfare states.[5] ### Experimental comparisons The strongest single piece of efficacy evidence is an experiment rather than a survey. China, Mills, and Cooper assigned pet dogs with known off-lead problems to training with remote electronic collars by industry-approved trainers, to the same trainers without collars, or to reward-based trainers. The reward-based group responded to "sit" and "come" more reliably and more quickly; the e-collar added no measurable benefit.[6] > **The honest caveats** > > Most of this evidence is correlational. Owners who choose punishment may already have more difficult dogs, and owners who choose reward-based methods may differ in other ways too. Surveys rely on self-report; clinic samples are not typical dogs; and school-based comparisons cannot fully separate the method from the trainer. What can be said is this: no controlled study has found aversive methods to be *more* effective, several have found reward-based methods equally or more effective, and the welfare and aggression findings point the same way across designs. That is enough for professional bodies to recommend reward-based training as the default.[7] [More on aversives and their side effects ›](https://operantconditioning.com/positive-punishment/#aversives-in-dog-training-what-the-evidence-shows) ## The dominance myth A great deal of popular dog training rests on the idea that dogs are trying to become the "alpha" of the household and must be shown their place. The idea traces to studies of unrelated wolves confined together in captivity in the mid-twentieth century, which showed constant fighting for rank. L. David Mech's 1970 book helped spread the "alpha wolf" concept; his own later fieldwork on wild wolves showed that a pack is simply a family — a breeding pair and their offspring — and that the parents lead the way parents do, without ritualized dominance contests. Mech published the correction in 1999, has spent years asking people to drop the term, and has said he asked his publisher to stop printing the 1970 book.[8][9] Domestic dogs are not wolves in any case, and studies of free-ranging dogs and of dog–human interactions find little support for the notion that problem behavior is a bid for status.[10] The dog that pulls on the leash is not staging a coup; pulling has simply been reinforced by getting where it wants to go. The American Veterinary Society of Animal Behavior's position statement on dominance theory recommends against confrontational "dominance" techniques, and its 2021 statement on humane dog training recommends reward-based methods and advises against aversive tools.[7][11] ## Marker and clicker training A treat that arrives three seconds after a sit reinforces whatever the dog was doing three seconds after it sat — usually standing up and sniffing your hand. A **marker** solves the timing problem. A clicker (or a short word like "yes") is paired with food until the sound itself becomes a **conditioned reinforcer**: through classical conditioning it comes to predict the treat, and through operant conditioning it can then strengthen whatever it follows.[12] Trainers call it a **bridge** because it spans the gap between the behavior and the primary reinforcer. [How classical and operant conditioning combine ›](https://operantconditioning.com/operant-vs-classical-conditioning/) The technique is older than most people think. Keller and Marian Breland, two of Skinner's early students, left academia in the 1940s to train animals commercially and used a hand-held clicker as a conditioned reinforcer across dozens of species.[13] Skinner himself described the method for a general audience in 1951, explaining how to train a dog with a conditioned reinforcer and successive approximations.[14] Marine-mammal trainers adopted a whistle as their bridge, and Karen Pryor, a former dolphin trainer, brought the approach to dog owners through *Don't Shoot the Dog* and the clicker-training movement that followed.[15] ### Charging the clicker Click, then treat; click, then treat — twenty or thirty times over a couple of short sessions, with the click always *preceding* the treat by about a second. When the dog's head whips toward you at the sound, the marker is charged. From then on the rule is simple: every click earns a treat, and the click marks the exact instant the behavior you want occurs. The treat can follow a moment later; the click has already done the teaching. ## Shaping, luring, and capturing There are three ways to get a behavior to happen so you can reinforce it. - **Capturing.** Wait for the dog to do the behavior on its own — lie down, make eye contact — and mark it. Slow for rare behaviors, excellent for common ones, and it produces behavior the dog "owns." - **Luring.** Use a treat in the hand to guide the dog into position: raise it over the nose and the rear drops into a sit. Fast, but the lure must be faded within a few repetitions or the dog learns that the behavior happens only when food is visible. - **Shaping.** Reinforce successive approximations — first a glance at the mat, then a step toward it, then a paw on it, then lying on it. Shaping builds behaviors that can't be lured and teaches the dog to experiment.[1] [How shaping works, step by step ›](https://operantconditioning.com/shaping/) ## What makes reinforcement work ### Timing: the one-second window Laboratory work on delayed reinforcement shows that a consequence loses much of its power within a few seconds of the behavior, and a delayed reinforcer tends to strengthen whatever happened just before *it* rather than the behavior you intended.[16] In practice: mark within about a second, and deliver the treat where you want the dog to be (feed a "down" on the floor between the paws, feed heel position at your left knee). ### Rate of reinforcement Early in learning, aim for many reinforced repetitions per minute — ten or fifteen is not unusual in a good session. A high rate keeps the dog engaged and out-competes distractions. ### Reinforcer value Reinforcers are not interchangeable. Most dogs rank kibble below cheese, cheese below roast chicken, and, for some, a tug toy above all of it. Match the reinforcer to the difficulty: kibble for a sit in the kitchen, chicken for a recall past a squirrel. And remember motivating operations: a dog trained right after dinner is a dog for whom food has stopped being a reinforcer. ### Life rewards and the Premack principle Anything the dog wants to do can reinforce something you want it to do — David Premack's [principle](https://operantconditioning.com/premack-principle/) that a higher-probability behavior reinforces a lower-probability one.[17] Sit, and the door opens. Look at me, and you are released to sniff that fascinating hydrant. Come when called, and you get sent back to play. ## Schedules of reinforcement in dog training While the dog is learning, reinforce *every* correct response — a continuous schedule, which produces the fastest acquisition. Once the behavior is reliable, thin to a **variable-ratio** schedule: reinforce most responses, then some, in an unpredictable pattern, while keeping the best responses on a "jackpot."[18] This is the step most owners skip, and skipping it in either direction causes trouble. Never thinning produces a dog that works only when food is visible. Thinning too fast, or reinforcing so rarely that the behavior stops paying, produces a dog that stops offering the behavior. > **Why occasional treats make behavior stronger, not weaker** > > Owners worry that "not rewarding every time" will erode the behavior. The opposite is true. Behavior maintained on an intermittent schedule is *more* resistant to extinction than behavior reinforced every time — the partial-reinforcement extinction effect that Ferster and Skinner documented in thousands of hours of cumulative records. A dog whose sit is reinforced unpredictably keeps sitting through long stretches without a treat, because long stretches without a treat are exactly what it has learned to expect.[18] The same principle works against you when counter-surfing pays off once a month. [Run the schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) ## Stimulus control, cues, and generalization ### Add the cue after the behavior is reliable A cue is a [discriminative stimulus](https://operantconditioning.com/glossary/#discriminative-stimulus): a signal that the behavior will now be reinforced. Saying "sit" to a dog that does not yet sit reliably teaches it nothing except that "sit" is background noise. The cleaner sequence is: get the behavior (by capturing, luring, or shaping), reinforce it until the dog offers it readily, and then say the cue *just as* the dog starts to do it. Within a few dozen pairings the word predicts the behavior and the reinforcement, and the dog starts responding to the word alone. A behavior is under stimulus control when it happens promptly on cue, does not happen without the cue, and does not happen to other cues.[1] ### Dogs don't generalize well — so train everywhere A dog that sits perfectly in the kitchen has learned "sit, in the kitchen, facing my owner, with the treat pouch on." Take it to the park and the behavior can vanish, not out of stubbornness but because none of the antecedent conditions match. Behavior analysts learned long ago that generalization has to be programmed rather than hoped for: train in many places, with many people, at many distances and levels of distraction, and reinforce across all of them.[19] Trainers summarize this as the three D's — distance, duration, distraction — and the rule that you raise only one at a time. ## Extinction and extinction bursts When a behavior that used to be reinforced stops working, it does not vanish quietly. It gets worse first. The dog whose jumping has always earned a hand, a voice, or a push-off will, when the attention stops, jump higher, faster, and with more mouth — an **extinction burst** — before the behavior declines.[20] The same happens with demand barking: ignore it and the dog will bark louder for a while. Two rules make extinction workable. Be completely consistent, because reinforcing the louder bark teaches the dog that louder is the new price. And always reinforce an alternative — four paws on the floor, a quiet sit — so the dog has a behavior that *does* pay. Pure extinction without a replacement is slow and unkind. [Extinction, bursts, and spontaneous recovery in depth ›](https://operantconditioning.com/extinction/) ## Management: control the antecedent before you train Every time a dog practices a behavior and it pays off, the behavior gets stronger. So before any training plan, arrange the environment so the unwanted behavior can't be rehearsed: baby gates at the front door, food off the counters, a leash on the dog when guests arrive, blinds down for the window barker, a long line for the dog whose recall isn't ready. This is antecedent control, the "A" of the [A-B-C model](https://operantconditioning.com/abc-model/), and it is what lets the reinforcement-based plan win: the old behavior stops being reinforced in the background while you teach the new one. ## How to teach a reliable recall, step by step Recall is the behavior that keeps dogs alive, and it is the one most often ruined by good intentions. The protocol below uses nothing but positive reinforcement, management, and schedules. 1. **Choose a fresh cue.** If "come" has been shouted at a dog that ignored it, or used to end walks, it is contaminated. Pick a new word — "here," a whistle — and protect it: never use it when you can't make it pay. 2. **Charge the cue.** Indoors, with the dog beside you, say the cue once and immediately feed three to five tiny pieces of something excellent, one after another. Ten repetitions, twice a day, for three or four days. The cue now predicts a jackpot before the dog has done anything. 3. **Add the behavior at short range.** Say the cue when the dog is a few feet away and likely to come anyway. When it reaches you, take the collar gently, *then* feed. The collar touch becomes part of the reinforced chain, so a hand reaching for the collar never becomes a signal to dodge. 4. **Build distance and distraction separately.** Increase distance indoors, then move to a quiet yard at short distance, then a long line in a park. Raise one variable at a time. The long line is management: it guarantees the cue is never disobeyed successfully. 5. **Send the dog back to the fun.** Most recalls in practice should end with "go play." If coming when called usually ends the walk, the recall is being punished by the loss of freedom — negative punishment of the very behavior you want. Use the Premack principle so that coming is the price of more freedom, not the end of it. 6. **Thin the schedule, keep the jackpots.** Once the recall is reliable, reinforce with food unpredictably, keep the surprise jackpots for fast responses through distractions, and lean on life rewards. Never let the schedule thin to zero. 7. **Never punish a recall.** If the dog comes slowly, reinforce anyway — a slow recall is still a recall, and scolding it teaches the dog that arriving is dangerous. If the dog doesn't come, go and get it calmly, then make the next repetition easier. 8. **Maintain for life.** A few surprise recalls on every walk, always paid, keep the behavior strong. ## Common mistakes in operant dog training - **Repeating the cue.** "Sit. Sit! SIT!" teaches the dog that the cue is "sit-sit-SIT." Say it once; if nothing happens, make the situation easier and try again. - **Poisoning the cue.** A cue that is sometimes followed by reinforcement and sometimes by a correction becomes ambiguous — Karen Pryor's term is a *poisoned cue* — and dogs respond to it slowly and with signs of stress. Keep each cue attached to one kind of consequence. - **Punishing the recall by ending the fun.** Calling the dog only to leave the park, get a bath, or be crated trains the dog to keep its distance. - **Treat dependence from never thinning.** If every sit for two years has produced a visible treat, the treat has become part of the cue. Fade the lure early and thin the schedule once the behavior is reliable. - **Bribing instead of reinforcing.** Showing the treat before the behavior is a lure or a bribe. Reinforcement comes *after*. - **Reinforcing the wrong moment.** A treat handed to a dog that sat and then stood reinforces standing. Mark the sit; feed in the sit. - **Training when the dog is over threshold.** A dog that is frantic about another dog across the street cannot learn. Add distance until it can eat and think, then train. ## Common problem behaviors: what maintains them, and a reinforcement-based plan The first question is never "how do I stop this?" but "what is this behavior getting?" Once you can name the reinforcer, the plan almost writes itself: manage so the old reinforcer stops arriving, and reinforce a behavior that can replace it. | Behavior | Likely maintaining reinforcer | Reinforcement-based plan | | --- | --- | --- | | Jumping on people | Attention: eye contact, voices, hands — even a push-off is contact R+ | Manage with a leash or gate at the door. All attention stops the instant paws leave the floor P−; four-on-the-floor or a sit earns the greeting, generously. Recruit guests, and expect a burst. | | Pulling on the leash | Forward progress toward smells and dogs R+ | A tight leash stops all forward motion (pulling no longer works); a loose leash makes the walk go on, plus frequent treats at your side. A front-clip harness for management while the new behavior builds. | | Barking at passersby from the window | The passerby always leaves R−, plus the arousal itself | Block the view (film on the glass, closed blinds). Teach and heavily reinforce "go to your mat" when someone passes; reinforce quiet glances at the window. | | Demand barking | Food, play, the door, or attention delivered to stop the noise R+ | Barking pays nothing, ever. Teach a quiet alternative request (a sit, a nose-touch) and pay it fast and often. Consistency from everyone in the house, and expect the burst. | | Counter-surfing | Food, on an intermittent schedule — the most durable kind VR | Management is non-negotiable: clear counters, closed kitchen. Reinforce lying on a mat in the kitchen while you cook. One sandwich a month will maintain the behavior indefinitely. | | Begging at the table | Scraps, from at least one family member, occasionally VR | Nobody feeds from the table, no exceptions. Give the dog a stuffed food toy on its bed during meals so lying there becomes the behavior that pays. | | Lunging and barking at dogs on leash | Distance: the other dog goes away, or the owner retreats R−, usually driven by fear or frustration | Work at a distance where the dog can eat and think. Pair the appearance of other dogs with excellent food (counterconditioning), and reinforce looking at the dog and back at you. Do this with a qualified reward-based trainer. | > **Aggression, fear, and resource guarding need a professional** > > Growling, snapping, biting, guarding food or objects, and severe fear are not obedience problems, and punishing them tends to suppress the warning while leaving the emotion intact — the dog that no longer growls may go straight to biting. Seek a certified reward-based trainer or, for aggression and anxiety, a veterinary behaviorist, who can also rule out pain and medical causes. ## Key takeaways - "Positive" and "negative" mean added and removed, not kind and cruel. R+ and P− require nothing unpleasant to be present; R− and P+ both depend on an aversive the trainer introduces, which is why evidence-based trainers work almost entirely in R+ and reach for P− when a behavior needs to shrink. - The research is smaller than the argument but points one way: no controlled study has found aversive methods more effective, reward-based methods were equally or more effective, and aversive methods are associated with stress, higher cortisol, and aggressive responses. The dominance idea rests on captive wolves and was retracted by the researcher who spread it. - A marker is a conditioned reinforcer that bridges the gap between the behavior and the treat; without it, a treat three seconds late reinforces whatever the dog was doing three seconds later. Mark within about a second and feed where you want the dog to be. - Reinforce every correct response while the dog is learning, then thin to a variable-ratio schedule and keep the jackpots. Intermittent reinforcement makes behavior more resistant to extinction, which is why a sit survives long stretches without treats and why one sandwich a month maintains counter-surfing. - Cues are discriminative stimuli added after the behavior is reliable, and dogs generalize poorly, so train everywhere and raise distance, duration, and distraction one at a time. Before training, manage the antecedent so the old behavior stops being rehearsed, and ask what the behavior is getting before asking how to stop it. ### Check yourself **An owner calls her dog at the park, and when it arrives she clips on the leash and goes home. She never scolds it, yet over the weeks the recall gets slower. Why is the recall weakening?** Coming when called reliably ends the dog's freedom, so the recall is being punished by the loss of play: negative punishment of the very behavior she wants. The fix is the Premack principle in reverse of what she has been doing: most recalls should end with "go play," so that coming is the price of more freedom rather than the end of it. **A trainer applies steady leash pressure and releases it the instant the dog steps toward her. A student calls this punishment, since leash pressure is unpleasant. Which quadrant is it, and what is the real concern?** It is negative reinforcement: something the dog dislikes is removed after the behavior, and stepping toward the handler increases. "Negative" means removed, not bad. The real concern is that R− requires the aversive to be present first, which is where the welfare cost lies and why it sits on the same side of the asymmetry as P+. **A dog jumps on every guest, and every guest pushes it off with a firm "No." Months later it still jumps. What is maintaining the behavior?** Attention: eye contact, voices, and hands, and even a push-off is contact. The intended correction is functioning as positive reinforcement, which the behavior proves by continuing. The plan is management at the door, all attention stopping the instant paws leave the floor, a generous greeting for four-on-the-floor or a sit, and an extinction burst to be expected first. **An owner worries that reinforcing sits only some of the time will erode the behavior, so she keeps treating every one. Is her worry justified?** No; the opposite is true. Behavior maintained on an intermittent, variable-ratio schedule is more resistant to extinction than behavior reinforced every time, the partial-reinforcement extinction effect. Reinforcing every sit for years makes the visible treat part of the cue. The errors to avoid are thinning too fast, or letting the schedule thin to zero. **Explain it to a friend.** Explain why a clicker works, without using the words "reinforcer," "conditioned," or "bridge." ## Frequently asked questions **What are the four quadrants of dog training?** Positive reinforcement (add something the dog wants — a treat for a sit), negative reinforcement (remove something unpleasant — leash pressure released when the dog moves), positive punishment (add something unpleasant — a leash pop), and negative punishment (remove something the dog wants — turning away when it jumps). "Positive" and "negative" mean added and removed, not good and bad. **Is positive reinforcement dog training effective?** Yes. Owner surveys associate reward-based methods with higher obedience and fewer problem behaviors, a controlled comparison found reward-based training more effective than electronic-collar training for recall and sit, and no controlled study has found aversive methods to be more effective. Reward-based methods also avoid the stress and aggression associated with confrontational techniques. **Do I have to give my dog treats forever?** No. Reinforce every correct response while the dog is learning, then shift to reinforcing unpredictably — a variable-ratio schedule — while replacing many treats with life rewards such as play, sniffing, and going through doors. Intermittent reinforcement makes behavior more durable, not less. What you should never do is let reinforcement stop entirely. **Is negative reinforcement bad for dogs?** Negative reinforcement is not punishment — it strengthens behavior by removing something unpleasant. But it requires the unpleasant thing to be present first, which is where the welfare cost lies. Mild, brief pressure released the instant the dog responds is used by many trainers; methods built on shock, prong collars, or sustained discomfort are associated with stress indicators and are discouraged by veterinary behavior organizations. **Is the alpha or dominance theory of dog training true?** No. The "alpha wolf" idea came from unrelated captive wolves forced to live together. Wild wolf packs are families led by the parents, as L. David Mech, who helped popularize the term, later showed and retracted. Dogs are not wolves, and studies of dog behavior do not support the idea that misbehavior is a bid for rank. Problem behavior is almost always a reinforcement problem, not a status problem. **What is clicker training and how does it work?** Clicker training uses a distinct sound that has been paired with food until it becomes a conditioned reinforcer. The click marks the precise instant the dog does the right thing and bridges the gap until the treat arrives. It solves the timing problem — you can click within a fraction of a second even if the treat takes longer — and it lets you shape complex behaviors in small steps. **Why does my dog only listen at home?** Because that is where the behavior was trained. Dogs learn cues together with the context — the room, the person, the pouch, the absence of distractions — and do not generalize well on their own. Retrain each behavior in new places, with new people, at greater distances and with more distraction, raising one variable at a time and reinforcing generously as you go. **Should I punish my dog for growling?** No. A growl is information — the dog is telling you it is uncomfortable. Punishing it can suppress the warning without changing the feeling, producing a dog that bites without growling first. Move the dog away from what is bothering it, and consult a certified reward-based trainer or veterinary behaviorist to address the underlying fear or guarding. ## References 1. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 2. Hiby, E. F., Rooney, N. J., & Bradshaw, J. W. S. (2004). Dog training methods: Their use, effectiveness and interaction with behaviour and welfare. *Animal Welfare, 13*(1), 63–69. 3. Herron, M. E., Shofer, F. S., & Reisner, I. R. (2009). Survey of the use and outcome of confrontational and non-confrontational training methods in client-owned dogs showing undesired behaviors. *Applied Animal Behaviour Science, 117*(1–2), 47–54. 4. Ziv, G. (2017). The effects of using aversive training methods in dogs — A review. *Journal of Veterinary Behavior, 19*, 50–60. 5. Vieira de Castro, A. C., Fuchs, D., Morello, G. M., Pastur, S., de Sousa, L., & Olsson, I. A. S. (2020). Does training method matter? Evidence for the negative impact of aversive-based methods on companion dog welfare. *PLoS ONE, 15*(12), e0225023. 6. China, L., Mills, D. S., & Cooper, J. J. (2020). Efficacy of dog training with and without remote electronic collars vs. a focus on positive reinforcement. *Frontiers in Veterinary Science, 7*, 508. 7. American Veterinary Society of Animal Behavior. (2021). *Position Statement on Humane Dog Training*. AVSAB. 8. Mech, L. D. (1999). Alpha status, dominance, and division of labor in wolf packs. *Canadian Journal of Zoology, 77*(8), 1196–1203. 9. Mech, L. D. (2008). Whatever happened to the term alpha wolf? *International Wolf, 18*(4), 4–8. 10. Bradshaw, J. W. S., Blackwell, E. J., & Casey, R. A. (2009). Dominance in domestic dogs — useful construct or bad habit? *Journal of Veterinary Behavior, 4*(3), 135–144. 11. American Veterinary Society of Animal Behavior. (2008). *Position Statement on the Use of Dominance Theory in Behavior Modification of Animals*. AVSAB. 12. Williams, B. A. (1994). Conditioned reinforcement: Experimental and theoretical issues. *The Behavior Analyst, 17*(2), 261–285. 13. Breland, K., & Breland, M. (1951). A field of applied animal psychology. *American Psychologist, 6*(6), 202–204. 14. Skinner, B. F. (1951). How to teach animals. *Scientific American, 185*(6), 26–29. 15. Pryor, K. (1984). *Don't Shoot the Dog! The New Art of Teaching and Training*. Simon & Schuster. 16. Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139. 17. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 18. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 19. Stokes, T. F., & Baer, D. M. (1977). An implicit technology of generalization. *Journal of Applied Behavior Analysis, 10*(2), 349–367. 20. Lerman, D. C., & Iwata, B. A. (1995). Prevalence of the extinction burst and its attenuation during treatment. *Journal of Applied Behavior Analysis, 28*(1), 93–94. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The quadrant modern dog training is built on — timing, contingency, and reinforcer types. - [Shaping](https://operantconditioning.com/shaping/): Successive approximations and chaining — how complex behaviors are built one step at a time. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): When to thin the treats, and why intermittent reinforcement makes behavior last. --- # How to Build Habits With Operant Conditioning: A Science-Based, Seven-Step Protocol > How to build habits with operant conditioning: what habit formation psychology shows (66 days, context cues), a 7-step protocol, and how to break bad habits. - Source: https://operantconditioning.com/habits/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Apply it · Self-management* Seventy years of behavioral science and a decade of habit research agree on the mechanics. Here is the evidence, a seven-step protocol built on the antecedent–behavior–consequence loop, and what to do when it breaks. > **Definition** > > A **habit** is a behavior that has come under strong stimulus control: it is cued automatically by a context, performed with little deliberation, and relatively insensitive to what you happen to want in the moment.[1] > > In operant terms a habit is a well-worn [three-term contingency](https://operantconditioning.com/abc-model/) — a cue (antecedent), a response (behavior), and the history of reinforcement that built and maintains it. Building a habit with operant conditioning means arranging all three terms on purpose, which is what most habit advice — and most habit apps — leave to chance. **In brief** - A habit is a behavior under strong stimulus control: cued automatically by context, performed with little deliberation, and insensitive to momentary wants. - Motivation is not a term in the [contingency](https://operantconditioning.com/abc-model/); build on an unmissable cue, a tiny behavior, and a consequence that arrives within seconds. - Automaticity took a median of 66 days (range 18–254) in the best real-world study, and one missed day made no material difference. ## What habit formation psychology actually shows The most-cited study of real-world habit formation followed 96 volunteers who each chose a new eating, drinking, or activity behavior and tied it to a once-daily cue — for example, a piece of fruit with lunch or a run before dinner. Self-reported automaticity rose along an asymptotic curve: fast gains early, then a plateau.[2] 66 median days to reach peak automaticity 18–254 range across individuals and behaviors 1 missed day: no material effect on the process Three findings matter more than the headline number. The range was enormous — drinking a glass of water plateaued fast, exercise slowly. Missing a single opportunity did not measurably disrupt the curve. And the "21 days" figure that circulates everywhere appears nowhere in the data.[2] Wendy Wood's research adds the mechanism: habits are cued by **context**, and they persist on context rather than on intention. Students who transferred universities kept their exercise, reading, and TV habits only when the new environment resembled the old one.[3] Habitual cinema popcorn-eaters ate stale, week-old popcorn as readily as fresh — in a cinema. In a meeting room, taste took over.[4] So a stable cue is not optional, and habits change most easily when context is already disrupted — a move, a new job, a new term. Temporal landmarks work the same way. Katy Milkman and colleagues found that searches for "diet," gym visits, and goal commitments all spike at the start of a week, a month, a year, and after birthdays — the **fresh start effect**.[5] A landmark separates the old self from a new one and changes the antecedent conditions under which the behavior is attempted. Use one, but do not wait for one. Finally, habit researchers measure habit as **automaticity** — behavior that is efficient, unintentional, and hard to control — rather than as frequency, because a behavior you do daily through gritted teeth is not yet a habit.[6] The test is not "did I do it?" but "did I have to decide to?" ## Why willpower and motivation are the wrong frame Notice what is missing from the three-term contingency: motivation. It is not a variable in the loop. The nearest thing to it — the [motivating operation](https://operantconditioning.com/abc-model/) — is deprivation or satiation, which changes how much a reinforcer is worth and which fluctuates hour to hour. Building a routine on how much you want it today is building on the one term guaranteed to be different tomorrow. Willpower fares no better. The influential "ego depletion" studies of the late 1990s reported that self-control is a limited resource that runs down with use; a preregistered replication across 23 laboratories in 2016 found an effect indistinguishable from zero.[7][8] Whatever willpower is, its footing is contested enough that no protocol should depend on it. The more robust finding cuts the other way: in experience-sampling studies, people high in self-control do not report resisting more temptations — they report *fewer* — and their advantage in life outcomes is carried largely by beneficial habits, because they have arranged their environments and routines so that the desired behavior runs automatically.[9] Self-control, in practice, is antecedent control. > **The reframe** > > You are not trying to want it more. You are trying to arrange a cue that is unmissable, a behavior that is small enough to be emitted, and a consequence that arrives fast enough to count. Motivation is what you feel while the arrangement is bad. ## The operant protocol for building a habit The seven steps below are the three-term contingency applied in order, with the schedule and failure-planning that the laboratory and the habit literature both insist on. 1. **Pick the behavior and make it tiny.** "Exercise" is an outcome; "put on running shoes and step outside" is a behavior. BJ Fogg's Tiny Habits method starts with a version that takes under a minute — two push-ups, one sentence, flossing one tooth — which is [shaping](https://operantconditioning.com/shaping/)'s first approximation by another name.[10] A tiny behavior gets emitted, and only emitted behavior can be reinforced. Once it is automatic, grow it. 2. **Choose a stable antecedent.** Anchor the new behavior to something that already happens reliably every day: after I pour the coffee, when I sit down at the desk. State it as an implementation intention — "when X happens, I will do Y" — which links cue to response in advance; a meta-analysis of 94 studies found a medium-to-large effect on goal attainment.[11] Then make the cue physically present: shoes by the door, book on the pillow, phone charging in the kitchen. Do not use a notification as the cue. It habituates, many people disable it, and it cues picking up the phone rather than the behavior you want. 3. **Arrange an immediate consequence.** The natural consequences of good habits arrive in weeks (fitness) or years (health), which is precisely why they fail to control behavior. So arrange one yourself, within seconds of finishing. Marking the behavior done works as a [conditioned reinforcer](https://operantconditioning.com/positive-reinforcement/) once it is paired with something that matters — visible progress, a small pleasure you allow only afterward, a message to someone watching. **Temptation bundling** pairs the behavior with something you already crave: gym-goers given audiobooks they could only hear at the gym went more often.[12] That is the **[Premack principle](https://operantconditioning.com/premack-principle/)** — a higher-probability behavior reinforces a lower-probability one — in a pair of headphones.[13] 4. **Reinforce every time at first.** During acquisition use continuous reinforcement: the consequence follows every occurrence. It is the fastest way to build a behavior, and it is where self-managers are too stingy, saving the reward for a "real" workout and never reinforcing the tiny one that had to come first.[14] 5. **Thin to an intermittent schedule once the behavior is stable.** Continuous reinforcement builds fragile behavior: stop the consequence and it extinguishes quickly. After a few weeks of reliable performance, shift to reinforcing unpredictably — a variable-ratio schedule — for the steadiest responding and the greatest resistance to [extinction](https://operantconditioning.com/extinction/). [How schedules work, with a simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) 6. **Track the behavior, not the outcome.** You cannot reinforce "lose ten pounds" on a Tuesday; you can reinforce "walked after lunch." Recording is itself an intervention: **self-monitoring is reactive** — observing and writing down your own behavior changes it, usually in the desired direction.[15] A meta-analysis of 138 experiments found that prompting people to monitor progress reliably improved goal attainment, more so when progress was physically recorded and reported to someone else.[16] 7. **Plan for extinction bursts, lapses, and resurgence.** When a new behavior stops being reinforced — you get sick, you travel — the old one it replaced tends to return; behavior analysts call this [resurgence](https://operantconditioning.com/extinction/), and it is the mechanism of most relapse. A lapse is not a failure; the 66-day data showed a missed day barely registered.[2] What turns a lapse into a collapse is the **abstinence violation effect**: the "I've blown it" judgment that follows one slip and licenses the next.[17] Decide the recovery response in advance — "if I miss a day, I do the tiny version tomorrow, no catching up" — and reinforce that. ### Writing the contingency down Whatever you use to track a habit — a notebook, a spreadsheet, an app — write all three terms for each habit: an antecedent (a real-world cue such as a time, a place, or a preceding routine), a behavior scoped as small as it needs to be, and a consequence you will actually deliver the moment the behavior is done. A record of only the middle term is a to-do list, not a contingency, and the value of writing the other two down is that it refuses to let you skip them. > **The app: Operant runs this loop for you.** A habit app for iPhone and Apple Watch from the publisher of this site. Each habit is set up as an antecedent, a behavior and a consequence — the three terms, not just the middle one. Free to download and try; a subscription unlocks the full app. [About the app](https://operantconditioning.com/app/) · [Download on the App Store](https://apps.apple.com/us/app/operant-behavior-change-app/id6802081776) ## Self-management in behavior analysis: what Skinner actually said Skinner devoted a chapter of *Science and Human Behavior* to self-control, and his position was simple: a person controls their own behavior the same way they control anyone else's — by manipulating the variables of which it is a function. One response (the *controlling* response) alters the conditions under which another (the *controlled* response) occurs.[18] His catalogue reads like a habit book written in 1953: - **Physical restraint and physical aid.** Walk out of the room; put the phone in a drawer; lay out the equipment the behavior needs. - **Changing the stimulus.** Remove the cues for the unwanted behavior and add cues for the wanted one — the whole of "environment design." - **Deprivation and satiation.** Eat before the party; build up an appetite for the reinforcer you plan to use. - **Manipulating emotional conditions and using aversive stimulation.** Set an alarm; make a public commitment you would be embarrassed to break. - **Operant conditioning and punishment of one's own behavior.** Self-administered reinforcers and penalties. - **Doing something else.** Emit an incompatible behavior — the seed of differential reinforcement. The last two items raised a debate that is still open: can you really reinforce yourself? Charles Catania argued in 1975 that a reinforcer you can take at any moment is not contingent on anything, so "self-reinforcement" is a misnomer for what is really rule-following.[19] A 1985 experiment sharpened the point: the benefit of a self-reward procedure disappeared when participants set their goals privately rather than publicly, suggesting the active ingredient was the social contingency.[20] The practical lesson survives whichever side wins: self-administered consequences work best when they are reliably withheld until the behavior occurs, externalized in something you cannot quietly waive — an app, a partner, a deposit — and backed by a social or financial contingency. Richard Malott's framing is the most useful: the natural consequences of most habits are too small, too delayed, or too improbable to control behavior, so the job of self-management is to add consequences that are sizable, immediate, and probable.[21] ## How to break a bad habit with operant conditioning A bad habit is a good contingency working for the wrong behavior. The steps mirror building one, in reverse. 1. **Identify the maintaining reinforcer.** Keep an [ABC record](https://operantconditioning.com/abc-model/) for three days. What does the behavior get (stimulation, food, attention) or escape (boredom, anxiety, an unpleasant task)? A habit maintained by escape needs a different replacement than one maintained by novelty. 2. **Make the cue unavailable.** The cheapest intervention is on the antecedent. No phone in the bedroom; no snacks on the counter; a route home that does not pass the bakery. Wood's popcorn study made the point dramatically: disrupting the habitual motor pattern — eating with the non-dominant hand — was enough to bring the behavior back under the control of taste.[4] 3. **Add friction and response cost.** Log out after every session so the feed opens to a password screen. Delete the app and reinstall it when you actually want it. Small increases in effort produce large decreases in a behavior that runs on automaticity, because automatic behavior stops at the first obstacle that requires a decision. 4. **Reinforce an alternative that serves the same function.** This is **differential reinforcement of alternative behavior**, and it is the step people skip. If scrolling was escape from boredom, queue a podcast; if the evening drink marked the boundary between work and rest, replace it with a walk that marks the same boundary. Remove a behavior without replacing it and the function goes unmet, and the old behavior resurges. 5. **Expect the burst and the recovery.** Withholding a reinforcer produces a temporary spike — the [extinction burst](https://operantconditioning.com/extinction/) — and an extinguished behavior can reappear after time away. Neither means the method failed. Giving in during the burst, however, reinforces a stronger version of the habit on an intermittent schedule, the worst of all outcomes. > **Do not rely on self-punishment** > > A penalty you administer to yourself is a penalty you can waive, and after the first waiver it is a penalty in name only. If you want an aversive consequence in the loop, hand it to a third party — a deposit contract, a friend who collects — so that the contingency is real. Even then, use it alongside reinforcement of the replacement behavior, not instead of it. [Why punishment is a poor first choice ›](https://operantconditioning.com/positive-punishment/) ## Common goals translated into tiny behaviors, antecedents, and consequences | Goal (outcome) | Tiny behavior (B) | Antecedent (A) | Immediate consequence (C) | | --- | --- | --- | --- | | Get fit | Put on running shoes and step outside | After the morning coffee is poured; shoes by the door | Mark it done; the podcast you only play while moving starts | | Read more | Read one page | When you get into bed; book on the pillow, phone in the kitchen | Mark it done; a moved bookmark is visible progress | | Meditate | Sit and take three slow breaths | After brushing teeth; cushion visible from the sink | Mark it done; the first sip of coffee follows the breaths | | Floss | Floss one tooth (the rest usually follows) | After putting the toothbrush down; floss pick on the brush | Mark it done; say "done" out loud — Fogg's celebration | | Journal | Write one sentence | After closing the laptop for the day; notebook open on the desk | Mark it done; the tea you make only after the sentence | | Use the phone less | Dock the phone at the kitchen charger | When you start the dishwasher after dinner | Mark it done; the evening show starts only once the phone is docked | | Study | Open the notes and answer one flashcard | When you sit down after the 4 p.m. class; deck open on the desk | Mark it done; a text to a study partner who replies | | Drink more water | Drink one glass | When the kettle is switched on; glass kept beside it | Mark it done; the coffee comes after the water | Notice the pattern in the consequence column. Almost every entry uses the Premack principle: a thing you were going to do anyway (the coffee, the show, the podcast) is made contingent on the tiny behavior. That costs nothing, it is immediate, and it is a reinforcer you already know works because you already do it. ## Commitment devices and behavioral economics A **commitment device** is an arrangement your present self makes to constrain your future self: a deadline you cannot move, money you forfeit if you fail, a membership that charges you whether or not you go. In operant terms it converts a consequence that is small, delayed, and cumulative into one that is large, certain, and near — Malott's prescription, implemented through a third party. The classic demonstration is Dan Ariely and Klaus Wertenbroch's deadline experiment. Students allowed to set their own binding deadlines for three papers set them earlier than they had to and performed better than students with a single end-of-term deadline — but worse than students given evenly spaced deadlines by the instructor. People know they procrastinate and will precommit to fight it; they just do it imperfectly.[22] Deposit contracts exploit **loss aversion**, the finding from prospect theory that a loss looms larger than an equivalent gain.[23] In a 16-week randomized trial, obese adults who put their own money at risk — refunded, with a match, only if they hit monthly weight targets — lost roughly three times as much weight as a control group given the same goals and weigh-ins; a lottery-incentive group did about as well. Much of the weight returned after the incentives ended, which is not an argument against the contingency so much as a demonstration of it: consequences control behavior while they are in force — the same pattern seen in [contingency management](https://operantconditioning.com/applications/#health) for addiction.[24] Plan the maintenance schedule before the acquisition schedule runs out. Every device on this list puts a consequence out of reach of your own leniency. That is the honest reason a partner, a coach, or an app that records the miss can outperform sheer resolve: none of them can be talked out of it at 10 p.m. ## Key takeaways - A habit is a well-worn three-term contingency: a cue, a response, and the reinforcement history that built it. Habits persist on context rather than intention, so a stable cue is not optional, and they change most easily when context is already disrupted. - Motivation and willpower are the wrong frame. Motivation is not a variable in the loop, ego depletion failed a 23-laboratory replication, and people high in self-control report fewer temptations rather than more resistance; self-control in practice is antecedent control. - The protocol runs the contingency in order: make the behavior tiny, anchor it to a stable cue stated as an implementation intention, arrange an immediate consequence (the Premack principle costs nothing), reinforce every time at first, then thin to an intermittent schedule, and track the behavior rather than the outcome. - Plan for lapses. A missed day barely registers; what turns a lapse into a collapse is the abstinence violation effect, so decide the recovery response in advance. When a new behavior stops being reinforced, the old one it replaced tends to resurge. - Breaking a habit mirrors building one: find the maintaining reinforcer, make the cue unavailable, add friction, and reinforce an alternative that serves the same function. Self-administered consequences work only when they cannot be quietly waived, which is what commitment devices and deposit contracts are for. ### Check yourself **Someone decides her morning coffee will be the reward for her morning run. On days she skips the run, she drinks the coffee anyway. Is the coffee reinforcing the run?** No. A reinforcer you can take at any moment is not contingent on anything, which was Catania's objection to "self-reinforcement." The coffee only works as a Premack consequence if it comes after the run and not otherwise; self-administered consequences work when they are withheld until the behavior occurs and externalized in something that cannot be quietly waived. **A person has reinforced a new habit every single time for six weeks and plans to keep doing so, reasoning that continuous reinforcement builds the strongest behavior. What does the schedule research say?** Continuous reinforcement is the fastest way to build a behavior, but it builds fragile behavior: stop the consequence and it extinguishes quickly. Once performance is stable, the protocol thins to an intermittent, variable-ratio schedule, which produces the steadiest responding and the greatest resistance to extinction. **Someone has done a behavior every day for two months but still has to force herself each time. Is it a habit yet?** Not by the researchers' definition. Habit is measured as automaticity, behavior that is efficient, unintentional, and hard to control, not as frequency; a behavior done daily through gritted teeth is not yet a habit. The test is not "did I do it?" but "did I have to decide to?" **A person who stopped late-night scrolling by charging the phone in the kitchen goes on a work trip, and in the hotel the scrolling returns. Did the method fail?** No. Habits are cued by context, and the antecedent arrangement that had been doing the work stayed at home; when a new behavior stops being reinforced, the old one it replaced resurges, which is the mechanism of most relapse. A lapse is not a collapse unless the "I've blown it" judgment makes it one, so the pre-decided recovery response, the tiny version the next day, is the behavior to reinforce. **Explain it to a friend.** Explain why a habit that keeps failing is usually an arrangement problem, using one habit you have tried and dropped, and without using the words "motivation" or "willpower." ## Frequently asked questions **How long does it take to form a habit?** In the best real-world study, the median was 66 days to reach peak automaticity, with a range of 18 to 254 days depending on the person and the behavior. Simple behaviors such as drinking a glass of water became automatic fastest; exercise took longest. The popular "21 days" figure has no basis in that data. **Does missing one day ruin a habit?** No. Missing a single opportunity had no material effect on habit formation in the 66-day study. What damages a habit is the "I've blown it" reaction that turns one lapse into a week of them. Decide in advance that after a miss you do the tiny version the next day, and treat that recovery as the behavior to reinforce. **What is the best reinforcer for building a habit?** One that is immediate, that you actually care about, and that you can reliably withhold until the behavior happens. The Premack principle is the easiest source: make something you already do daily (the coffee, the show, the podcast) contingent on the tiny behavior. A check-off works as a conditioned reinforcer once it is paired with visible progress. **Why don't habit-app notifications work as cues?** Three reasons. Repeated notifications habituate, so they stop grabbing attention. Many people disable them. And a notification is a cue for picking up the phone, not for the behavior you want, so it puts phone-checking under stimulus control instead. A stable real-world cue — a time, a place, a routine you already have — is a far stronger antecedent. **Can you really reinforce yourself?** Behavior analysts have argued about this since the 1970s. A reward you can take any time is not truly contingent, and experiments suggest public goal-setting often does the real work. In practice, self-administered consequences work when they are withheld until the behavior occurs, externalized in something you cannot quietly waive (an app, a partner, a deposit), and backed by a social or financial contingency. **How do you break a bad habit with operant conditioning?** Find what is reinforcing it, remove or change its cue, add friction so it can no longer run automatically, and reinforce an alternative behavior that serves the same function. Expect a temporary increase (the extinction burst) and occasional reappearances; neither means the method is failing. **What is habit stacking, and does it work?** Habit stacking (BJ Fogg calls it anchoring) attaches a new behavior to an existing routine: "after I pour my coffee, I will write one sentence." It works because the existing routine is a stable, reliable antecedent, and because stating the plan as an if-then implementation intention links cue to response in advance. It is antecedent control in plain language. **Are streaks a good idea?** A streak turns a habit into an avoidance contingency: the reinforcer becomes not losing the count. That can help while the streak is intact and hurt badly when it breaks, because the loss often triggers the abstinence violation effect. If you use streaks, decide in advance that a miss resets nothing but the number, and reinforce the return rather than mourning the run. ## References 1. Wood, W., & Rünger, D. (2016). Psychology of habit. *Annual Review of Psychology, 67*, 289–314. 2. Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. *European Journal of Social Psychology, 40*(6), 998–1009. 3. Wood, W., Tam, L., & Witt, M. G. (2005). Changing circumstances, disrupting habits. *Journal of Personality and Social Psychology, 88*(6), 918–933. 4. Neal, D. T., Wood, W., Wu, M., & Kurlander, D. (2011). The pull of the past: When do habits persist despite conflict with motives? *Personality and Social Psychology Bulletin, 37*(11), 1428–1437. 5. Dai, H., Milkman, K. L., & Riis, J. (2014). The fresh start effect: Temporal landmarks motivate aspirational behavior. *Management Science, 60*(10), 2563–2582. 6. Gardner, B. (2015). A review and analysis of the use of 'habit' in understanding, predicting and influencing health-related behaviour. *Health Psychology Review, 9*(3), 277–295. 7. Baumeister, R. F., Bratslavsky, E., Muraven, M., & Tice, D. M. (1998). Ego depletion: Is the active self a limited resource? *Journal of Personality and Social Psychology, 74*(5), 1252–1265. 8. Hagger, M. S., Chatzisarantis, N. L. D., Alberts, H., et al. (2016). A multilab preregistered replication of the ego-depletion effect. *Perspectives on Psychological Science, 11*(4), 546–573. 9. Hofmann, W., Baumeister, R. F., Förster, G., & Vohs, K. D. (2012). Everyday temptations: An experience sampling study of desire, conflict, and self-control. *Journal of Personality and Social Psychology, 102*(6), 1318–1335. See also Galla, B. M., & Duckworth, A. L. (2015). More than resisting temptation: Beneficial habits mediate the relationship between self-control and positive life outcomes. *Journal of Personality and Social Psychology, 109*(3), 508–525. 10. Fogg, B. J. (2020). *Tiny Habits: The Small Changes That Change Everything*. Houghton Mifflin Harcourt. 11. Gollwitzer, P. M. (1999). Implementation intentions: Strong effects of simple plans. *American Psychologist, 54*(7), 493–503. See also Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. *Advances in Experimental Social Psychology, 38*, 69–119. 12. Milkman, K. L., Minson, J. A., & Volpp, K. G. M. (2014). Holding the Hunger Games hostage at the gym: An evaluation of temptation bundling. *Management Science, 60*(2), 283–299. 13. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 14. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 15. Nelson, R. O., & Hayes, S. C. (1981). Theoretical explanations for reactivity in self-monitoring. *Behavior Modification, 5*(1), 3–14. 16. Harkin, B., Webb, T. L., Chang, B. P. I., et al. (2016). Does monitoring goal progress promote goal attainment? A meta-analysis of the experimental evidence. *Psychological Bulletin, 142*(2), 198–229. 17. Marlatt, G. A., & Gordon, J. R. (Eds.). (1985). *Relapse Prevention: Maintenance Strategies in the Treatment of Addictive Behaviors*. Guilford Press. 18. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. (Chapter 15, "Self-control.") 19. Catania, A. C. (1975). The myth of self-reinforcement. *Behaviorism, 3*(2), 192–199. 20. Hayes, S. C., Rosenfarb, I., Wulfert, E., Munt, E. D., Korn, Z., & Zettle, R. D. (1985). Self-reinforcement effects: An artifact of social standard setting? *Journal of Applied Behavior Analysis, 18*(3), 201–214. 21. Malott, R. W. (1989). The achievement of evasive goals: Control by rules describing contingencies that are not direct acting. In S. C. Hayes (Ed.), *Rule-Governed Behavior: Cognition, Contingencies, and Instructional Control* (pp. 269–322). Plenum. 22. Ariely, D., & Wertenbroch, K. (2002). Procrastination, deadlines, and performance: Self-control by precommitment. *Psychological Science, 13*(3), 219–224. 23. Kahneman, D., & Tversky, A. (1979). Prospect theory: An analysis of decision under risk. *Econometrica, 47*(2), 263–291. 24. Volpp, K. G., John, L. K., Troxel, A. B., Norton, L., Fassbender, J., & Loewenstein, G. (2008). Financial incentive-based approaches for weight loss: A randomized trial. *JAMA, 300*(22), 2631–2637. ## Related - [The ABC model](https://operantconditioning.com/abc-model/): Antecedent, behavior, consequence — the loop every habit runs on. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): When to reinforce every time and when to thin — with a live simulator. - [Extinction](https://operantconditioning.com/extinction/): Bursts, resurgence, and why relapse is predictable. --- # 50+ Operant Conditioning Examples: Everyday Life, Classroom, Work, and Animals > 50+ operant conditioning examples from everyday life, the classroom, work, and animal training — each labeled by quadrant, plus schedules and tricky cases. - Source: https://operantconditioning.com/examples/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Examples* More than fifty examples of operant conditioning in everyday life, at home, in the classroom, at work, in animal training, in apps, and in nature — each one labeled by quadrant and explained, plus the schedules and the tricky cases that fool people. > **The four quadrants, in two sentences** > > **Operant conditioning** is learning from consequences: a behavior that is followed by [positive reinforcement](https://operantconditioning.com/positive-reinforcement/) (something added) or [negative reinforcement](https://operantconditioning.com/negative-reinforcement/) (something removed) becomes more frequent, and a behavior followed by [positive punishment](https://operantconditioning.com/positive-punishment/) (something added) or [negative punishment](https://operantconditioning.com/negative-punishment/) (something removed) becomes less frequent. "Positive" and "negative" mean added and removed, not good and bad, and a consequence counts as reinforcement or punishment only by its actual effect on the behavior.[1] ## How to read an example of operant conditioning Every example on this page has the same three parts, the [A-B-C](https://operantconditioning.com/abc-model/) of behavior analysis: an **antecedent** (the situation), a **behavior** (what the person or animal does), and a **consequence** (what happens next). To classify the consequence, ask two questions in order: 1. **Did the behavior become more or less likely afterward?** More likely means reinforcement. Less likely means punishment. 2. **Was a stimulus added or removed?** Added means "positive." Removed means "negative."  *The four quadrants. The column answers "added or removed?"; the row answers "more or less likely?" Positive and negative are arithmetic signs, not judgments.* Two more labels cover what the quadrants leave out. Extinction means a behavior that used to be reinforced no longer is, and it fades.[11] Schedule marks examples where the interesting part is not *what* the consequence is but *how often* it arrives.[2] > **The label depends on the effect, not the intention** > > Every row below states the effect on behavior ("barking increases," "swearing decreases") because the label is only correct if that effect actually happens. A scolding that makes a child act out *more* is reinforcement, whatever the parent meant by it. When you analyze your own examples, watch the behavior over time before you name the quadrant.[1] ## Examples of operant conditioning in everyday life | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | You hold the door for a stranger | They smile and thank you; you hold doors more | Positive reinforcement | Social approval is added; the behavior increases. | | You try a new restaurant | The meal is excellent; you go back | Positive reinforcement | A good meal is added after the choice; returning increases. | | You tell a joke at dinner | Everyone laughs; you tell more jokes | Positive reinforcement | Laughter is added; joke-telling increases. | | You put on noise-cancelling headphones on the train | The noise disappears; you reach for them every ride | Negative reinforcement | An aversive stimulus (noise) is removed; the behavior increases. This is escape. | | You feed the parking meter | No ticket; feeding the meter becomes automatic | Negative reinforcement | The behavior prevents an aversive event. This is avoidance — the ticket never has to happen for the habit to hold. | | You mute a group chat that buzzes constantly | The buzzing stops; you mute chats faster in future | Negative reinforcement | An aversive stimulus is removed contingent on the behavior; muting increases. | | You skip sunscreen at the beach | Painful sunburn; you skip it less often | Positive punishment | An aversive stimulus is added; the behavior decreases. A natural punisher — no one had to deliver it. | | You text while walking | You walk into a lamppost; you text-and-walk less | Positive punishment | Pain is added; the behavior decreases. | | You leave your bike unlocked | The bike is stolen; you never leave one unlocked again | Negative punishment | A valued item is removed; the behavior decreases. One trial was enough. | | You show up late to a friend's dinners, repeatedly | The invitations stop; you become punctual with other friends | Negative punishment | A reinforcer (invitations) is withdrawn contingent on lateness; lateness decreases. | | You wave at a neighbor who never waves back | Nothing; after a few weeks you stop waving | Extinction | The behavior used to produce a wave back. With the reinforcer gone, it fades. | | You buy a lottery ticket every week | Occasional small wins; you keep buying | Schedule · VR | Reinforcement after an unpredictable number of responses — variable ratio, the most persistent pattern of all. | Notice how many everyday examples are natural consequences that nobody arranged. Sunburn, a stolen bike, and a good meal shape behavior exactly the way a food pellet shapes a rat's lever press. [The complete guide to operant conditioning ›](https://operantconditioning.com/) ## Examples at home and in parenting | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Child shares a toy with a sibling | Parent notices and says exactly what was good; sharing increases | Positive reinforcement | Specific attention is added right after the behavior. "Catch them being good." | | Child brushes teeth without a fuss | Then comes the bedtime story; brushing gets easier | Positive reinforcement | A preferred activity follows a less-preferred one — the Premack principle.[3] | | Teenager texts "home safe" | Parent's stream of check-in calls stops; texting increases | Negative reinforcement | The aversive stimulus (repeated calls) is removed; the behavior increases. | | Child eats the agreed three bites of broccoli | Excused from the table; bites happen faster | Negative reinforcement | Sitting at the table is aversive; the behavior ends it. Escape. | | Child whines for a snack before dinner | Parent gives in "just this once" — every few days; whining intensifies | Positive reinforcement | The snack is added, on a variable-ratio schedule. Giving in occasionally builds more persistent whining than giving in every time.[2] | | Child rides a bike without a helmet | The bike is put away for the rest of the day; helmet-less riding decreases | Negative punishment | A reinforcer (the bike) is removed; the behavior decreases. | | Child swears at dinner | A brief, sharp reprimand; swearing at dinner decreases | Positive punishment | An aversive stimulus is added; the behavior decreases. If swearing had *increased*, the reprimand would be attention — see the tricky cases below. | | Child calls out from bed for a fifth glass of water | Parent answers the first request only, then stops responding; after a noisy week the calling fades | Extinction | A behavior maintained by attention no longer produces it. Expect a burst before the decline.[4] | Families are full of *reciprocal* contingencies: the child's behavior is shaped by the parent's response, and the parent's response is shaped by what stops the child. Gerald Patterson called the escalating version of this the coercive family process.[5] [Operant conditioning in parenting ›](https://operantconditioning.com/applications/#parenting) ## Operant conditioning examples in the classroom | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Student hands in homework on time | Written feedback returned the next day; on-time work increases | Positive reinforcement | Feedback is added soon after the behavior. Prompt feedback reinforces; feedback three weeks later mostly doesn't. | | Class lines up quietly after recess | Teacher adds a marble to the jar; a full jar earns a class party | Positive reinforcement | A token — a conditioned, generalized reinforcer — is added. The basis of every token economy.[6] | | Student finishes every problem in class | Tonight's homework is waived; finishing in class increases | Negative reinforcement | An aversive task is removed contingent on the behavior. "Homework passes" work this way. | | Student who dreads reading aloud acts silly when called on | Teacher sighs and moves to the next student; acting silly increases | Negative reinforcement | The demand is removed. This is escape-maintained problem behavior, and it is extremely common.[7] | | Student turns in sloppy, illegible work | Has to redo it during free time; sloppy work decreases | Positive punishment | Extra effort is added contingent on the behavior; the behavior decreases. | | Student uses a phone during a lesson | Phone is kept at the desk until the bell; phone use decreases | Negative punishment | A reinforcer is removed for a set time; the behavior decreases. | | Student makes wisecracks that used to get laughs | Classmates stop laughing; the wisecracks dry up | Extinction | The reinforcer (peer attention) is no longer delivered. Peer attention is often the reinforcer teachers can't control. | | Students stay on task during independent work | Teacher circulates and gives praise at unpredictable moments | Schedule · VI | The first on-task moment after an unpredictable interval is reinforced — variable interval, which produces steady, sustained work. | Classroom studies in the 1960s established that teacher attention is a powerful reinforcer, that praise for on-task behavior plus planned ignoring of mild disruption reduces disruption, and that reprimands can backfire when the attention they carry is what the student is working for.[8] [Operant conditioning in education ›](https://operantconditioning.com/applications/#education) ## Examples at work | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Employee proposes a process fix | Manager credits her by name in the team meeting; proposals increase | Positive reinforcement | Public recognition is added. For many people it outperforms a small bonus. | | Employee completes the mandatory compliance module | The nagging pop-up at every login disappears; completion happens sooner | Negative reinforcement | An aversive stimulus is removed. Most corporate "reminders" are negative-reinforcement systems. | | New hire offers ideas in meetings | Nobody responds; after a month the ideas stop | Extinction | No reinforcer follows the behavior, so it fades. Silence trains people as surely as criticism does. | | Employee interrupts colleagues | A colleague calls it out in front of the group; interruptions drop | Positive punishment | An aversive stimulus is added; the behavior decreases — with the usual side effect of resentment toward the punisher. | | Employee expenses a personal dinner | Loses bonus eligibility for the quarter; padding expenses stops | Negative punishment | A reinforcer is removed contingent on the behavior. This is response cost. | | Worker assembles units on a production line | Paid a fixed amount per unit | Schedule · FR | Piece-rate pay is a fixed-ratio schedule: high rates of work with a pause after each payout. | [How organizational behavior management applies these ideas ›](https://operantconditioning.com/applications/#workplace) ## Animal training examples | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Dolphin touches a target pole | Whistle, then a fish; target-touching increases | Positive reinforcement | The whistle is a conditioned reinforcer that bridges the gap to the fish. Same logic as a clicker. | | Horse steps forward | Rider's leg pressure is released; the horse moves off pressure more readily | Negative reinforcement | Pressure is removed contingent on the behavior. Most traditional horsemanship is negative reinforcement. | | Cat jumps onto the kitchen counter | A motion-triggered puff of air; counter-jumping decreases | Positive punishment | An aversive stimulus is added. Because the device delivers it, the cat doesn't learn to avoid the owner. | | Parrot screams for attention | Owner leaves the room; screaming decreases | Negative punishment | The reinforcer (the owner's presence) is removed contingent on the behavior. | | Rat presses a lever that used to deliver food | Food no longer comes; pressing surges, then fades | Extinction | The laboratory original. The surge is the extinction burst.[4] | | Pigeon pecks a lighted key | Food after an average of 50 pecks | Schedule · VR | Variable ratio: pigeons on this schedule peck at very high, steady rates for long stretches.[2] | Modern dog training leans almost entirely on the first quadrant, for reasons the evidence supports. [Operant conditioning in dog training ›](https://operantconditioning.com/dog-training/) ## Examples in technology and apps | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | You complete a lesson in a language app | Chime, confetti, points; lesson completion increases | Positive reinforcement | Conditioned reinforcers are added immediately — the timing is the whole trick. | | You open a loot box in a game | Occasionally a rare item; opening boxes increases | Schedule · VR | Variable ratio, same as a slot machine. This is why regulators treat loot boxes as gambling-adjacent. | | You pay a bill in the banking app | The red "overdue" badge disappears; you pay sooner next month | Negative reinforcement | An aversive stimulus is removed contingent on the behavior. | | You hit "reply all" on a company-wide email | Two hundred "please remove me" replies; you never reply-all again | Positive punishment | An aversive flood is added; the behavior decreases sharply. | | You post spam in a forum | Post deleted and account suspended for 24 hours; spamming decreases | Negative punishment | Access (a reinforcer) is removed for a period. Time-out, in software. | | You open an app's notifications | They never contain anything useful; you start swiping them away unread | Extinction | Opening was reinforced by relevant content. Without it, the behavior fades — which is why irrelevant notifications destroy an app's reach. | [Operant design in technology ›](https://operantconditioning.com/applications/#technology) ## Examples in health and therapy | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Patient in addiction treatment submits a drug-negative urine sample | Voucher whose value rises with each consecutive negative sample; abstinence increases | Positive reinforcement | Contingency management — one of the best-supported treatments for stimulant use disorders.[9] | | Person afraid of flying drives twelve hours instead | The dread evaporates; driving instead of flying becomes the rule | Negative reinforcement | Avoidance removes fear and is reinforced by that relief. The fear itself never gets tested, so it never fades. | | Client in exposure therapy stays in the feared situation | Nothing bad happens; fear declines across sessions | Extinction | The avoidance response is blocked, and the conditioned fear extinguishes. [More on extinction ›](https://operantconditioning.com/extinction/) | | Patient takes a new medication | Nausea within the hour; doses get skipped | Positive punishment | An aversive stimulus is added after the behavior; taking the pill decreases. A major, under-recognized cause of non-adherence. | | Patient on a hospital ward shouts at staff | Loses tokens that buy privileges; shouting decreases | Negative punishment | Response cost within a token economy.[6] | [Operant conditioning in health and clinical practice ›](https://operantconditioning.com/applications/#health) ## Examples in nature and evolution | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Crow drops a walnut onto a road | A passing car cracks it open; dropping nuts on roads increases | Positive reinforcement | Food access is added after the behavior. No trainer required. | | Toad snaps at a bumblebee | Stung on the tongue; snapping at striped insects decreases | Positive punishment | Pain is added; the behavior decreases, often after a single trial. | | Gull begs at a picnic table where nobody feeds it | Nothing; it stops coming to that table | Extinction | Begging that was reinforced at other tables is not reinforced here, so it fades — and comes under stimulus control of the tables that do pay. | Skinner argued that operant conditioning is a second kind of selection: natural selection picks traits across generations, and reinforcement picks behaviors within a lifetime.[10] Both work by consequences, and neither needs a plan. ## Examples for yourself: self-management | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | You write 200 words | You cross off today's box on a paper tracker; writing sessions increase | Positive reinforcement | A conditioned reinforcer, delivered immediately by you. Small and instant beats big and delayed. | | You wash the dishes right after dinner | No crusted pile waiting in the morning; the habit sticks | Negative reinforcement | The behavior prevents an aversive state — avoidance, working in your favor. | | You eat a huge lunch | A 3 p.m. slump; big lunches decrease, slowly | Positive punishment | An aversive state is added — but two hours late, which is why natural punishers for eating are so weak. | | You miss a goal in a commitment contract | Money you pledged goes to a cause you dislike; missing decreases | Negative punishment | Response cost, arranged in advance by you against your future self. | | You get a generic "time to work out!" notification | Nothing happens whether you obey it or not; within two weeks you swipe it away unread | Extinction | A cue with no contingent consequence loses control over behavior. Most habit-app reminders die this way. | [The full guide to building habits with operant conditioning ›](https://operantconditioning.com/habits/) ## Examples of schedules of reinforcement in daily life Ferster and Skinner spent a decade cataloguing what happens when reinforcement arrives only some of the time.[2] Four schedules cover most of daily life. | Schedule | Reinforcer arrives… | Three everyday examples | Characteristic pattern | | --- | --- | --- | --- | | **Fixed ratio (FR)** | after a set number of responses | A coffee card stamped every purchase, free on the tenth · Piece-rate pay for each unit sewn · "After every 25 flashcards, a five-minute break" (self-set) | Fast, steady work with a pause right after each payout — the "just got my free coffee, no hurry" lull. | | **Variable ratio (VR)** | after an unpredictable number of responses | A slot machine · Cold-call sales, where roughly one call in thirty closes · Loot boxes and gacha games | Very high, very steady rates and extreme resistance to extinction. The hardest schedule to walk away from. | | **Fixed interval (FI)** | for the first response after a set time | Studying that ramps up the night before a weekly Friday quiz · Peering down the street for a bus that runs every 15 minutes · Checking the washing machine as its 45-minute cycle nears the end | The "scallop": almost nothing right after reinforcement, accelerating as the interval runs out. | | **Variable interval (VI)** | for the first response after an unpredictable time | Checking email · Redialing a busy customer-service line · A surfer paddling for waves that arrive at irregular intervals | Moderate, remarkably steady rates. Checking never quite stops because the next one might be the one. | The practical rule in every setting: reinforce continuously while a behavior is being learned, then thin to an intermittent schedule so it lasts. [Run the interactive schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) ## Tricky examples that fool people Textbook examples are clean. Real ones usually contain two contingencies at once, or a consequence whose effect is the opposite of its intent. Five that trip up students, parents, and trainers: ### 1. The candy that ends the tantrum A child screams for candy in the checkout line; the parent, mortified, hands it over; the screaming stops. Two things were learned. The child's tantrum was followed by candy — **positive reinforcement of tantrums**. The parent's giving-in was followed by the end of the screaming — **negative reinforcement of giving in**. Each person trained the other, and each will do it faster next time.[5] The fix is not a bigger punishment for the tantrum; it is making sure tantrums stop working (extinction) while *asking politely* starts working (reinforcement). ### 2. The scolding that is really attention A teacher tells a student to sit down every time he wanders; the wandering gets worse. Students almost always label this "positive punishment that failed." It isn't. The behavior increased, so the consequence was a reinforcer: the reprimand delivered attention, and attention was what the wandering was for. Classroom studies from the 1960s onward found exactly this — reprimands can increase the behavior they target.[8] In applied behavior analysis, this is called attention-maintained behavior, and it is identified with a functional analysis rather than guessed.[7] ### 3. The snooze button "The alarm is positive punishment for sleeping in." No: the alarm is an antecedent, not a consequence, and the behavior that matters is *pressing snooze*. Pressing snooze removes the alarm instantly, every single time — **negative reinforcement on a continuous schedule with zero delay**, which is about as strong as a contingency gets. That is why the only alarm that reliably works is the one across the room: it makes standing up, not pressing a button, the escape response. ### 4. The grounding that removes nothing A teenager is "grounded for a week" for a bad grade. She spends the week in her room, texting her friends, exactly as she would have anyway. Nothing that functions as a reinforcer was removed, so this is not **negative punishment** no matter what it is called — and if being grounded also excuses her from family dinners she dislikes, it may be negative *reinforcement* of whatever produced the grade. The test for negative punishment is that a reinforcer the person actually contacts is withdrawn and the behavior decreases. [How negative punishment actually works ›](https://operantconditioning.com/negative-punishment/) ### 5. The seat-belt chime People often call the chime "punishment for not wearing a seat belt." But "not buckling" is not a behavior that can be punished; it is the absence of one. The chime is present until you buckle; buckling ends it — **negative reinforcement of buckling** (escape). After a few weeks most drivers buckle before the chime ever starts, which is **avoidance**: the behavior now prevents the aversive stimulus rather than ending it. Engineers who design these systems are doing behavior analysis whether they call it that or not. > **A fast test for the hard cases** > > Name the behavior first — the specific thing the person or animal *did*. Then ask what changed in the environment right after it, and whether the behavior went up or down over the following days. If you find two behaviors (the child's and the parent's, the dog's and the owner's), analyze each one separately. Most "tricky" examples are just two ordinary examples stacked on top of each other. ## Classify it yourself Take any behavior from your own day and run it through the two questions. If you want a check, the [interactive quadrant finder on the home page](https://operantconditioning.com/#which-quadrant-is-it-an-interactive-check) asks the same two questions and names the quadrant for you, and the [20-question quiz](https://operantconditioning.com/quiz/) tests you on scenarios like the ones above with instant explanations. For the vocabulary, see the [glossary](https://operantconditioning.com/glossary/). ## Frequently asked questions **What are some examples of operant conditioning in everyday life?** Holding a door and getting a thank-you (positive reinforcement), putting on headphones to cut out noise (negative reinforcement), getting sunburned after skipping sunscreen (positive punishment), and having a bike stolen after leaving it unlocked (negative punishment). Lottery tickets and slot machines are variable-ratio schedules, and you stop waving at a neighbor who never waves back through extinction. **What is an example of operant conditioning in the classroom?** A teacher adds a marble to a jar when the class lines up quietly, and a full jar earns a party — positive reinforcement with tokens. A student who finishes all the problems in class has homework waived — negative reinforcement. A phone used during a lesson is kept until the bell — negative punishment. A student who acts silly to escape reading aloud, and gets skipped, is being negatively reinforced for acting silly. **What are the four types of operant conditioning with examples?** Positive reinforcement: a dolphin touches a target and gets a fish, so touching increases. Negative reinforcement: a horse steps forward and leg pressure is released, so stepping forward increases. Positive punishment: a cat jumps on the counter and gets a puff of air, so jumping decreases. Negative punishment: a parrot screams and the owner leaves the room, so screaming decreases. **Is a speeding ticket positive or negative punishment?** It depends which part you analyze, and textbooks disagree. Being pulled over and handed a citation adds an aversive event (positive punishment). Paying the fine and losing license points removes money and privileges (negative punishment). If speeding decreases, both descriptions are correct; if it doesn't, neither is, because punishment is defined by its effect. **What is an example of negative reinforcement that is not punishment?** Feeding a parking meter. Nothing unpleasant happens, and no behavior decreases; the behavior of feeding the meter increases because it prevents a ticket. Negative reinforcement always strengthens a behavior by removing or preventing something aversive. Punishment weakens a behavior. **What is an example of a variable-ratio schedule in everyday life?** A slot machine pays out after an unpredictable number of pulls; a salesperson closes roughly one cold call in thirty; a loot box occasionally contains a rare item. In each case reinforcement depends on the number of responses but the number varies, which produces high, steady rates of behavior that are very hard to extinguish. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 3. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 4. Lerman, D. C., & Iwata, B. A. (1995). Prevalence of the extinction burst and its attenuation during treatment. *Journal of Applied Behavior Analysis, 28*(1), 93–94. 5. Patterson, G. R. (1982). *Coercive Family Process*. Castalia. 6. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 7. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 8. Madsen, C. H., Becker, W. C., & Thomas, D. R. (1968). Rules, praise, and ignoring: Elements of elementary classroom control. *Journal of Applied Behavior Analysis, 1*(2), 139–150. 9. Higgins, S. T., Budney, A. J., Bickel, W. K., Foerg, F. E., Donham, R., & Badger, G. J. (1994). Incentives improve outcome in outpatient behavioral treatment of cocaine dependence. *Archives of General Psychiatry, 51*(7), 568–576. 10. Skinner, B. F. (1981). Selection by consequences. *Science, 213*(4507), 501–504. 11. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. ## Related - [Can you spot it?](https://operantconditioning.com/quiz/): Twenty scenarios, instant explanations. Can you name the quadrant every time? - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Fixed, variable, ratio, interval — with a live cumulative-record simulator. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): The quadrant behind the snooze button, the seat-belt chime, and most of the tricky cases. --- # Operant Conditioning Glossary: 128 Behavior Analysis Terms Defined > Operant conditioning glossary: 128 behavior analysis terms, from abolishing operation to variable ratio, defined precisely and linked to full explanations. - Source: https://operantconditioning.com/glossary/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-09 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reference* Every term you are likely to run into — in a paper, a therapy session, a dog-training class, or an argument on the internet — defined precisely, with links to the full explanations. - **ABC model**: The antecedent–behavior–consequence framework for reading any operant: what set the occasion (A), what the organism did (B), and what followed (C). It is another name for the three-term contingency and the basis of functional behavior assessment. [The ABC model in depth ›](https://operantconditioning.com/abc-model/) - **Abolishing operation (AO)**: A motivating operation that temporarily *decreases* the effectiveness of a reinforcer (or punisher) and decreases the current frequency of behavior that has produced it. A large meal is an abolishing operation for food: food reinforces less, and food-seeking drops. The opposite of an establishing operation. [Motivating operations ›](https://operantconditioning.com/abc-model/) - **Acquisition**: The phase of learning in which a new behavior is being established and its rate rises from baseline as reinforcement takes hold. Continuous reinforcement produces the fastest acquisition; behavior is then usually shifted to an intermittent schedule so that it persists. [Schedules of reinforcement ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Antecedent**: Any stimulus or condition present before a behavior that influences whether it occurs: a cue, an instruction, a place, a time of day, a bodily state. The "A" in A-B-C. The two most important kinds are discriminative stimuli, which signal that reinforcement is available, and motivating operations, which change how much the reinforcer is worth. [Antecedents explained ›](https://operantconditioning.com/abc-model/) - **Applied behavior analysis (ABA)**: The discipline that applies the principles of operant conditioning to socially significant behavior and measures whether the change is real. Baer, Wolf, and Risley defined the field in 1968 by seven dimensions: applied, behavioral, analytic, technological, conceptually systematic, effective, and showing generality. Used in autism intervention, education, organizational management, and behavioral medicine. [ABA in practice ›](https://operantconditioning.com/applications/#aba) - **Automatic reinforcement**: Reinforcement produced directly by the behavior itself rather than delivered by another person — scratching an itch, humming, rocking, twirling hair. Behavior that persists when the person is alone is often automatically reinforced, which is why "alone" is one of the conditions in a functional analysis. - **Autoshaping (sign-tracking)**: The emergence of a response — a pigeon pecking a lit key, a rat nosing a lever — when a stimulus is repeatedly followed by a reinforcer *whether or not the animal responds*. Discovered by Brown and Jenkins in 1968, it persists even when responding cancels the reinforcer, so it is not maintained by reinforcement; it is classical conditioning producing behavior directed at the signal. [Autoshaping explained ›](https://operantconditioning.com/operant-vs-classical-conditioning/#behavior-without-reinforcement-autoshaping-and-contrafreeloading) - **Aversive stimulus**: A stimulus an organism will work to escape or avoid. Functionally, one whose removal reinforces behavior (negative reinforcement) and whose presentation punishes it (positive punishment). Like reinforcers, aversive stimuli are identified by their effect on behavior, not by how unpleasant they seem to an observer. - **Avoidance**: Behavior that prevents an aversive stimulus from occurring at all, maintained by negative reinforcement: buckling up before the chime sounds, paying a bill before the late fee. Contrast escape, which ends an aversive stimulus that is already present. [Avoidance learning ›](https://operantconditioning.com/avoidance-learning/#discriminated-signaled-avoidance) - **Backward chaining**: Teaching a behavior chain by starting with the *last* step and adding earlier steps one at a time, so every training trial ends with the chain's natural reinforcer. A child learning to put on a shirt first does only the final tug down, then the last two steps, and so on. [Forward, backward and total-task chaining ›](https://operantconditioning.com/chaining/) - **Baseline**: Measurement of a behavior before any intervention, used as the comparison for judging whether the intervention worked. In single-case research designs the baseline is the "A" phase of an A-B or A-B-A-B design (unrelated to the "A" for antecedent). - **Behavior**: Anything an organism does that can be observed and measured — walking, talking, pressing, and, in radical behaviorism, private events such as thinking. A useful target behavior is active and specific, and passes the dead man's test. - **Behavior analysis**: The natural science of behavior founded on Skinner's work, with three branches: the experimental analysis of behavior (basic laboratory research), applied behavior analysis, and the conceptual branch, radical behaviorism. [B. F. Skinner ›](https://operantconditioning.com/bf-skinner/) - **Behavioral contrast**: A change in the rate of a behavior in one situation caused by a change in reinforcement in *another*. When reinforcement is reduced in the presence of one stimulus, responding often increases in the presence of a second stimulus where reinforcement has not changed (Reynolds, 1961). Practically: put a behavior on extinction at school and it may rise at home. - **Behavioral momentum**: John Nevin's metaphor for the persistence of behavior under disruption — extinction, distraction, satiation, or free reinforcers. The "mass" of a behavior depends on the rate of reinforcement obtained in the presence of a stimulus, so behavior from a richly reinforced context is harder to disrupt than behavior from a lean one. [Momentum and resistance to extinction ›](https://operantconditioning.com/extinction/#resistance-to-extinction-and-behavioral-momentum) - **Behaviorism**: The position that psychology should be a science of behavior. **Methodological behaviorism**, associated with John B. Watson, restricts the science to publicly observable events. **Radical behaviorism**, Skinner's version, treats private events — thoughts, feelings — as behavior subject to the same principles. [The history of behaviorism ›](https://operantconditioning.com/history/) - **Bridge (marker signal)**: A conditioned reinforcer, such as a clicker or the word "yes," delivered the instant the target behavior occurs to "bridge" the delay until the primary reinforcer arrives. The bridge is what makes precise timing possible when the treat is still in your pocket. [Clicker training ›](https://operantconditioning.com/dog-training/) - **Chaining**: Teaching a sequence of behaviors in which each response produces the discriminative stimulus for the next, with a reinforcer at the end of the chain. Built from a task analysis and taught forward, backward, or as a whole task. Contrast shaping, which builds a single new response. [Chaining in depth ›](https://operantconditioning.com/chaining/) - **Classical conditioning**: Pavlov's form of learning, in which a neutral stimulus paired with an unconditioned stimulus comes to elicit a reflexive response on its own: bell, then food, until the bell alone produces salivation. Also called respondent or Pavlovian conditioning. The key event comes *before* the response, whereas in operant conditioning it comes after. [Operant vs. classical conditioning ›](https://operantconditioning.com/operant-vs-classical-conditioning/) - **Compulsion loop**: A game-design pattern — anticipation, activity, reward, repeat — built on a variable schedule of reinforcement to keep players engaged. The term comes from the industry; the mechanism is the variable-ratio schedule. [Technology and product design ›](https://operantconditioning.com/applications/#technology) - **Conditioned punisher**: A previously neutral stimulus that has acquired punishing function by being paired with other punishers — the word "no," a frown, a warning light. Also called a secondary punisher. Like all punishers, it is defined by the fact that it decreases the behavior it follows. - **Conditioned reinforcer**: A previously neutral stimulus that has acquired reinforcing function through pairing with existing reinforcers: money, praise, grades, tokens, the click of a clicker. Also called a secondary reinforcer. Conditioned reinforcers make delayed real-world consequences workable. [Conditioned reinforcers in depth ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Consequence**: The stimulus change that follows a behavior — the "C" in A-B-C. A consequence that increases the behavior is reinforcement; one that decreases it is punishment; the absence of a consequence that used to arrive produces extinction. How much a consequence matters depends on contiguity and contingency. - **Contiguity**: Closeness in time between a behavior and its consequence. The shorter the delay, the stronger the effect; delays of even a few seconds sharply weaken operant learning, which is why a paycheck at month's end reinforces very little of any particular Tuesday's work. [Why timing matters ›](https://operantconditioning.com/positive-reinforcement/) - **Contingency**: The dependency between a behavior and a consequence: *if* this behavior, *then* this consequence. Learning requires that the consequence actually depend on the behavior; reinforcers that arrive regardless produce superstitious behavior instead. The word is also used for the whole antecedent–behavior–consequence relation, the three-term contingency. - **Contingency management**: A treatment, chiefly for substance use disorders, in which tangible reinforcers such as vouchers or prize draws are delivered contingent on objectively verified behavior — most often drug-negative urine samples or attendance. It is one of the best-supported psychosocial treatments for stimulant use disorders. [Operant conditioning in health care ›](https://operantconditioning.com/applications/#health) - **Continuous reinforcement (CRF)**: A schedule in which every occurrence of the behavior is reinforced. It produces the fastest acquisition and the fastest extinction, so it is used to build a behavior and then thinned to an intermittent schedule to make the behavior durable. [Continuous vs. intermittent reinforcement ›](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) - **Contrafreeloading**: Working for a reinforcer that is also freely available: a rat with a dish of food will still press a lever for the same food. Reported by Jensen in 1963 and found in most species tested, it shows that the opportunity to explore and manipulate is reinforcing in its own right. - **Cumulative record**: The graph produced by Skinner's cumulative recorder: paper feeds at a constant speed while a pen steps up with every response, so the slope of the line is the rate of responding. The signature patterns of each schedule — the fixed-interval scallop, the fixed-ratio break-and-run — were discovered on cumulative records. [The Skinner box and its recorder ›](https://operantconditioning.com/skinner-box/) - **Dead man's test**: A rule of thumb proposed by Ogden Lindsley: if a dead man can do it, it isn't behavior. "Not hitting" and "staying quiet" fail the test; "keeping hands on the desk" and "raising a hand" pass. It keeps target behaviors active and reinforceable. - **Delay discounting**: The decline in the present value of a reinforcer as the delay to receiving it increases. Steep discounting — preferring a small reward now to a larger one later — is associated with impulsivity and addiction, and it is the reason immediate consequences beat delayed ones in habit change. [Self-control and delay discounting ›](https://operantconditioning.com/matching-law/#self-control-when-the-alternatives-differ-in-time) - **Deprivation**: Going without a reinforcer for a period, which increases its effectiveness and increases behavior that has produced it. Hours without food make food a stronger reinforcer. Deprivation is the classic establishing operation; its opposite is satiation. - **Differential reinforcement**: Reinforcing one response or response class while withholding reinforcement from others. It is the procedure inside shaping and discrimination training, and the basis of the "DR" family of interventions: DRA (alternative behavior), DRI (incompatible behavior), DRO (other behavior), DRL (low rates, reinforcing only when responding is infrequent), and DRH (high rates). [Differential reinforcement in depth ›](https://operantconditioning.com/differential-reinforcement/) - **Differential reinforcement of alternative behavior (DRA)**: Reinforcing a desirable alternative to a problem behavior while placing the problem behavior on extinction — reinforcing "may I have a break?" instead of screaming. The most widely used function-based treatment; when the alternative is a communicative response, it is called functional communication training. [DRA, DRI, DRO and DRL ›](https://operantconditioning.com/differential-reinforcement/) - **Differential reinforcement of incompatible behavior (DRI)**: A form of DRA in which the reinforced alternative physically cannot occur at the same time as the problem behavior. Sitting is incompatible with running around the room; hands in pockets are incompatible with hitting. [DRA, DRI, DRO and DRL ›](https://operantconditioning.com/differential-reinforcement/) - **Differential reinforcement of other behavior (DRO)**: Delivering a reinforcer when the target behavior has *not* occurred for a set interval — a token for every five minutes without calling out. Despite the name, it reinforces the absence of a behavior rather than any specific alternative. Also called omission training. [DRA, DRI, DRO and DRL ›](https://operantconditioning.com/differential-reinforcement/) - **Discrete trial training (DTT)**: A teaching format in which learning is broken into short, clearly bounded trials: an instruction, a prompt if needed, the response, a consequence, and a brief pause. Associated with early intensive behavioral intervention for autism. Contrast the free operant, in which the learner can respond at any time. [ABA methods ›](https://operantconditioning.com/applications/#aba) - **Discrimination**: Responding differently in the presence of different stimuli, the result of reinforcement being available under one stimulus and not another. The rat presses when the light is on and not when it is off; you swear with friends and not with your grandmother. The opposite of generalization. [Stimulus control ›](https://operantconditioning.com/abc-model/) - **Discriminative stimulus (S^D)**: An antecedent stimulus in whose presence a behavior has been reinforced and in whose absence it has not, so that it now raises the probability of the behavior. A ringing phone is an S^D for answering; a green light for driving on. Its counterpart, the S-delta (S^Δ), signals that reinforcement is *not* available. [The discriminative stimulus explained ›](https://operantconditioning.com/abc-model/) - **Errorless learning**: Discrimination training arranged so the learner rarely or never responds to the wrong stimulus — for example, by introducing the SΔ so faintly and briefly at first that it is never responded to, then fading it in. Terrace (1963) showed it avoids the frustration and side effects of trial-and-error discrimination; it underlies prompting and fading in applied work. [Stimulus control ›](https://operantconditioning.com/abc-model/) - **Escape**: Behavior that terminates an aversive stimulus already present, maintained by negative reinforcement: taking an aspirin for a headache, leaving a loud room, handing a screaming child the candy. Contrast avoidance, which prevents the aversive stimulus from starting. [Escape, on the avoidance page ›](https://operantconditioning.com/avoidance-learning/#escape-the-easy-case) - **Establishing operation (EO)**: A motivating operation that temporarily *increases* the effectiveness of a reinforcer and increases the current frequency of behavior that has produced it. Food deprivation is the classic example; so are thirst, cold, and pain. The term was introduced by Jack Michael in 1982. [Motivating operations ›](https://operantconditioning.com/abc-model/) - **Extinction**: The procedure of no longer reinforcing a previously reinforced behavior, and the resulting decline in that behavior. Extinction is not one of the four quadrants: nothing is added or removed, the consequence that used to arrive simply stops. Extinguished behavior can return through spontaneous recovery, renewal, and resurgence. [Extinction in depth ›](https://operantconditioning.com/extinction/) - **Extinction burst**: A temporary increase in the frequency, intensity, or variability of a behavior when its reinforcer is first withheld, before the behavior declines. Pressing the elevator button harder and faster when it fails to respond. Giving in during the burst reinforces the more intense version of the behavior. [What is an extinction burst? ›](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) - **Fading**: Gradually removing prompts, or gradually changing a stimulus, so that the behavior comes under the control of the natural antecedent. Full physical guidance is faded to a light touch, then a gesture, then the spoken instruction alone. [Prompting and fading ›](https://operantconditioning.com/shaping/) - **Fixed interval (FI)**: A schedule in which the first response after a fixed period of time is reinforced (FI 60 s: the first press after a minute has elapsed). Produces the FI scallop: a pause after each reinforcer, then accelerating responding as the interval ends. [Fixed interval schedules in depth ›](https://operantconditioning.com/fixed-interval-schedule/) - **Fixed ratio (FR)**: A schedule in which reinforcement follows a set number of responses (FR 10: every tenth response). Produces a high, steady "break-and-run" rate with a post-reinforcement pause. Piecework pay and "buy ten, get one free" are fixed-ratio schedules. [Fixed ratio schedules in depth ›](https://operantconditioning.com/fixed-ratio-schedule/) - **Free operant**: A procedure in which the organism can respond at any time and at any rate — the lever in the operant chamber is always available — so that rate of responding becomes the primary measure. Skinner's central methodological innovation, in contrast to discrete-trial procedures such as mazes and puzzle boxes, where the experimenter starts each trial. - **Function of behavior**: What a behavior accomplishes for the organism: the reinforcer that maintains it. Functional assessment sorts problem behavior into four common functions — attention, escape or avoidance, access to tangibles, and automatic reinforcement — and treatment is matched to the function, not to what the behavior looks like. - **Functional analysis**: An experimental method for identifying the function of a behavior by systematically arranging conditions — attention, demand, alone, and a play control — and measuring which one produces the most behavior. Developed by Brian Iwata and colleagues (1982, reprinted 1994), it is the gold standard of assessment in applied behavior analysis. [ABA in practice ›](https://operantconditioning.com/applications/#aba) - **Functional behavior assessment (FBA)**: The broader process of identifying the antecedents and consequences maintaining a behavior, through interviews, rating scales, direct A-B-C observation and, when needed, a functional analysis. Widely required in schools before a behavior intervention plan is written. [A-B-C observation ›](https://operantconditioning.com/abc-model/) - **Generalization**: The spread of a learned behavior beyond the conditions in which it was trained. **Stimulus generalization**: responding to stimuli similar to the training stimulus. **Response generalization**: untrained but related responses also change. Generalization across settings, people, and time is a central goal of applied work (Stokes & Baer, 1977). The opposite of discrimination. - **Generalization gradient**: The curve relating response strength to how similar a test stimulus is to the training stimulus. A pigeon reinforced for pecking a 550 nm light pecks less and less as the color moves away from 550 nm (Guttman & Kalish, 1956). A steep gradient means sharp discrimination; a flat one means broad generalization. - **Generalized reinforcer**: A conditioned reinforcer that has been paired with many different reinforcers — money, tokens, social approval — so that it works regardless of any single deprivation state. Generalized reinforcers are why token economies work. [Generalized reinforcers in depth ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Habit**: In behavioral terms, a well-practiced operant under tight stimulus control, emitted with little deliberation when its cue appears. Habits are built by pairing a stable antecedent with a small behavior and an immediate consequence, and they persist because of intermittent reinforcement and behavioral momentum. In the neuroscience literature a habit is behavior that continues even after its outcome has lost value. [Building habits with operant conditioning ›](https://operantconditioning.com/habits/) - **Habituation**: The decline in responding to a stimulus that is simply repeated — you stop noticing the ticking clock. It is a form of non-associative learning and involves no consequences, so it is not operant conditioning. Distinct from extinction, which requires a history of reinforcement, and from satiation. - **Incentive salience**: The "wanting" that the brain's dopamine system attaches to reinforcers and to the cues that predict them, making them attention-grabbing and worth working for. Berridge and Robinson showed it is separate from "liking": animals without dopamine still enjoy sugar but will not seek it. [The neuroscience of operant conditioning ›](https://operantconditioning.com/neuroscience/) - **Instinctive drift**: The tendency of a trained behavior to drift toward the species-typical behavior the reinforcer naturally evokes, even at the cost of reinforcement. Keller and Marian Breland's raccoons, taught to drop coins in a box for food, began rubbing the coins together and refusing to let go — food-washing intruding on a trained response. Reported in "The misbehavior of organisms" (1961), it showed that reinforcement works within biological limits. [The Brelands in the history of the field ›](https://operantconditioning.com/history/) - **Instrumental conditioning**: The older term, from the Thorndike tradition, for what Skinner called operant conditioning: learning in which behavior is "instrumental" in producing an outcome. Some authors reserve "instrumental" for discrete-trial procedures such as mazes and puzzle boxes and "operant" for free-operant ones; in most modern usage the terms are interchangeable. [From Thorndike to Skinner ›](https://operantconditioning.com/history/) - **Intermittent reinforcement**: Any schedule in which only some occurrences of a behavior are reinforced — the fixed and variable ratio and interval schedules and their many variants. Also called partial reinforcement. Behavior on intermittent schedules is markedly more resistant to extinction than behavior on continuous reinforcement, the partial reinforcement extinction effect. [Continuous vs. intermittent reinforcement ›](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) - **Interval schedule**: A schedule in which reinforcement depends on the passage of time: the first response after an interval — fixed or variable — is reinforced. Responding faster does not bring reinforcement sooner, so interval schedules produce lower rates than ratio schedules. [Interval vs. ratio schedules ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Law of effect**: Edward Thorndike's principle, from his puzzle-box experiments with cats (1898; stated in full in 1911): responses followed by satisfaction are "stamped in" and become more likely, while responses followed by discomfort are "stamped out." The direct ancestor of reinforcement and punishment. [Thorndike and the law of effect ›](https://operantconditioning.com/history/) - **Learned helplessness**: The finding, first reported by Seligman and Maier in 1967, that animals exposed to inescapable aversive events later fail to escape even when escape becomes possible — as if they had learned that responding and outcomes are unrelated. Fifty years on, the same authors reframed it: passivity is the default response to prolonged aversive events, and what is learned is control. [Learned helplessness in depth ›](https://operantconditioning.com/learned-helplessness/) - **Learned industriousness**: Eisenberger's (1992) finding that reinforcing high effort in one task increases effort in unrelated tasks: the sensation of effort itself acquires conditioned reinforcing value. The counterpart of learned helplessness, and one reason to reinforce trying rather than only succeeding. - **Matching law**: Richard Herrnstein's 1961 finding that when two responses are available, the relative rate of each matches the relative rate of reinforcement it produces: pigeons on two keys distribute their pecks in proportion to the reinforcers each key delivers. It explains choice, and why a problem behavior persists when it is reinforced more richly than its alternative. [The matching law in depth ›](https://operantconditioning.com/matching-law/) - **Motivating operation (MO)**: An environmental variable that (1) alters the effectiveness of a stimulus as a reinforcer or punisher and (2) alters the current frequency of behavior that has produced that stimulus. Establishing operations increase both; abolishing operations decrease both. The concept was developed by Jack Michael. [Motivating operations ›](https://operantconditioning.com/abc-model/) - **Near miss**: A losing outcome that resembles a win — two cherries and a lemon on a slot machine. Near misses increase the urge to keep playing and recruit some of the same brain circuitry as wins, which is why machines are designed to produce them more often than chance would. [Gambling and games ›](https://operantconditioning.com/applications/#technology) - **Negative punishment**: The process in which a behavior is followed by the *removal* of a stimulus and decreases in future frequency as a result. Time-out, response cost, and losing privileges are the standard forms. "Negative" means removed, not bad. [Negative punishment in depth ›](https://operantconditioning.com/negative-punishment/) - **Negative reinforcement**: The process in which a behavior is followed by the *removal* (or avoidance) of a stimulus and increases in future frequency as a result. Buckling up to silence the chime; taking a painkiller to end a headache. It is not punishment: the behavior goes up. [Negative reinforcement in depth ›](https://operantconditioning.com/negative-reinforcement/) - **Noncontingent reinforcement (NCR)**: Delivering a reinforcer on a time-based schedule regardless of behavior, so the behavior no longer has to occur to obtain it. Used as a treatment for attention-maintained behavior — attention is given freely every few minutes — it breaks the contingency and acts as an abolishing operation for the reinforcer. Because no behavior is being strengthened, some analysts argue the procedure should not be called "reinforcement" at all (Poling & Normand, 1999). - **Operant**: A class of responses defined by its effect on the environment rather than by its form: "lever-pressing" includes pressing with the left paw, the right paw, or the nose. As an adjective, behavior that is emitted by the organism and controlled by its consequences. Skinner introduced the term in 1937. [Skinner's concept of the operant ›](https://operantconditioning.com/bf-skinner/) - **Operant conditioning**: Learning in which the future probability of a behavior is changed by the consequences that follow it: behaviors followed by reinforcement become more likely, behaviors followed by punishment less likely. Also called instrumental conditioning. [The complete guide to operant conditioning ›](https://operantconditioning.com/) - **Operant conditioning chamber (Skinner box)**: Skinner's apparatus for studying free-operant behavior: an enclosed space with a manipulandum (a lever or key), a device for delivering reinforcers (a food hopper or water dipper), optional stimuli such as lights and tones, and automatic recording. "Skinner box" was never Skinner's own term. [The Skinner box, part by part ›](https://operantconditioning.com/skinner-box/) - **Operant hoarding**: Letting earned reinforcers accumulate instead of consuming each one immediately. Cole (1990) found that rats on a schedule in which collecting pellets triggered a pause in reinforcement learned to let pellets pile up and collect them in batches — a form of self-control that simple impulsivity accounts do not predict. [Choice and self-control ›](https://operantconditioning.com/matching-law/) - **Operant variability**: Variability in behavior treated as a dimension that reinforcement controls. Page and Neuringer (1985) reinforced pigeons only for response sequences that differed from recent ones, and the birds became more variable; reinforcing repetition makes behavior stereotyped. New behavior in shaping comes from this variability. [Shaping ›](https://operantconditioning.com/shaping/) - **Overcorrection**: A positive punishment procedure developed by Richard Foxx and Nathan Azrin with two forms: **restitution** (repairing the environment beyond its original state — cleaning the whole table, not just the spill) and **positive practice** (repeatedly performing the correct form of the behavior). [Reprimands and overcorrection ›](https://operantconditioning.com/positive-punishment/#overcorrection) - **Partial reinforcement extinction effect (PREE)**: The finding that behavior reinforced intermittently is more resistant to extinction than behavior reinforced every time, first demonstrated systematically by Lloyd Humphreys in 1939. Paradoxically, less reinforcement produces more persistence. [The partial reinforcement extinction effect ›](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) - **Pavlovian-instrumental transfer (PIT)**: The increase in the rate of an operant behavior when a separately trained Pavlovian cue for the same or a similar reinforcer is presented — a rat presses faster for food while a tone that predicts food plays, although the tone was never part of the lever-press contingency. First reported by Estes in 1948; a laboratory model of how drug and food cues energize seeking. [How the two kinds of learning interact ›](https://operantconditioning.com/operant-vs-classical-conditioning/) - **Peak shift**: After discrimination training between an SD and a similar SΔ, the peak of the generalization gradient moves *away* from the SΔ: pigeons reinforced at 550 nm and extinguished at 555 nm respond most at about 540 nm (Hanson, 1959). Evidence that the SΔ carries an inhibitory gradient of its own. - **Positive punishment**: The process in which a behavior is followed by the *addition* of a stimulus and decreases in future frequency as a result. A burn after touching the stove; a reprimand after an interruption, if interruptions then decrease. "Positive" means added, not good. [Positive punishment in depth ›](https://operantconditioning.com/positive-punishment/) - **Positive reinforcement**: The process in which a behavior is followed by the *addition* of a stimulus and increases in future frequency as a result. A treat after a sit; praise after a chore; a like after a post. The most-used procedure in the field. [Positive reinforcement in depth ›](https://operantconditioning.com/positive-reinforcement/) - **Post-reinforcement pause**: The pause in responding that follows each reinforcer on fixed schedules (FR and FI). Its length grows with the size of the ratio or interval — the larger the requirement, the longer the break before the organism starts the next run. [The post-reinforcement pause ›](https://operantconditioning.com/fixed-ratio-schedule/) - **Premack principle**: David Premack's 1959 principle that a higher-probability behavior can reinforce a lower-probability behavior when access to it is made contingent: "first homework, then video games." Sometimes called Grandma's rule. [The Premack principle in depth ›](https://operantconditioning.com/premack-principle/) - **Primary reinforcer**: A stimulus that reinforces without any learning history because of its biological significance: food, water, warmth, sexual contact, relief from pain. Its effectiveness depends on deprivation. Also called an unconditioned reinforcer. [Primary and secondary reinforcers ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Prompt**: A supplementary antecedent stimulus added to evoke a correct response that the natural discriminative stimulus does not yet control: a verbal hint, a gesture, a model, physical guidance, or a highlighted stimulus. Prompts are meant to be faded. [Prompting ›](https://operantconditioning.com/shaping/) - **Prompt hierarchy**: An ordered set of prompts from least to most intrusive, used systematically in teaching. **Least-to-most** prompting waits for an independent response, then escalates — verbal, gestural, model, physical. **Most-to-least** begins with full guidance and fades it as the learner succeeds. - **Punisher**: A stimulus change that, when it follows a behavior, decreases the future frequency of that behavior. Defined entirely by effect: a consequence that fails to decrease the behavior is not a punisher, however unpleasant it looks. - **Punishment**: The process in which a behavior is followed by a stimulus change — the addition of a stimulus (positive punishment) or the removal of one (negative punishment) — and decreases in future frequency as a result. Punishment is defined by the decrease, not by intent or severity. [Punishment: types, evidence, alternatives ›](https://operantconditioning.com/punishment/) - **Radical behaviorism**: Skinner's philosophy of the science of behavior: private events such as thoughts and feelings are behavior, subject to the same principles as public behavior, and explanations should be sought in the organism's environmental history rather than in mental causes. "Radical" means thoroughgoing, not extreme. [Skinner's radical behaviorism ›](https://operantconditioning.com/bf-skinner/) - **Rate of response**: Responses per unit of time — Skinner's preferred measure of behavior, because it is continuous, sensitive, and reflects the probability of responding. It is read directly from the slope of a cumulative record. - **Ratio schedule**: A schedule in which reinforcement depends on the number of responses made, fixed or variable. Because faster responding produces reinforcement sooner, ratio schedules generate higher rates than interval schedules. [Ratio vs. interval schedules ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Ratio strain**: The breakdown of responding — long pauses, erratic rates, quitting — that occurs when a ratio schedule is thinned too quickly or set too high. The remedy is to lower the requirement and thin more gradually. [Ratio strain in depth ›](https://operantconditioning.com/fixed-ratio-schedule/) - **Reinforcement**: The process in which a behavior is followed by a stimulus change — the addition of a stimulus (positive reinforcement) or the removal of one (negative reinforcement) — and increases in future frequency as a result. Reinforcement is defined by the increase, not by whether the consequence seems rewarding. [Reinforcement: types, reinforcers, what makes it work ›](https://operantconditioning.com/reinforcement/) - **Reinforcer**: A stimulus change that, when it follows a behavior, increases the future frequency of that behavior. Not a synonym for "reward": a reward is something the giver thinks is nice, while a reinforcer is something that demonstrably strengthens the behavior it follows. [What counts as a reinforcer ›](https://operantconditioning.com/positive-reinforcement/) - **Renewal**: The return of an extinguished behavior when the context changes — a behavior extinguished at the clinic reappears at home. Extinction learning is unusually tied to the setting in which it happened, so extinction must be carried out in every context that matters. [Renewal ›](https://operantconditioning.com/extinction/#renewal) - **Resistance to extinction**: How long, and how much, a behavior persists once reinforcement stops. It is increased by intermittent schedules, by a rich history of reinforcement, and by a long history of the behavior paying off. [Resistance to extinction ›](https://operantconditioning.com/extinction/#resistance-to-extinction-and-behavioral-momentum) - **Respondent behavior**: Behavior elicited by a prior stimulus: reflexes and conditioned reflexes such as salivation, the startle response, the eye-blink, and conditioned fear. Skinner's term, chosen to contrast with operant behavior, which is emitted and controlled by its consequences. [Operant vs. respondent ›](https://operantconditioning.com/operant-vs-classical-conditioning/) - **Respondent conditioning**: Skinner's name for classical conditioning: a neutral stimulus paired with a stimulus that elicits a reflex comes to elicit the reflex itself. Both respondent and operant conditioning show extinction, spontaneous recovery, and renewal. [Full comparison ›](https://operantconditioning.com/operant-vs-classical-conditioning/) - **Response**: A single instance of behavior — one lever press, one spoken word, one glance at the phone. "Response" and "behavior" are used almost interchangeably; operant and response class refer to the group of responses that share a function. - **Response class**: A group of responses that share a function — the same effect on the environment — even when they look different. All the ways of getting a door open form one response class. This is why suppressing one form of a problem behavior often produces another form with the same function. - **Response cost**: A negative punishment procedure in which a specified amount of a reinforcer — tokens, points, money, minutes of screen time — is removed contingent on a behavior. Fines, penalties, and losing points are response cost. [Response cost ›](https://operantconditioning.com/negative-punishment/#response-cost) - **Response deprivation hypothesis**: Timberlake and Allison's (1974) rule for when one behavior will reinforce another: a contingency works if it restricts the contingent behavior below the level the organism performs when free. It generalizes the Premack principle and explains why a less-preferred activity can sometimes reinforce a more-preferred one. [Premack and response deprivation ›](https://operantconditioning.com/premack-principle/) - **Resurgence**: The reappearance of a previously reinforced behavior when a more recently reinforced behavior is placed on extinction. A child taught to ask politely instead of screaming goes back to screaming when polite requests stop being answered. The remedy is to keep the replacement behavior reliably reinforced. [Resurgence ›](https://operantconditioning.com/extinction/#resurgence) - **Reward prediction error**: The difference between the reinforcer received and the reinforcer predicted. Midbrain dopamine neurons fire a burst for a positive error, stay quiet for a fully predicted reward, and dip below baseline when an expected reward fails to arrive — the brain's teaching signal for operant learning and the quantity reinforcement-learning algorithms compute. [The neuroscience of operant conditioning ›](https://operantconditioning.com/neuroscience/) - **Satiation**: The reduced effectiveness of a reinforcer after it has been consumed or contacted in quantity, and the reduced behavior that follows: kibble is a weak reinforcer after dinner. Satiation is an abolishing operation; its opposite is deprivation. - **Scallop (fixed-interval scallop)**: The curved pattern a fixed-interval schedule leaves on a cumulative record: little responding just after a reinforcer, then a gradual acceleration to a high rate as the interval ends. Students who study little after an exam and cram before the next one trace the same curve. [The fixed-interval scallop ›](https://operantconditioning.com/fixed-interval-schedule/) - **Schedule of reinforcement**: The rule specifying which occurrences of a behavior will be reinforced: continuous, fixed ratio, variable ratio, fixed interval, variable interval, and many compound schedules built from them. Ferster and Skinner catalogued their effects in *Schedules of Reinforcement* (1957). [Schedules of reinforcement, with a simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Secondary reinforcer**: Another name for a conditioned reinforcer: a stimulus that acquired its reinforcing power through pairing with other reinforcers, such as money, praise, or a clicker's sound. [Primary and secondary reinforcers ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Self-management**: Applying the principles of behavior to one's own behavior: arranging antecedents, defining small target behaviors, self-monitoring, and delivering one's own consequences. Skinner devoted a chapter of *Science and Human Behavior* (1953) to it, and it is the basis of evidence-based habit change. [Self-management and habits ›](https://operantconditioning.com/habits/) - **Setting event**: A broader condition, often distant in time, that alters how an antecedent–behavior–consequence relation plays out: poor sleep, illness, hunger, an argument earlier in the morning. Closely related to, and often treated as a kind of, motivating operation. [Antecedents and setting events ›](https://operantconditioning.com/abc-model/) - **Shaping**: Building a new behavior by reinforcing successive approximations toward it while withholding reinforcement from earlier approximations. Skinner used it to teach pigeons to play ping-pong; trainers use it to teach a dog to spin; speech therapists use it to build words from sounds. [How shaping works ›](https://operantconditioning.com/shaping/) - **Sidman avoidance (free-operant avoidance)**: Avoidance without a warning signal: shocks arrive on a timer unless the organism responds, and each response postpones the next shock for a set interval. Introduced by Murray Sidman in 1953; rats learn a steady response rate that keeps shocks rare. [Avoidance learning ›](https://operantconditioning.com/avoidance-learning/) - **Spontaneous recovery**: The reappearance of an extinguished behavior after a period away from the extinction setting, usually weaker than before and weaker with each recurrence if reinforcement is still withheld. First described by Pavlov for conditioned reflexes; it occurs just as reliably for operant behavior. [Spontaneous recovery ›](https://operantconditioning.com/extinction/#spontaneous-recovery) - **Stimulus**: Any event or energy change in the environment that can affect behavior: a sound, a light, a word, a touch, a food pellet. Stimuli that precede behavior are antecedents; stimuli that follow it are consequences. - **Stimulus control**: The condition in which a behavior occurs reliably in the presence of a particular stimulus and less often in its absence, established by discrimination training. Strong stimulus control is what makes a habit feel automatic and makes changing the cue an easier route to change than willpower. [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) - **Successive approximations**: The series of increasingly close versions of a target behavior that are reinforced, one after another, during shaping: any movement toward the lever, then touching it, then pressing it. [Shaping step by step ›](https://operantconditioning.com/shaping/) - **Superstitious behavior**: Behavior maintained by accidental reinforcement — a consequence that happened to follow the behavior without depending on it. In Skinner's 1948 experiment, pigeons fed at regular intervals regardless of what they did developed rituals such as turning and head-bobbing. Lucky socks work the same way. [Superstitious behavior in depth ›](https://operantconditioning.com/superstitious-behavior/) - **Target behavior**: The specific, observable, measurable behavior chosen for change in an intervention — "raises hand before speaking," not "is respectful." A good target behavior passes the dead man's test and can be counted. - **Task analysis**: Breaking a complex skill into its component steps, in order, as the basis for chaining. Hand-washing becomes eight discrete steps; each is taught and reinforced until the whole chain runs on its own. [Task analysis and chaining ›](https://operantconditioning.com/chaining/) - **Three-term contingency**: Skinner's basic unit of analysis: a discriminative stimulus sets the occasion, a response occurs, and a consequence follows. Written A → B → C, it is the same thing as the ABC model and the frame behind every entry in this glossary. [The three-term contingency ›](https://operantconditioning.com/abc-model/) - **Time-out**: Short for *time-out from positive reinforcement*: a negative punishment procedure in which access to reinforcement is removed for a brief period, contingent on a behavior. It works only when the "time-in" environment is actually reinforcing; a child sent from a hard lesson to a comfortable hallway has been negatively reinforced, not punished. [Time-out done correctly ›](https://operantconditioning.com/negative-punishment/#time-out-from-positive-reinforcement) - **Token economy**: A system in which tokens — points, stars, chips — are delivered contingent on target behaviors and later exchanged for backup reinforcers. Tokens are generalized conditioned reinforcers. Ayllon and Azrin's 1968 program on a psychiatric ward is the founding example. [Token economies in depth ›](https://operantconditioning.com/token-economy/) - **Topography**: The physical form of a behavior — what it looks like — as distinct from its function. A raised hand and a wave have similar topography and different functions; hitting and kicking have different topographies and may share one function. - **Two-factor theory of avoidance**: Mowrer's account of avoidance learning: a warning signal is first classically conditioned to elicit fear, then the avoidance response is operantly reinforced by escape from that fear. It explains why avoidance is so persistent — the animal never stays to learn the shock has stopped — but struggles with unsignaled avoidance and with the calm of well-trained avoiders. [Avoidance learning ›](https://operantconditioning.com/avoidance-learning/) - **Unconditioned reinforcer**: Another name for a primary reinforcer: a stimulus such as food, water, or warmth that reinforces without any prior learning. Contrast conditioned reinforcer. [Primary and secondary reinforcers ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Variable interval (VI)**: A schedule in which the first response after an unpredictable interval, averaging some value, is reinforced (VI 60 s: intervals that average a minute). Produces a moderate, steady rate. Checking email is the everyday example: messages arrive on their own schedule, and the first check afterward is the one that pays off. [Variable interval schedules in depth ›](https://operantconditioning.com/variable-interval-schedule/) - **Variable ratio (VR)**: A schedule in which reinforcement follows an unpredictable number of responses, averaging some value (VR 10). Produces the highest, steadiest rate of any simple schedule and the greatest resistance to extinction — the schedule behind slot machines and infinite scroll. [Variable ratio schedules in depth ›](https://operantconditioning.com/variable-ratio-schedule/) - **Verbal behavior**: Skinner's 1957 analysis of language as operant behavior reinforced through the mediation of other people. It classifies verbal operants by function — the mand (request), tact (label), echoic (repetition), and intraverbal (reply to someone else's words), among others — and underlies most language programs in applied behavior analysis. Noam Chomsky's 1959 review of the book became a founding document of cognitive psychology. [Verbal Behavior and its critics ›](https://operantconditioning.com/history/) No terms match that search. ## Questions about operant conditioning terms **What's the difference between a reinforcer and reinforcement?** A reinforcer is the stimulus; reinforcement is the process. The treat is the reinforcer; the fact that giving the treat after a sit makes sitting more frequent is reinforcement. The same distinction separates a punisher (the stimulus) from punishment (the process). In both cases the stimulus earns its name only by its effect on behavior. **Why is it called positive punishment if it's bad?** Because "positive" in this vocabulary means *added*, like a plus sign, not pleasant. Positive punishment adds a stimulus after a behavior and the behavior decreases; negative punishment removes a stimulus and the behavior decreases. The same logic applies to reinforcement: positive reinforcement adds something, negative reinforcement removes something, and both make the behavior more likely. **What does S^D mean?** S^D, pronounced "ess-dee," stands for discriminative stimulus: an antecedent in whose presence a behavior has been reinforced, so that its presence now makes the behavior more likely. Its counterpart is S^Δ ("ess-delta"), a stimulus in whose presence the behavior has not been reinforced. The ringing phone is an S^D for answering; a phone that is switched off is an S^Δ. **Is a habit the same as an operant?** A habit is a kind of operant, but not every operant is a habit. An operant is any behavior controlled by its consequences. A habit is an operant that has been practiced so often in the presence of a stable cue that the cue alone triggers it with little deliberation — an operant under very strong stimulus control. That is why habits are built by fixing the antecedent and the consequence, not by relying on motivation. ## References Definitions follow the usage of the standard texts and primary sources below. 1. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 4. Catania, A. C. (2013). *Learning* (5th ed.). Sloan Publishing. 5. Michael, J. (1982). Distinguishing between discriminative and motivational functions of stimuli. *Journal of the Experimental Analysis of Behavior, 37*(1), 149–155. 6. Michael, J. (1993). Establishing operations. *The Behavior Analyst, 16*(2), 191–206. 7. Laraway, S., Snycerski, S., Michael, J., & Poling, A. (2003). Motivating operations and terms to describe them: Some further refinements. *Journal of Applied Behavior Analysis, 36*(3), 407–414. 8. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 9. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 10. Skinner, B. F. (1957). *Verbal Behavior*. Appleton-Century-Crofts. 11. Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. 12. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97. 13. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 14. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 15. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 16. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 17. Reynolds, G. S. (1961). Behavioral contrast. *Journal of the Experimental Analysis of Behavior, 4*(1), 57–71. 18. Seligman, M. E. P., & Maier, S. F. (1967). Failure to escape traumatic shock. *Journal of Experimental Psychology, 74*(1), 1–9. 19. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 20. Stokes, T. F., & Baer, D. M. (1977). An implicit technology of generalization. *Journal of Applied Behavior Analysis, 10*(2), 349–367. 21. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158. ## Related - [The ABC model](https://operantconditioning.com/abc-model/): Antecedent, behavior, consequence — the frame behind every term on this page. - [50+ examples](https://operantconditioning.com/examples/): See the terms in action, sorted by quadrant and setting. - [Can you spot it?](https://operantconditioning.com/quiz/): Twenty questions with instant explanations. Can you tell the quadrants apart? --- # Operant Conditioning Quiz: 20 Practice Questions With Instant Explanations > Free 20-question operant conditioning quiz with instant explanations: reinforcement and punishment scenarios, schedules, extinction, shaping, and more. - Source: https://operantconditioning.com/quiz/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-08 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Practice* Twenty questions, from a dog and a treat to motivating operations, with an explanation the moment you answer. Most people get the negative-reinforcement questions wrong. Will you? > **The two-question test** > > For any scenario, ask, in this order: **(1) Did the behavior become more or less likely?** More likely means reinforcement; less likely means punishment. **(2) Was something added or removed?** Added means positive; removed means negative. Ignore whether the consequence seems pleasant, ignore what anyone intended, and look only at what happened to the behavior. ## What this operant conditioning quiz covers This operant conditioning quiz has 20 multiple-choice practice questions that get harder as you go. The first eight ask you to identify the quadrant — positive or negative reinforcement, positive or negative punishment — from everyday scenarios that are written to catch the classic mistakes: the seat-belt chime, the candy that ends a tantrum, the scolding that turns out to be attention, the time-out that is really an escape. The rest cover [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/), the [extinction burst and spontaneous recovery](https://operantconditioning.com/extinction/), [shaping versus chaining](https://operantconditioning.com/shaping/), [operant versus classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/), a little [history](https://operantconditioning.com/history/), and the two ideas that separate people who understand this subject from people who have memorized it: the motivating operation and the functional definition of a reinforcer. ## How the quiz is scored One point per question, no penalty for wrong answers, and a short explanation after every answer whether you got it right or not. Your score appears at the end, and you can retake the quiz as often as you like. Nothing is recorded or sent anywhere. ## Answer key and explanations Every question from the quiz, with the correct answer and the reasoning. Open any question to check your thinking. **Q1. A dog sits when told to and immediately gets a treat. Over the next week, the dog sits on cue more reliably. Which process is this?** **Answer:** Positive reinforcement. Ask the two questions. Did the behavior increase? Yes, so it is reinforcement. Was something added or removed? A treat was added, so it is positive reinforcement. **Q2. You start the car and an irritating chime sounds until you buckle your seat belt. Over time you buckle up faster and more consistently. What is maintaining the buckling?** **Answer:** Negative reinforcement. Buckling increased, so this is reinforcement, and it increased because an aversive stimulus (the chime) was removed, so it is negative reinforcement. Nothing here is punishment: no behavior decreased. **Q3. A teenager comes home after curfew and loses the car keys for a week. Curfew violations become less frequent. Which quadrant is this?** **Answer:** Negative punishment. The behavior decreased, so it is punishment. It decreased because access to something (the car) was removed, so it is negative punishment. Grounding and losing privileges are the classic examples. **Q4. A teacher scolds a student every time he calls out without raising his hand. Over the month, calling out increases. What has happened?** **Answer:** Positive reinforcement, because the behavior increased. Consequences are defined by their effect, not by how they look. Scolding added a stimulus (attention) and the behavior went up, so the scolding functioned as a positive reinforcer. For a student who gets little attention, even a reprimand can be the payoff. **Q5. A child screams for candy in the checkout line. The parent hands over the candy and the screaming stops. On later shopping trips, screaming becomes more frequent. What is the candy, for the child?** **Answer:** A positive reinforcer for screaming. Candy was added after the screaming and the screaming increased: positive reinforcement of the tantrum. Call it a bribe if you like, but the contingency is still doing its work, and it is working on the child. **Q6. In the same checkout-line scenario, the parent finds herself handing over the candy faster and faster on each trip. What is happening to the parent’s behavior?** **Answer:** It is being negatively reinforced by the end of the screaming. The parent’s candy-giving increased because an aversive stimulus (the screaming) stopped when she did it. That is escape, which is negative reinforcement. Two behaviors were strengthened in one exchange, and neither person intended either of them. **Q7. A student hits the snooze button, the alarm stops, and she drifts back to sleep. Over the semester she hits snooze more and more often. Which process best explains the snooze-pressing?** **Answer:** Negative reinforcement: the alarm is removed. The immediate consequence of pressing the button is that the alarm stops: an aversive stimulus is removed and the behavior increases. That is negative reinforcement by escape. Extra sleep may add positive reinforcement on top, but the defining event is the removal of the noise. Being late is delayed and inconsistent, which is exactly why it fails to punish. **Q8. A boy who dislikes math worksheets is sent to the hallway for five minutes every time he swears during math. Swearing during math increases. What is the hallway time-out actually doing?** **Answer:** Negatively reinforcing swearing by removing the worksheet. Time-out is a punishment procedure only when the behavior decreases. Here swearing increased, and what changed when he swore was that the aversive task went away, so the time-out is functioning as negative reinforcement (escape). Time-out only works when the time-in environment is more reinforcing than the time-out. **Q9. A garment worker is paid a fixed amount for every 20 shirts she sews. Which schedule of reinforcement is this?** **Answer:** Fixed ratio (FR). Reinforcement depends on a set number of responses, so it is a ratio schedule, and the number is always the same, so it is fixed: FR 20. Fixed-ratio schedules produce a high rate with a brief pause after each reinforcer. **Q10. New email arrives at unpredictable times. Checking your inbox more often does not make messages arrive sooner, but the first check after a message lands is the one that pays off. Which schedule is this?** **Answer:** Variable interval (VI). Reinforcement depends on time passing, not on how many checks you make, so it is an interval schedule, and the time is unpredictable, so it is variable: VI. Variable-interval schedules produce moderate, steady responding. **Q11. Which schedule of reinforcement produces behavior that is most resistant to extinction?** **Answer:** Variable ratio. On a variable-ratio schedule the organism can never tell whether the next response will pay off, so responding persists long after reinforcement stops. It is the schedule behind slot machines and infinite scroll. Continuous reinforcement is at the other extreme: fast to learn, fast to extinguish. **Q12. On a cumulative record, one schedule shows a pause after each reinforcer followed by a gradually accelerating rate of responding as the next reinforcer approaches, a pattern called a scallop. Which schedule is it?** **Answer:** Fixed interval (FI). The FI scallop appears because responding early in a fixed interval is never reinforced, so the organism waits, then speeds up as the interval ends. Students who study little after an exam and cram before the next one show the same curve. Fixed ratio produces a pause too, but it is followed by an abrupt switch to a high, steady rate, not a gradual acceleration. **Q13. For months, a toddler’s bedtime crying has always brought a parent into the room. The parents decide to stop going in. On the first night the crying is louder, longer, and more varied than ever before. What is this?** **Answer:** An extinction burst. When a reinforcer is first withheld, behavior often becomes more frequent, more intense, and more variable before it declines. That is the extinction burst. Giving in at this point reinforces the louder version of the crying on an intermittent schedule, which makes it harder to extinguish later. **Q14. The parents hold firm and the crying stops within a week. Ten days later, with nothing else changed, the crying briefly returns one night at a lower intensity, then fades again. What is this called?** **Answer:** Spontaneous recovery. Spontaneous recovery is the reappearance of an extinguished behavior after time away from extinction, usually weaker than before and weaker each time if reinforcement is still withheld. It is not a sign that extinction failed. Resurgence is different: an old behavior returning when a newer replacement behavior stops being reinforced. **Q15. A trainer teaching a dog to spin in a circle first rewards a slight head turn, then a quarter turn, then a half turn, and finally a full spin, withholding rewards for earlier versions as each new step is learned. Which procedure is this?** **Answer:** Shaping. Reinforcing successive approximations of a target behavior while extinguishing earlier ones is shaping. Chaining is different: it links separate, already-learned behaviors into a sequence in which each step cues the next, usually built from a task analysis. **Q16. A dog starts to salivate when it hears the treat bag rustle. Which kind of learning best describes the salivation?** **Answer:** Classical (respondent) conditioning. Salivation is a reflex elicited by a stimulus, not a voluntary behavior strengthened by its consequences. The rustle was paired with food and now elicits the response on its own: classical conditioning. The dog’s running to the kitchen when it hears the bag, by contrast, is operant behavior, reinforced by getting the treat. **Q17. Which statement correctly distinguishes operant conditioning from classical conditioning?** **Answer:** Classical conditioning involves reflexive responses elicited by antecedent stimuli; operant conditioning involves emitted behavior controlled by its consequences. Classical conditioning is stimulus-stimulus learning: a neutral stimulus paired with an unconditioned stimulus comes to elicit a reflexive response, and the key event comes before the response. Operant conditioning is behavior-consequence learning: the organism emits a behavior and what follows changes its future probability. Both apply across species. **Q18. Whose puzzle-box experiments with cats produced the law of effect, the principle Skinner later developed into the concept of reinforcement?** **Answer:** Edward Thorndike. Thorndike timed cats escaping from latched boxes and found that responses followed by a satisfying outcome were “stamped in.” He published the law of effect in 1898. Skinner coined the term “operant” in 1937 and formalized the science in *The Behavior of Organisms* (1938); Pavlov studied reflexes; Watson launched behaviorism in 1913. **Q19. A rat has learned to press a lever for food pellets only when a light is on. One day the experimenter feeds the rat to fullness before the session. The light comes on, but the rat barely presses. What is the pre-session feeding?** **Answer:** An abolishing operation that reduces the value of food as a reinforcer. The light is the discriminative stimulus: it signals that pressing will pay off, and it still does. What changed is the rat’s motivation. Satiation is an abolishing operation, a motivating operation that lowers the effectiveness of a reinforcer and reduces the behavior that produces it. Discriminative stimuli signal availability; motivating operations change value. **Q20. A manager begins praising employees publicly whenever they submit reports early. Over the next month, early submissions decrease. Which statement is correct?** **Answer:** The praise was not a reinforcer for this behavior; whether a stimulus is a reinforcer is determined by its effect on behavior. A reinforcer is defined by its effect: a stimulus that follows a behavior and increases it. If early submissions fell after public praise was added, the praise was not reinforcing them; if the drop was caused by the praise, it functioned as a positive punisher, which is common when public attention is embarrassing. The only way to know what reinforces a behavior is to watch what happens to the behavior. ## Study guide: where each topic is explained Missed a few? Each question maps to one page on this site. Read the page, retake the quiz, and the same scenarios should feel obvious. | Questions | Topic | Where to study it | | --- | --- | --- | | 1, 4, 5, 20 | Positive reinforcement and the functional definition of a reinforcer | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | | 2, 6, 7, 8 | Negative reinforcement, escape, and avoidance | [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) | | 4, 20 | Positive punishment and why attention is not always punishment | [Positive punishment](https://operantconditioning.com/positive-punishment/) | | 3, 8 | Negative punishment, grounding, and time-out done correctly | [Negative punishment](https://operantconditioning.com/negative-punishment/) | | 9–12 | Fixed and variable, ratio and interval schedules; the scallop; resistance to extinction | [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) (with a simulator) | | 13, 14 | Extinction burst, spontaneous recovery, resurgence, renewal | [Extinction](https://operantconditioning.com/extinction/) | | 15 | Shaping, successive approximations, chaining | [Shaping](https://operantconditioning.com/shaping/) | | 16, 17 | Respondent vs. operant behavior | [Operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/) | | 18 | Thorndike, Watson, Skinner | [B. F. Skinner](https://operantconditioning.com/bf-skinner/) · [History](https://operantconditioning.com/history/) | | 19 | Discriminative stimuli and motivating operations | [The ABC model](https://operantconditioning.com/abc-model/) | For quick definitions of any term that appeared in the quiz, use the [operant conditioning glossary](https://operantconditioning.com/glossary/); for more scenarios to practice on, see the [50+ examples sorted by quadrant](https://operantconditioning.com/examples/). The [complete guide](https://operantconditioning.com/) covers everything on this page in one place. ## Frequently asked questions **Is this quiz free?** Yes. Nothing is recorded, you can take it as many times as you like, and the full answer key with explanations is printed on this page. **How many questions are in the operant conditioning quiz?** Twenty multiple-choice questions, each with four options and an instant explanation. Eight cover identifying the four quadrants from scenarios, four cover schedules of reinforcement, two cover extinction, and the rest cover shaping, operant versus classical conditioning, history, motivating operations, and the functional definition of a reinforcer. **Can I use this to study for AP Psychology or an intro psych exam?** Yes. The quiz covers the operant conditioning material that appears in AP Psychology and most introductory psychology courses: the four quadrants, the schedules and their response patterns, extinction and spontaneous recovery, shaping, and the difference between operant and classical conditioning. The scenario questions are written in the same style as exam items, and the study guide above links each topic to a full explanation. **What score is good?** Fourteen or more out of twenty (70%) indicates a solid grasp of the fundamentals; eighteen or more means you could teach it. If you score below ten, the quadrant pages and the schedules page will fix most of the gaps, because those topics account for more than half the questions. Most first-time mistakes are the negative-reinforcement items, which is exactly what the quiz is designed to expose. ## Related - [Glossary](https://operantconditioning.com/glossary/): 128 terms from "abolishing operation" to "variable ratio," defined. - [50+ examples](https://operantconditioning.com/examples/): More scenarios to practice on, sorted by quadrant and setting. - [The complete guide](https://operantconditioning.com/): Operant conditioning from definition to application, on one page. --- # Operant Conditioning Diagrams to Download and Reuse > Every diagram on this site — four quadrants, extinction curve, four schedules, shaping, escape vs. avoidance, peak shift, Skinner box — as SVG and PNG, CC BY. - Source: https://operantconditioning.com/diagrams/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-10 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reference · Diagram library* Every diagram on this site, as a crisp vector file for slides and handouts and as a PNG for anything else. All of them are free to reuse with attribution. Each image below is the diagram exactly as it appears on its page, redrawn as a standalone file with a white background and no external dependencies. Click a diagram to open the SVG, which scales to any size; use the PNG where a bitmap is required. The text under each one gives the page it comes from, and the description that screen readers hear. > **License: Creative Commons Attribution 4.0 (CC BY 4.0)** > > You may copy, adapt, and redistribute these diagrams, including in slides, handouts, worksheets, videos, and commercial textbooks, provided you credit **operantconditioning.com** and link to the diagram's page. A line such as "Diagram: operantconditioning.com (CC BY 4.0)" is enough. The license text is at [creativecommons.org/licenses/by/4.0](https://creativecommons.org/licenses/by/4.0/).  **The four quadrants of operant conditioning** — The four quadrants. The column answers "added or removed?"; the row answers "more or less likely?" Positive and negative are arithmetic signs, not judgments. — From [Examples](https://operantconditioning.com/examples/) · [SVG](https://operantconditioning.com/assets/diagrams/the-four-quadrants-of-operant-conditioning.svg) · [PNG](https://operantconditioning.com/assets/diagrams/the-four-quadrants-of-operant-conditioning.png)  **Schematic of an operant conditioning chamber** — An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a food pellet on any schedule; the light signals when pressing will pay off; the recorder draws responses over time. — From [The complete guide](https://operantconditioning.com/) · [SVG](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber.svg) · [PNG](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber.png)  **Schematic of an operant conditioning chamber for a rat** — An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a pellet on any schedule; the cue light signals when pressing will pay off; the recorder draws responses against time. — From [Skinner Box](https://operantconditioning.com/skinner-box/) · [SVG](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber-for-a-rat.svg) · [PNG](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber-for-a-rat.png)  **An annotated cumulative record** — Reading a cumulative record: slope is rate, ticks are reinforcers, flat stretches are pauses, and the curved "scallop" is the fixed-interval pattern. The line only ever goes up; a behavior that has stopped draws a horizontal line. — From [Skinner Box](https://operantconditioning.com/skinner-box/) · [SVG](https://operantconditioning.com/assets/diagrams/an-annotated-cumulative-record.svg) · [PNG](https://operantconditioning.com/assets/diagrams/an-annotated-cumulative-record.png)  **Idealized cumulative records for the four basic schedules** — Idealized cumulative records. Each upward step is a response; green ticks mark reinforcers. Patterns after Ferster & Skinner (1957). — From [Schedules of Reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) · [SVG](https://operantconditioning.com/assets/diagrams/idealized-cumulative-records-for-the-four-basic-schedules.svg) · [PNG](https://operantconditioning.com/assets/diagrams/idealized-cumulative-records-for-the-four-basic-schedules.png)  **The extinction curve** — The shape of extinction. When reinforcement stops, responding briefly rises (the burst), then declines; after a rest it returns at a lower level (spontaneous recovery) and fades again. — From [Extinction](https://operantconditioning.com/extinction/) · [SVG](https://operantconditioning.com/assets/diagrams/the-extinction-curve.svg) · [PNG](https://operantconditioning.com/assets/diagrams/the-extinction-curve.png)  **Shaping as a rising criterion** — Shaping as a moving criterion. Reinforcing only the upper tail of what the animal currently does shifts the whole distribution; the criterion is then raised again. Variability (the width of each curve) is what makes the next step possible. — From [Shaping](https://operantconditioning.com/shaping/) · [SVG](https://operantconditioning.com/assets/diagrams/shaping-as-a-rising-criterion.svg) · [PNG](https://operantconditioning.com/assets/diagrams/shaping-as-a-rising-criterion.png)  **Peak shift after discrimination training** — Schematic of Hanson's result. After discrimination training against a nearby S−, the peak of the gradient moves away from the S− (here from 550 to about 540 nm) and rises above the control gradient. — From [Stimulus Control](https://operantconditioning.com/stimulus-control/) · [SVG](https://operantconditioning.com/assets/diagrams/peak-shift-after-discrimination-training.svg) · [PNG](https://operantconditioning.com/assets/diagrams/peak-shift-after-discrimination-training.png)  **Escape versus avoidance** — Escape ends an aversive stimulus that is present; avoidance prevents one that would have come. Both increase the behavior, so both are negative reinforcement. — From [Negative Reinforcement](https://operantconditioning.com/negative-reinforcement/) · [SVG](https://operantconditioning.com/assets/diagrams/escape-versus-avoidance.svg) · [PNG](https://operantconditioning.com/assets/diagrams/escape-versus-avoidance.png)  **Classical versus operant conditioning as sequences** — The two kinds of learning as sequences. In classical conditioning a stimulus comes to elicit a reflex; in operant conditioning a consequence changes how often a voluntary behavior recurs. — From [Operant vs. Classical](https://operantconditioning.com/operant-vs-classical-conditioning/) · [SVG](https://operantconditioning.com/assets/diagrams/classical-versus-operant-conditioning-as-sequences.svg) · [PNG](https://operantconditioning.com/assets/diagrams/classical-versus-operant-conditioning-as-sequences.png)  **The matching relation** — Schematic of the matching relation. Real data cluster around the diagonal; systematic departures from it are captured by the generalized matching law below. — From [Matching Law](https://operantconditioning.com/matching-law/) · [SVG](https://operantconditioning.com/assets/diagrams/the-matching-relation.svg) · [PNG](https://operantconditioning.com/assets/diagrams/the-matching-relation.png)  **The Premack principle as a reversal of baseline probabilities** — Premack's reversal. Whichever behavior is more probable at baseline can reinforce the other; deprivation changes which one that is. — From [Premack Principle](https://operantconditioning.com/premack-principle/) · [SVG](https://operantconditioning.com/assets/diagrams/the-premack-principle-as-a-reversal-of-baseline-probabilities.svg) · [PNG](https://operantconditioning.com/assets/diagrams/the-premack-principle-as-a-reversal-of-baseline-probabilities.png) ## Using the diagrams in teaching - **Slides.** Drag the SVG into Keynote, PowerPoint, or Google Slides; it stays sharp at any projector size and inherits none of the site's fonts. - **Handouts and exams.** The extinction curve, the escape-versus-avoidance timeline, and the schedules diagram work as unlabeled prompts: crop the labels and ask students to supply them. - **Learning management systems.** Use the PNG where the platform will not render SVG. Every SVG carries an accessible description in its `