# Operant Conditioning — full text > The complete guide to how consequences shape behavior. Generated 2026-09-13 from https://operantconditioning.com. Each section below is one page; the URL is given under its title. # Operant Conditioning: The Complete Guide > Operant conditioning is learning through consequences. The definition, Skinner's four quadrants, schedules of reinforcement, examples, and how to apply it. - Source: https://operantconditioning.com/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *The complete guide* Operant conditioning is learning from consequences: a behavior becomes more likely if it’s followed by reinforcement and less likely if it’s followed by punishment. - **350 BC** — [Aristotle](https://operantconditioning.com/history/#aristotle-on-habit): "Pleasure and pain are also the standards by which we all, in a greater or less degree, regulate our actions." - **1898** — [Thorndike](https://operantconditioning.com/history/)'s law of effect - **1937** — [Skinner](https://operantconditioning.com/bf-skinner/) coins the term "operant conditioning" - **Now** — [Reinforcement learning](https://operantconditioning.com/applications/#economics) is used to improve AI ## What is operant conditioning? > **Definition** > > **Operant conditioning** is a form of learning in which the future frequency of a behavior is changed by the consequences that follow it. Behaviors followed by [reinforcement](https://operantconditioning.com/positive-reinforcement/) become more likely; behaviors followed by [punishment](https://operantconditioning.com/positive-punishment/) become less likely. > > The term was introduced by American psychologist [B. F. Skinner](https://operantconditioning.com/bf-skinner/) in 1937, building on Edward Thorndike's law of effect (1898). It is also called *instrumental conditioning*, and it is the foundation of applied behavior analysis, modern animal training, and most evidence-based habit-change methods.[1][2] **Pick a door.** - [I can't put my phone down.](https://operantconditioning.com/#the-machine-you-are-in) — Checking and scrolling run on two different schedules. One of them is the strongest of the five. - [I want a habit that actually sticks.](https://operantconditioning.com/#how-to-use-operant-conditioning-on-yourself) — How motivated you feel is the part you cannot arrange. Three other things you can. - [My kid keeps doing the thing.](https://operantconditioning.com/parenting/) — What the evidence says about time-out, rewards, and why consistency beats severity. - [My dog is ignoring me.](https://operantconditioning.com/dog-training/) — Marker timing, clicker mechanics, and the method that replaced dominance theory. - [Just explain it properly.](https://operantconditioning.com/#operant-conditioning-in-one-paragraph) — The definition, the four quadrants, the evidence, the five schedules, and where the theory runs out. Teaching this, or studying it? The quiz, glossary, diagrams, citation formats and discussion questions all live on [one page](https://operantconditioning.com/for-teachers/). ## The machine you're already in Sometime in the last hour you picked up your phone without deciding to. You weren't looking for anything in particular. You just checked. Then, probably, you kept scrolling. Those are two different mechanisms, and behavior analysts have a name for each. **Checking** pays off on a time basis — either something arrived since you last looked or it didn't, and pressing the button twice as often will not make messages arrive twice as fast. That is a [variable-interval schedule](https://operantconditioning.com/schedules-of-reinforcement/), and it produces a moderate, stubbornly steady rate of responding. It is why you check again four minutes later. **Scrolling** is the other kind. Each swipe is a response, and whether it pays depends on how many swipes you make rather than on how long you wait. That is a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/), and of the five basic schedules — catalogued, with dozens of combinations, in a 700-page book Ferster and Skinner published in 1957 — it is the one that produces the highest and steadiest rate of responding, and the one that keeps behavior going longest after the payoffs stop. It is what a slot machine is. It is what a loot box is. It is what an infinite feed is. None of which is a character flaw, and in most cases it is not addiction in any clinical sense. It is a schedule doing what seventy years of data say a schedule will do. Which is the genuinely useful part. A mechanism strong enough to keep you swiping against your own stated wishes is strong enough to aim somewhere else, and aiming it is a skill rather than a personality trait. The rest of this page is that mechanism from the ground up: what a consequence has to do to count as one, the four things it can be, the five ways to time it, and what happens when it stops. > **One honest caveat before you go further.** Schedules explain a great deal about behavior and not all of it. People also learn by watching someone else, by building a map of a situation before any reward arrives, and — most awkwardly for the theory — through language. Where this runs out, the page says so rather than papering over it. ## Watch: operant vs. classical conditioning in four minutes Start here if you are new to the topic. This TED-Ed lesson by Peggy Andover walks through Pavlov's dogs, then shows how Skinner's operant conditioning differs — behavior first, consequence second — and how reinforcement and punishment change what an animal (or a person) does next. *Figure: "The difference between classical and operant conditioning," a TED-Ed lesson by Peggy Andover (2013). Embedded from YouTube's privacy-enhanced player.* > **What to watch for** > > Two things the video makes vivid: in classical conditioning the animal is *passive* — the bell and the food arrive whether or not it does anything — while in operant conditioning the animal's own action is what produces the consequence. And "negative" never means "bad": it means something was *taken away*. The rest of this page builds on both ideas. ## Operant conditioning in one paragraph Every organism that can learn is constantly running the same experiment: *do something, notice what happens next, adjust.* A rat presses a lever and a food pellet drops, so it presses again. A toddler says "please" and gets the cookie, so "please" becomes a habit. You check your phone, a notification rewards you, and the checking becomes automatic. In each case the behavior **operates** on the environment (hence "operant") and the environment answers back with a consequence. Operant conditioning is the study of how those consequences select which behaviors survive and which fade away.[3] Three ideas do most of the work: - **Consequences are defined by their effect, not their appearance.** A "reward" that doesn't increase behavior is not a [reinforcer](https://operantconditioning.com/glossary/#reinforcer). A "punishment" that doesn't decrease behavior is not a punisher. This is a functional definition, and it is the single most important thing to understand about the whole field. - **Behavior is selected over time, the way evolution selects traits.** Skinner called this "selection by consequences." Reinforcement doesn't teach a rule; it shifts probabilities.[4] - **The unit of analysis is the three-term contingency:** an antecedent sets the occasion, a behavior occurs, a consequence follows. Learn to see the A-B-C pattern and you can read almost any behavior. ## How operant conditioning works Skinner's central insight was that behavior is not just triggered by what comes *before* it (as in Pavlov's reflexes), it is shaped by what comes *after* it. He formalized this as the **[three-term contingency](https://operantconditioning.com/glossary/#three-term-contingency)**, often written A → B → C:[3] Two more variables determine how much a consequence matters: - **[Contiguity](https://operantconditioning.com/glossary/#contiguity) (timing).** Consequences that follow within seconds are far more effective than delayed ones. In laboratory studies the strength of learning drops off steeply as the delay grows to even a few seconds — one reason a paycheck at the end of the month is a poor reinforcer for any specific behavior on a Tuesday morning.[5] - **[Contingency](https://operantconditioning.com/glossary/#contingency) (dependability).** The consequence must actually depend on the behavior. If food arrives whether or not the rat presses the lever, lever-pressing doesn't get learned — and, as Skinner showed in his famous "superstition" experiment, pigeons fed on a timer developed odd ritual behaviors that happened to precede the food by chance.[6] > **Operant vs. respondent behavior** > > Skinner distinguished **[respondent](https://operantconditioning.com/glossary/#respondent-behavior)** behavior (reflexes elicited by a prior stimulus — salivation, the eye-blink, the startle response) from **[operant](https://operantconditioning.com/glossary/#operant)** behavior (actions emitted by the organism and controlled by their consequences). Pavlov studied the first; Skinner studied the second. Most of what people mean by "behavior" — walking, talking, working, scrolling — is operant. [Full comparison ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## Inside the Skinner box The apparatus that made all of this measurable was Skinner's **operant conditioning chamber** — the "Skinner box." Thorndike had timed cats escaping from puzzle boxes one trial at a time; Skinner's innovation was a box the animal never had to leave, so it could respond whenever it liked and the *rate* of responding could be recorded continuously.[2] [The Skinner box in full: every part, and why it was built that way ›](https://operantconditioning.com/skinner-box/) ![Schematic of an operant conditioning chamber](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber.svg) *An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a food pellet on any schedule; the light signals when pressing will pay off; the recorder draws responses over time.* A typical experiment runs in four steps. The rat is kept mildly hungry so food works as a reinforcer. It first learns that the click of the food dispenser means a pellet has arrived (a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer)). Then the experimenter [shapes](https://operantconditioning.com/shaping/) lever-pressing, reinforcing closer and closer approximations until the rat presses on its own. Finally the schedule is thinned — every press, then every fifth, then an unpredictable number — and the recorder shows how the pattern of responding changes. Pigeons peck a lit key instead of pressing a lever; the logic is identical. [More on Skinner and the box ›](https://operantconditioning.com/bf-skinner/) ## Run your own Skinner box At the top of this page you were the rat. Here you hold the pellet button. Reading about shaping is one thing; doing it is another. This lab puts a hungry, untrained rat in a chamber and hands you the pellet button. Magazine-train it, shape a lever press one approximation at a time, put the press on a schedule, then take the food away and watch extinction — all in about three minutes. The counters under the chamber track every pellet you deliver and every press the rat makes, from the first step on. The rat is a stylized model — its tendencies shift with what you reinforce, drift back when you don't, and follow the schedule patterns Ferster and Skinner documented — not a replay of real data. Real shaping takes longer and real rats are more surprising. ## The four quadrants: reinforcement and punishment Every consequence can be sorted along two questions. **Did the behavior increase or decrease?** (That tells you whether it was reinforcement or punishment.) **Was a stimulus added or removed?** (That tells you whether it was "positive" or "negative.") Crucially, in this vocabulary *positive* and *negative* mean plus and minus — added and removed — not good and bad. - **Positive reinforcement** — A pleasant stimulus is added after the behavior. The dog sits; the dog gets a treat. You finish a task; you feel a hit of satisfaction. (https://operantconditioning.com/positive-reinforcement/) - **Negative reinforcement** — An aversive stimulus is removed after the behavior. You buckle up; the seat-belt chime stops. You take an aspirin; the headache goes away. (https://operantconditioning.com/negative-reinforcement/) - **Positive punishment** — An aversive stimulus is added after the behavior. You touch a hot pan; it burns. You speed; you get a ticket. (https://operantconditioning.com/positive-punishment/) - **Negative punishment** — A pleasant stimulus is removed after the behavior. A teenager breaks curfew; the car keys are gone for a week. A player fouls; they sit out. (https://operantconditioning.com/negative-punishment/) > **The two-question test** > > Ask, in this order: **(1) Did the behavior become more or less likely?** More likely means reinforcement; less likely means punishment. **(2) Was something added or removed?** Added means positive; removed means negative. Those two answers name the quadrant every time. ### The four types of operant conditioning at a glance | Type | What happens after the behavior | Effect on the behavior | Example | | --- | --- | --- | --- | | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | A stimulus is added | Increases | A dog sits and gets a treat; sitting becomes more frequent | | [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) | A stimulus is removed | Increases | You buckle up and the seat-belt chime stops; buckling up becomes faster and more reliable | | [Positive punishment](https://operantconditioning.com/positive-punishment/) | A stimulus is added | Decreases | You touch a hot pan and get burned; touching hot pans becomes rarer | | [Negative punishment](https://operantconditioning.com/negative-punishment/) | A stimulus is removed | Decreases | A teenager breaks curfew and loses the car keys; breaking curfew becomes rarer | A fifth process, [extinction](https://operantconditioning.com/extinction/), is not a quadrant: the reinforcer that used to follow the behavior simply stops arriving, and the behavior fades. ### Which quadrant is it? An interactive check Almost everyone gets one pair backwards the first time, and it is nearly always *negative reinforcement* mistaken for *punishment*. Answer the two questions about any scenario and the tool will classify it. ### Five mistakes almost everyone makes 1. **Treating "negative reinforcement" as a polite word for punishment.** It is the opposite: negative reinforcement makes a behavior *more* likely by taking something unpleasant away. Taking an aspirin to end a headache is negative reinforcement of aspirin-taking. 2. **Reading "positive" as good and "negative" as bad.** They mean added and removed. A spray of water in a dog's face is *positive* punishment. 3. **Calling something a reinforcer because it seems nice.** A reinforcer is anything that increases the behavior it follows — it is defined by its effect. Scolding that a bored child finds attention-worthy is a reinforcer; a sticker a teenager finds embarrassing is not. 4. **Mixing up operant and classical conditioning.** If the key event comes *before* the response and the response is a reflex (salivating, flinching), it is classical. If the key event comes *after* a voluntary behavior, it is operant. 5. **Confusing extinction with punishment.** Extinction means the reinforcer simply stops arriving; nothing is added or taken away as a consequence. The behavior fades — usually after a brief burst — rather than being suppressed. ### Try three ## Reinforcement versus punishment: what the evidence says Skinner believed punishment was a poor way to change behavior, in part because an early experiment by his student W. K. Estes suggested that punishment only temporarily suppressed responding.[7] Later research complicated that picture: punishment *can* produce lasting decreases when it is immediate, consistent, and sufficiently intense from the outset.[8] But those same studies documented why practitioners still prefer reinforcement: - Punishment teaches what *not* to do without teaching what to do instead. Reinforcement builds a replacement. - Punishment tends to produce escape and avoidance — of the punisher as much as the behavior. (The child learns not to get caught.) - It can elicit aggression and emotional side effects, and it models the use of aversive control. - It works best at intensities and consistencies that are ethically or practically unavailable in most human settings. In parenting specifically, a large body of research links corporal punishment to worse, not better, long-term behavioral outcomes.[9] Modern applied behavior analysis therefore treats reinforcement-based procedures as the default and reserves punishment for narrow, supervised cases where reinforcement alone has failed and the behavior is dangerous.[10] [Reinforcement: the two types, kinds of reinforcers, and what makes it work ›](https://operantconditioning.com/reinforcement/) [Punishment: the two types, the side effects, and the alternatives ›](https://operantconditioning.com/punishment/) ## Schedules of reinforcement Once a behavior is learned, *how often* it gets reinforced changes both how fast the organism responds and how long the behavior persists when reinforcement stops. Ferster and Skinner catalogued these patterns in a 700-page 1957 volume, and the basic findings have held up for seventy years.[11] | Schedule | Reinforcer delivered… | Typical response pattern | Everyday example | | --- | --- | --- | --- | | **Continuous (CRF)** | after every response | Fast learning; fast extinction | A vending machine | | **Fixed ratio (FR)** | after a set number of responses | High rate with a pause after each reinforcer | Paid per piece; "buy 10, get 1 free" | | **Variable ratio (VR)** | after an unpredictable number of responses | Highest, steadiest rate; most resistant to extinction | Slot machines; social-media feeds | | **Fixed interval (FI)** | for the first response after a set time | "Scallop": slow after a reinforcer, accelerating as the interval ends | Checking the oven as the timer nears zero | | **Variable interval (VI)** | for the first response after an unpredictable time | Moderate, steady rate | Checking email | The practical rule: **use continuous reinforcement to build a behavior, then thin to an [intermittent schedule](https://operantconditioning.com/glossary/#intermittent-reinforcement) to make it durable.** The variable-ratio schedule is why gambling and infinite-scroll apps are so hard to quit — and why a behavior you reinforce only sometimes can end up stronger than one you reinforce every time. [Run the interactive schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) · [How organisms choose between schedules: the matching law ›](https://operantconditioning.com/matching-law/) ## Extinction, shaping, and stimulus control ### Extinction When a previously reinforced behavior stops producing reinforcement, it gradually declines. But not immediately: there is usually an **[extinction burst](https://operantconditioning.com/glossary/#extinction-burst)** — a temporary spike in the frequency, intensity, and variability of the behavior — before it fades. (Push the elevator button; nothing happens; you push it harder and faster before giving up.) Behavior that has been extinguished can also show **[spontaneous recovery](https://operantconditioning.com/glossary/#spontaneous-recovery)** after a rest period. [More on extinction ›](https://operantconditioning.com/extinction/) ### Shaping Complex behavior is rarely emitted fully formed, so it can't simply be reinforced. **Shaping** solves this by reinforcing *[successive approximations](https://operantconditioning.com/glossary/#successive-approximations)* — first any movement toward the lever, then touching it, then pressing it. Skinner used shaping to teach pigeons to play ping-pong; trainers use it to teach dolphins to jump through hoops; speech therapists use it to build words from sounds. [How shaping works ›](https://operantconditioning.com/shaping/) ### Stimulus control and the antecedent A behavior reinforced in one context and not in another comes under **[stimulus control](https://operantconditioning.com/glossary/#stimulus-control)**: it appears when the "discriminative stimulus" (S^D) is present and not otherwise. The rat presses only when the light is on. You swear with friends and not with your grandmother. This is the "A" in A-B-C, and it is the most under-used lever in self-improvement — changing the cue is often easier than willing a new response. [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) · [The ABC model ›](https://operantconditioning.com/abc-model/) ## Operant vs. classical conditioning The two great forms of associative learning are often confused. The clean distinction is *what gets associated with what*: | Aspect | Classical (Pavlovian) conditioning | Operant (instrumental) conditioning | | --- | --- | --- | | **Association** | Stimulus ↔ stimulus (bell → food) | Behavior ↔ consequence (press → food) | | **Behavior type** | Involuntary, reflexive (salivation, fear, nausea) | Voluntary, "emitted" (pressing, speaking, working) | | **Organism's role** | Passive; the stimulus is presented regardless | Active; the consequence depends on what it does | | **Timing of key event** | Stimulus comes *before* the response | Consequence comes *after* the response | | **Founders** | Ivan Pavlov (1890s–1927) | Edward Thorndike (1898); B. F. Skinner (1937–38) | In real life the two run together. The sound of the treat bag classically conditions excitement in a dog *and* operantly reinforces running to the kitchen. [Full comparison with examples ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## Can you spot it? Eight scenarios, instant explanations, no score kept. When you can call all eight without hesitating, you have it — and there is a [longer set of twenty](https://operantconditioning.com/quiz/). ## What happens in the brain Reinforcement has a physical address. In 1953 James Olds and Peter Milner found that a rat would press a lever thousands of times an hour for a pulse of electricity to its own brain, and in 1997 Wolfram Schultz and colleagues showed what the relevant neurons are doing: midbrain **dopamine** cells fire when a reinforcer is *better than expected*, fall silent when it is exactly as expected, and dip when an expected one fails to arrive.[21][16] That **reward prediction error** is the brain's teaching signal, and it explains why unpredictable reinforcers hold behavior so well and why a fully predictable one stops teaching. It is not a pleasure signal: Kent Berridge and Terry Robinson showed that animals without dopamine still *like* sugar but no longer *want* it.[22] [The neuroscience of operant conditioning, in depth ›](https://operantconditioning.com/neuroscience/) ## A brief history - **1898** — Thorndike's puzzle boxes Edward Thorndike times cats escaping from latched boxes. Escapes get faster, and he proposes the **[law of effect](https://operantconditioning.com/glossary/#law-of-effect)**: responses followed by satisfaction are "stamped in"; those followed by discomfort are "stamped out."[1] - **1930–1938** — Skinner's operant chamber At Harvard, as a graduate student and then a junior fellow, Skinner builds the apparatus later nicknamed the "Skinner box" and invents the cumulative recorder; at Minnesota he coins "operant" (1937) and lays out the science of operant behavior in *The Behavior of Organisms* (1938).[12][2] - **1948–1957** — Superstition, schedules, and Verbal Behavior The pigeon "superstition" study (1948), *Science and Human Behavior* (1953), *Schedules of Reinforcement* with Ferster (1957), and *Verbal Behavior* (1957) extend the analysis to society and language. - **1968** — Applied behavior analysis is born Baer, Wolf, and Risley publish the founding paper of ABA in the first issue of the *Journal of Applied Behavior Analysis*.[15] - **1997** — The dopamine connection Schultz, Dayan, and Montague show that midbrain dopamine neurons encode a *reward prediction error* — a biological implementation of the learning signal operant theory had assumed.[16] [The full history of operant conditioning ›](https://operantconditioning.com/history/) · [B. F. Skinner: life, work, and the Skinner box ›](https://operantconditioning.com/bf-skinner/) · [The original books, full text, in the library ›](https://operantconditioning.com/library/) ## Examples of operant conditioning in everyday life Once you know the pattern you see it everywhere: - **Positive reinforcement:** A barista is thanked warmly for a latte-art heart and starts making them on every cup. A student's essay earns praise and she writes more. Your phone lights up with a like. - **Negative reinforcement:** You clean the kitchen to end your partner's nagging. A student finishes homework early to escape the anxiety of a deadline. A driver takes the side street to avoid the traffic jam. - **Positive punishment:** A dog gets sprayed with water for jumping on the couch. A late invoice draws a fee. You bite into a moldy strawberry. - **Negative punishment:** A child loses screen time for hitting a sibling. A driver loses points from their license. A team member loses a project after missing deadlines. - **Schedules in the wild:** Slot machines (variable ratio). A bakery loyalty card (fixed ratio). Watching the clock in the last minutes of a shift (fixed interval). Fishing (variable interval). [50+ examples, sorted by quadrant and setting ›](https://operantconditioning.com/examples/) ## Operant conditioning in pop culture Two scenes worth watching a second time. Both are embedded from YouTube and load only when you press play. *Figure: **What to look for:** — every time Penny does something Sheldon likes, a chocolate appears — — [positive reinforcement](https://operantconditioning.com/positive-reinforcement/) — , delivered immediately, on a continuous schedule. When Leonard objects, Sheldon reaches for a spray bottle — that would be — [positive punishment](https://operantconditioning.com/positive-punishment/) — . Sheldon even names the procedure.* *Figure: **What almost everyone gets wrong here:** — Jim calls it Pavlov, but is it? Dwight's hand reaching out is a voluntary behavior that has been reinforced with a mint whenever the chime sounds — the chime is working as a — [discriminative stimulus](https://operantconditioning.com/glossary/#discriminative-stimulus) — , which makes the reaching — *operant* — . The dry mouth he notices is the — [classical](https://operantconditioning.com/operant-vs-classical-conditioning/) — part. Most real learning is both at once.* Clips are uploaded by third parties and may disappear; NBC hosts [the official Office clip](https://www.nbc.com/the-office/video/jims-pavlovian-prank-on-dwight-the-office/4141507). ## Key terms to know The twelve terms that do most of the work, defined the way behavior analysts define them. Hover or tap any underlined term anywhere on this page for its definition; the [full glossary](https://operantconditioning.com/glossary/) has 128 entries. - **Operant**: A class of behavior defined by its effect on the environment (what it accomplishes), not by its exact form. Lever-pressing with the left paw or the right paw is the same operant. - **Reinforcer**: Any consequence that increases the future frequency of the behavior it follows. *Positive* reinforcers are added; *negative* reinforcers are removed. - **Punisher**: Any consequence that decreases the future frequency of the behavior it follows. - **Primary vs. secondary reinforcer**: Primary reinforcers work without learning (food, water, warmth). Secondary (conditioned) reinforcers acquire their power by being paired with primary ones — money, praise, grades, a clicker. - **Three-term contingency**: Antecedent → Behavior → Consequence: the basic unit of analysis. [The ABC model ›](https://operantconditioning.com/abc-model/) - **Discriminative stimulus (S^D)**: A cue that signals a behavior will be reinforced. When a behavior reliably occurs in its presence and not otherwise, the behavior is under **stimulus control**. - **Schedule of reinforcement**: The rule for which responses get reinforced: continuous, fixed ratio, variable ratio, fixed interval, or variable interval. [Schedules ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Extinction**: Withholding the reinforcer that maintained a behavior, so the behavior declines — often after an **extinction burst**, a temporary increase. [Extinction ›](https://operantconditioning.com/extinction/) - **Shaping**: Building a new behavior by reinforcing successive approximations of it. [Shaping ›](https://operantconditioning.com/shaping/) - **Motivating operation**: A condition such as deprivation or satiation that changes how effective a reinforcer is. Food reinforces a hungry rat, not a full one. - **Premack principle**: A more probable behavior can reinforce a less probable one — "finish your homework, then you can play." [Premack ›](https://operantconditioning.com/premack-principle/) - **Law of effect**: Thorndike's 1898 principle that responses followed by satisfying consequences are strengthened and those followed by discomfort are weakened — the ancestor of operant conditioning. [The law of effect in depth ›](https://operantconditioning.com/law-of-effect/) *In practice* ## Where operant conditioning is used. The same three-term contingency runs a therapy session, a classroom, a family dinner, a factory floor, and a habit tracker. - [Applied behavior analysis](https://operantconditioning.com/applications/#aba): The clinical discipline built on operant principles, used in autism intervention, developmental disability support, and behavioral medicine. - [Education](https://operantconditioning.com/classroom/): Token economies, positive behavioral supports, immediate feedback, and the teaching machines Skinner pioneered in the 1950s. - [Parenting](https://operantconditioning.com/parenting/): Catch them being good, planned ignoring, time-out done correctly, and why consistency beats severity. - [Animal training](https://operantconditioning.com/dog-training/): Clicker training, marker signals, and the reinforcement-based methods that replaced dominance theory. - [Workplace & management](https://operantconditioning.com/applications/#workplace): Organizational behavior management, safety programs, and why annual reviews fail to change daily behavior. - [Habits & self-management](https://operantconditioning.com/habits/): Using antecedents, tiny behaviors, and immediate consequences to build habits that stick — on yourself. ## Criticisms and limitations An honest account includes what operant conditioning does *not* explain well. - **Biological constraints.** Organisms are not blank slates. The Brelands found that raccoons trained to deposit coins would instead "wash" them — an instinctive food-handling behavior that intruded even though it delayed reinforcement, a drift toward species-typical behavior they called [instinctive drift](https://operantconditioning.com/glossary/#instinctive-drift).[14] Reinforcement works with an animal's evolved tendencies, not against them. - **Cognition and language.** Chomsky argued that reinforcement cannot explain how children acquire grammar from limited input;[13] Tolman's rats learned mazes without obvious reinforcement, and Bandura's children learned by watching.[17] Operant learning is one powerful process among several, not a complete theory of mind. - **Rewards and intrinsic motivation.** Expected, tangible rewards for an activity someone already enjoys can reduce their interest once the rewards stop — the overjustification effect. A large meta-analysis found it; a rival meta-analysis found it small and narrow.[18][19] The fair reading: praise and feedback rarely undermine motivation; paying people for things they already love sometimes does. - **Ethics of control.** Skinner's *Beyond Freedom and Dignity* (1971) argued that since behavior is always controlled by its environment, we should design that environment deliberately. Critics saw a road to manipulation; the ethics codes of behavior analysis answer with consent, least-restrictive procedures, and the client's own goals.[10] What survived every critique is the core: the law of effect, schedule effects, extinction bursts, stimulus control, and shaping replicate across species and remain the working toolkit of clinicians, teachers, and trainers. Reinforcement learning — the branch of AI behind game-playing systems and the tuning of language models — is a mathematical descendant of the same ideas.[20] ## How to use operant conditioning on yourself The same contingency that trains a pigeon can be turned inward, and most self-improvement advice ignores two-thirds of it: it obsesses over motivation, which is not a term in the equation, and neglects antecedents and consequences, which are. The protocol is short. Attach the new behavior to a cue that already happens every day; shrink the behavior until it is almost embarrassing; deliver a small reinforcer within seconds, every time at first; then thin the schedule so the habit survives missed days. [The full seven-step protocol, with the evidence on how long habits take ›](https://operantconditioning.com/habits/) ### Design your own contingency Fill in the three terms for a habit you actually want. The tool checks each against the science — is the cue stable, is the behavior really a behavior, is the consequence immediate — and gives you a card to print. Nothing you type leaves your browser. [Full guide to building habits ›](https://operantconditioning.com/habits/) > **The app: Operant runs this loop for you.** A habit app for iPhone and Apple Watch from the publisher of this site. Each habit is set up as an antecedent, a behavior and a consequence — the three terms, not just the middle one. Free to download and try; a subscription unlocks the full app. [About the app](https://operantconditioning.com/app/) · [Download on the App Store](https://apps.apple.com/us/app/operant-behavior-change-app/id6802081776) ## Before you leave: can you answer these without opening them? **What is operant conditioning in simple terms?** Operant conditioning is learning from consequences. When a behavior is followed by something good (or the removal of something bad), it happens more often. When it is followed by something bad (or the loss of something good), it happens less often. The organism "operates" on its environment and the results shape what it does next. **Who discovered operant conditioning?** The underlying principle — the law of effect — was discovered by Edward Thorndike in 1898 through his puzzle-box experiments with cats. B. F. Skinner named it "operant" conditioning in 1937, developed the experimental methods to study it, and built the field around it beginning with *The Behavior of Organisms* in 1938. **What are the four types of operant conditioning?** Positive reinforcement (add something, behavior increases), negative reinforcement (remove something, behavior increases), positive punishment (add something, behavior decreases), and negative punishment (remove something, behavior decreases). "Positive" and "negative" refer to adding and removing a stimulus, not to whether the outcome is good or bad. **What is Skinner's theory of operant conditioning?** Skinner's theory is that behavior is selected by its consequences, much as species are selected by their environments. A behavior that is followed by reinforcement becomes more frequent; one followed by punishment or by no reinforcement at all becomes less frequent. He distinguished this *operant* behavior, which acts on the environment, from *respondent* behavior, the reflexes studied by Pavlov, and he showed that the schedule on which reinforcement arrives controls how fast and how persistently an organism responds. The theory deliberately explains behavior by its history of consequences rather than by inner states such as wants or intentions, which Skinner treated as behavior to be explained rather than as causes. [B. F. Skinner: the theory, the experiments, and the critiques ›](https://operantconditioning.com/bf-skinner/) **Why is it called "operant" conditioning?** Skinner chose the word in 1937 because the behavior *operates* on the environment to produce a consequence: the rat's press operates the lever, the lever delivers food, and the food changes future pressing. He contrasted operant behavior with respondent behavior, which is elicited by a stimulus that comes before it, as a puff of air elicits a blink. The older name, *instrumental* conditioning, makes the same point from the other side: the behavior is instrumental in producing the outcome. **What are the three components of operant conditioning?** The antecedent, the behavior, and the consequence, usually written A-B-C and called the three-term contingency. The antecedent is the situation or cue that sets the occasion for the behavior; the behavior is what the organism does; the consequence is what follows, and it is the consequence that changes how likely the behavior is next time. Every example on this site can be broken into those three parts. [The ABC model in depth ›](https://operantconditioning.com/abc-model/) **Is negative reinforcement the same as punishment?** No — this is the most common mistake in the whole subject. Negative reinforcement *increases* a behavior by removing something unpleasant (taking a painkiller to end a headache). Punishment *decreases* a behavior. "Negative" only means something was taken away. The two forms of negative reinforcement, escape and avoidance, have [their own page](https://operantconditioning.com/avoidance-learning/). **Which schedule of reinforcement is most resistant to extinction?** The variable-ratio schedule, in which reinforcement follows an unpredictable number of responses. Because the organism can never tell whether the next response will pay off, responding persists long after reinforcement has stopped. This is why gambling and social-media checking are hard to extinguish. ## References 1. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. See also Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 4. Skinner, B. F. (1981). Selection by consequences. *Science, 213*(4507), 501–504. 5. Grice, G. R. (1948). The relative effects of delay of reinforcement in the discrimination learning of rats. *Journal of Experimental Psychology, 38*(1), 1–16. See also Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139. 6. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 7. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 8. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 9. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 10. Behavior Analyst Certification Board. (2020). *Ethics Code for Behavior Analysts*. BACB. See also Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 11. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 12. Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279. 13. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58. 14. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 15. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97. 16. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 17. Tolman, E. C. (1948). Cognitive maps in rats and men. *Psychological Review, 55*(4), 189–208; Bandura, A. (1977). *Social Learning Theory*. Prentice-Hall. 18. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 19. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423. 20. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. 21. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427; Olds, J. (1958). Self-stimulation of the brain. *Science, 127*(3294), 315–324. 22. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? *Brain Research Reviews, 28*(3), 309–369. *Keep going* ## Go deeper. - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The most-used tool in the kit — and the most misunderstood word in it. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Fixed, variable, ratio, interval — with a live cumulative-record simulator. - [B. F. Skinner](https://operantconditioning.com/bf-skinner/): The man, the box, the pigeons, the controversies. - [50+ examples](https://operantconditioning.com/examples/): Everyday, classroom, workplace, and animal examples for every quadrant. - [Glossary](https://operantconditioning.com/glossary/): Every term from "abolishing operation" to "variable interval," defined. - [Can you spot it?](https://operantconditioning.com/quiz/): A 20-question quiz with instant explanations. Can you tell the quadrants apart? - [The library](https://operantconditioning.com/library/): Thorndike, Morgan, James, Yerkes, Darwin: the public-domain sources in full text, for readers and for AI. - [The Operant app](https://operantconditioning.com/app/): A habit app for iPhone and Apple Watch that runs the A-B-C loop for each habit. --- # Reinforcement: Definition, the Two Types, and What Makes It Work > Reinforcement is any consequence that makes a behavior more likely. Positive vs. negative reinforcement, types of reinforcers, what makes it work, and more. - Source: https://operantconditioning.com/reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · The hub* Half of operant conditioning is about making behavior more likely. This page defines reinforcement, sorts its two types and its kinds of reinforcers, and points you to the deep pages on each. > **Definition** > > **Reinforcement** is the process in which a consequence that follows a behavior makes that behavior *more likely* in the future. The consequence is called a **reinforcer**. Whether something is a reinforcer is decided by its effect on behavior, never by how it looks or feels.[1] > > A "reward" that does not increase the behavior it follows is not a reinforcer. A scolding that does increase it is. **In brief** - Reinforcement is any consequence that makes a behavior more likely; it is defined by that effect, never by how it looks or feels. - Positive reinforcement adds a stimulus and negative reinforcement removes one; both increase behavior, and the signs are arithmetic, not judgments. - Most failures of reinforcement are failures of implementation: late, non-contingent, too small, delivered to a satiated person, or on the wrong [schedule](https://operantconditioning.com/schedules-of-reinforcement/). ## The two types of reinforcement Both types make behavior more likely. They differ only in what happens to the stimulus: it is **added** (positive, +) or **removed** (negative, −). "Positive" and "negative" are arithmetic signs, not judgments. - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): A stimulus is **added** after the behavior and the behavior increases. The dog sits and gets a treat; you finish a task and feel the satisfaction. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): A stimulus is **removed** after the behavior and the behavior increases. You buckle up and the chime stops; you take an aspirin and the headache goes. > **The two-question test** > > Ask, in this order: **(1) Did the behavior become more or less likely?** More likely means reinforcement. **(2) Was something added or removed?** Added means positive; removed means negative. [Try it on any scenario with the quadrant checker ›](https://operantconditioning.com/#the-four-quadrants-reinforcement-and-punishment) ## Kinds of reinforcers | Kind | What it is | Examples | | --- | --- | --- | | Primary (unconditioned) | Reinforcing without any learning, because of biology | Food, water, warmth, sleep, sex, relief from pain | | Secondary (conditioned) | Acquires its power by being paired with other reinforcers | Praise, grades, a clicker, a "like," a check mark | | Generalized | A conditioned reinforcer paired with many others, so it works almost regardless of the person's state | Money, tokens, attention, approval | | Activity | The opportunity to do something more probable than the target behavior | Play after homework, the walk after the coffee — the [Premack principle](https://operantconditioning.com/premack-principle/) | | Social | Delivered by other people | Smiles, thanks, being listened to | | Automatic | Produced by the behavior itself, with no one delivering it | The feel of scratching an itch, the sound of your own humming | ## What makes reinforcement work - **Immediacy.** Reinforcers that arrive within seconds teach; delayed ones mostly strengthen whatever happened just before they arrived. Bridge a delay with a conditioned reinforcer (a word, a click, a check-off).[2] - **Contingency.** The reinforcer has to depend on the behavior. Reinforcement that arrives anyway teaches nothing — or teaches superstition.[3] - **Magnitude and quality.** Bigger and better reinforcers work better, with diminishing returns; a cut in magnitude is felt as a loss.[4] - **Motivating operations.** Deprivation makes a reinforcer stronger; satiation makes it weaker. Food does not reinforce a full rat. - **Schedule.** Reinforce every occurrence while a behavior is being learned, then thin to an intermittent schedule to make it durable. [Schedules of reinforcement ›](https://operantconditioning.com/schedules-of-reinforcement/) - **The alternatives.** A behavior's strength depends on what everything else pays. Enrich the alternatives and the behavior weakens without any punishment. [The matching law ›](https://operantconditioning.com/matching-law/) ## Reinforcement is not bribery, and not a reward A bribe is offered *before* a behavior to induce it; a reinforcer follows the behavior. A reward is something a person gives because it seems nice; a reinforcer is defined afterward, by the fact that the behavior went up. Most failures of "positive reinforcement" in classrooms and homes are failures of one of the five factors above — the reinforcer was late, non-contingent, too small, delivered to a satiated person, or on the wrong schedule — rather than failures of the principle. ## Key takeaways - A reinforcer is defined afterward, by the fact that the behavior went up. A "reward" that does not increase the behavior is not a reinforcer; a scolding that does increase it is. - Ask two questions, in order: did the behavior become more or less likely, and was something added or removed? More likely means reinforcement; added means positive and removed means negative. - Reinforcers come in kinds: primary (biological), conditioned (learned by pairing), generalized (money, tokens, approval), activity (a more probable behavior), social, and automatic (produced by the behavior itself). - Reinforcement depends on immediacy, contingency, magnitude, motivating operations, and schedule. Reinforce every occurrence while a behavior is being learned, then thin to an intermittent schedule to make it durable. - A bribe is offered before a behavior; a reinforcer follows it. A behavior's strength also depends on what everything else pays, so enriching the alternatives weakens it without any punishment. ### Check yourself **A teacher scolds a student every time he calls out, and calling out becomes more frequent. Was the scolding a punisher?** No. A consequence is classified by its effect on behavior, never by how it looks or feels. The behavior became more likely, so the scolding was a reinforcer: attention added after the behavior, which makes it positive reinforcement. **Food is delivered to a pigeon on a timer, regardless of what the pigeon does. Why might it end up repeating some odd movement?** Because reinforcement that arrives anyway strengthens whatever happened just before it arrived. Without contingency, the reinforcer does not teach the intended behavior; it teaches superstition. **You take an aspirin and your headache fades, and you reach for aspirin sooner next time. Is this reinforcement or punishment, positive or negative?** Reinforcement, because the behavior became more likely; negative, because a stimulus (the headache) was removed. Negative reinforcement is not punishment: the sign only says whether something was added or taken away. **A trainer wants to reinforce a dog's recall, but the treats are in a bag across the yard. What should happen the instant the dog arrives?** A conditioned reinforcer, such as a word or a click, should mark the behavior immediately. Reinforcers that arrive within seconds teach, while delayed ones mostly strengthen whatever happened just before they arrived; the conditioned reinforcer bridges the delay to the treat. **Explain it to a friend.** Explain what makes a consequence a reinforcer, using one example involving a person and one involving an animal. ## Go deeper - [Schedules](https://operantconditioning.com/schedules-of-reinforcement/): Fixed and variable, ratio and interval — with a live simulator. - [Shaping](https://operantconditioning.com/shaping/): Building behavior that does not yet exist by reinforcing approximations. - [Premack principle](https://operantconditioning.com/premack-principle/): When a behavior is the reinforcer. - [Avoidance learning](https://operantconditioning.com/avoidance-learning/): Negative reinforcement's strangest and most persistent form. - [The matching law](https://operantconditioning.com/matching-law/): How reinforcement divides behavior between options. - [Examples](https://operantconditioning.com/examples/): Fifty-plus scenarios sorted by quadrant. ## Frequently asked questions **What is reinforcement in psychology?** Any consequence that makes the behavior it follows more likely in the future. It is defined by its effect: if the behavior does not increase, no reinforcement occurred, whatever the consequence looked like. **What is the difference between positive and negative reinforcement?** Both increase behavior. Positive reinforcement adds a stimulus (a treat, praise); negative reinforcement removes one (a chime stops, a headache ends). "Negative" does not mean bad, and negative reinforcement is not punishment. **What are the types of reinforcers?** Primary (biological: food, water), secondary or conditioned (learned: praise, money, a clicker), generalized (paired with many reinforcers: money, tokens), activity (a preferred behavior), social, and automatic (produced by the behavior itself). **Why does reinforcement sometimes fail?** Usually because of timing (too late), contingency (delivered regardless of behavior), magnitude (too small), motivating operations (the person is satiated), or schedule (thinned too fast). The principle rarely fails; its implementation often does. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 3. Skinner, B. F. (1948). "Superstition" in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 4. Hutt, P. J. (1954). Rate of bar pressing as a function of quality and quantity of food reward. *Journal of Comparative and Physiological Psychology, 47*(3), 235–239. ## Related - [Punishment](https://operantconditioning.com/punishment/): The other half: making behavior less likely, and what the evidence says. - [Extinction](https://operantconditioning.com/extinction/): What happens when reinforcement stops. - [The complete guide](https://operantconditioning.com/): Everything on one page, in order. --- # Punishment: Definition, the Two Types, and What the Evidence Says > Punishment is any consequence that makes a behavior less likely. Positive vs. negative punishment, why it is not extinction, side effects, and alternatives. - Source: https://operantconditioning.com/punishment/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Punishment · The hub* The other half of operant conditioning: making behavior less likely. This page defines punishment precisely, separates its two types from extinction and from negative reinforcement, and gives an honest reading of whether it works. > **Definition** > > **Punishment** is the process in which a consequence that follows a behavior makes that behavior *less likely* in the future. The consequence is called a **punisher**. As with reinforcement, the definition is functional: a consequence that does not reduce the behavior is not punishment, however unpleasant it was meant to be.[1] > > In everyday speech "punishment" means a penalty someone intended. In behavior analysis it means a consequence that worked. **In brief** - Punishment is any consequence that makes the behavior it follows less likely; a consequence that does not reduce the behavior is not punishment. - Positive punishment adds a stimulus, negative punishment removes one; both decrease behavior, unlike [negative reinforcement](https://operantconditioning.com/negative-reinforcement/), which increases it. - It works only when immediate, consistent, intense from the outset, and paired with a reinforced alternative — conditions rarely met outside a laboratory. ## The two types of punishment - [Positive punishment](https://operantconditioning.com/positive-punishment/): A stimulus is **added** after the behavior and the behavior decreases. You touch the hot pan and it burns; you speed and get a ticket. - [Negative punishment](https://operantconditioning.com/negative-punishment/): A stimulus is **removed** after the behavior and the behavior decreases. A teenager misses curfew and loses the car keys; a player fouls and sits out. ## Three things punishment is not - **It is not negative reinforcement.** Negative reinforcement *increases* a behavior by removing something aversive. Punishment decreases behavior. The word "negative" is shared; the direction is opposite. [Why students mix these up ›](https://operantconditioning.com/negative-reinforcement/) - **It is not extinction.** In extinction nothing is added or removed as a consequence; the reinforcer that used to follow the behavior simply stops arriving. Ignoring a tantrum is extinction (if attention was the reinforcer); taking away the tablet for the tantrum is negative punishment. [Extinction ›](https://operantconditioning.com/extinction/) - **It is not defined by intent.** A parent who "punishes" a child by yelling, and finds the behavior increasing, has reinforced it — attention was the reinforcer. The behavior, not the parent, decides what the consequence was. ## Does punishment work? Skinner thought not. An early experiment by his student W. K. Estes found that punishing rats' lever pressing suppressed it only temporarily; when the punishment stopped, the pressing returned, and the total number of responses to extinction was about the same as for unpunished rats.[2] Skinner concluded that punishment merely suppresses, and argued against it for the rest of his career. Later research complicated the picture. Azrin and Holz's 1966 review, still the reference work, found that punishment can produce large and lasting decreases when it is **immediate**, **delivered every time**, **intense enough from the outset** (rather than escalated gradually), and combined with reinforcement of an alternative behavior.[3] A review of the applied literature three decades later reached the same conclusion for clinical settings.[4] Punishment, in other words, works — under conditions that are rarely met outside a laboratory, and at a price. ## The side effects - **It teaches nothing new.** Punishment says what not to do. The gap is filled by whatever else the person can do, which may be worse. Reinforcement of an alternative fills the gap deliberately. - **Escape and avoidance.** The punished organism learns to avoid the punisher as much as the behavior — the child learns not to get caught, the employee stops reporting mistakes. [Avoidance learning ›](https://operantconditioning.com/avoidance-learning/) - **Emotional and aggressive responding.** Aversive stimulation produces fear and aggression, sometimes toward bystanders.[3] - **It models aversive control.** People treated with punishment use punishment; the punisher's own behavior is negatively reinforced by the brief pause in the problem, which is why punishing tends to escalate. - **Corporal punishment specifically.** A 2016 meta-analysis of 75 studies covering more than 160,000 children found spanking associated with worse, not better, outcomes — more aggression, more antisocial behavior, more mental-health problems — with no evidence of benefit.[5] ## What to do instead | Alternative | How it works | Where it is explained | | --- | --- | --- | | Differential reinforcement of an alternative | Reinforce a behavior that does the same job for the person; the problem behavior loses its function | [Glossary: DRA](https://operantconditioning.com/glossary/#differential-reinforcement-of-alternative-behavior) | | Extinction | Identify the reinforcer maintaining the behavior and stop it arriving; expect a burst first | [Extinction](https://operantconditioning.com/extinction/) | | Antecedent change | Remove the cue, add a prompt, change the setting so the behavior is never triggered | [The ABC model](https://operantconditioning.com/abc-model/) | | Time-out done correctly | Brief, calm removal from reinforcement — which only works if the situation left was reinforcing | [Negative punishment](https://operantconditioning.com/negative-punishment/) | | Response cost | A predictable, proportionate loss of a token or privilege inside a system that also pays for good behavior | [Negative punishment](https://operantconditioning.com/negative-punishment/) | Modern applied behavior analysis treats reinforcement-based procedures as the default and reserves punishment for narrow, supervised cases where reinforcement alone has failed and the behavior is dangerous, with consent and the least restrictive procedure as ethical requirements.[6] ## Key takeaways - In behavior analysis, punishment means a consequence that worked: the behavior became less likely. The behavior, not the punisher's intent, decides what the consequence was. - Punishment is not negative reinforcement, which increases behavior by removing something aversive, and it is not extinction, in which the reinforcer simply stops arriving. Ignoring a tantrum is extinction if attention was the reinforcer; taking away the tablet for it is negative punishment. - Punishment can produce large, lasting decreases when it is immediate, delivered every time, intense enough from the outset, and combined with reinforcement of an alternative. Those conditions are rarely met outside a laboratory. - Its side effects: it teaches nothing new, it produces escape and avoidance of the punisher, it elicits fear and aggression, and it models aversive control. Spanking is associated with worse outcomes and no evidence of benefit. - Reinforcement-based procedures are the default. Differential reinforcement of an alternative, extinction, antecedent change, and time-out or response cost done correctly come first; punishment is reserved for narrow, supervised cases where the behavior is dangerous. ### Check yourself **A parent yells at a child for whining, and the whining increases over the following weeks. What kind of consequence was the yelling?** A reinforcer. Punishment is not defined by intent: the behavior, not the parent, decides what the consequence was. Attention was added and the behavior increased, so the yelling was positive reinforcement. **A child throws tantrums for attention. One parent stops responding to tantrums entirely; the other takes away the tablet each time. Which one is punishment?** Taking the tablet is negative punishment: a stimulus is removed as a consequence. Ignoring the tantrum is extinction, because nothing is added or removed; the reinforcer that used to follow the behavior simply stops arriving. Expect a burst before extinction works. **A manager waits until the monthly review to reprimand an employee for skipping safety steps, and mentions it only sometimes. Why is this unlikely to reduce the behavior?** Punishment produces lasting decreases only when it is immediate and delivered every time. A delayed, occasional reprimand meets neither condition, and it teaches nothing about what to do instead. **A teenager loses her phone for a day each time she misses curfew, and missed curfews decrease. A friend calls this negative reinforcement, because something was removed. What is wrong with that?** The direction. Negative reinforcement increases a behavior by removing something aversive; here a behavior decreased when something wanted was removed, which is negative punishment. The two share the word "negative" and nothing else. **Explain it to a friend.** Explain why behavior analysts reach for reinforcement before punishment, using an example from your own life in which a penalty failed to change what someone did. ## Go deeper - [Positive punishment](https://operantconditioning.com/positive-punishment/): Examples, the spanking evidence, alternatives. - [Negative punishment](https://operantconditioning.com/negative-punishment/): Time-out and response cost, done right. - [Extinction](https://operantconditioning.com/extinction/): The burst, spontaneous recovery, and how to use it. ## Frequently asked questions **What is punishment in psychology?** Any consequence that makes the behavior it follows less likely in the future. Positive punishment adds a stimulus (a burn, a fine); negative punishment removes one (losing privileges). It is defined by its effect on behavior, not by anyone's intention. **Is negative reinforcement a type of punishment?** No. Negative reinforcement increases behavior by removing something unpleasant; punishment decreases behavior. They share the word "negative" and nothing else. **Does punishment work?** It can reduce behavior when it is immediate, consistent, sufficiently intense from the start, and paired with reinforcement of an alternative. Those conditions are rarely met in daily life, and punishment carries side effects — avoidance, aggression, no new learning — which is why behavior analysts use reinforcement first. **What is the difference between punishment and extinction?** Punishment adds or removes a stimulus as a consequence. Extinction changes nothing except that the reinforcer stops coming. Both reduce behavior; extinction usually produces a temporary burst first. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), 1–40. 3. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 4. Lerman, D. C., & Vorndran, C. M. (2002). On the status of knowledge for using punishment: Implications for treating behavior disorders. *Journal of Applied Behavior Analysis, 35*(4), 431–464. 5. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 6. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. ## Related - [Reinforcement](https://operantconditioning.com/reinforcement/): The other half: making behavior more likely. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): The idea most often confused with punishment. - [The complete guide](https://operantconditioning.com/): Everything on one page, in order. --- # Positive Reinforcement: Definition, Examples, and How to Use It > Positive reinforcement adds a stimulus after a behavior to make it more likely. Definition, examples, types of reinforcers, what makes it work, and mistakes. - Source: https://operantconditioning.com/positive-reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Adding something* The workhorse of operant conditioning: add something after a behavior, and the behavior grows. Here is exactly how it works, what counts as a reinforcer, and why so many attempts at it fail. > **Definition** > > **Positive reinforcement** is the process in which a behavior is followed by the *addition* of a stimulus, and as a result the behavior becomes more frequent, intense, or likely in the future. The added stimulus is called a **positive reinforcer**. > > "Positive" means something is added (think plus sign), not that the outcome is pleasant — although positive reinforcers usually are. Whether a stimulus is a reinforcer is determined only by its effect on behavior.[1] **In brief** - Positive reinforcement adds a stimulus after a behavior, and the behavior becomes more frequent, intense, or likely in the future. - A reinforcer is defined by its effect on behavior, not by whether the giver thinks it is nice. - Reinforcement must be immediate and contingent; reinforce every occurrence while learning, then thin to an [intermittent schedule](https://operantconditioning.com/schedules-of-reinforcement/) so the behavior persists. ## How positive reinforcement works The sequence is always the same: an [antecedent](https://operantconditioning.com/abc-model/) sets the occasion, a behavior occurs, and immediately afterward a stimulus appears that was not there before. If the behavior then happens more often under similar conditions, positive reinforcement has taken place. - A rat presses a lever → a food pellet drops → lever-pressing increases. - A child says "thank you" → a parent smiles and says "you're welcome" → "thank you" increases. - You post a photo → likes appear → posting increases. Notice that in each case the consequence is delivered *because of* the behavior (contingency) and *right after* it (contiguity). Remove either and the effect weakens sharply. A bonus paid in December for effort in March reinforces very little of that effort; the pellet that arrives ten seconds after the press teaches the rat almost nothing about pressing.[2] > **The functional definition, again** > > A reinforcer is not a "reward." A reward is something the giver thinks is nice. A reinforcer is something that demonstrably increases the behavior it follows. Praise that a teenager finds embarrassing is not a reinforcer for them. Attention — even scolding — often *is* a reinforcer for a child who gets little of it. You find out what reinforces a behavior by watching what happens to the behavior. ### Check any scenario in two questions Not sure whether a consequence is this quadrant or a neighbor? Answer the two questions and the checker names it. ## Examples of positive reinforcement | Setting | Behavior | Stimulus added | Result | | --- | --- | --- | --- | | Home | Toddler uses the potty | Sticker and enthusiastic praise | Uses the potty more often | | Home | Child clears the table without being asked | "That was really helpful — thank you." | Clears the table more often | | Classroom | Student raises hand instead of calling out | Teacher calls on them | Hand-raising increases | | Classroom | Class transitions quietly | Marble added to the class jar (token) | Quiet transitions increase | | Workplace | Employee submits a report early | Public recognition in the team meeting | Early submissions increase | | Workplace | Salesperson closes a deal | Commission | Closing behavior increases | | Dog training | Dog sits on cue | Click, then a treat | Sitting on cue increases | | Dog training | Dog returns when called | Play with a favorite toy | Recall improves | | Technology | User opens the app | New content, likes, badges | App-opening increases | | Health | Patient attends a treatment session | Voucher (contingency management) | Attendance increases | | Self-management | You put on running shoes at 6:30 a.m. | Marking the habit done; a satisfying check | Shoes go on more mornings | | Nature | Bee visits a flower | Nectar | Visits to that flower type increase | Not every example of "adding something nice" is positive reinforcement. If a parent gives a child candy to *stop* a tantrum in the store, the candy positively reinforces the tantrum (and the end of the tantrum negatively reinforces the parent's candy-giving). Both people learned something; neither learned what they intended. ## Types of positive reinforcers ### Primary (unconditioned) reinforcers Stimuli that reinforce without any learning history because they relate to biological needs: food, water, warmth, sexual contact, relief from pain, and — for social species — physical contact. Their effectiveness depends on deprivation: food is a powerful reinforcer for a hungry rat and a weak one for a full one. ### Secondary (conditioned) reinforcers Neutral stimuli that acquire reinforcing power by being paired with existing reinforcers. Money is the classic example; so are grades, praise, a clicker's sound, and a green checkmark. Conditioned reinforcers are what make delayed real-world consequences workable: the click bridges the gap between the dog's sit and the treat that follows a second later.[3] ### Generalized reinforcers Conditioned reinforcers paired with *many* different reinforcers, so they work regardless of the organism's current state. Money, tokens, and social approval are generalized reinforcers; that is why token economies work in classrooms and psychiatric wards.[4] ### Social, tangible, activity, and sensory reinforcers Practitioners often sort reinforcers by category: **social** (attention, praise, a smile), **tangible** (a sticker, a toy, a bonus), **activity** (getting to play, choosing the music), and **sensory** (a pleasant sound or texture). The **Premack principle** is the rule for activity reinforcers: a higher-probability behavior can reinforce a lower-probability one. "You can play video games after you finish your homework" is Premack in action.[5] [The Premack principle in depth ›](https://operantconditioning.com/premack-principle/) ## What makes positive reinforcement effective 1. **Immediacy.** Deliver the reinforcer within seconds. If you can't, deliver a conditioned reinforcer (a click, a word, a checkmark) immediately and the real one later. 2. **Contingency.** The reinforcer follows the behavior and only the behavior. Non-contingent goodies are nice, but they don't teach. 3. **Magnitude and quality.** Bigger and better reinforcers work better up to a point: rats press faster for larger and for sweeter food rewards, with diminishing returns as amount grows.[8] But the effect is relative. Crespi found that rats switched from a large reward to a small one ran *slower* than rats that had only ever had the small one — a negative contrast effect — and rats shifted upward briefly ran faster than rats always given the large amount.[9] Cutting a reinforcer is felt as a loss, not as a smaller gain. And small, immediate reinforcers usually beat large, delayed ones. 4. **Motivating operations.** Deprivation makes a reinforcer stronger; satiation makes it weaker. Training a dog after dinner with kibble is a losing game. 5. **Schedule.** Reinforce every occurrence while the behavior is being learned, then shift to an intermittent schedule so it persists. [Schedules of reinforcement ›](https://operantconditioning.com/schedules-of-reinforcement/) 6. **Individualization.** Reinforcers are personal. What works is discovered by preference assessment and by watching the data, not assumed. ## Positive reinforcement vs. negative reinforcement vs. bribery Three things get confused constantly: | Aspect | Positive reinforcement | Negative reinforcement | Bribery | | --- | --- | --- | --- | | **What happens** | Stimulus is added after the behavior | Stimulus is removed after the behavior | Reward is offered *before* the behavior, often to stop misbehavior | | **Effect** | Behavior increases | Behavior increases | Usually reinforces the misbehavior that prompted the bribe | | **Example** | Praise after homework is done | Nagging stops when homework is done | "If you stop screaming, I'll buy you the toy" | The difference between reinforcement and bribery is timing and target. Reinforcement follows the desired behavior. Bribery precedes it and is typically triggered by an undesired behavior, so the undesired behavior is what gets strengthened. [More on negative reinforcement ›](https://operantconditioning.com/negative-reinforcement/) ## Common mistakes - **Reinforcing too late.** "Great job on the presentation" three days later is pleasant, not reinforcing. - **Reinforcing the wrong behavior.** Attention given during a tantrum reinforces the tantrum. Reinforcement must be contingent on the behavior you want. - **Assuming the reinforcer.** Stickers do not reinforce every child; public praise does not reinforce every employee. - **Reinforcing every time forever.** Continuous reinforcement builds behavior quickly but leaves it fragile. Thin the schedule. - **Reinforcing outcomes you can't control.** "Lose weight" can't be reinforced in the moment; "walked after lunch" can. - **Pairing praise with criticism.** "Good, but…" turns the praise into a warning signal. ## Does positive reinforcement undermine intrinsic motivation? This is the most serious research-based objection, and the honest answer is "sometimes, under specific conditions." A meta-analysis of 128 studies found that *expected, tangible* rewards for doing an already-interesting activity reduced free-choice engagement with it afterward; verbal praise and unexpected rewards did not, and informational feedback generally increased intrinsic motivation.[6] A competing meta-analysis found the negative effect small and limited to a narrow set of conditions.[7] The practical guidance that follows from both: - Use social and informational reinforcers (specific praise, feedback on progress) freely. - Use tangible rewards for behaviors the person would not otherwise do, and fade them as the behavior contacts natural reinforcers. - Avoid paying people for things they already love doing, especially with rewards contingent merely on engaging rather than on quality. ## How to use positive reinforcement: a six-step protocol 1. **Define the behavior so a camera could see it.** "Raises her hand before speaking," not "is respectful." If you cannot count it, you cannot reinforce it. 2. **Find a reinforcer that works for this individual.** Watch what the person chooses when free, ask, or test. A reinforcer is proven by its effect, not by your intentions; a sticker that a teenager finds embarrassing is not one. 3. **Deliver it immediately and every time — at first.** Within seconds, and after every occurrence, until the behavior is reliable. If the real reinforcer must wait, mark the behavior instantly with a conditioned reinforcer: a word, a click, a check. 4. **Make it contingent and only contingent.** The reinforcer follows the behavior and nothing else. Free access to the same reinforcer at other times drains its power. 5. **Thin the schedule.** Once the behavior is steady, reinforce it only sometimes, unpredictably. Intermittent reinforcement is what makes the behavior survive days when no one is watching. [How to thin a schedule ›](https://operantconditioning.com/schedules-of-reinforcement/) 6. **Measure, and fade to natural reinforcers.** Count the behavior before and after. Then hand the job to consequences the world already provides — the finished essay, the dog's walk, the satisfaction of the check mark — so the behavior no longer depends on you. ### Praise, done properly Praise is the cheapest positive reinforcer available, and most of it is wasted. Jere Brophy's analysis of teacher praise found that it changed behavior only when it was *contingent* (delivered for the behavior, not as a reflex), *specific* (named what was done well), and *credible* (sincere, and not inflated).[10] A later review added a fourth condition: praise for effort and strategy supports motivation, while praise for ability or for merely finishing can undermine it.[11] "You kept the paragraph to one idea — that made it clear" reinforces; "great job" does not. [Praise, token economies, and the Good Behavior Game in the classroom ›](https://operantconditioning.com/classroom/) ### A worked example A second-grade teacher wants a student, Maya, to start her worksheet without a reminder. **Behavior:** pencil on paper within one minute of the worksheet landing on her desk. **Reinforcer:** the teacher notices that Maya lights up when asked to hand out materials, so "hand out the next set" becomes the consequence. **Delivery:** the moment the pencil moves, a quiet "you started on your own — you're handing out the readers at ten." Every time, for a week. **Thinning:** in week two, the job goes to two starts out of three, then to unpredictable ones. **Result:** starts rose from one in five worksheets to almost all of them, and by the end of the month the teacher's occasional nod was enough. Every step is on this page; none of it required a sticker chart. ### Quick check ## Positive reinforcement in practice Positive reinforcement is the default procedure of [applied behavior analysis](https://operantconditioning.com/applications/#aba), the core of reward-based [dog training](https://operantconditioning.com/dog-training/), the basis of classroom systems like PBIS and token economies, the engine of contingency management in addiction treatment, and the mechanism behind most successful [habit-formation](https://operantconditioning.com/habits/) methods. The through-line: identify the behavior, find a reinforcer that actually works for that individual, deliver it immediately and contingently, and thin the schedule over time. ## Key takeaways - "Positive" means a stimulus is added, not that the outcome is pleasant. A reinforcer is not a reward: it is defined only by its effect on the behavior it follows. - Reinforcers are personal and are discovered by watching the behavior, not assumed. Praise a teenager finds embarrassing is not a reinforcer; attention during a tantrum, even scolding, often is. - Contingency and contiguity are both required: the reinforcer arrives because of the behavior and right after it. Bribery is offered before the behavior, usually to stop misbehavior, so the misbehavior is what gets strengthened. - Reinforce every occurrence while the behavior is being learned, then thin to an intermittent schedule and fade to the natural reinforcers the world already provides. - Expected, tangible rewards for an activity a person already enjoys can reduce intrinsic motivation. Specific praise, unexpected rewards, and informational feedback do not. **Explain it to a friend.** Explain why a reinforcer is not the same thing as a reward, using an example from your own week and without using the word "reward" itself. ## Frequently asked questions **What is positive reinforcement in simple terms?** Adding something after a behavior so the behavior happens more often. A dog sits, gets a treat, and sits more often. The treat is the positive reinforcer. **What is an example of positive reinforcement?** A child puts their toys away and a parent says, "Thank you for cleaning up — that was a big help." If the child cleans up more often afterward, the praise was a positive reinforcer. Other examples: a paycheck for hours worked, a like on a post, a treat for a dog's trick. **Is positive reinforcement the same as a reward?** Not exactly. A reward is something intended to be pleasant. A positive reinforcer is defined by its effect: it must actually increase the behavior it follows. Many rewards fail to reinforce, and some unpleasant things (like scolding, which is attention) do reinforce. **What is the difference between positive and negative reinforcement?** Both increase behavior. Positive reinforcement adds a stimulus (a treat appears). Negative reinforcement removes one (an alarm stops). "Positive" and "negative" mean added and removed, not good and bad. **Who developed positive reinforcement?** The principle traces to Edward Thorndike's law of effect (1898). B. F. Skinner developed the concept of reinforcement, the term "positive reinforcement," and the experimental science around it, beginning with *The Behavior of Organisms* (1938). **Is positive reinforcement effective for adults?** Yes. Organizational behavior management uses it to improve safety and productivity, contingency management uses it in addiction treatment with strong evidence, and it is the mechanism behind effective self-management and habit apps. Adults simply have more complex reinforcers (money, status, autonomy, feedback) than treats. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139. 3. Williams, B. A. (1994). Conditioned reinforcement: Experimental and theoretical issues. *The Behavior Analyst, 17*(2), 261–285. 4. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 5. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 6. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 7. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423. 8. Hutt, P. J. (1954). Rate of bar pressing as a function of quality and quantity of food reward. *Journal of Comparative and Physiological Psychology, 47*(3), 235–239. 9. Crespi, L. P. (1942). Quantitative variation of incentive and performance in the white rat. *American Journal of Psychology, 55*(4), 467–517. 10. Brophy, J. (1981). Teacher praise: A functional analysis. *Review of Educational Research, 51*(1), 5–32. 11. Henderlong, J., & Lepper, M. R. (2002). The effects of praise on children's intrinsic motivation: A review and synthesis. *Psychological Bulletin, 128*(5), 774–795. ## Related - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): Removing something to strengthen behavior — and why it isn't punishment. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): How often to reinforce, with a live simulator. - [Reinforcement: the hub](https://operantconditioning.com/reinforcement/): Both types, the kinds of reinforcers, and what makes reinforcement work. --- # Negative Reinforcement: Definition, Examples, and Why It Isn't Punishment > Negative reinforcement removes an aversive stimulus so a behavior increases. Definition, 17 examples, escape vs. avoidance, and why it is not punishment. - Source: https://operantconditioning.com/negative-reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Taking something away* Not punishment — relief. Negative reinforcement is what keeps you buckling your seat belt, hitting snooze, and avoiding what you fear, and it is the most misread idea in operant conditioning. > **Definition** > > **Negative reinforcement** is the process in which a behavior is followed by the *removal* (or prevention) of an aversive stimulus, and as a result the behavior becomes more frequent or more likely in the future. The stimulus that is removed is called a **negative reinforcer**. > > "Negative" means something is subtracted (think minus sign), not that the outcome is bad. Like all reinforcement, negative reinforcement *strengthens* behavior; that is what separates it from punishment, which weakens it.[1] **In brief** - Negative reinforcement removes or prevents an aversive stimulus after a behavior, and the behavior becomes more frequent or more likely. - "Negative" means subtracted, not bad: negative reinforcement strengthens behavior, which is exactly what separates it from punishment. - Escape ends an aversive stimulus that is present; [avoidance](https://operantconditioning.com/avoidance-learning/) prevents one that would have come, and avoidance is extremely persistent. ## How negative reinforcement works In [positive reinforcement](https://operantconditioning.com/positive-reinforcement/), the world is quiet before the behavior and something good appears after it. In negative reinforcement the order is reversed: something unpleasant is already present (or about to arrive), the behavior makes it go away, and the going-away is what teaches. The reinforcer is the *change* — from aversive to not-aversive. - A rat is receiving a mild shock → it presses a lever → the shock stops → pressing increases. - The seat-belt chime is beeping → you buckle up → the chime stops → you buckle faster next time. - Your head is pounding → you take an aspirin → the pain fades → aspirin-taking increases. The same rules apply as for any operant in the [antecedent–behavior–consequence](https://operantconditioning.com/abc-model/) loop: the relief must be *contingent* on the behavior and *immediate*. A chime that stops on its own after 30 seconds teaches you to wait, not to buckle. Skinner devoted whole chapters of *Science and Human Behavior* to aversive control, and argued that governments, schools, and other institutions lean on it far too much.[1] > **The two-question test** > > To classify any consequence, ask only two things. **Did the behavior go up or down?** Up is reinforcement; down is punishment. **Was something added or taken away?** Added is "positive"; taken away is "negative." Negative reinforcement is "behavior went up, something was taken away." How unpleasant the stimulus felt is not part of the test — if removing it increased the behavior, it was aversive by definition. ### Check any scenario in two questions Not sure whether a consequence is this quadrant or a neighbor? Answer the two questions and the checker names it. ## Escape and avoidance: the two kinds of negative reinforcement In **escape**, the aversive stimulus is already present and the behavior ends it: the shock is on, the rat presses, the shock goes off; the baby is crying, the parent picks her up, the crying stops. Escape is learned quickly because the relief is felt directly. ![Escape versus avoidance](https://operantconditioning.com/assets/diagrams/escape-versus-avoidance.svg) *Escape ends an aversive stimulus that is present; avoidance prevents one that would have come. Both increase the behavior, so both are negative reinforcement.* In **avoidance**, the behavior comes first and prevents or postpones the aversive stimulus. Dogs trained with a warning signal in a shuttle box learned to jump the barrier within seconds of the signal and kept jumping for hundreds of trials without ever being shocked again.[2] That persistence is the puzzle of avoidance — how can the *absence* of a shock reinforce anything? Mowrer's two-factor answer was that the signal first acquires fear by classical conditioning and the response is then reinforced by escape from that fear; later work showed that avoidance can be learned with no signal at all (Sidman avoidance) and that a reduction in the overall rate of aversive events can itself be the reinforcer.[3][4][5][6] [Avoidance learning in depth: the paradox, the theories, and learned helplessness ›](https://operantconditioning.com/avoidance-learning/) | Aspect | Escape | Avoidance | | --- | --- | --- | | **Aversive stimulus** | Already present | Not yet present; the behavior prevents or postpones it | | **What the behavior does** | Ends it | Prevents it (or turns off its signal) | | **Example** | Taking a painkiller for a headache | Leaving early to miss the traffic | | **How it is learned** | Directly; the relief is felt | Usually grows out of escape as the response comes earlier | | **Persistence** | Fades when the behavior stops working | Extremely persistent; the threat is never tested | ## Examples of negative reinforcement In every row, an aversive stimulus is present or imminent, a behavior removes or prevents it, and the behavior becomes more likely. | Setting | Aversive stimulus | Behavior | Result | | --- | --- | --- | --- | | Home | Partner nagging about the dishes | Do the dishes; nagging stops | Dishes get done sooner (and the nagging is reinforced too) | | Home | Baby crying | Parent picks the baby up; crying stops | Parent picks up faster next time | | Classroom | Looming pop quiz | Class works quietly all week; quiz is cancelled | Quiet work increases | | Classroom | Hard math worksheet | Student acts out; is sent to the hall | Acting out during math increases (escape-maintained) | | Workplace | Manager's reminder emails | Submit the report; emails stop | Reports go in sooner | | Workplace | Machine noise | Put on ear protection | Ear protection worn more reliably | | Driving | Seat-belt chime | Buckle up; chime stops | Buckling is faster and more reliable | | Driving | Predicted traffic jam | Take the side street | Side street becomes habitual (avoidance) | | Technology | Alarm blaring | Hit snooze; alarm stops | Snoozing increases | | Health | Headache | Take an aspirin; pain fades | Aspirin at the first twinge increases | | Health | Social anxiety at a party | Leave early; anxiety drops | Leaving (then not going) increases | | Dog training | Leash pressure | Dog steps toward handler; pressure released | Dog follows pressure more readily | | Dog training | Mail carrier at the door | Dog barks; carrier leaves (as always) | Barking at the door increases | | Weather | Rain | Open an umbrella | Umbrella use increases; carrying one becomes avoidance | | Self-management | Dread of writing the essay | Clean the apartment instead; dread fades | Procrastination increases | | Self-management | Dread of writing the essay | Write one sentence; dread lifts | Starting increases — the healthy version of the same loop | | Nature | Midday heat | Lizard moves into shade | Shade-seeking increases | Notice how many rows involve *two* people reinforcing each other. When a baby cries and a parent picks her up, the parent's behavior is negatively reinforced (the crying stops) and the baby's crying is positively reinforced (contact arrives). These reciprocal traps explain an enormous amount of family life, and they are covered below. [More examples for every quadrant ›](https://operantconditioning.com/examples/) ## Negative reinforcement vs. punishment This is the confusion the quadrant vocabulary exists to prevent. "Negative" sounds like "bad," so people hear "negative reinforcement" and picture a scolding. But a scolding that reduces a behavior is *positive punishment* (something added, behavior down). Negative reinforcement always *increases* behavior. | Aspect | Positive reinforcement | Negative reinforcement | Positive punishment | Negative punishment | | --- | --- | --- | --- | --- | | **Stimulus** | Added | Removed | Added | Removed | | **Effect on behavior** | Increases | Increases | Decreases | Decreases | | **Kind of stimulus** | Wanted | Unwanted (aversive) | Unwanted (aversive) | Wanted | | **Before the behavior** | Reinforcer absent | Aversive present or imminent | Aversive absent | Reinforcer present | | **Example** | Dog sits → treat | Dog sits → leash pressure released | Dog jumps → sprayed with water | Dog jumps → person turns away | | **Page** | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | This page | [Positive punishment](https://operantconditioning.com/positive-punishment/) | [Negative punishment](https://operantconditioning.com/negative-punishment/) | The two often use the *same* aversive stimulus for opposite jobs. A shock that stops when the rat presses negatively reinforces pressing; a shock that starts when the rat presses punishes it. A parent who nags until the room is clean is using negative reinforcement; one who starts nagging on seeing the mess is using positive punishment. ## The negative reinforcement trap: how it maintains problem behavior Reinforcement does not care whether anyone wants the behavior it strengthens. Negative reinforcement quietly maintains a large share of the behaviors people go to therapists, behavior analysts, and trainers to get rid of. ### Escape-maintained behavior In applied behavior analysis, a **functional analysis** tests which consequence maintains a problem behavior by presenting each candidate in turn.[7] In the *demand* condition, tasks are presented and briefly withdrawn if the problem behavior occurs; behavior that peaks there is **escape-maintained**. In the largest published series of functional analyses of self-injury, escape from demands was the single most common function identified.[8] The everyday version needs no clinic. A child is asked to put on shoes; she screams; the parent, late for work, drops the demand and carries her to the car. Screaming was negatively reinforced (the demand vanished) and so was giving up (the screaming stopped). Gerald Patterson called this the **coercive family process**: parent and child train each other, in escalating rounds, to use aversive behavior to get their way.[9] The remedy is not more punishment but **escape extinction** — following through so the behavior no longer works — plus reinforcement for compliance and easier, better-prompted demands.[10] [How escape extinction differs from ignoring ›](https://operantconditioning.com/extinction/#extinction-is-not-ignoring) ### Avoidance and anxiety Two-factor theory turned out to describe human anxiety well. A person who fears elevators takes the stairs; the anxiety drops; stair-taking is negatively reinforced. Because they never ride the elevator, they never learn that nothing bad happens, so the fear never extinguishes. The same loop maintains compulsions (checking relieves doubt), social avoidance (leaving relieves dread), and panic-related avoidance.[11] Exposure-based treatment is built on *blocking the avoidance response* long enough for fear to fade — extinction applied to a behavior that negative reinforcement has protected for years. ### Procrastination Procrastination is escape, not laziness. The thought of the task is aversive, doing something else makes the feeling go away, and that relief reinforces the diversion — researchers describe it as short-term mood repair at the expense of long-term goals.[12] The fix follows from the analysis: make the first step so small that the relief of having started outweighs the dread of starting, so that *starting* becomes the escape response. That is why "open the document and write one sentence" works when "write the essay" does not, and it is the logic behind the [tiny-behavior approach to habits](https://operantconditioning.com/habits/). ### Nagging, whining, and the snooze button Nagging persists because it works intermittently, and intermittent reinforcement builds the most durable behavior of all (see [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/)). Whining persists the same way: hold out four times, give in on the fifth, and you have trained a five-whine behavior. The snooze button is the purest case — snooze removes the alarm instantly, every morning, which is why the alarm across the room is the only one that reliably works. ## How to use negative reinforcement ethically Negative reinforcement is not inherently unethical — the seat-belt chime saves lives. But it requires an aversive stimulus to be present, so it carries the side effects of aversive control (stress, escape, avoidance), and if you have to *create* the aversive yourself, you become the thing the learner wants to get away from.[13] 1. **Prefer positive reinforcement.** Use negative reinforcement when an aversive is already unavoidably present — pain, noise, a deadline, a task that must be done. 2. **Use natural aversives, not manufactured ones.** Finishing homework ending the homework is fine. Inventing an unpleasant condition so you can remove it is coercion. 3. **Make the desired behavior the fastest route to relief.** If tantrums end demands faster than compliance does, tantrums win. 4. **Close the other exits.** If the learner can also escape by lying, hiding, or quitting, those get reinforced too. 5. **Pair it with positive reinforcement, then fade it.** The goal is behavior maintained by natural positive consequences, not by relief from something you hold over the learner. 6. **Watch the side effects.** If the learner starts avoiding *you*, the procedure has failed whatever the target behavior is doing. In [dog training](https://operantconditioning.com/dog-training/), methods built on negative reinforcement — leash corrections that stop when the dog complies, electronic collars whose stimulation ends on recall — are associated with more stress behaviors and no better obedience than reward-based methods, which is why veterinary behavior organizations recommend against them.[14] ## Common mistakes - **Calling it punishment.** If the behavior increased, it was reinforcement, whatever it felt like. - **Missing your own reinforcement.** Parents, teachers, and managers are negatively reinforced every time giving in makes an unpleasant situation stop. Ask what *your* behavior is escaping from. - **Reinforcing the escape instead of the task.** Sending a disruptive student out of class may be exactly the consequence the student was working for. - **Waiting for avoidance to extinguish on its own.** It won't; successful avoidance never contacts the change. Something has to block the response. - **Confusing it with negative punishment.** Both remove something; one strengthens behavior (removing an aversive), the other weakens it (removing a reinforcer). [Negative punishment explained ›](https://operantconditioning.com/negative-punishment/) ## Is the positive/negative distinction even real? Some behavior analysts think not. Jack Michael argued in 1975 that the two cannot be told apart in many real cases — is a warm room after a cold one the addition of warmth or the removal of cold? — and that the field should simply say "reinforcement."[15] The distinction survived because it maps onto clearly different laboratory procedures and forces students to notice what was present *before* the behavior.[16] Don't agonize over borderline cases; ask what the behavior accomplishes and whether that makes it more likely. ## Key takeaways - Negative reinforcement is "behavior went up, something was taken away." How unpleasant the stimulus felt is not part of the test; if removing it increased the behavior, it was aversive by definition. - It is not punishment: a scolding that reduces a behavior is positive punishment, while nagging that stops when the room is cleaned is negative reinforcement. The same aversive stimulus can do either job, depending on whether the behavior ends it or starts it. - Escape ends an aversive stimulus that is already present and is learned directly. Avoidance prevents one that would have come, and it persists because the threat is never tested. - Negative reinforcement quietly maintains much of the behavior people try to get rid of: escape-maintained tantrums, procrastination, and anxious avoidance. The fix is to stop the escape from working and make the desired behavior the fastest route to relief, not to add punishment. - Prefer positive reinforcement. Use negative reinforcement only when an aversive is already unavoidably present, and never invent an unpleasant condition so you can remove it. ### Check yourself **A teacher sends a student to the hallway whenever he acts out during math. Acting out during math becomes more frequent. Is the teacher punishing the behavior?** No. The behavior increased, so the consequence was reinforcement, whatever the teacher intended. Being sent out removed the hard worksheet, so the acting out is escape-maintained: negatively reinforced by the removal of the demand. **A parent starts nagging the moment she sees a messy room, and messes become less frequent. A friend calls this negative reinforcement. Is the friend right?** No. The behavior went down and something was added, so this is positive punishment. Negative reinforcement would be nagging that stops once the room is cleaned, which makes cleaning more likely. **Someone who fears elevators has taken the stairs for years, and the fear has never faded. Why not?** Taking the stairs is avoidance: the anxiety drops, so stair-taking is negatively reinforced, and the elevator is never ridden. Because the threat is never tested, the fear never gets the chance to extinguish; something has to block the avoidance response. **You open an umbrella when rain starts. Later you begin carrying one whenever rain is forecast. Which is escape, and which is avoidance?** Opening the umbrella in the rain is escape: the aversive stimulus is present and the behavior ends it. Carrying one so you never get wet is avoidance: the behavior comes first and prevents the stimulus. Both increase the behavior, so both are negative reinforcement. **Explain it to a friend.** Explain why negative reinforcement is not punishment, without using the words "positive" or "negative." ## Frequently asked questions **What is negative reinforcement in simple terms?** Removing something unpleasant after a behavior so the behavior happens more often. The seat-belt chime stops when you buckle up, so you buckle up faster. The unpleasant thing that goes away is the negative reinforcer. **What is an example of negative reinforcement?** Taking an aspirin to end a headache: the headache is present, you take the pill, the pain fades, and you take aspirin sooner next time. Others: hitting snooze, doing chores to stop nagging, opening an umbrella, leaving a party to escape anxiety. **Is negative reinforcement the same as punishment?** No. Negative reinforcement increases a behavior by removing something unpleasant; punishment decreases a behavior. "Negative" only means something was taken away. A scolding that reduces a behavior is positive punishment; nagging that stops when the room is cleaned is negative reinforcement. **What is the difference between escape and avoidance?** In escape, the unpleasant stimulus is already present and the behavior ends it (a painkiller for a headache). In avoidance, the behavior comes first and prevents the stimulus (leaving early to miss traffic). Avoidance is persistent because the person never finds out whether the threat is still real. **Is negative reinforcement bad?** Not inherently — pain relief and seat-belt chimes depend on it. But it requires an aversive stimulus, so it shares punishment's side effects, and it silently maintains tantrums, procrastination, and anxiety. Positive reinforcement is usually the better tool when you have a choice. **What is negative reinforcement in the classroom?** Any case where a behavior removes something a student finds aversive. Useful: finishing work early ends the work period. Problematic: a student acts out during hard tasks and is sent out of the room, which removes the task and strengthens acting out — escape-maintained behavior. **Who came up with negative reinforcement?** B. F. Skinner introduced the reinforcement vocabulary in *The Behavior of Organisms* (1938) and elaborated it in *Science and Human Behavior* (1953). Avoidance research was advanced by O. H. Mowrer's two-factor theory (1947) and Murray Sidman's unsignaled avoidance experiments (1953). ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Solomon, R. L., & Wynne, L. C. (1953). Traumatic avoidance learning: Acquisition in normal dogs. *Psychological Monographs, 67*(4), 1–19. 3. Mowrer, O. H. (1947). On the dual nature of learning — a re-interpretation of "conditioning" and "problem-solving." *Harvard Educational Review, 17*, 102–148. 4. Mowrer, O. H. (1960). *Learning Theory and Behavior*. Wiley. 5. Sidman, M. (1953). Avoidance conditioning with brief shock and no exteroceptive warning signal. *Science, 118*(3058), 157–158. 6. Herrnstein, R. J. (1969). Method and theory in the study of avoidance. *Psychological Review, 76*(1), 49–69. 7. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 8. Iwata, B. A., Pace, G. M., Dorsey, M. F., Zarcone, J. R., Vollmer, T. R., Smith, R. G., et al. (1994). The functions of self-injurious behavior: An experimental-epidemiological analysis. *Journal of Applied Behavior Analysis, 27*(2), 215–240. 9. Patterson, G. R. (1982). *Coercive Family Process*. Castalia. 10. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 11. Barlow, D. H. (2002). *Anxiety and Its Disorders: The Nature and Treatment of Anxiety and Panic* (2nd ed.). Guilford Press. 12. Sirois, F., & Pychyl, T. (2013). Procrastination and the priority of short-term mood regulation: Consequences for future self. *Social and Personality Psychology Compass, 7*(2), 115–127. 13. Iwata, B. A. (1987). Negative reinforcement in applied behavior analysis: An emerging technology. *Journal of Applied Behavior Analysis, 20*(4), 361–378. 14. Ziv, G. (2017). The effects of using aversive training methods in dogs — A review. *Journal of Veterinary Behavior, 19*, 50–60. See also Vieira de Castro, A. C., Fuchs, D., Morello, G. M., Pastur, S., de Sousa, L., & Olsson, I. A. S. (2020). Does training method matter? Evidence for the negative impact of aversive-based methods on companion dog welfare. *PLoS ONE, 15*(12), e0225023. 15. Michael, J. (1975). Positive and negative reinforcement, a distinction that is no longer necessary; or a better way to talk about bad things. *Behaviorism, 3*(1), 33–44. 16. Baron, A., & Galizio, M. (2005). Positive and negative reinforcement: Should the distinction be preserved? *The Behavior Analyst, 28*(2), 85–98. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): Adding something to strengthen behavior — the tool to reach for first. - [Positive punishment](https://operantconditioning.com/positive-punishment/): Adding an aversive to weaken behavior, and what the evidence says about it. - [Reinforcement: the hub](https://operantconditioning.com/reinforcement/): Both types, the kinds of reinforcers, and what makes reinforcement work. --- # Positive Punishment: Definition, Examples, Side Effects, and What the Evidence Says > Positive punishment adds an aversive stimulus after a behavior to reduce it. Definition, 16 examples, side effects, the spanking evidence, and alternatives. - Source: https://operantconditioning.com/positive-punishment/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Punishment · Adding something* Add something unpleasant after a behavior and the behavior shrinks. It works — under conditions that are rarely met — and it comes with a bill. Here is what the research shows, and what to do instead. > **Definition** > > **Positive punishment** is the process in which a behavior is followed by the *addition* of a stimulus, and as a result the behavior becomes less frequent or less likely in the future. The added stimulus is called a **punisher**. > > "Positive" means something is added, not that the procedure is good. And a stimulus is a punisher only if it actually reduces the behavior it follows; something that feels unpleasant but leaves the behavior unchanged is not punishment in the technical sense.[1] **In brief** - Positive punishment adds a stimulus after a behavior, and the behavior becomes less frequent or less likely in the future. - It can suppress behavior, but only when it is immediate, consistent, intense from the start, inescapable, and paired with a reinforced alternative. - It brings side effects — escape, avoidance, aggression, fear — and teaches nothing to do instead, so reinforcing an alternative is usually the better tool. ## How positive punishment works Positive punishment is the mirror image of [positive reinforcement](https://operantconditioning.com/positive-reinforcement/): the behavior occurs, a stimulus appears that was not there before, and the future rate of the behavior drops. - A rat presses a lever → a brief shock follows → pressing decreases. - A toddler touches the stove → it burns → stove-touching decreases (a natural punisher). - A driver runs a red light → a ticket arrives → red-light running decreases, at least at that intersection. Like reinforcement, punishment depends on *contingency* and *contiguity*. A ticket that arrives three weeks later is a weak punisher for a specific act of speeding, which is why cameras and points systems aim to make the consequence certain rather than severe. The same [A-B-C analysis](https://operantconditioning.com/abc-model/) applies: antecedent, behavior, and a consequence that changes the future probability of that behavior in that context. > **Punishers are defined by their effect, not their intent** > > A teacher who yells at a class clown intends to punish. But if the yelling is the most attention the student gets all day, clowning goes *up* — the yelling was a positive reinforcer. The only way to know whether a consequence is punishing is to measure the behavior afterward. "I punished him and it didn't work" is a contradiction in terms. ### Check any scenario in two questions Not sure whether a consequence is this quadrant or a neighbor? Answer the two questions and the checker names it. ## Examples of positive punishment | Setting | Behavior | Stimulus added | Result | | --- | --- | --- | --- | | Home | Child touches a hot pan | Burn | Touching hot pans decreases (natural punisher) | | Home | Child hits a sibling | Firm verbal reprimand | Hitting decreases — if the reprimand functions as a punisher | | Home | Teenager comes home late | Extra chores | Lateness decreases | | Classroom | Student talks during instruction | Teacher's pointed look and name said aloud | Talking decreases | | Classroom | Student writes on the desk | Must clean every desk in the room (overcorrection) | Desk-writing decreases | | Workplace | Employee skips a safety step | Written warning | Skipped steps decrease | | Driving | Driver speeds past a camera | Fine in the mail | Speeding at that spot decreases | | Driving | Driver drifts out of the lane | Rumble-strip vibration | Lane drifting decreases | | Dog training | Dog jumps on the couch | Spray of water | Couch-jumping decreases (when the owner is present) | | Dog training | Dog pulls on the leash | Leash jerk | Pulling decreases, with stress side effects | | Technology | User mistypes a password five times | Lockout and warning | Careless attempts decrease | | Health | Person skips sunscreen at the beach | Painful sunburn | Skipping sunscreen decreases (natural punisher) | | Health | Person drinks alcohol while taking disulfiram | Flushing, nausea, headache | Drinking decreases | | Sports | Player commits a foul | Yellow card and coach's rebuke | Fouling decreases | | Self-management | Person bites their nails | Bitter-tasting nail polish | Nail-biting decreases (while the taste is present) | | Nature | Young dog pounces on a bee | Sting on the nose | Pouncing on buzzing insects decreases | Notice that the natural punishers — burns, nausea — produce fast, durable learning, while many human examples carry a qualifier: "when the owner is present," "at that spot." Contrived punishment tends to suppress behavior in the punisher's presence rather than eliminate it. [Examples for every quadrant ›](https://operantconditioning.com/examples/) ## What the research says about punishment ### Thorndike and Skinner: the early doubts Thorndike's original law of effect (1898) gave punishment equal billing: responses followed by discomfort were "stamped out." By 1932 he had concluded that annoying consequences did not weaken connections nearly as directly as satisfying ones strengthened them.[2] Skinner reached the same position from the laboratory. His student W. K. Estes trained rats to press a lever for food, then briefly punished pressing with shock during extinction. Pressing dropped sharply while the shock was in effect, then recovered; by the end of extinction the punished rats had made about as many responses as unpunished ones.[3] Skinner concluded that punishment merely *suppresses* behavior temporarily, and argued against it for the rest of his career.[1] ### Azrin and Holz: when punishment does work The picture changed in the 1960s, when Nathan Azrin and colleagues ran punishment through the same parametric analysis Skinner had applied to reinforcement. Their landmark 1966 chapter concluded that punishment can produce complete and lasting suppression — but only under specific conditions.[4][5] | Condition | What the research found | Why it is hard to meet | | --- | --- | --- | | **Immediacy** | The punisher should follow within seconds; delayed punishment is far weaker. | Most human punishment is delayed by hours or weeks. | | **Intensity from the start** | A punisher introduced at full intensity suppresses behavior; one that starts mild and is gradually increased produces adaptation. | Parents and trainers naturally escalate, which trains tolerance. | | **Consistency** | Every occurrence should be punished; intermittent punishment is much weaker. | Nobody catches every instance; the behavior is reinforced on the rest. | | **No unauthorized escape** | If the organism can avoid the punisher by hiding, lying, or leaving, it learns the escape instead. | Humans are very good at finding escape routes. | | **An alternative response** | Suppression is far greater when another behavior produces the same reinforcer. | Punishment is usually applied without teaching a replacement. | | **Reduced motivation** | Punishment works better when the maintaining reinforcer is weak or unavailable. | The maintaining reinforcer is often unknown. | The same research documented the costs: escape and avoidance of the punishing agent, aggression elicited by aversive stimulation, and the odd finding that a punisher can become a signal for reinforcement, so that punishment paradoxically *increases* behavior when it predicts reward.[4] Russell Church's review reached a similar verdict: punishment is a lawful process, but its effects depend on parameters most users never control.[6] ### The applied literature In 2002, Dorothea Lerman and Christina Vorndran reviewed what applied behavior analysts actually knew about punishment: laboratory findings on intensity, immediacy, and schedule had rarely been tested with people in clinical settings; punishment sometimes remained necessary when reinforcement-based treatments failed for dangerous behavior; and the field needed better data on using it with the least intensity and fewest side effects.[7] Punishment is neither the reliable tool folk wisdom assumes nor the useless one Estes's rats suggested. ## Side effects of positive punishment Even when punishment suppresses a target behavior, it produces collateral effects that reinforcement does not.[4][7] - **Escape and avoidance.** The learner avoids the punisher — and the person, place, or task associated with it. The child punished for spilling hides spills; the employee stops reporting mistakes; the dog steals food only when no one is in the room. - **Aggression.** Aversive stimulation elicits fighting. Rats shocked together attack each other, and the same reflexive aggression appears across species.[8] Punished children and animals lash out, often at an unrelated target. - **Emotional responding.** Fear, crying, freezing, and disruption of behavior you wanted to keep. A punished student may stop talking out of turn and also stop participating. - **Modeling.** Punishment teaches that adding aversive consequences is how you change people. Children who are hit are more likely to hit. - **Suppression, not elimination.** The behavior returns when the punisher is absent or the motivation rises. Punishment does not teach what to do instead. - **The punisher becomes a conditioned aversive stimulus.** Anything reliably paired with punishment — the parent's raised voice, the classroom, the training collar — acquires aversive properties of its own, spreading escape and avoidance further. ## Corporal punishment: what the evidence shows Spanking is the most studied form of positive punishment in humans. A 2002 meta-analysis of 88 studies found corporal punishment associated with immediate compliance and with worse outcomes on essentially every other measure — aggression, antisocial behavior, mental health, the parent–child relationship.[9] A 2016 meta-analysis by Elizabeth Gershoff and Andrew Grogan-Kaylor addressed the objection that earlier work had lumped spanking together with abuse. Restricting the analysis to ordinary open-handed spanking across 75 studies and more than 160,000 children, they found spanking significantly associated with 13 of the 17 outcomes examined, all in the harmful direction; none favored spanking.[10] The evidence is largely correlational, and it is fair to ask whether difficult children simply get spanked more. But longitudinal studies that control for prior behavior still find spanking predicting later increases in problem behavior, and no comparable evidence shows benefits. In 2018 the American Academy of Pediatrics recommended that parents not spank, hit, or otherwise physically punish children, nor use verbal abuse that shames or humiliates, and instead rely on positive reinforcement, limit-setting, and brief time-out.[11] The operant analysis predicts exactly this: spanking as actually practiced — inconsistent, delayed, escalating, delivered by a person the child depends on — violates every condition on Azrin and Holz's list. ## Aversives in dog training: what the evidence shows The same question has been studied in companion dogs, with the same answer. An early owner survey found reward-based methods associated with higher obedience and punishment-based methods with more problem behaviors.[12] A survey at a veterinary behavior clinic found that confrontational techniques — hitting, alpha rolls, leash jerks — frequently provoked an aggressive response from the dog.[13] A 2017 review concluded that aversive methods carry welfare risks and no advantage in effectiveness.[14] And a 2020 study of dogs at aversive-based versus reward-based training schools found the aversively trained dogs showed more stress behaviors, higher salivary cortisol, and a more pessimistic bias in a cognitive test outside training.[15] Positive punishment trades short-term suppression for long-term stress, which is why the reward-based methods on our [dog training page](https://operantconditioning.com/dog-training/) are the standard recommendation of veterinary behavior organizations. ## When is punishment used in applied behavior analysis? Rarely, and only under conditions. In 1988 a task force of the Association for Behavior Analysis affirmed that clients have a right to *effective* treatment, which may in rare cases include restrictive procedures when less intrusive ones have failed and the behavior is dangerous.[16] The Behavior Analyst Certification Board's ethics code operationalizes that position: prioritize reinforcement-based procedures; recommend punishment only when the severity of the behavior or the failure of less intrusive procedures warrants it; always include reinforcement of alternative behavior; and monitor data and side effects continuously.[17] In practice a punishment procedure in [ABA](https://operantconditioning.com/applications/#aba) follows a functional assessment, sits inside a plan that teaches a replacement, uses the least intrusive procedure that works, requires informed consent and oversight, and is faded as soon as possible.[18] The contrast with everyday punishment — improvised, delayed, escalating, standalone — could hardly be sharper. ## Natural vs. contrived punishers A **natural punisher** is produced by the behavior itself: the burn from the stove, the fall from tipping a chair. A **contrived** punisher is delivered by another person: the reprimand, the fine, the leash jerk. Natural punishers teach efficiently because they are immediate, perfectly consistent, impossible to escape, and there is no punishing agent to become a conditioned aversive stimulus. "Let natural consequences happen" is sound advice whenever the consequence is safe: a child who forgets a coat gets cold, and cold is a better teacher than a lecture. ## Reprimands and overcorrection ### Reprimands A **reprimand** is a brief, firm verbal statement delivered immediately after the behavior. Classroom research found that *soft* reprimands, audible only to the target student, reduced disruptive behavior more than loud public ones, which can function as attention and entertainment for the class.[19] Reprimands work better delivered up close, with eye contact, than shouted across the room.[20] They lose their effect when used constantly or when the student is reinforced by the reaction. ### Overcorrection **Overcorrection** requires the person to repair the effects of the behavior beyond the original state (*restitutional* — clean the whole table, not just your spill) or to practice a correct alternative repeatedly (*positive practice* — walk calmly to the door five times after running). Foxx and Azrin developed it in the early 1970s for severe behavior in institutional settings.[21] It is effortful, can provoke aggression when the person is physically guided, and is used far less today than reinforcement-based procedures. ## Alternatives to positive punishment Because punishment does not teach a replacement, the most effective way to reduce a behavior is usually to build something else in its place. | Procedure | What you do | Example | | --- | --- | --- | | **DRA** — differential reinforcement of alternative behavior | Reinforce an appropriate behavior that serves the same function; withhold reinforcement for the problem behavior | A child who grabs toys is taught to ask, and asking gets the toy | | **DRI** — differential reinforcement of incompatible behavior | Reinforce a behavior that cannot happen at the same time | A dog that jumps on guests is reinforced for sitting when they arrive | | **DRO** — differential reinforcement of other behavior | Reinforce each interval that passes without the problem behavior | A student earns a point for each five minutes with no shouting | | **Extinction** | Identify the maintaining reinforcer and stop delivering it | Whining for candy never produces candy again | | **Antecedent changes** | Alter the setting so the behavior is less likely to be triggered or needed | Move the cookie jar; break the hard worksheet into pieces; offer a choice | | **Negative punishment** | Remove a reinforcer contingent on the behavior | Brief time-out from play after hitting | Differential reinforcement paired with [extinction](https://operantconditioning.com/extinction/) is the default treatment package in ABA and the core of evidence-based parent training. [Negative punishment](https://operantconditioning.com/negative-punishment/) — removing a reinforcer rather than adding an aversive — carries fewer side effects and is what pediatricians mean by time-out. Antecedent strategies are the most under-used option: the easiest way to reduce a behavior is often to change the situation that occasions it, as our [guide to the ABC model](https://operantconditioning.com/abc-model/) explains. ## Positive punishment vs. negative punishment vs. negative reinforcement Positive punishment *adds* an aversive to *decrease* behavior (a fine for speeding). Negative punishment *removes* a reinforcer to *decrease* behavior (losing the car keys for speeding). [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) *removes* an aversive to *increase* behavior (slowing down ends the passenger's complaints). "Positive" and "negative" only ever mean added and removed; whether it is punishment or reinforcement depends on which way the behavior moves. ## Key takeaways - "Positive" means added, not good. A stimulus is a punisher only if it reduces the behavior it follows; "I punished him and it didn't work" is a contradiction in terms. - Punishment can produce complete suppression, but only when it is immediate, at full intensity from the start, consistent, impossible to escape, and paired with a reinforced alternative. Everyday punishment is usually delayed, escalating, and inconsistent, which is why it mostly suppresses behavior in the punisher's presence. - Its side effects are escape and avoidance of the punisher, aggression, fear, modeling of aversive control, and a punisher that becomes a conditioned aversive stimulus. Punishment does not teach what to do instead. - Spanking is associated with worse outcomes on nearly every measure studied and with no long-term benefits, and pediatricians recommend positive reinforcement, limit-setting, and brief time-out instead. Aversive dog-training methods show the same pattern: more stress and no advantage in effectiveness. - The most effective way to reduce a behavior is usually to build something else in its place: differential reinforcement, extinction, antecedent changes, or negative punishment, which carries fewer side effects. ### Check yourself **A teacher yells at a class clown after every joke. Over the month, the joking increases. What was the yelling?** A positive reinforcer. Consequences are classified by their effect, not their intent: the behavior went up and something was added. If the yelling is the most attention the student gets all day, it reinforces the clowning it was meant to stop. **A dog is sprayed with water whenever it jumps on the couch, and it stops jumping — while the owner is home. What has the dog learned?** The behavior has been suppressed in the punisher's presence, not eliminated. Contrived punishment tends to work only where the punisher is, and the dog has found an escape route: jumping when no one is in the room. Nothing has taught it what to do instead. **A parent scolds mildly, then louder, then more firmly as a behavior continues. Why does this escalation tend to fail?** A punisher that starts mild and is gradually increased produces adaptation; each step trains tolerance of the next. Punishment suppresses behavior when it is introduced at full intensity from the start, which is one reason it is so hard to use well outside a laboratory. **One child forgets a coat and gets cold. Another is lectured about forgetting coats. Which consequence teaches better, and why?** The cold. It is a natural punisher: immediate, perfectly consistent, impossible to escape, and with no punishing agent to become a conditioned aversive stimulus. The lecture is a contrived punisher delivered by a person, and that person can become something the child learns to avoid. **Explain it to a friend.** Explain why punishment that seems to work in the moment usually fails in the long run, in two sentences a twelve-year-old would follow. ## Frequently asked questions **What is positive punishment in simple terms?** Adding something unpleasant after a behavior so the behavior happens less often. A dog jumps on the couch, gets sprayed with water, and jumps less. The spray is the positive punisher. "Positive" means added, not good. **What is an example of positive punishment?** A driver runs a red light and gets a ticket; red-light running decreases. Others: a child touches a hot stove and is burned, a student receives a reprimand for talking, an employee gets a written warning for skipping a safety check, a bird eats a bitter insect and vomits. **What is the difference between positive and negative punishment?** Both decrease behavior. Positive punishment adds an aversive stimulus (a reprimand, a fine, a shock). Negative punishment removes a pleasant one (a privilege, tokens, attention — as in time-out). Negative punishment has fewer side effects and is preferred when punishment is used at all. **Does positive punishment work?** It can reduce behavior when it is immediate, consistent, intense enough from the start, impossible to escape, and paired with a reinforced alternative. Those conditions are rarely met outside a laboratory. In everyday use it suppresses behavior mainly in the punisher's presence, with side effects: avoidance, aggression, fear, and imitation of aversive control. **Is spanking positive punishment?** Yes — an aversive stimulus is added after a behavior to reduce it. Meta-analyses find spanking associated with worse outcomes across the board (more aggression, more behavior problems, worse mental health) and no long-term benefits, and the American Academy of Pediatrics recommends against it. **Is positive punishment ever used in ABA therapy?** Rarely, and only within strict limits. Ethics guidelines require reinforcement-based procedures first, reserve punishment for dangerous behavior that has not responded to less intrusive treatment, require reinforcement of an alternative behavior alongside it, and require consent and monitoring of side effects. **Why is punishment less effective than reinforcement?** Punishment says what not to do but not what to do instead, so the motivation finds another outlet. It also creates escape and avoidance (including sneaking and lying), elicits aggression and fear, turns the punisher into something to avoid, and usually suppresses behavior only while the punisher is around. Reinforcement builds a replacement that persists on its own. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Thorndike, E. L. (1932). *The Fundamentals of Learning*. Teachers College, Columbia University. 3. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 4. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 5. Azrin, N. H. (1960). Effects of punishment intensity during variable-interval reinforcement. *Journal of the Experimental Analysis of Behavior, 3*(2), 123–142. 6. Church, R. M. (1963). The varied effects of punishment on behavior. *Psychological Review, 70*(5), 369–402. 7. Lerman, D. C., & Vorndran, C. M. (2002). On the status of knowledge for using punishment: Implications for treating behavior disorders. *Journal of Applied Behavior Analysis, 35*(4), 431–464. 8. Ulrich, R. E., & Azrin, N. H. (1962). Reflexive fighting in response to aversive stimulation. *Journal of the Experimental Analysis of Behavior, 5*(4), 511–520. 9. Gershoff, E. T. (2002). Corporal punishment by parents and associated child behaviors and experiences: A meta-analytic and theoretical review. *Psychological Bulletin, 128*(4), 539–579. 10. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 11. Sege, R. D., Siegel, B. S., Council on Child Abuse and Neglect, & Committee on Psychosocial Aspects of Child and Family Health. (2018). Effective discipline to raise healthy children. *Pediatrics, 142*(6), e20183112. 12. Hiby, E. F., Rooney, N. J., & Bradshaw, J. W. S. (2004). Dog training methods: Their use, effectiveness and interaction with behaviour and welfare. *Animal Welfare, 13*(1), 63–69. 13. Herron, M. E., Shofer, F. S., & Reisner, I. R. (2009). Survey of the use and outcome of confrontational and non-confrontational training methods in client-owned dogs showing undesired behaviors. *Applied Animal Behaviour Science, 117*(1–2), 47–54. 14. Ziv, G. (2017). The effects of using aversive training methods in dogs — A review. *Journal of Veterinary Behavior, 19*, 50–60. 15. Vieira de Castro, A. C., Fuchs, D., Morello, G. M., Pastur, S., de Sousa, L., & Olsson, I. A. S. (2020). Does training method matter? Evidence for the negative impact of aversive-based methods on companion dog welfare. *PLoS ONE, 15*(12), e0225023. 16. Van Houten, R., Axelrod, S., Bailey, J. S., Favell, J. E., Foxx, R. M., Iwata, B. A., & Lovaas, O. I. (1988). The right to effective behavioral treatment. *Journal of Applied Behavior Analysis, 21*(4), 381–384. 17. Behavior Analyst Certification Board. (2020). *Ethics Code for Behavior Analysts*. BACB. 18. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 19. O'Leary, K. D., Kaufman, K. F., Kass, R. E., & Drabman, R. S. (1970). The effects of loud and soft reprimands on the behavior of disruptive students. *Exceptional Children, 37*(2), 145–155. 20. Van Houten, R., Nau, P. A., MacKenzie-Keating, S. E., Sameoto, D., & Colavecchia, B. (1982). An analysis of some variables influencing the effectiveness of reprimands. *Journal of Applied Behavior Analysis, 15*(1), 65–83. 21. Foxx, R. M., & Azrin, N. H. (1973). The elimination of autistic self-stimulatory behavior by overcorrection. *Journal of Applied Behavior Analysis, 6*(1), 1–14. ## Related - [Negative punishment](https://operantconditioning.com/negative-punishment/): Time-out and response cost — removing a reinforcer to weaken behavior. - [Extinction](https://operantconditioning.com/extinction/): Reducing behavior by withholding its reinforcer, and what to expect when you do. - [Punishment: the hub](https://operantconditioning.com/punishment/): Both types, what the evidence says, and the alternatives. --- # Negative Punishment: Definition, Examples, Time-Out, and Response Cost > Negative punishment removes a reinforcer after a behavior so it decreases. Definition, examples, time-out and response cost, and how it differs from extinction. - Source: https://operantconditioning.com/negative-punishment/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Punishment · Taking something away* Take something good away after a behavior and the behavior shrinks. Time-out, fines, lost privileges, and penalty points all work this way — when they are done right, which is less often than you would think. > **Definition** > > **Negative punishment** is the process in which a behavior is followed by the *removal* of a reinforcing stimulus, and as a result the behavior becomes less frequent or less likely in the future. Its two main forms are **time-out from positive reinforcement** and **response cost**. > > "Negative" means something is subtracted; "punishment" means the behavior goes down. Whether a removal is actually punishing is decided by its effect on the behavior, not by how much the person seems to mind.[1] **In brief** - Negative punishment removes a reinforcing stimulus after a behavior, and the behavior becomes less frequent; its main forms are time-out and response cost. - Time-out only removes what time-in provides; if the situation the child leaves was not rewarding, removal is an escape, not a punishment. - It is preferred over [positive punishment](https://operantconditioning.com/positive-punishment/) because it adds no aversive stimulus, but it teaches no replacement and is weak when delayed. ## How negative punishment works Negative punishment requires that a reinforcer be present or available *before* the behavior, so that it can be taken away after. A teenager who has the car keys can lose them; a child who is playing can be removed from play; a driver with a clean license can lose points. The behavior costs the person something they had, and its future rate drops. - A child hits a sibling → is removed from the game for two minutes → hitting decreases. - A player fouls → sits out for two minutes → fouling decreases. - A driver is caught speeding → loses three points → speeding decreases. The usual rules apply. The loss must follow the behavior closely (a privilege revoked on Friday for something done on Monday teaches little), be contingent on the behavior and nothing else, and be consistent. The [A-B-C analysis](https://operantconditioning.com/abc-model/) is the same as for reinforcement, except that the consequence is a subtraction. > **The reinforcer has to be real** > > You cannot remove what the person does not value. Taking away dessert from a child who does not like dessert is not negative punishment. Sending a bored student out of a boring classroom is not either — it removes nothing the student wants, and it may remove something they wanted to escape, which makes it [negative reinforcement](https://operantconditioning.com/negative-reinforcement/) of the behavior that got them sent out. ### Check any scenario in two questions Not sure whether a consequence is this quadrant or a neighbor? Answer the two questions and the checker names it. ## Time-out from positive reinforcement ### What time-out actually is The full name matters: **time-out from positive reinforcement**. Time-out is a period during which, contingent on a behavior, the person loses access to the reinforcers that were available a moment earlier. The chair, the corner, and the bedroom are incidental; the procedure is the *removal of access to reinforcement*. It is not "go and think about what you did," and it works only when the environment the child leaves — the **time-in** — is actually rewarding. Time-out entered clinical practice in the early 1960s. In one of the first applied case studies, a three-and-a-half-year-old boy whose tantrums and self-injury had made him impossible to treat was placed briefly in his room contingent on tantrums while appropriate behavior was reinforced; the tantrums declined until he could wear his glasses, eat at the table, and eventually attend school.[2] Time-out has since become one of the most researched discipline techniques in existence and a standard component of evidence-based parent training.[3] ### Exclusionary vs. non-exclusionary time-out | Type | Procedure | Example | | --- | --- | --- | | **Non-exclusionary** (person stays in the setting) | Planned ignoring | Adult withdraws attention for a set period | | Withdrawal of a specific reinforcer | The toy or tablet is removed for two minutes | | | Contingent observation ("sit and watch") | Child sits at the edge of the activity, watches, then rejoins | | | Time-out ribbon | A ribbon signals that reinforcement is available; it is removed after misbehavior and returned after a short interval[4] | | | **Exclusionary** (person leaves the reinforcing setting) | Partition or corner | Child sits facing away from the group in the same room | | Hallway or another room | Child sits in the hallway or a quiet room without toys | | | Seclusion | An isolated time-out room — a restrictive procedure requiring oversight, reserved for dangerous behavior | | The least restrictive version that works is the right one. For most children in most homes, that is a brief non-exclusionary or corner time-out, delivered calmly. ### Why time-out fails when time-in isn't reinforcing The most common reason time-out does not work is that the child was not losing anything. A classic pair of experiments made the point. In one, time-out reduced a teenager's tantrums when the time-in environment was enriched with activities and attention, but not when it was impoverished. In the other, time-out actually *increased* a young child's problem behavior, because the time-out period gave her uninterrupted opportunity for self-stimulatory behavior she found reinforcing.[5] Time-out is only punishing relative to what it interrupts; for a child in a dull or demanding situation, being removed is an escape or an opportunity — hence the child happy to be sent to a bedroom full of toys, and the student who acts out to be sent from a hard lesson to a comfortable office. Before using time-out, ask what time-in offers. If the answer is "not much," fix that first. ### What the evidence says Decades of research, mostly with young children and their parents, support time-out for reducing aggression, noncompliance, and tantrums when it is brief, consistent, and embedded in a home rich in positive attention.[3][6] It is a component of every major evidence-based parent-management program, and the American Academy of Pediatrics recommends it as an alternative to physical punishment.[7] A 2019 review examined the claim that time-out harms attachment and concluded that, implemented as designed, the evidence does not support that concern — and that discouraging time-out risks steering parents away from one of the few discipline strategies with a strong evidence base.[8] One caution: the time-out parameters recommended in popular books and websites often diverge from what the research supports, so the details below matter.[9] ## Response cost **Response cost** is the removal of a specified amount of a reinforcer, contingent on a behavior — a fine, in other words. The term comes from laboratory work in the early 1960s in which people working for points lost some of them for responding under certain conditions, which reduced responding.[10] In applied settings it usually means losing tokens, points, money, minutes of a privilege, or the privilege itself.[11] - A student in a token economy loses two tokens for leaving her seat without permission. - A driver loses points from his license — and eventually the license — for moving violations. - A teenager loses thirty minutes of screen time for each unfinished chore. - A gym member forfeits a deposit for each missed class (a commitment contract). Response cost fits naturally inside a token economy, where reinforcers are earned for desired behavior and lost for undesired behavior. The practical rules: the person must have something to lose (a child at zero tokens has no reason to behave); the cost should matter without wiping out the day's earnings; and earnings should outpace losses. A system in which people lose more than they earn becomes an aversive environment everyone tries to escape. ## Examples of negative punishment | Setting | Behavior | Reinforcer removed | Result | | --- | --- | --- | --- | | Home | Child throws a toy at a sibling | Two-minute time-out from play | Toy-throwing decreases | | Home | Teenager misses curfew | Car keys for the weekend | Curfew violations decrease | | Home | Child interrupts repeatedly at dinner | Parent stops talking to them for a minute | Interrupting decreases | | School | Student shouts out answers | One token from the day's total | Shouting out decreases | | School | Student pushes in line | Place in line (sent to the back) | Pushing decreases | | School | Student misuses lab equipment | Lab privileges for a week | Misuse decreases | | Sports | Hockey player trips an opponent | Two minutes of play (penalty box) | Tripping decreases | | Sports | Soccer player keeps arguing after a warning | Sent off: loses the rest of the match | Arguing decreases | | Driving | Driver runs a red light | Points on the license | Red-light running decreases | | Driving | Driver parks illegally | Money (fine) | Illegal parking decreases | | Workplace | Employee is repeatedly late | Flexible-hours privilege | Lateness decreases | | Workplace | Salesperson skips compliance training | Eligibility for the quarter's bonus | Skipped trainings decrease | | Games and apps | Player attacks a teammate online | Ranking points; temporary ban | Team-killing decreases | | Games and apps | User skips a day in a streak-based app | The streak (reset to zero) | Skipped days decrease, for those who value the streak | | Self-management | You check social media during a focus block | A dollar into a jar you don't get back | Checking decreases | | Dog training | Puppy nips during play | Play; the person leaves for 30 seconds | Nipping decreases | The dog-training row is negative punishment done well: the reinforcer removed (play) is exactly the one maintaining the nipping, the removal is immediate, and play resumes as soon as the dog is calm, so the contrast is unmistakable. It teaches without adding anything aversive, which is why it is the recommended way to handle puppy biting. [Reward-based dog training ›](https://operantconditioning.com/dog-training/) · [Examples for every quadrant ›](https://operantconditioning.com/examples/) ## Negative punishment vs. extinction This is the distinction students most often get wrong, because both procedures involve a reinforcer failing to arrive and both reduce behavior. The difference is *which* reinforcer. | Aspect | Extinction | Negative punishment | | --- | --- | --- | | **What happens** | The reinforcer that has been *maintaining* the behavior is no longer delivered for it | A reinforcer the person already has is removed *because* the behavior occurred | | **Is anything taken away?** | No — a reinforcer is withheld | Yes — the person loses something | | **Which reinforcer** | The one maintaining the behavior | Usually a *different* one | | **Typical course** | Gradual decline, often after an initial burst | Usually a faster decrease | | **Child whines for candy** | Whining no longer ever produces candy | Each whine costs a token from the sticker chart | | **Dog jumps up for attention** | Jumping is never followed by attention | Jumping ends the walk for one minute | The test question: *was the removed reinforcer the one paying for the behavior?* If whining simply stops working, that is extinction — you declined to give something. If the whine costs a token, that is response cost — the token was never the reason for whining, and you removed it. The two behave differently: [extinction](https://operantconditioning.com/extinction/) produces an extinction burst and requires you to control the maintaining reinforcer, while response cost can be applied even when you cannot control what maintains the behavior (peer laughter, for instance). ## Negative punishment vs. negative reinforcement Both begin with "negative," so both involve removing something. That is where the similarity ends. | Aspect | Negative reinforcement | Negative punishment | | --- | --- | --- | | **What is removed** | An aversive stimulus | A reinforcing stimulus | | **Effect on behavior** | Increases | Decreases | | **Example** | Buckling up silences the seat-belt chime | Speeding costs points on a license | | **What the learner feels** | Relief | Loss | A single event can be both, for different people. When a parent carries a screaming toddler out of a restaurant, the toddler may be negatively punished (loses the crayons and the attention) while the parent is negatively reinforced (the embarrassment ends). And if a "punishment" removes something the person wanted to escape anyway, it has flipped into negative reinforcement. ## Effectiveness and side effects Negative punishment is the preferred form of punishment when punishment is used at all. It adds no aversive stimulus, so it produces less of the fear, reflexive aggression, and conditioned aversiveness that follow [positive punishment](https://operantconditioning.com/positive-punishment/); it is easier to apply consistently without escalation; and it fits naturally with reinforcement-based systems as the debit side of a ledger.[12][13] But it is still punishment, and it shares punishment's limits: - **It does not teach a replacement.** A time-out says what not to do. The child still needs a reinforced way to get what hitting was getting. - **Emotional responding and escape.** Children resist going to time-out, argue about fines, and sometimes escalate; the struggle to enforce the procedure can become more aversive than the procedure. - **Response cost can provoke aggression and collapse.** Losing a great deal at once elicits anger, and once a person is at zero there is nothing left to lose. - **Overuse turns a rich environment into a poor one.** A classroom where tokens are constantly taken away is one children want to leave, which brings escape-maintained behavior with it. - **Delay kills it.** "You're grounded next weekend" is a weak consequence for a Tuesday-night behavior. ## How to do time-out correctly Most failed time-outs fail on the details. The research-supported procedure looks like this:[3][9] 1. **Make time-in rich first.** Time-out only removes what time-in provides. Catch the child being good many times a day. 2. **Decide in advance which behaviors earn it.** A short list — hitting, throwing, deliberate destruction. Minor behavior is better handled by planned ignoring and reinforcing the alternative. 3. **Deliver it immediately and calmly.** One brief statement ("No hitting. Time-out."), no lecture, no negotiation. Anger and explanation are attention, and attention is a reinforcer. 4. **Keep it brief.** A few minutes is enough; about one minute per year of age is the common rule of thumb for young children. Longer is not more effective and is harder to enforce. 5. **End it on calm, not the clock alone.** Release when the interval has passed *and* the child has been quiet for a few seconds, so release does not reinforce protesting. 6. **Return to time-in without a debrief.** Back to the activity, and reinforce the first good behavior you see. The point is the contrast: good behavior gets warmth and access; hitting gets a brief, boring pause. 7. **Be consistent.** Every listed behavior, every time, from every caregiver. Intermittent time-out teaches that the behavior sometimes pays. 8. **Track it.** If the behavior is not decreasing within a couple of weeks, the procedure is not functioning as punishment. Check whether time-in is rewarding and whether time-out offers an escape. The same principles transfer to response cost: small, immediate, consistent, and embedded in a system where the person earns far more than they lose. ## Common mistakes - **Time-out from nothing.** Removing a child from a situation they wanted to leave is negative reinforcement, and it will increase the behavior. - **Time-out with a lecture.** Explaining, scolding, and checking in during the interval deliver attention and turn the procedure into a conversation. - **Making it long.** A thirty-minute time-out is less effective than a two-minute one and more likely to end in a fight. - **Fining people into the red.** Once tokens hit zero the system has no leverage. - **Confusing it with extinction.** Withholding the maintaining reinforcer is extinction; removing a different reinforcer is negative punishment. They need different planning. - **Using it instead of teaching.** Punishment of any kind supplements reinforcement of the behavior you want; it never substitutes for it. [Positive parenting strategies ›](https://operantconditioning.com/applications/#parenting) ## Key takeaways - Negative punishment needs a reinforcer that is present or available before the behavior, so that it can be taken away after. You cannot remove what the person does not value, and whether a removal is punishing is decided by its effect on the behavior. - Time-out means time-out from positive reinforcement. The chair and the corner are incidental; the procedure works only when time-in is rewarding, and for a child in a dull or demanding situation, being removed is an escape or an opportunity. - Response cost is a fine: a set amount of a reinforcer is removed each time the behavior occurs. It works only when the person has something to lose and earns more than they lose. - Extinction withholds the reinforcer that has been maintaining the behavior; negative punishment removes a reinforcer the person already has, usually a different one. Ask whether the removed reinforcer was the one paying for the behavior. - Negative punishment is the preferred form of punishment because it adds no aversive stimulus, but it still teaches no replacement, weakens with delay, and turns a rich environment poor when overused. Punishment supplements reinforcement of the behavior you want; it never substitutes for it. ### Check yourself **A bored student acts out during a hard lesson and is sent to sit in the office. Acting out increases. Was this negative punishment that failed?** No. It was negative reinforcement. The student lost nothing they wanted and escaped something they wanted to leave, so the removal strengthened the behavior. A consequence is classified by its effect, not by its label. **A child whines for candy. Parent A never gives candy for whining again. Parent B takes a token off the sticker chart for each whine. Which one is negative punishment?** Parent B. The token was never the reason for whining, and it was removed because the behavior occurred: response cost. Parent A is using extinction, in which the maintaining reinforcer is withheld and nothing is taken away, and should expect an extinction burst. **A parent puts a child in a two-minute time-out for hitting, explains during the interval why hitting is wrong, and checks in every thirty seconds. Hitting does not decrease. What went wrong?** The time-out was delivering attention. Explaining, scolding, and checking in are attention, and attention is a reinforcer, so the procedure became a conversation rather than a removal of reinforcement. One brief statement, no lecture, then back to time-in and reinforce the first good behavior. **A classroom token system fines students for every minor infraction, and several students are at zero tokens by mid-morning. Why does the system stop working?** Once a person is at zero there is nothing left to lose, so response cost has no leverage. A room where tokens are constantly taken away also becomes an aversive environment children want to leave, which brings escape-maintained behavior with it. Earnings must outpace losses. **Explain it to a friend.** Explain why a time-out only works when the situation it interrupts is rewarding, using an example from a sport or a game rather than from parenting. ## Frequently asked questions **What is negative punishment in simple terms?** Taking something good away after a behavior so the behavior happens less often. A child hits, loses two minutes of playtime, and hits less. The removed playtime is the negative punisher. "Negative" means removed, not bad. **What is an example of negative punishment?** A teenager comes home after curfew and loses the car for the weekend; curfew violations decrease. Others: time-out for hitting, losing license points for speeding, a penalty-box stint in hockey, losing tokens in a classroom system, a parking fine. **Is time-out negative punishment?** Yes, when done properly. Time-out from positive reinforcement removes access to reinforcers (play, attention, activities) for a brief period contingent on a behavior, and the behavior decreases. If the child was not enjoying the situation they were removed from, time-out is not functioning as punishment and may act as an escape instead. **What is the difference between negative punishment and negative reinforcement?** Both remove something. Negative punishment removes a reinforcer and decreases behavior (losing screen time for hitting). Negative reinforcement removes an aversive stimulus and increases behavior (buckling up to stop the seat-belt chime). Ask whether the behavior went up or down. **What is the difference between negative punishment and extinction?** Extinction withholds the specific reinforcer that has been maintaining the behavior — whining no longer produces candy. Negative punishment removes a reinforcer the person already has, usually a different one — each whine costs a token. Extinction produces an extinction burst; response cost usually works faster but does not address why the behavior occurs. **What is response cost?** A form of negative punishment in which a set amount of a reinforcer — tokens, points, money, minutes of a privilege — is removed each time a behavior occurs. Fines, penalty points, and losing tokens in a classroom are all response cost. It works best inside a system where the person earns more than they lose. **How long should a time-out be?** Brief. A few minutes is enough for young children; one minute per year of age is a common guideline. Longer time-outs are not more effective and are harder to enforce. End it when the interval has passed and the child has been calm for a few seconds. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Wolf, M. M., Risley, T., & Mees, H. (1964). Application of operant conditioning procedures to the behaviour problems of an autistic child. *Behaviour Research and Therapy, 1*(2–4), 305–312. 3. Everett, G. E., Hupp, S. D. A., & Olmi, D. J. (2010). Time-out with parents: A descriptive analysis of 30 years of research. *Education and Treatment of Children, 33*(2), 235–259. 4. Foxx, R. M., & Shapiro, S. T. (1978). The timeout ribbon: A nonexclusionary timeout procedure. *Journal of Applied Behavior Analysis, 11*(1), 125–136. 5. Solnick, J. V., Rincover, A., & Peterson, C. R. (1977). Some determinants of the reinforcing and punishing effects of timeout. *Journal of Applied Behavior Analysis, 10*(3), 415–424. 6. Kazdin, A. E. (2005). *Parent Management Training: Treatment for Oppositional, Aggressive, and Antisocial Behavior in Children and Adolescents*. Oxford University Press. 7. Sege, R. D., Siegel, B. S., Council on Child Abuse and Neglect, & Committee on Psychosocial Aspects of Child and Family Health. (2018). Effective discipline to raise healthy children. *Pediatrics, 142*(6), e20183112. 8. Dadds, M. R., & Tully, L. A. (2019). What is it to discipline a child: What should it be? A reanalysis of time-out from the perspective of child mental health, attachment, and trauma. *American Psychologist, 74*(7), 794–808. 9. Corralejo, S. M., Jensen, S. A., Greathouse, A. D., & Ward, L. E. (2018). Parameters of time-out: Research update and comparison to parenting programs, books, and online recommendations. *Behavior Therapy, 49*(1), 99–112. 10. Weiner, H. (1962). Some effects of response cost upon human operant behavior. *Journal of the Experimental Analysis of Behavior, 5*(2), 201–208. 11. Kazdin, A. E. (1972). Response cost: The removal of conditioned reinforcers for therapeutic change. *Behavior Therapy, 3*(4), 533–546. 12. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 13. Lerman, D. C., & Vorndran, C. M. (2002). On the status of knowledge for using punishment: Implications for treating behavior disorders. *Journal of Applied Behavior Analysis, 35*(4), 431–464. ## Related - [Positive punishment](https://operantconditioning.com/positive-punishment/): Adding an aversive stimulus — and why the evidence favors alternatives. - [Extinction](https://operantconditioning.com/extinction/): Withholding the maintaining reinforcer: bursts, recovery, and how to do it. - [Punishment: the hub](https://operantconditioning.com/punishment/): Both types, what the evidence says, and the alternatives. --- # Schedules of Reinforcement: Fixed, Variable, Ratio, and Interval — With an Interactive Simulator > Schedules of reinforcement decide how often a behavior pays: continuous vs. intermittent, fixed and variable ratio and interval, with a live simulator. - Source: https://operantconditioning.com/schedules-of-reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Timing and rate* A behavior's strength depends not just on *whether* it is reinforced but on *how often* and *on what rule*. Those rules are schedules, and they explain everything from slot machines to why you keep checking your phone. > **Definition** > > A **schedule of reinforcement** is the rule that determines which occurrences of a behavior are followed by a reinforcer. A **continuous** schedule reinforces every response; an **intermittent** (partial) schedule reinforces only some, based on the number of responses (*ratio* schedules) or the passage of time (*interval* schedules), on a fixed or variable basis. > > Each schedule produces a characteristic pattern of responding and a characteristic resistance to [extinction](https://operantconditioning.com/extinction/). The systematic study of schedules was the work of Charles Ferster and B. F. Skinner, whose 1957 book *Schedules of Reinforcement* reported hundreds of cumulative records from pigeons and rats.[1] **In brief** - A schedule of reinforcement is the rule deciding which responses earn a reinforcer: every response (continuous) or only some (intermittent). - Intermittent schedules are based on a count of responses (ratio) or on time elapsed (interval), with a fixed or variable requirement. - Continuous reinforcement teaches a behavior fastest; variable ratio maintains the highest, steadiest rate and is the hardest to [extinguish](https://operantconditioning.com/extinction/). ## Try it: the schedule simulator Below is a cumulative record — the same kind of graph Skinner's recorder drew on a rolling strip of paper. Time runs left to right; every response steps the line up by one; a small green tick marks each reinforcer. Pick a schedule and let the simulated subject run to watch the classic patterns appear — or switch the mode to "I'll respond" and press **Respond** yourself. The simulated subject is a stylized model of the response patterns Ferster and Skinner documented — post-reinforcement pauses on fixed ratio, steady high rates on variable ratio, the fixed-interval "scallop," and a moderate steady rate on variable interval — not a replay of real data. Try FR 20 to see a longer post-reinforcement pause, or FI 15 to see the scallop. ## Continuous versus intermittent reinforcement On a **continuous reinforcement** schedule (CRF), every response is reinforced. It is the fastest way to establish a new behavior: the relationship between behavior and consequence is unmistakable. It is also the fastest schedule to extinguish, because the first unreinforced response is immediately informative — something has changed. On an **intermittent** schedule, only some responses are reinforced. Learning is slower, but the behavior becomes far more durable. This is the **partial reinforcement extinction effect**: behavior maintained on an intermittent schedule persists much longer once reinforcement stops than behavior maintained on CRF.[2] A vending machine that fails once loses you immediately; a slot machine that fails a hundred times in a row is doing exactly what it always does. > **The practical rule** > > Reinforce continuously while a behavior is being learned. Once it is reliable, thin gradually to an intermittent schedule — ideally a variable-ratio schedule — to make it resistant to extinction. Thin too fast and you get **ratio strain**: pausing, erratic responding, and eventually extinction. ## The four basic intermittent schedules Two questions define the basic schedules. Is reinforcement based on the *number of responses* (ratio) or on *time elapsed* (interval)? And is the requirement *fixed* or *variable*? ![Idealized cumulative records for the four basic schedules](https://operantconditioning.com/assets/diagrams/idealized-cumulative-records-for-the-four-basic-schedules.svg) *Idealized cumulative records. Each upward step is a response; green ticks mark reinforcers. Patterns after Ferster & Skinner (1957).* ### Fixed ratio (FR) Reinforcement follows every *n*th response. FR 1 is continuous reinforcement; FR 10 pays after every tenth response. Fixed-ratio schedules produce a high rate of responding with a distinctive **post-reinforcement pause** after each reinforcer — the larger the ratio, the longer the pause — followed by a rapid, steady run to the next one. Piece-rate pay, "buy ten coffees, get one free," and a set of ten push-ups before a rest are fixed ratios. ### Variable ratio (VR) Reinforcement follows an unpredictable number of responses that averages *n*. VR 10 might pay after 3, then 15, then 8, then 14 responses. Because the very next response might be the one that pays, there is little pausing; the result is the highest, steadiest response rate of any basic schedule and the greatest resistance to extinction. Slot machines are the textbook example. So are fishing casts, sales calls, refreshing a social feed, and opening loot boxes.[1] ### Fixed interval (FI) The first response after a fixed time has elapsed is reinforced; responses during the interval earn nothing. FI 60 s pays the first response after a minute. Animals on FI schedules produce the famous **scallop**: almost no responding right after a reinforcer, then an accelerating rate as the end of the interval approaches. Checking the oven as the timer runs down, studying more as the exam nears, and glancing at the mailbox around delivery time follow this pattern. (A salary is *not* a good example — it is paid on a time basis regardless of the number of responses, which makes it closer to a fixed-time schedule than a fixed-interval one.) ### Variable interval (VI) The first response after an unpredictable interval that averages *t* is reinforced. VI schedules produce a moderate, remarkably steady rate — the organism keeps checking, because the reinforcer could become available at any moment, but doesn't race, because responding faster doesn't make it come sooner. Checking email, waiting for a reply to a text, and a manager's random walk-throughs of the shop floor all maintain behavior on variable intervals. Because VI produces such stable baselines, it is the workhorse schedule in laboratory research on choice and drugs. | Schedule | Reinforcer delivered | Response pattern | Resistance to extinction | Everyday examples | | --- | --- | --- | --- | --- | | **Continuous (CRF)** | After every response | Steady while satiation is far off | Low | Light switch, vending machine, a new skill being taught | | **Fixed ratio (FR)** | After a fixed number of responses | High rate; post-reinforcement pause | Moderate–high | Piece-rate pay, loyalty punch cards, sets of reps | | **Variable ratio (VR)** | After a variable number of responses | Very high, steady rate; little pausing | Highest | Slot machines, sales calls, social-media feeds, fishing | | **Fixed interval (FI)** | First response after a fixed time | Scallop: pause then acceleration | Moderate | Watching the oven timer, cramming before a scheduled exam | | **Variable interval (VI)** | First response after a variable time | Moderate, steady rate | High | Checking email or texts, surprise inspections | ## Which schedule is it? A two-question rule Students confuse the four schedules almost as often as they confuse the quadrants, and the same trick works: two questions, in order. **(1) Does the reinforcer depend on how many responses were made, or on how much time has passed?** Count means *ratio*; time means *interval*. **(2) Is the requirement always the same, or does it vary around an average?** Same means *fixed*; varies means *variable*. (If every single response is reinforced, it is continuous reinforcement, and if none are, it is extinction.) A detail that catches people: on an interval schedule, responding faster does not bring the reinforcer sooner — only the *first* response after the interval counts. | Basis | Requirement fixed | Requirement varies | | --- | --- | --- | | Depends on a count of responses | Fixed ratio (FR) | Variable ratio (VR) | | Depends on time elapsed | Fixed interval (FI) | Variable interval (VI) | ### Practice: name the schedule Decide, then open the answer. **A factory pays a worker $2 for every 50 shirts sewn.** **Fixed ratio (FR 50).** The reinforcer depends on a count, and the count is always the same. Expect a brief pause after each payment, then a run. **A slot machine pays out after an unpredictable number of pulls, averaging about one in twenty.** **Variable ratio (VR 20).** Count-based, requirement varies around an average. High, steady responding; the hardest schedule to extinguish. **Your paycheck arrives every other Friday, provided you have worked that fortnight.** **Fixed interval (FI two weeks) — approximately.** Time-based and fixed. Real paychecks are a poor example of pure FI because the response requirement is not a single response after the interval, but the classic textbook classification is FI, and the "scallop" of end-of-period effort is real. **You check your phone for messages that arrive at unpredictable times; the first check after a message arrives finds it.** **Variable interval.** A message becomes available after a varying time, and only a check after that point is reinforced. Checking faster does not make messages arrive sooner — which is why VI produces a moderate, steady rate rather than a frantic one. **A dog gets a treat for every single sit during the first week of training.** **Continuous reinforcement (CRF, or FR 1).** Every response pays. Fast learning, fast extinction; the right schedule for building a behavior, the wrong one for keeping it. **A quality inspector checks a machine that jams at random intervals; she can only fix a jam once it has happened.** **Variable interval.** The opportunity to be reinforced (finding and fixing a jam) becomes available after a varying time; checks before a jam cannot be reinforced. Time-based, variable. ## Why variable ratio is so powerful — and so dangerous Three features make VR the schedule of choice for anyone who wants behavior to persist: no signal ever tells the organism that the next response is pointless; the average payout can be made arbitrarily lean once the behavior is established; and the occasional early payout after a long dry run is itself a powerful reinforcer of persistence. Casinos, mobile games, and social platforms have converged on variable-ratio designs for exactly these reasons.[3] The same mechanism that makes a behavior hard to quit makes it hard to *build* deliberately — you cannot start on VR. The order is always CRF, then a gradual stretch. ## Beyond the basic four | Schedule | Rule | What it's used for | | --- | --- | --- | | **Fixed / variable time (FT, VT)** | Reinforcer delivered after time passes, regardless of behavior | Noncontingent reinforcement; the procedure behind Skinner's "superstition" study[4] and a treatment for attention-maintained problem behavior | | **Differential reinforcement of low rate (DRL)** | Reinforced only if at least *t* seconds have passed since the last response | Slowing behavior that is fine at a low rate (talking in class, eating speed) | | **Differential reinforcement of high rate (DRH)** | Reinforced only if *n* responses occur within *t* seconds | Speeding up fluent skills (math facts, typing) | | **Differential reinforcement of other behavior (DRO)** | Reinforced if the target behavior does *not* occur for an interval | Reducing problem behavior without punishment; see [alternatives to punishment](https://operantconditioning.com/positive-punishment/) | | **Progressive ratio (PR)** | Ratio increases after each reinforcer until the organism stops (the "breakpoint") | Measuring how hard an organism will work for a reinforcer — a standard index of reinforcer value and drug abuse liability[5] | | **Concurrent schedules** | Two or more schedules available at once on different responses | Studying choice; produced Herrnstein's [matching law](https://operantconditioning.com/matching-law/): relative response rate matches relative reinforcement rate[6] | | **Chained schedules** | Completing one schedule produces the stimulus for the next; only the last delivers the primary reinforcer | Analyzing sequences of behavior; conditioned reinforcement | | **Multiple and mixed schedules** | Schedules alternate, with (multiple) or without (mixed) a signal | Studying stimulus control and behavioral contrast | ## Do humans follow these schedules? Mostly, with an important caveat. Human performance on simple schedules is often shaped by verbal rules and self-instruction as much as by the schedule itself. Adults told (or who figure out) that reinforcement is time-based may respond at a very low rate or a very high steady rate on FI rather than producing the clean scallop seen in pigeons, and instructions can override contingencies for a surprisingly long time.[7] Infants and young children, who have less verbal behavior to lean on, look more like the animal data. The lesson for anyone applying schedules to people: the *description* of the contingency is itself a variable, and rules that don't match the real schedule eventually lose to it. ## Using schedules deliberately 1. **Start continuous.** While a behavior is new — a child's first attempts at a chore, a dog's first sits, your first week of a habit — reinforce every occurrence. 2. **Stretch the ratio slowly.** Move to reinforcing two out of three, then every other, then unpredictably around one in three. Watch for ratio strain and back off if the behavior falters. 3. **Prefer variable over fixed.** Variable schedules produce steadier behavior with fewer pauses and greater persistence. 4. **Prefer ratio over interval when you want rate.** If you want more of a behavior, tie reinforcement to responses. If you want steady monitoring, an interval schedule is fine. 5. **Let natural reinforcers take over.** The goal of any contrived schedule is to hand the behavior off to the consequences the world already provides — the clean kitchen, the finished chapter, the dog that comes when called. [How to do this for your own habits ›](https://operantconditioning.com/habits/) ## Schedules and problem behavior The same principles explain why unwanted behavior is so persistent. A parent who gives in to a tantrum "only occasionally" has put tantrums on a lean variable-ratio schedule — the most extinction-resistant schedule there is. A manager who answers after-hours emails "only when urgent" has done the same for after-hours emailing. When you are trying to reduce a behavior, the first question is always: *what schedule is currently maintaining it, and who is delivering the reinforcer?* [See extinction and the extinction burst ›](https://operantconditioning.com/extinction/) ## Key takeaways - Continuous reinforcement, in which every response is reinforced, establishes a new behavior fastest and extinguishes fastest. Intermittent reinforcement teaches more slowly but makes behavior far more durable: the partial reinforcement extinction effect. - Two questions identify any basic schedule. Does the reinforcer depend on a count of responses (ratio) or on time elapsed (interval), and is the requirement always the same (fixed) or does it vary around an average (variable)? - Each schedule has a signature pattern. Fixed ratio produces a post-reinforcement pause followed by a run; variable ratio a high, steady rate; fixed interval the scallop; variable interval a moderate, steady rate. - Variable ratio is the most resistant to extinction because no signal ever tells the organism that the next response is pointless. Slot machines, social feeds, and tantrums that are given in to "only occasionally" all run on it. - To use schedules deliberately, reinforce continuously while a behavior is new, then thin gradually to an intermittent, ideally variable-ratio, schedule. Thin too fast and you get ratio strain: pausing, erratic responding, and eventually extinction. **Explain it to a friend.** Explain why a slot machine keeps people playing long after a broken vending machine would have lost them, without using the words schedule, ratio, or interval. ## Frequently asked questions **What are the four schedules of reinforcement?** Fixed ratio (reinforcement after a set number of responses), variable ratio (after an unpredictable number of responses), fixed interval (first response after a set time), and variable interval (first response after an unpredictable time). Continuous reinforcement — every response reinforced — is the fifth, simplest case. **Which schedule of reinforcement is most effective?** It depends on the goal. Continuous reinforcement is best for teaching a new behavior. Variable ratio produces the highest, steadiest rate and the greatest resistance to extinction, so it is best for maintaining an established behavior. Variable interval produces the most stable moderate rate. **Which schedule is most resistant to extinction?** Variable ratio. Because reinforcement has always been unpredictable, a long run without it is not a signal that anything has changed, so responding persists. This is why gambling and social-media checking are so hard to stop. **Is a salary a fixed interval schedule?** Not really. Fixed interval requires a response after the interval; a salary is paid on a time basis regardless of the amount of behavior, which makes it more like a fixed-time (noncontingent) schedule. Piece-rate pay and commissions are ratio schedules. This is one reason salaries do little to reinforce specific daily behaviors. **What is the fixed-interval scallop?** The characteristic pattern on a cumulative record under a fixed-interval schedule: little or no responding right after a reinforcer, then an accelerating rate as the end of the interval approaches, producing a curve that looks like a series of scallops. **What is a post-reinforcement pause?** The pause in responding that follows each reinforcer on a fixed-ratio (and fixed-interval) schedule. On fixed ratio, the pause grows with the size of the ratio — the "break" in break-and-run responding. **What is the partial reinforcement extinction effect?** The finding that behavior reinforced only some of the time persists longer during extinction than behavior that was reinforced every time. Intermittent reinforcement makes the transition to no reinforcement harder to detect, so the behavior keeps going. **What is ratio strain?** The breakdown in responding — long pauses, erratic bursts, eventually extinction — that occurs when a ratio schedule is thinned too quickly or set too high for the value of the reinforcer. The cure is to drop back to a richer schedule and stretch more gradually. ## References 1. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 2. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158. See also Mowrer, O. H., & Jones, H. (1945). Habit strength as a function of the pattern of reinforcement. *Journal of Experimental Psychology, 35*(4), 293–311. 3. Schüll, N. D. (2012). *Addiction by Design: Machine Gambling in Las Vegas*. Princeton University Press. 4. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 5. Hodos, W. (1961). Progressive ratio as a measure of reward strength. *Science, 134*(3483), 943–944. 6. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 7. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): What a reinforcer is and what makes it work. - [Extinction](https://operantconditioning.com/extinction/): What happens when reinforcement stops — bursts, recovery, resurgence. - [B. F. Skinner](https://operantconditioning.com/bf-skinner/): The man who built the box and the cumulative recorder. --- # Extinction in Operant Conditioning: Definition, the Extinction Burst, Spontaneous Recovery, and Examples > Extinction in psychology is the fading of a behavior when its reinforcer is withheld. The extinction burst, spontaneous recovery, resurgence, and how to use it. - Source: https://operantconditioning.com/extinction/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Core process* Stop paying for a behavior and it fades — but not quietly, not immediately, and not for good. What extinction in psychology really means, why behavior gets worse before it gets better, and how to use it well. > **Definition** > > **Extinction** in operant conditioning is the process in which a behavior that was previously reinforced is no longer followed by its reinforcer, and as a result the behavior decreases and eventually stops. The procedure is withholding the reinforcer; the effect is the decline in behavior. > > Extinction is not one of the four quadrants — nothing is added or removed; the consequence that used to arrive simply stops. Nor is it forgetting: extinguished behavior can return through spontaneous recovery, renewal, and resurgence, so extinction is best understood as new learning laid over the old rather than erasure.[1][2] **In brief** - Extinction is what happens when a previously reinforced behavior stops producing its reinforcer: the behavior declines and eventually stops. - Behavior can get worse before it fades — the extinction burst — and giving in at that point reinforces a more intense version, intermittently. - Extinction is not ignoring: it withholds whichever reinforcer maintains the behavior, and behavior reinforced [intermittently](https://operantconditioning.com/schedules-of-reinforcement/) resists extinction far longer. ## How extinction works Every reinforced behavior exists because of a contingency: do this, get that. Extinction breaks the contingency. The rat that has learned to press a lever for food presses, and nothing happens. It presses again. Over the next hour its rate climbs, wobbles, and slides toward zero, tracing the **extinction curve** Skinner recorded in 1938.[1] - A rat presses a lever → no pellet, ever again → pressing declines over the session. - A child whines for a cookie → whining is never again followed by a cookie → whining declines over days. - You text a friend who has stopped replying → no replies → you text less and eventually stop. ![The extinction curve](https://operantconditioning.com/assets/diagrams/the-extinction-curve.svg) *The shape of extinction. When reinforcement stops, responding briefly rises (the burst), then declines; after a rest it returns at a lower level (spontaneous recovery) and fades again.* Two things distinguish extinction from other ways of reducing behavior. It works only on the reinforcer that has been *maintaining* the behavior; withholding something else is not extinction. And it is slow and bumpy, because a history of the behavior paying off does not vanish when the payoff stops — it is tested, protested, and eventually overwritten. How long that takes depends on the reinforcement history, as the partial reinforcement extinction effect explains. ### Watch it happen The simulator below starts in extinction: a subject with a history of reinforcement, and no more reinforcers. Run it and watch the burst, the decline, and the flattening record. Switch the schedule to compare how quickly behavior built on continuous versus intermittent reinforcement gives up. ## Operant vs. respondent extinction "Extinction" is used in both branches of conditioning. In **respondent (classical) extinction**, a conditioned stimulus is presented repeatedly without the unconditioned stimulus — the bell rings, no food follows — and the conditioned response fades. In **operant extinction**, a behavior occurs repeatedly without its reinforcer and the behavior fades. Both show extinction curves, spontaneous recovery, and renewal, and both are now understood as new inhibitory learning rather than unlearning.[2] The difference is what is disconnected: a stimulus from a stimulus, or a behavior from its consequence. [Operant vs. classical conditioning in full ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## What is an extinction burst? An **extinction burst** is a temporary increase in the frequency, intensity, or variability of a behavior immediately after its reinforcer is withheld, before the behavior declines. Press the elevator button; nothing happens; you press again, harder, three times, then hold it, then try the other button. More, harder, different — that is the burst. The organism's history says the behavior *should* work, and vigorous, varied responding is the evolved answer to a missing payoff. The burst is also the engine of [shaping](https://operantconditioning.com/shaping/): when a trainer stops reinforcing the current approximation, the burst of variability produces the slightly better response to reinforce next. But in everyday life the burst is where almost everyone gives up. The child whose whining has stopped producing candy whines louder, then cries, then throws himself on the floor — and a parent who gives in at that point has reinforced the floor-throwing. How common is it? A review of 113 published data sets in which extinction was applied to problem behavior found a burst in about a quarter of cases, more often when extinction was used alone than when combined with procedures such as differential reinforcement.[3] A later analysis of 41 cases of self-injury treated with extinction found bursts in 39% and extinction-induced aggression in 22%, again less often when extinction was combined with other treatments.[4] The burst is worth planning for but not inevitable, and the way to make it less likely is also the way to make extinction work better: reinforce something else at the same time. > **The burst is the moment that decides everything** > > If you give in during the extinction burst, you have not returned to where you started. You have reinforced a more intense version of the behavior, on an intermittent schedule — which makes it *more* resistant to extinction than before. Do not start extinction unless you can outlast the burst. If you cannot, use a different procedure. ## Extinction-induced aggression and emotional responding Withholding an expected reinforcer also elicits emotional behavior: agitation, frustration, and aggression, often aimed at whatever is nearby. In a landmark experiment, pigeons pecking a key for grain were placed on extinction, and when the grain stopped they attacked a restrained pigeon at the other end of the chamber — a target that had nothing to do with the reinforcer.[5] The human version is familiar: the vending machine gets kicked, the customer-service line gets yelled at, a child on extinction hits a sibling. Emotional responding is the second main side effect of extinction, after the burst, and a central reason it is rarely used alone for serious behavior.[4][6] ## Spontaneous recovery, resurgence, and renewal Extinguished behavior comes back in at least three distinct ways, each with its own trigger. Knowing which is which tells you what to do about it. ### Spontaneous recovery **Spontaneous recovery** is the reappearance of an extinguished behavior after time away from the extinction setting. The rat that stopped pressing by the end of Monday's session starts again on Tuesday, at a lower rate, and stops sooner. First described by Pavlov for conditioned reflexes, it occurs just as reliably for operant behavior, and each recovery is smaller than the last if reinforcement stays withheld.[7] A behavior that seems gone on Friday may be back on Monday; that is not a sign extinction failed. ### Resurgence **Resurgence** is the reappearance of a *previously* reinforced behavior when a *more recently* reinforced behavior is placed on extinction. Teach a pigeon to peck key A for food, extinguish it, teach it to peck key B, then extinguish B — and the pigeon starts pecking A again.[8][9] This explains a common form of relapse: a child learns to ask politely instead of screaming, then in a busy classroom polite requests go unanswered, and the screaming returns.[10] The remedy is to keep the replacement behavior reliably reinforced, especially where the old behavior once worked. ### Renewal **Renewal** is the return of an extinguished behavior when the context changes. Mark Bouton's research showed that extinction learning is unusually tied to the place and cues in which it happened, whereas the original learning transfers broadly — which is why a fear extinguished in a therapist's office returns in the parking garage, and a behavior extinguished at the clinic returns at home.[2][11] The remedy is to conduct extinction in every context that matters. | Phenomenon | Trigger | Example | What to do | | --- | --- | --- | --- | | **Spontaneous recovery** | Time since the last extinction session | Bedtime crying that stopped last week returns after a weekend away | Keep withholding; each recovery is weaker | | **Resurgence** | A newer behavior stops being reinforced | Polite requests go unanswered and tantrums return | Keep the replacement reliably reinforced | | **Renewal** | A change of context | Behavior extinguished at school returns at home | Extinguish in every relevant setting | ## The partial reinforcement extinction effect The most important variable in how long extinction takes is the [schedule](https://operantconditioning.com/schedules-of-reinforcement/) on which the behavior was reinforced. Behavior reinforced **intermittently** is far more resistant to extinction than behavior reinforced every time. This is the **partial reinforcement extinction effect**, first demonstrated systematically by Lloyd Humphreys in 1939: responses reinforced only half the time persisted much longer without reinforcement than responses reinforced every time.[12] The effect is paradoxical — *less* reinforcement produces *more* persistence — and two theories survived decades of testing. Abram Amsel's **frustration theory** holds that organisms on intermittent schedules learn to keep responding *through* the frustration of nonreward, so nonreward during extinction is nothing new.[13] E. J. Capaldi's **sequential theory** holds that the memory of unreinforced trials becomes a cue for responding.[14] In plain language: an organism on a continuous schedule can tell instantly that the rules have changed; one on a variable schedule cannot, because unrewarded stretches were always part of the deal. The consequences are everywhere. Slot-machine players keep pulling through long droughts. A parent who gives in to whining "only sometimes" has built whining on a variable schedule and faces a long extinction. And a habit you want to *keep* should be moved from continuous to intermittent reinforcement once established, so it survives the days the reinforcer does not show up. ## Resistance to extinction and behavioral momentum Resistance to extinction is one case of what John Nevin called **behavioral momentum**. In physics, momentum depends on velocity and mass; in Nevin's analogy, a behavior's velocity is its rate and its mass is its resistance to disruption — by extinction, satiation, distraction, or free food. His central finding is that mass depends on the *rate of reinforcement obtained in the presence of a stimulus*, largely independent of response rate: behavior from a richly reinforced context is harder to disrupt than behavior from a lean one, even at the same rate.[15][16] This has a counterintuitive implication for treatment. Reinforcing an alternative behavior makes the alternative more likely, but it also enriches the context, which can make the *problem* behavior more resistant to extinction when its reinforcement is later withheld.[16] Practitioners deliver the alternative reinforcement in clearly distinct situations and expect the persistence. For the rest of us, momentum is why a well-established habit, good or bad, shrugs off disruption. ## Extinction is not ignoring "Just ignore it" is the folk version of extinction, and it is wrong often enough to be dangerous. Ignoring is extinction only when the reinforcer maintaining the behavior is *your attention*. Extinction means withholding whichever reinforcer is doing the maintaining; a functional analysis identifies it, and the procedure must match it.[17][18] | What maintains the behavior | What extinction requires | Example | What "ignoring" would do | | --- | --- | --- | --- | | **Attention** (social positive reinforcement) | *Planned ignoring*: no eye contact, comment, or reaction | A child interrupts for attention; the parent no longer responds | Works — this is the case ignoring was made for | | **Escape from demands** (social negative reinforcement) | *Escape extinction*: the demand stays in place; the behavior no longer ends it | A student shoves the worksheet away; the teacher calmly keeps it there and prompts the next step | Makes it worse — ignoring lets the student escape the task | | **Tangible items** | The item is never delivered after the behavior | Grabbing for the phone never produces the phone | Irrelevant unless the adult was handing over the item | | **Automatic (sensory) reinforcement** | *Sensory extinction*: the sensory consequence is blocked | A child spins objects for the visual effect; the effect is removed or altered | Does nothing — the behavior reinforces itself | The escape row is the one that catches parents and teachers. A child who tantrums to get out of putting on shoes is not seeking attention, and ignoring the tantrum while the shoes stay off is not extinction; it is the reinforcer, delivered. Escape extinction means the shoes still go on — planned carefully, with calm, safety, and often an easier demand to begin with. [More on escape-maintained behavior ›](https://operantconditioning.com/negative-reinforcement/) ## Examples of extinction | Situation | Reinforcer withheld | Burst | Outcome | | --- | --- | --- | --- | | Elevator button | Elevator arriving | Repeated, harder presses; holding the button | You stop pressing and take the stairs | | Vending machine that eats a coin | Snack dropping | Pressing every button; shaking; a kick | You walk away (and avoid that machine) | | Crying at bedtime | Parent returning to the room | Louder, longer crying on the first nights | Crying declines within about ten nights if parents hold firm | | Dog begging at the table | Scraps | Whining, pawing, intense staring | Begging fades — unless one family member keeps slipping food | | Checking a phone with notifications off | New messages and likes | More frequent checks for a few days | Checking declines | | Slot machine on a losing streak | Payout | Faster play, bigger bets | Very slow decline: the variable-ratio history resists extinction | | App that has shut down | Content, updates, responses | A few repeated opens and refreshes | You stop opening it within days | | Light switch during a power cut | Light | Flipping it several times | You stop within a day — and flip it again entering a new room (renewal) | The bedtime example is a real case, published in 1959. A 21-month-old boy screamed when his parents left the room, keeping one of them there for up to two hours a night. The parents put him to bed, left, and did not return. He cried for about 45 minutes the first night, less on subsequent nights, and by the tenth night not at all. A week later an aunt put him to bed, returned when he cried, and the tantrums came back at full strength — requiring a second extinction, which again succeeded.[19] The case contains the whole story: burst, curve, accidental intermittent reinforcement, recovery, and the persistence needed to see it through. ## How to use extinction 1. **Identify the maintaining reinforcer.** Attention? Escape? An item? A sensation? If you cannot tell, the procedure cannot be designed. 2. **Make sure you control that reinforcer.** If a classroom behavior is maintained by peers' laughter, a teacher's ignoring changes nothing. 3. **Decide whether you can withstand the burst.** If the behavior is dangerous at its worst — self-injury, aggression, bolting — extinction alone is not appropriate, and a professional should design the plan. 4. **Reinforce a replacement at the same time.** Extinction says what no longer works; differential reinforcement says what does. The combination reduces bursts and aggression. 5. **Get everyone on the same page.** One family member or one weekend with the grandparents can reinforce the behavior intermittently and undo weeks of progress. 6. **Expect recovery, resurgence, and renewal.** Returns after a break, under stress, and in new settings are normal, and each is weaker if the plan holds. 7. **Track it.** Extinction curves are bumpy; without data a bad week looks like failure. ### When not to use extinction - When the behavior is dangerous and the burst could cause harm. - When you cannot control the maintaining reinforcer — peer attention, sensory consequences you cannot block, reinforcers delivered by others. - When you cannot be consistent; intermittent extinction is intermittent reinforcement. - When an antecedent change would remove the need — moving the cookie jar is easier than extinguishing cookie-seeking. [Antecedent strategies in the ABC model ›](https://operantconditioning.com/abc-model/) ## Combining extinction with differential reinforcement Extinction is rarely used alone in modern practice: by itself it produces the burst and the emotional side effects and leaves a hole where the behavior was. The standard package pairs it with reinforcement of something else.[6][20] - **DRA (alternative behavior):** reinforce an appropriate behavior that gets the same result. The child who screamed for help is taught to tap the adult's arm; tapping is answered every time, screaming never. - **DRI (incompatible behavior):** reinforce a behavior that cannot happen at the same time. Sitting is reinforced; jumping up cannot coexist with it. - **DRO (other behavior):** reinforce the passage of time without the behavior. Every five minutes without shouting earns a point. - **Noncontingent reinforcement:** deliver the maintaining reinforcer freely on a schedule, so the behavior has nothing to earn. With a replacement in place, extinction becomes a redirection rather than a standoff: the reinforcer the learner wants is still available, through the door you chose. The same applies when the learner is you — the reliable way to extinguish a habit you do not want is to make its payoff available from one you do. [Building habits with operant conditioning ›](https://operantconditioning.com/habits/) · [Positive reinforcement in depth ›](https://operantconditioning.com/positive-reinforcement/) ## Key takeaways - Extinction breaks the contingency: the behavior occurs and its reinforcer no longer arrives. Nothing is added or removed, so it is not one of the four quadrants, and it is not forgetting but new learning laid over the old. - The extinction burst is a temporary rise in the frequency, intensity, or variability of the behavior. Giving in during the burst reinforces a more intense version on an intermittent schedule, so do not start extinction unless you can outlast it. - Extinguished behavior returns in three ways: spontaneous recovery after time away, resurgence when a newer behavior stops being reinforced, and renewal when the context changes. Each is weaker if the plan holds, and none means extinction failed. - Behavior reinforced intermittently is far more resistant to extinction than behavior reinforced every time: the partial reinforcement extinction effect. Giving in "only sometimes" builds whining on a variable schedule and guarantees a long extinction. - Ignoring is extinction only when attention is the maintaining reinforcer; escape-maintained behavior requires escape extinction, in which the demand stays in place. In practice, pair extinction with reinforcement of a replacement, which reduces bursts and aggression. ### Check yourself **A child tantrums when told to put on shoes. The parent ignores the tantrum, and the shoes stay off. Is this extinction?** No. The tantrum is maintained by escape from the demand, not by attention, so ignoring it while the shoes stay off delivers the reinforcer. Escape extinction means the shoes still go on; the behavior no longer ends the demand. **A child whines for cookies. Parent A never gives a cookie for whining again. Parent B takes away screen time for each whine. Which one is extinction?** Parent A. Extinction stops delivering the reinforcer the behavior used to produce; nothing is added or removed. Parent B removes a reinforcer the child already has, which is negative punishment, not extinction. **Bedtime crying that stopped last week returns after a weekend at the grandparents' house. Has extinction failed?** Not necessarily. A return after time away from the extinction setting is spontaneous recovery, which is normal and weaker each time if reinforcement stays withheld. The real risk is that someone responded to the crying over the weekend: one intermittent reinforcement can bring the behavior back at full strength. **Two children whined for candy. One was always given candy for whining; the other got it only sometimes. Both families now stop entirely. Whose whining fades faster, and why?** The child who was always given candy. Behavior reinforced every time is far easier to extinguish than behavior reinforced intermittently, which is the partial reinforcement extinction effect. The child on the continuous schedule can tell instantly that the rules have changed; the other cannot, because unrewarded stretches were always part of the deal. **Explain it to a friend.** Explain why a behavior can get worse before it fades once its payoff stops, using an example that has nothing to do with children or pets. ## Frequently asked questions **What is extinction in psychology?** The weakening and eventual disappearance of a learned behavior when the consequence that maintained it stops occurring. In operant conditioning, a behavior is no longer followed by its reinforcer and declines. In classical conditioning, a conditioned stimulus is presented without the unconditioned stimulus and the conditioned response fades. **What is an extinction burst?** A temporary increase in the frequency, intensity, or variability of a behavior right after its reinforcer is withheld — pressing the elevator button harder and faster before giving up. Bursts occur in a substantial minority of cases when extinction is used alone and are less likely when it is combined with reinforcement of another behavior. **What is an example of extinction in operant conditioning?** A child who used to get a cookie by whining stops getting cookies for whining; after an initial increase, the whining fades over days. Others: a dog stops begging when scraps stop, you stop pressing a broken elevator button, a toddler stops crying at bedtime once parents reliably stop returning. **Is extinction the same as ignoring?** Only when attention is the reinforcer maintaining the behavior. Extinction means withholding the specific reinforcer that keeps the behavior going. If a behavior is maintained by escape from a task, ignoring it lets the escape happen and makes things worse; escape extinction means keeping the task in place. **Is extinction a form of punishment?** No. Punishment adds an aversive stimulus or removes a reinforcer the person already has. Extinction simply stops delivering the reinforcer the behavior used to produce. Both reduce behavior, but extinction is slower, produces a burst, and is not one of the four quadrants. **What is spontaneous recovery?** The return of an extinguished behavior after time away from the extinction setting — the rat that stopped pressing on Monday presses again briefly on Tuesday. Each recovery is weaker than the last if reinforcement continues to be withheld. **Why is behavior on a variable schedule harder to extinguish?** Because unrewarded stretches were always part of the experience, so the organism cannot tell that reinforcement has stopped. This is the partial reinforcement extinction effect. It is why slot machines and social media are hard to quit and why "giving in sometimes" produces the most persistent whining. **How long does extinction take?** It depends on the reinforcement history. Behavior reinforced every time may fade within a session or a few days; behavior reinforced intermittently can persist for hundreds of unreinforced responses. In the classic bedtime-crying case, tantrums disappeared within about ten nights. Consistency matters more than anything else. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Bouton, M. E. (2004). Context and behavioral processes in extinction. *Learning & Memory, 11*(5), 485–494. 3. Lerman, D. C., & Iwata, B. A. (1995). Prevalence of the extinction burst and its attenuation during treatment. *Journal of Applied Behavior Analysis, 28*(1), 93–94. 4. Lerman, D. C., Iwata, B. A., & Wallace, M. D. (1999). Side effects of extinction: Prevalence of bursting and aggression during the treatment of self-injurious behavior. *Journal of Applied Behavior Analysis, 32*(1), 1–8. 5. Azrin, N. H., Hutchinson, R. R., & Hake, D. F. (1966). Extinction-induced aggression. *Journal of the Experimental Analysis of Behavior, 9*(3), 191–204. 6. Lerman, D. C., & Iwata, B. A. (1996). Developing a technology for the use of operant extinction in clinical settings: An examination of basic and applied research. *Journal of Applied Behavior Analysis, 29*(3), 345–382. 7. Rescorla, R. A. (2004). Spontaneous recovery. *Learning & Memory, 11*(5), 501–509. See also Pavlov, I. P. (1927). *Conditioned Reflexes* (G. V. Anrep, Trans.). Oxford University Press. 8. Epstein, R. (1983). Resurgence of previously reinforced behavior during extinction. *Behaviour Analysis Letters, 3*, 391–397. 9. Epstein, R. (1985). Extinction-induced resurgence: Preliminary investigations and possible applications. *The Psychological Record, 35*(2), 143–153. 10. Lattal, K. A., & St. Peter Pipkin, C. (2009). Resurgence of previously reinforced responding: Research and application. *The Behavior Analyst Today, 10*(2), 254–266. 11. Bouton, M. E. (2002). Context, ambiguity, and unlearning: Sources of relapse after behavioral extinction. *Biological Psychiatry, 52*(10), 976–986. 12. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158. 13. Amsel, A. (1958). The role of frustrative nonreward in noncontinuous reward situations. *Psychological Bulletin, 55*(2), 102–119. 14. Capaldi, E. J. (1966). Partial reinforcement: A hypothesis of sequential effects. *Psychological Review, 73*(5), 459–477. 15. Nevin, J. A. (1974). Response strength in multiple schedules. *Journal of the Experimental Analysis of Behavior, 21*(3), 389–408. 16. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 17. Iwata, B. A., Pace, G. M., Cowdery, G. E., & Miltenberger, R. G. (1994). What makes extinction work: An analysis of procedural form and function. *Journal of Applied Behavior Analysis, 27*(1), 131–144. 18. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 19. Williams, C. D. (1959). The elimination of tantrum behavior by extinction procedures. *Journal of Abnormal and Social Psychology, 59*(2), 269. 20. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Why intermittent schedules resist extinction — with a live simulator. - [Negative punishment](https://operantconditioning.com/negative-punishment/): Time-out and response cost, and how they differ from extinction. - [Shaping](https://operantconditioning.com/shaping/): How the variability of the extinction burst builds new behavior. --- # Shaping in Operant Conditioning: Successive Approximations, Step by Step > Shaping is differential reinforcement of successive approximations to a target behavior. How it works, a step-by-step protocol, examples, shaping vs. chaining. - Source: https://operantconditioning.com/shaping/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Core process* You cannot reinforce a behavior that never happens. Shaping is how operant conditioning builds behavior that does not yet exist — one small approximation at a time, in pigeons, children, patients, and yourself. > **Definition** > > **Shaping** is the **differential reinforcement of successive approximations** to a target behavior. Responses that come closer to the goal are reinforced; earlier, cruder forms are no longer reinforced; and the criterion for reinforcement is moved step by step until the target behavior appears. > > The term is B. F. Skinner's, and so is the analogy: "Operant conditioning shapes behavior as a sculptor shapes a lump of clay." Shaping is the standard method for teaching a new response in [applied behavior analysis](https://operantconditioning.com/applications/#aba), animal training, rehabilitation, and skill learning of every kind.[1] **In brief** - Shaping is differential reinforcement of successive approximations: reinforce responses closer to the goal, extinguish cruder ones, and raise the criterion step by step. - Reinforcement acts only on behavior that occurs; shaping builds behavior that does not yet exist by selecting from the learner's natural variation. - Mark each approximation immediately, keep the steps small enough that reinforcement stays frequent, and [thin the schedule](https://operantconditioning.com/schedules-of-reinforcement/) once the terminal behavior is reliable. ## Why shaping is needed [Reinforcement](https://operantconditioning.com/positive-reinforcement/) can only act on behavior that occurs. A rat in an operant chamber will not press the lever by chance for a very long time; a child who has never said "water" cannot be praised for saying it; a stroke patient who cannot lift an arm cannot be rewarded for lifting it. If you wait for the finished behavior, you wait forever. Shaping solves the problem by reinforcing whatever the organism *can* do that resembles the goal, then demanding a little more. It relies on a fact that is easy to miss: no response is ever repeated exactly. Every lever press has a slightly different force, every attempt at a word a slightly different sound. Shaping selects from that natural variation, the way breeding selects from variation in a population.[1] ## How shaping works: the mechanism Shaping is two procedures running together — reinforcement of the current approximation and [extinction](https://operantconditioning.com/extinction/) of everything else — and it cycles through four phases: ![Shaping as a rising criterion](https://operantconditioning.com/assets/diagrams/shaping-as-a-rising-criterion.svg) *Shaping as a moving criterion. Reinforcing only the upper tail of what the animal currently does shifts the whole distribution; the criterion is then raised again. Variability (the width of each curve) is what makes the next step possible.* 1. **Reinforce a starting behavior.** The rat turns toward the lever; a pellet drops. Turning toward the lever increases in frequency. 2. **Variability appears.** As the behavior is repeated it varies: some turns are closer, some include a step forward, some a raised paw. Extinction increases variability further — when a response stops paying off, organisms do it in new ways.[2][3] 3. **Raise the criterion.** Once turns are reliable, only turns that include a step forward are reinforced. Plain turns go on extinction; steps forward increase. 4. **Repeat.** Steps toward, touching, pawing, pressing. Each stage is built on the reinforced behavior of the previous one and the extinction-driven variability that behavior produces. The variability step is the one people forget. Allen Neuringer's research showed that variability is not noise around a "true" response but a dimension of behavior that reinforcement controls: reinforce variety and organisms become more variable; reinforce sameness and they become stereotyped.[2] Good shaping keeps the current form reinforced often enough to persist but not so often that it hardens into a fixed habit. > **Differential reinforcement is the engine** > > "Differential" means some forms of the response are reinforced and others are not. Without the extinction half you are not shaping — you are reinforcing whatever happens, and the behavior settles at the easiest form that still pays. Without the reinforcement half, the behavior disappears. The skill is holding both at once and moving the line between them. ## Skinner's discovery: the pigeon that learned to bowl Skinner dated his understanding of shaping to a single day in 1943. He, Keller Breland, and Norman Guttman were working on [Project Pigeon](https://operantconditioning.com/bf-skinner/) on the top floor of a flour mill in Minneapolis and decided, for amusement, to teach a pigeon to swipe a small wooden ball down a miniature alley with its beak. Waiting for a full swipe went nowhere. So they reinforced any response that resembled it — a look at the ball, a move toward it, a touch — and within minutes the bird was bowling.[4][5] Skinner later wrote that the day changed how he thought about behavior: complex acts could be built rapidly from nothing by hand-delivered reinforcement. He showed the public how in a 1951 *Scientific American* article, "How to Teach Animals," which walks a reader through establishing a conditioned reinforcer (a sound paired with food) and using it to shape a dog or pigeon to a new behavior in a single session — the recipe clicker trainers still follow.[6] His pigeons went on to play a version of ping-pong, pecking a ball back and forth across a table.[7] ## How to shape a behavior: step-by-step 1. **Define the terminal behavior.** Be exact about the finished form: "sits on the mat with all four paws for 10 seconds," "says 'water' clearly," "raises the affected arm to shoulder height." A vague goal cannot be approximated. 2. **Find the starting behavior.** Something the learner already does, at least occasionally, that is on the road to the goal. Glancing at the mat. Saying "wa." Moving the arm two inches. If it never occurs, you need an easier start. 3. **Plan the approximations.** Write down the steps you expect, but hold the plan loosely — the learner's variability will suggest better ones. 4. **Reinforce immediately and every time.** Use a conditioned reinforcer (a click, a "yes," a checkmark) to mark the exact instant the criterion is met, then deliver the real reinforcer. A second's delay reinforces whatever happened in that second instead. 5. **Raise the criterion in small steps.** Move on when the current approximation is reliable — Karen Pryor's rule of thumb is when the learner succeeds most of the time — and keep each step small enough that reinforcement stays frequent.[8] 6. **Don't stay too long; don't move too fast.** Staying reinforces a plateau until it resists change. Moving too fast puts the learner on extinction: variability spikes, then the behavior collapses. If that happens, drop back a step. 7. **Thin the schedule at the end.** Once the terminal behavior is reliable, shift to [intermittent reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) and bring the behavior under the cue you want it to answer to. ## Examples of shaping across settings | Setting | Terminal behavior | Successive approximations | | --- | --- | --- | | Laboratory | Rat presses a lever | Faces lever → approaches → touches → rests paw on it → presses | | Speech / early intervention | Child says "water" | Any vocalization → "wa" → "wa-wa" → "wa-ter" → "water" clearly; the same method Lovaas used to build first words[9] | | Parenting | Toilet training | Sits on potty clothed → sits unclothed → sits at scheduled times → urinates on potty → initiates independently[10] | | Classroom | Shy student answers in class | Nods → one-word answer to a direct question → a sentence → volunteers an answer → asks a question | | Dog training | "Go to mat" | Looks at mat → steps toward → one paw on → all four paws → lies down → stays as handler moves away | | Rehabilitation | Stroke patient uses affected arm | Small movements reinforced, then larger range, then functional tasks — the core of constraint-induced movement therapy[11] | | Psychiatric care | Mute patient speaks | Eye movement toward gum → lip movement → any sound → a word → answering questions, in a classic 1960 case[12] | | Self-management | Writing 500 words a day | Open the document → one sentence → one paragraph → 100 words → 250 → 500 | | Business (analogy) | A finished product | Minimum viable version → customer feedback → iterate; a loose analogy, since the "reinforcer" is market data rather than an immediate consequence for a specific response | ## Shaping vs. chaining Shaping builds a *new form* of a single response. **Chaining** links *existing* responses into a sequence, where each step produces the cue for the next and the whole chain ends in a reinforcer. Making a bed, brushing teeth, and a dog's retrieve are chains. The stimulus each step produces does two jobs at once: it is the discriminative stimulus for the next response and a conditioned reinforcer for the one just completed, which is why chains hold together — and why a link that stops paying early in a chain lets everything after it fall apart.[16] The first step in chaining is a **task analysis**: breaking the sequence into teachable components.[13] | Aspect | Shaping | Forward chaining | Backward chaining | Total-task chaining | | --- | --- | --- | --- | --- | | **What is taught** | A new response form | A sequence, first step first | A sequence, last step first | Whole sequence every trial | | **How** | Reinforce closer approximations; extinguish earlier ones | Teach step 1 to mastery, then 1+2, and so on | Trainer does all but the last step; learner completes it and is reinforced; then last two… | Learner attempts all steps with prompts where needed | | **Reinforcer** | After each approximation | After the last mastered step | Always at the natural end of the chain | At the end, plus prompts faded | | **Best for** | Behavior that does not yet occur in any form | Learners who can already do the early steps | Learners who benefit from finishing every trial with success | Learners who can do most steps already | | **Example** | Teaching a first word | Putting on a coat | Zipping a jacket (start with the last inch) | Making a sandwich with a picture schedule | In practice the two combine: a step within a chain that the learner cannot yet perform is shaped. Backward chaining has a special advantage — the learner completes the chain and contacts the terminal reinforcer on every trial, and each newly added step is reinforced by the chance to perform the steps already mastered.[13] ## Shaping vs. prompting and fading A **prompt** is help added before or during the response to make it occur: a verbal instruction, a gesture, a model to imitate, or physical guidance. Prompting produces the behavior now; shaping waits for it to emerge. Both end in the same place — the behavior under the control of its natural cue — but prompting requires **fading**: removing the help gradually so the learner does not become dependent on it.[13] - **Most-to-least prompting** starts with full physical guidance and fades to lighter prompts; useful for learners who make many errors. - **Least-to-most prompting** gives the learner a chance to respond alone and adds help only as needed; it lets independent responding happen sooner. - **Time delay** keeps the prompt but waits progressively longer before giving it, so the learner has room to beat the prompt. The rule of thumb: prompt when the behavior is physically possible but the learner does not know what is wanted; shape when the behavior does not yet exist in the repertoire. Many programs do both — prompt a rough form, fade the prompt, then shape toward fluency. [More on cues, prompts, and stimulus control ›](https://operantconditioning.com/abc-model/) ## Shaping vs. luring in animal training Trainers distinguish **free shaping** — waiting for the animal to offer approximations and marking them — from **luring**, in which food is used to steer the animal into position (a treat moved over a dog's head produces a sit). Luring is faster for simple positions but has a cost: the dog learns to follow the food, and the lure must itself be faded or the behavior never happens without it. Shaped behavior tends to be more durable and produces animals that actively experiment; luring is a prompt and should be treated like one. [Reinforcement-based dog training ›](https://operantconditioning.com/dog-training/) ## Common shaping mistakes - **Steps that are too big.** The learner rarely meets the new criterion, reinforcement drops, and the behavior extinguishes. The fix is always to split the step. - **Staying too long on one step.** A heavily reinforced approximation becomes stereotyped and the learner stops varying. Move on while variability is still present. - **Inconsistent criteria.** Reinforcing a below-criterion response "because they tried" teaches that the criterion is negotiable. - **Late reinforcement.** A marker delivered a second late reinforces whatever followed the target — the head-turn after the sit. This is how [superstitious behaviors](https://operantconditioning.com/glossary/#superstitious-behavior) get built into a shaped response. - **Lumping instead of splitting.** Karen Pryor's term for demanding two improvements at once — a longer stay *and* a straighter sit. Raise one criterion at a time, and relax the others temporarily when you introduce a new one.[8] - **Not planning the endgame.** Shaping ends with the terminal behavior on continuous reinforcement, which is fragile. Thin the schedule and add the cue, or the behavior will not survive real life. ## Shaping yourself: technology and self-management Most successful [habit-building](https://operantconditioning.com/habits/) is shaping in disguise. The fitness watch that suggests a slightly higher daily goal after a week of hitting the old one, the running program that adds a minute each week, the language app that lengthens sessions as you improve — all are raising criteria on a behavior they reinforce immediately. The ones that fail usually fail by lumping: they ask for the terminal behavior on day one. To shape your own behavior, start with an approximation you will certainly emit — one push-up, one sentence, shoes on by the door — and reinforce it on the spot, even if only by marking it done. When the small behavior is happening on most days, raise the criterion a little. Real-world habit research suggests automaticity takes weeks to months and varies widely between people and behaviors, so raise the bar on the evidence of your own record, not a calendar.[14] A missed step is data: the criterion was raised too far, and the correct response is to split it, not to try harder. ### The percentile schedule: shaping as a formula Because human shapers drift — too generous on good days, too strict on bad ones — researchers formalized the procedure. In a **percentile schedule** the criterion is set from the learner's own recent performance: a response is reinforced if it beats a fixed percentage of the last several responses (say, half of the last ten). The criterion rises exactly as fast as the learner improves, falls if performance drops, and keeps the rate of reinforcement roughly constant. Gregory Galbicka's 1994 review proposed bringing the method into applied settings; it remains the clearest statement of what a good shaper does intuitively.[15] ## Key takeaways - Shaping is two procedures running together: reinforcement of the current approximation and extinction of everything else. Without the extinction half you are reinforcing whatever happens; without the reinforcement half the behavior disappears. - Variability is what makes each step possible. No response is ever repeated exactly, extinction increases variability further, and reinforcement itself can make behavior more variable or more stereotyped. - Define the terminal behavior exactly, start with something the learner already does, mark the instant the criterion is met with a conditioned reinforcer, and raise the criterion in small steps. Steps that are too big put the learner on extinction; staying too long hardens a plateau. - Shaping builds a new form of a single response; chaining links existing responses into a sequence; prompting adds help that must then be faded. Shape when the behavior does not yet exist, and prompt when it is possible but the learner does not know what is wanted. - Shaping ends with the terminal behavior on continuous reinforcement, which is fragile. Thin the schedule and bring the behavior under its cue, or it will not survive real life. ### Check yourself **You are shaping yourself toward writing 500 words a day. After a week at 100 words you jump to 400 and miss three days in a row. What went wrong, and what is the fix?** The criterion was raised too far, so reinforcement dropped and the behavior went on extinction. A missed step is data, not a character flaw: the fix is always to split the step, not to try harder. **A trainer shaping a dog to lie on its mat sometimes clicks when only two paws are on the mat, "because it tried." Why is this a problem?** Reinforcing a below-criterion response teaches that the criterion is negotiable. Differential reinforcement is the engine of shaping, with some forms reinforced and others not; without the extinction half, the behavior settles at the easiest form that still pays. **A trainer moves a treat over a dog's head so that it sits, and repeats this for a week. The dog now sits only when a treat is in the hand. Was the sit shaped?** No. That is luring, in which food steers the animal into position; it is a prompt, and it must be faded or the behavior never happens without it. Shaping waits for the animal to offer approximations and marks them, which produces more durable behavior. **A child can put each arm into a coat sleeve, pull the coat up, and zip it, but never does them in order without help. Shaping or chaining?** Chaining. The responses already exist; what is missing is the sequence, in which each step produces the cue for the next and the whole chain ends in a reinforcer. Shaping is for a behavior that does not yet occur in any form. **Explain it to a friend.** Explain how shaping builds a behavior that has never happened yet, without using the words "step," "approximation," or "criterion." ## Frequently asked questions **What is shaping in psychology?** Shaping is a procedure in operant conditioning for teaching a behavior that does not yet occur. The trainer reinforces successive approximations — responses that come progressively closer to the target — while withholding reinforcement from earlier, cruder forms, until the target behavior is performed. **What are successive approximations?** The series of intermediate behaviors between the learner's starting point and the goal, each a little closer to the goal than the last. Teaching a rat to press a lever, the approximations might be facing the lever, approaching it, touching it, and pressing it. **What is an example of shaping?** A parent teaching a toddler to say "water" praises "wa," then only "wa-wa," then only the full word. A dog trainer teaching "go to your mat" clicks and treats for looking at the mat, then a step toward it, then a paw on it, then all four paws, then lying down. **What is the difference between shaping and chaining?** Shaping creates a new form of a single behavior by reinforcing closer approximations. Chaining links behaviors the learner can already do into a sequence, such as the steps of brushing teeth. Chaining can be done forward, backward, or as a total task, and a step the learner cannot perform is often shaped separately. **Who invented shaping?** B. F. Skinner, with Keller Breland and Norman Guttman, discovered shaping in 1943 while teaching a pigeon to "bowl" during Project Pigeon. Skinner described the method in *Science and Human Behavior* (1953) and for a general audience in "How to Teach Animals" (*Scientific American*, 1951). **What is differential reinforcement of successive approximations?** It is the technical definition of shaping. "Differential reinforcement" means some responses are reinforced and others are placed on extinction; "successive approximations" means the reinforced responses are, in sequence, progressively closer to the target behavior. **How is shaping used in ABA therapy?** Behavior analysts use shaping to build first words and speech sounds, motor skills, feeding and self-care behaviors, social responses such as eye contact, and tolerance of medical or dental procedures. It is usually combined with prompting, fading, and chaining within a task-analyzed program. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Neuringer, A. (2002). Operant variability: Evidence, functions, and theory. *Psychonomic Bulletin & Review, 9*(4), 672–705. See also Page, S., & Neuringer, A. (1985). Variability is an operant. *Journal of Experimental Psychology: Animal Behavior Processes, 11*(3), 429–452. 3. Antonitis, J. J. (1951). Response variability in the white rat during conditioning, extinction, and reconditioning. *Journal of Experimental Psychology, 42*(4), 273–281. 4. Peterson, G. B. (2004). A day of great illumination: B. F. Skinner's discovery of shaping. *Journal of the Experimental Analysis of Behavior, 82*(3), 317–328. 5. Skinner, B. F. (1958). Reinforcement today. *American Psychologist, 13*(3), 94–99. 6. Skinner, B. F. (1951). How to teach animals. *Scientific American, 185*(6), 26–29. 7. Skinner, B. F. (1962). Two "synthetic social relations." *Journal of the Experimental Analysis of Behavior, 5*(4), 531–533. 8. Pryor, K. (1984). *Don't Shoot the Dog! The New Art of Teaching and Training*. Simon & Schuster. 9. Lovaas, O. I., Berberich, J. P., Perloff, B. F., & Schaeffer, B. (1966). Acquisition of imitative speech by schizophrenic children. *Science, 151*(3711), 705–707. 10. Azrin, N. H., & Foxx, R. M. (1971). A rapid method of toilet training the institutionalized retarded. *Journal of Applied Behavior Analysis, 4*(2), 89–99. 11. Taub, E., Crago, J. E., Burgio, L. D., Fleming, W. C., Nepomuceno, C. S., Connell, J. S., & Miller, N. E. (1994). An operant approach to rehabilitation medicine: Overcoming learned nonuse by shaping. *Journal of the Experimental Analysis of Behavior, 61*(2), 281–293. 12. Isaacs, W., Thomas, J., & Goldiamond, I. (1960). Application of operant conditioning to reinstate verbal behavior in psychotics. *Journal of Speech and Hearing Disorders, 25*(1), 8–12. 13. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 14. Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. *European Journal of Social Psychology, 40*(6), 998–1009. 15. Galbicka, G. (1994). Shaping in the 21st century: Moving percentile schedules into applied settings. *Journal of Applied Behavior Analysis, 27*(4), 739–760. 16. Kelleher, R. T., & Gollub, L. R. (1962). A review of positive conditioned reinforcement. *Journal of the Experimental Analysis of Behavior, 5*(4), 543–597. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The consequence that does the work in every approximation. - [Extinction](https://operantconditioning.com/extinction/): The other half of differential reinforcement — bursts, variability, and recovery. - [Habits](https://operantconditioning.com/habits/): Shaping your own behavior with antecedents, tiny steps, and immediate consequences. --- # Stimulus Control: Discrimination, Generalization, Peak Shift, and Errorless Learning > Stimulus control means a behavior's probability depends on an antecedent. SD vs. S-delta, discrimination, generalization, peak shift, and errorless learning. - Source: https://operantconditioning.com/stimulus-control/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Antecedents · Core concept* Consequences decide whether a behavior is learned; antecedents decide when it shows up. Here is how a stimulus comes to govern behavior without forcing it, how far that control spreads, why it can shift in strange directions, and how it is used on everything from pigeons to insomnia. > **Definition** > > A behavior is under **stimulus control** when its probability — how often it occurs, how quickly, or in what form — depends on whether a particular antecedent stimulus is present. The stimulus does not force the behavior; it changes the odds, because in the past the behavior has been reinforced in its presence and not in its absence. > > Skinner introduced the analysis, and the symbols S^D and S^Δ, in *The Behavior of Organisms* (1938); the classic experimental review is Terrace (1966).[1][2] **In brief** - A behavior is under stimulus control when its probability depends on an antecedent stimulus, because it has been reinforced in that stimulus's presence before. - A discriminative stimulus (S^D) signals that a response will be reinforced; an S-delta signals that it will not, and discrimination training alternates the two. - Generalization spreads responding to similar stimuli along a gradient; discrimination training steepens that gradient and can shift its peak away from the S-delta. ## What is stimulus control? Put a rat in a chamber where lever presses produce food only while a light is on. At first it presses at the same rate light or dark. After a few sessions it presses briskly the moment the light comes on and hardly at all when it goes off. Nothing about the light compels the press — a rat that has just eaten will ignore it — but the light has become the condition under which pressing is worth doing. Pressing is under stimulus control, and the gap between the two rates measures how tightly.[1] Human life is dense with the same relation. A ringing phone, a green light, a colleague's raised eyebrow, an "Open" sign, the first bars of a song you know: each raises the probability of a specific behavior because that behavior has paid off in its presence before. This is the antecedent term of the [three-term contingency](https://operantconditioning.com/abc-model/) — in the presence of S^D, response R produces reinforcer S^R. Reinforcement builds a behavior; stimulus control decides where and when it appears.[3] > **Control is not compulsion** > > "Control" is a technical word here, closer to the way a thermostat controls a furnace than the way a puppeteer controls a puppet. A stimulus controls an operant by changing its probability, and only because of a history of consequences; change the consequences and the same stimulus stops working. A stimulus that *forces* a response — the puff of air that makes you blink — is eliciting a reflex, a different process. [Operant vs. classical conditioning ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## S^D and S-delta: how discrimination training works A **discriminative stimulus**, S^D (pronounced "ess-dee"), is a stimulus in whose presence a response has been reinforced. An **S-delta**, S^Δ, is a stimulus in whose presence the same response has gone unreinforced. In the generalization literature the pair is often written S+ and S−. **Discrimination training** is simply reinforcement in one and extinction in the other, alternated until responding diverges. The organism is said to *discriminate* when it responds differently to the two.[1] | S^D (respond) | S^Δ (don't) | Behavior | What the S^D signals | | --- | --- | --- | --- | | Light on | Light off | Rat presses the lever | Presses will produce food | | "Sit" | Any other word | Dog sits | A treat is available for sitting now | | Teacher looking at the class | Teacher writing on the board | Student raises a hand | Hand-raising will get called on | | Friends at a bar | Grandmother at dinner | Telling a crude joke | Laughter, not a frown | | Green "Walk" signal | Red hand | Crossing the street | Crossing will be safe and unfined | | Bed, dark room, 11 p.m. | Desk, daylight | Falling asleep | Sleep will come — unless the bed has also become a cue for scrolling | Stimulus control is not automatic. A stimulus that is always present, with no difference in consequences attached, may acquire no control at all. Jenkins and Harrison trained pigeons to peck a key with a 1000-hertz tone playing throughout; tested later with tones of other pitches, the birds responded equally at every pitch — a flat gradient. The tone had been there all along, but they had never had a reason to attend to it. Birds trained with the tone on during reinforcement and off during extinction produced a sharply peaked gradient centered on 1000 hertz.[4] A cue has to *predict a difference* to gain control, which is why "sit" said to a dog that will be fed anyway teaches nothing. ## Generalization and the generalization gradient **Stimulus generalization** is the mirror image of discrimination: responding to stimuli that resemble the S^D without having been trained on them. The more similar the stimulus, the more responding it evokes — a relation that, plotted, is a **generalization gradient**. The textbook demonstration is Guttman and Kalish's. They reinforced pigeons for pecking a key lit with a single wavelength — separate groups at 530, 550, 580, and 600 nanometers — and then, in extinction so the test itself taught nothing, presented a range of wavelengths on either side. Responding peaked at the training wavelength and fell away smoothly on both sides.[5] The gradient's shape is not fixed: discrimination training steepens it, and what the organism learns is less the stimulus itself than how it differs from what surrounds it.[6] | Aspect | Discrimination | Generalization | | --- | --- | --- | | **What it is** | Responding differently to different stimuli | Responding similarly to similar stimuli | | **Laboratory sign** | A steep gradient; near-zero responding in S^Δ | A flat, wide gradient | | **Everyday example** | Answering your own ringtone, not a stranger's | Braking for a stop sign you have never seen before | | **When it fails you** | A skill that works only in the room it was taught in | A fear of one dog that spreads to every dog | | **How to get more of it** | Differential reinforcement: pay in S^D, never in S^Δ | Train with many examples in many settings | Both are adaptive and both can misfire. Generalization lets a toddler call the neighbor's terrier a dog; discrimination stops her calling the cat one. In applied work generalization is usually the scarce commodity: a behavior taught in one place with one person by one method tends to stay there unless generalization is deliberately programmed.[7] ## Peak shift: when discrimination training moves the peak Discrimination training does something stranger than sharpen the gradient. In 1937 Kenneth Spence proposed that an S+ builds up a gradient of excitation and an S− a gradient of inhibition, and that the two add algebraically. If the S− sits close to the S+ on the same dimension, the sum should peak not at the S+ but a little beyond it, on the side away from the S−.[8] The prediction sat on paper for twenty years. H. M. Hanson tested it in 1959. He reinforced pigeons for pecking at 550 nm, then extinguished pecking at a slightly yellower 555 nm (other groups had S− at 560, 570, or 590 nm; a control group had no S− at all). In the generalization test, the control birds peaked at 550, as Guttman and Kalish's had. The birds trained against 555 peaked at 540 — a wavelength they had never been reinforced for, displaced away from the S− — and pecked far more overall than the controls. The nearer the S− had been to the S+, the larger the shift.[9] This is **peak shift**: evidence that the S^Δ does not merely switch responding off but exerts an inhibitory gradient of its own, which subtracts most on the side nearest it. ![Peak shift after discrimination training](https://operantconditioning.com/assets/diagrams/peak-shift-after-discrimination-training.svg) *Schematic of Hanson's result. After discrimination training against a nearby S−, the peak of the gradient moves away from the S− (here from 550 to about 540 nm) and rises above the control gradient.* ## Errorless discrimination learning In ordinary discrimination training the learner meets the S^Δ at full strength, responds to it, and is extinguished: it learns by making errors. Herbert Terrace showed in 1963 that the errors are optional. He trained pigeons to peck a red key and not a green one, but introduced the green key from the first session so dimly and briefly that the birds never pecked it, then raised its brightness and duration in small steps. Pigeons trained this way acquired the discrimination with few or no errors, where birds trained conventionally made thousands.[10] In a companion study he transferred the discrimination from colors to vertical and horizontal lines by superimposing the lines on the colors and fading the colors out.[11] The errorless birds differed in more than their error count. They showed none of the agitation that conventionally trained birds displayed during the S−, and in a wavelength test they showed no peak shift: the S− had acquired no inhibitory gradient, because it had never been responded to and extinguished.[10][12] Whether a stimulus becomes aversive depends on how it was learned. The legacy is everywhere in teaching. **Prompting** — a gesture, a model, a highlighted answer, physical guidance — gets the right response before an error can occur, and **fading** transfers control from the prompt to the natural S^D in graded steps.[13] In neuropsychology, Baddeley and Wilson found that people with amnesia learned word lists better when prevented from guessing than by trial and error: without explicit memory they could not correct their mistakes, so each error was simply practiced.[14] Errorless learning has costs — a learner who never meets the S^Δ may be thrown when it finally appears, and a prompt that is never faded produces a learner who waits for it — but for fragile or fearful learners it is usually the right default. [Prompting and fading in depth ›](https://operantconditioning.com/shaping/) ## Contextual control and renewal Stimulus control belongs not only to discrete cues but to the background: the room, the time of day, the people present. Contexts acquire control more weakly than a lit key, but reliably, and their influence is sharpest at the moment of extinction. Mark Bouton's research showed that when a behavior is learned in one context and extinguished in another, returning to the first context brings it back — **renewal**. The original learning transfers across contexts; the extinction learning largely does not. Extinction does not erase what was learned but adds a second, inhibitory lesson whose retrieval depends on the context in which it was learned, so that the stimulus becomes ambiguous and the setting decides which meaning wins.[15] The practical consequences are large. A fear reduced in a therapist's office can return in the parking garage; a tantrum extinguished at the clinic can return at home; a habit dropped on holiday returns on the first morning back at work. The remedies follow directly: extinguish in every context that matters, and teach the replacement behavior where the old one used to pay. [Renewal, resurgence, and spontaneous recovery ›](https://operantconditioning.com/extinction/#renewal) ## Stimulus control in practice ### Stimulus control therapy for insomnia The most direct clinical use of the concept treats a bed that has stopped working. For a good sleeper, bed, darkness, and bedtime are S^Ds for falling asleep. For a chronic insomniac they have become cues for lying awake, worrying, checking the time, and watching television, because that is what has repeatedly happened there. Richard Bootzin's **stimulus control treatment**, introduced in 1972, re-establishes the bed as a cue for sleep and nothing else.[16] 1. **Go to bed only when sleepy**, not merely tired or because it is late. 2. **Use the bed only for sleep.** No reading, eating, screens, or worrying in bed. (Sex is the conventional exception.) 3. **If you cannot fall asleep within about 10 to 20 minutes, get up**, go to another room, and return only when sleepy. Repeat as needed. 4. **Get up at the same time every morning**, however little you slept. 5. **Do not nap** during the day. The rules are uncomfortable for a week or two and then, for most people, they work. Stimulus control is a core component of cognitive behavioral therapy for insomnia (CBT-I), which the American College of Physicians recommends as the first-line treatment for chronic insomnia in adults, ahead of medication.[17] ### "Train it everywhere": dog training A dog that sits perfectly in the kitchen has learned "sit, in the kitchen, facing my owner, treat pouch on." At the park none of those stimuli are present, and neither is the sit. The dog is not stubborn; it is discriminating exactly as its training taught it to. The fix is to program generalization: train with many examples — rooms, people, distances, distractions — until the word alone controls the behavior. Stokes and Baer's "train sufficient exemplars" is the same advice in the language of applied behavior analysis.[7] [Cues and generalization in dog training ›](https://operantconditioning.com/dog-training/) ### Habit design: choose the cue Habits are behaviors under tight stimulus control: the context evokes the response with little deliberation, and the response survives on the context rather than on intention.[18] That gives you two levers. To build a habit, attach the behavior to a specific, stable cue that already occurs — after the coffee is poured, when the front door closes — and reinforce it there until the cue does the work. To break one, remove or change the cue; a phone charging in the kitchen is not an S^D for scrolling in bed. Skinner listed "changing the stimulus" among the basic techniques of self-control for exactly this reason.[3] [The operant protocol for building habits ›](https://operantconditioning.com/habits/) ### Study in one place The classic study-skills prescription is Bootzin's rule applied to a desk. Work in one place used for nothing else; when you stop working, leave it. Over a few weeks the desk becomes an S^D for working and stops being one for daydreaming, snacking, and messaging, because those behaviors are never reinforced there. "I'll study on the couch" so rarely produces studying because the couch already controls something else. ## Which is it: discrimination or generalization? Five scenarios. Decide which process each shows before opening the answer. **1. A toddler who has learned "doggie" for the family beagle says it to a neighbor's Labrador, then to a goat.** **Generalization.** The response spreads to stimuli that resemble the training stimulus. The Labrador is a useful generalization; the goat is an overgeneralization that discrimination training — "no, that's a goat" — will correct. **2. A rat presses the lever rapidly when the light is on and almost never when it is off.** **Discrimination.** Responding differs sharply between S^D and S^Δ; the behavior is under tight stimulus control. **3. A dog that sits reliably in the kitchen ignores "sit" at the park.** **Discrimination — more than the trainer wanted.** The behavior is controlled by the kitchen's stimuli, not by the word. The trainer's job is to build generalization to the cue alone by training across settings. **4. You reach for your phone whenever you hear a notification chime, including other people's.** **Generalization.** The response evoked by your own chime spreads to similar chimes. If you later stop reaching for others' phones because only yours has ever paid off, that is discrimination developing. **5. A student swears freely with friends and never at the dinner table.** **Discrimination.** The same behavior occurs in one social setting (where it has been reinforced with laughter) and not in another (where it has met disapproval). Friends are the S^D; family dinner is the S^Δ. ## Common confusions ### Discriminative stimulus vs. conditioned stimulus | Aspect | Discriminative stimulus (S^D) | Conditioned stimulus (CS) | | --- | --- | --- | | **What it does** | *Evokes* an operant: raises its probability | *Elicits* a respondent: triggers a reflex | | **How it got its power** | The behavior was reinforced in its presence | It was paired with an unconditioned stimulus, whatever the organism did | | **Does the behavior matter?** | Yes: the consequence depends on responding | No: the US arrives regardless | | **Example** | A lit key: pecking now produces grain | A tone that precedes food: salivation follows the tone | The two often ride on the same event. The click of the food magazine is a CS (it elicits approach and salivation), a conditioned reinforcer (it strengthens the press it follows), and an S^D for going to the tray. Asking which function a stimulus is serving, rather than what it "is," is the behavior analyst's habit. ### Stimulus control vs. elicitation An S^D changes probabilities; it never guarantees a response, and a motivating operation can cancel it entirely (the lit key evokes nothing in a stuffed pigeon). Elicitation is closer to a guarantee: the reflex follows the stimulus whether or not the organism is hungry, tired, or busy. When a stimulus seems to "make" someone do something — a craving at the sight of a bar, a flinch at a raised hand — the elicited, Pavlovian component is often doing more of the work than the operant one. ### Three more - **An S^D is not a motivating operation.** The S^D signals that a reinforcer is *available*; a motivating operation changes how much it is *worth*. [Motivating operations explained ›](https://operantconditioning.com/abc-model/) - **An S^Δ is not a punisher.** It signals that responding will go unreinforced, not that it will be punished — though, as Terrace's pigeons showed, a stimulus learned through extinction can become mildly aversive in its own right. - **"Discrimination" here has no social meaning.** It is the technical word for responding differently to different stimuli, as in a discriminating palate. ## Key takeaways - Stimulus control means an antecedent changes the probability of a behavior; it does not force it. An S^D evokes an operant only because of a history of consequences, and a stimulus that forces a response is eliciting a reflex, a different process. - Discrimination training is reinforcement in the presence of the S^D and extinction in the presence of the S^Δ, alternated until responding diverges. A cue has to predict a difference in consequences to gain any control at all. - Generalization is responding to untrained stimuli that resemble the S^D, and the generalization gradient shows how responding falls off with similarity. Discrimination training steepens the gradient and, when the S− is close to the S+, shifts its peak away from the S−: peak shift. - Errors are optional. Introducing the S^Δ so faintly that it is never responded to, then fading it in, produces a discrimination with few or no errors, no agitation, and no peak shift; prompting and fading apply the same idea in teaching. - Contexts control behavior too, and extinction learning is more context-bound than the original learning, so an extinguished behavior renews when the setting changes. Stimulus control therapy for insomnia, training a dog in many settings, and choosing a cue for a habit all apply the same principle. **Explain it to a friend.** Explain why a dog that sits perfectly in the kitchen ignores "sit" at the park, in two sentences a twelve-year-old would follow. ## Frequently asked questions **What is stimulus control in psychology?** Stimulus control is the condition in which a behavior occurs more often, faster, or more reliably in the presence of a particular stimulus than in its absence, because the behavior has been reinforced in that stimulus's presence in the past. The stimulus is called a discriminative stimulus. Stimulus control is the antecedent side of operant conditioning: consequences build a behavior, and antecedents decide when it appears. **What is an example of stimulus control?** A rat that presses a lever only when a light is on; a driver who brakes at a red light and not a green one; a dog that sits when it hears "sit"; a student who swears with friends but not at the family table. In each case the behavior is possible at any time but is reliably evoked by one stimulus and not by others. **What is the difference between SD and S-delta?** An S^D (discriminative stimulus) is a stimulus in whose presence a behavior has been reinforced, so the behavior becomes more likely when it appears. An S^Δ (S-delta) is a stimulus in whose presence the same behavior has not been reinforced, so the behavior becomes less likely. Discrimination training alternates the two — reinforcement in S^D, extinction in S^Δ — until responding diverges. **What is the difference between discrimination and generalization?** Discrimination is responding differently to different stimuli: pressing when the light is on but not when it is off. Generalization is responding similarly to similar stimuli: pecking a slightly different color that was never trained. They are two ends of one continuum, and the generalization gradient — how responding falls off as a test stimulus becomes less like the training stimulus — measures where an organism sits on it. **What is peak shift?** After discrimination training between an S+ and a nearby S− on the same dimension, the peak of the generalization gradient moves away from the S−. Hanson's pigeons, reinforced at 550 nm and extinguished at 555 nm, responded most to 540 nm, a color they had never been reinforced for. Spence predicted the effect in 1937 from the idea that excitatory and inhibitory gradients add. **What is stimulus control therapy for insomnia?** A behavioral treatment, developed by Richard Bootzin in 1972, that makes the bed a cue for sleep and nothing else: go to bed only when sleepy, use the bed only for sleep, get up if you cannot sleep and return only when sleepy, rise at the same time every day, and do not nap. It is a core component of cognitive behavioral therapy for insomnia, the recommended first-line treatment for chronic insomnia. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Terrace, H. S. (1966). Stimulus control. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 271–344). Appleton-Century-Crofts. 3. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 4. Jenkins, H. M., & Harrison, R. H. (1960). Effect of discrimination training on auditory generalization. *Journal of Experimental Psychology, 59*(4), 246–253. 5. Guttman, N., & Kalish, H. I. (1956). Discriminability and stimulus generalization. *Journal of Experimental Psychology, 51*(1), 79–88. 6. Honig, W. K., & Urcuioli, P. J. (1981). The legacy of Guttman and Kalish (1956): Twenty-five years of research on stimulus generalization. *Journal of the Experimental Analysis of Behavior, 36*(3), 405–445. 7. Stokes, T. F., & Baer, D. M. (1977). An implicit technology of generalization. *Journal of Applied Behavior Analysis, 10*(2), 349–367. 8. Spence, K. W. (1937). The differential response in animals to stimuli varying within a single dimension. *Psychological Review, 44*(5), 430–444. 9. Hanson, H. M. (1959). Effects of discrimination training on stimulus generalization. *Journal of Experimental Psychology, 58*(5), 321–334. 10. Terrace, H. S. (1963). Discrimination learning with and without "errors." *Journal of the Experimental Analysis of Behavior, 6*(1), 1–27. 11. Terrace, H. S. (1963). Errorless transfer of a discrimination across two continua. *Journal of the Experimental Analysis of Behavior, 6*(2), 223–232. 12. Terrace, H. S. (1964). Wavelength generalization after discrimination learning with and without "errors." *Science, 144*(3615), 78–80. 13. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 14. Baddeley, A., & Wilson, B. A. (1994). When implicit learning fails: Amnesia and the problem of error elimination. *Neuropsychologia, 32*(1), 53–68. 15. Bouton, M. E. (2004). Context and behavioral processes in extinction. *Learning & Memory, 11*(5), 485–494. 16. Bootzin, R. R. (1972). Stimulus control treatment for insomnia. *Proceedings of the 80th Annual Convention of the American Psychological Association, 7*, 395–396. 17. Qaseem, A., Kansagara, D., Forciea, M. A., Cooke, M., & Denberg, T. D. (2016). Management of chronic insomnia disorder in adults: A clinical practice guideline from the American College of Physicians. *Annals of Internal Medicine, 165*(2), 125–133. 18. Wood, W., & Rünger, D. (2016). Psychology of habit. *Annual Review of Psychology, 67*, 289–314. ## Related - [The ABC model](https://operantconditioning.com/abc-model/): Antecedent, behavior, consequence — the three-term contingency that stimulus control is the front half of. - [Extinction](https://operantconditioning.com/extinction/): Bursts, spontaneous recovery, resurgence, and the renewal that context brings. - [Dog training](https://operantconditioning.com/dog-training/): Adding the cue, proofing it everywhere, and why a kitchen "sit" vanishes at the park. --- # The ABC Model of Behavior: Antecedent, Behavior, and Consequence Explained > The ABC model of behavior (antecedent, behavior, consequence) explained: the three-term contingency, discriminative stimuli, motivating operations, ABC data. - Source: https://operantconditioning.com/abc-model/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Core concept* Every operant behavior sits between what came before it and what came after it. Learn to read that three-part pattern and you can explain almost any behavior — and change it at three points instead of one. > **Definition** > > The **ABC model of behavior** — antecedent, behavior, consequence — is the **three-term contingency**, the basic unit of analysis in operant conditioning. An *antecedent* sets the occasion for a *behavior*; a *consequence* follows the behavior and changes how likely it is to occur again under similar antecedents. > > B. F. Skinner made the three-term contingency the centerpiece of *Science and Human Behavior* (1953), building on his laboratory work on discriminated operants. Each term names an observable event, and the relations between them can be measured.[1][2] **In brief** - The ABC model, or three-term contingency, is the basic unit of operant analysis: an antecedent sets the occasion, a behavior occurs, a consequence follows. - An antecedent evokes a behavior rather than causing it, and only because of what has followed that behavior there before. - A consequence is defined by its effect on future behavior, not by its appearance, and the model gives you three points to intervene. ## How the ABC model of behavior works The word doing the work is *contingency*. The consequence happens because the behavior happened, and the behavior is emitted in the presence of the antecedent because that is where the consequence has followed before. A rat whose lever-presses produce food only while a light is on comes to press when the light is on and to ignore the lever when it is off.[2] Nothing about the light forces the press; it matters only because of the history attached to it. The antecedent is the front of the loop, but the consequence is what loads it. ## What is an antecedent? An antecedent is anything in place before the behavior: the immediate stimuli, the wider context, and the state of the organism. Behavior analysts distinguish several kinds, because each is changed in a different way. ### The discriminative stimulus (S^D) and S-delta A **discriminative stimulus**, written S^D, is a stimulus in whose presence a behavior has been reinforced. Its counterpart, **S-delta** (S^Δ), is one in whose presence the same behavior has gone unreinforced. The lit "Open" sign is an S^D for pulling the door; the dark sign is an S^Δ. A friend's grin is an S^D for a crude joke; your grandmother's presence is an S^Δ for the same joke. The S^D does not *elicit* behavior the way a puff of air elicits a blink. It *evokes* it — raises its probability — and only because of what has followed the behavior there before. Pavlov's stimulus produces the response; Skinner's merely signals that a response will pay. [Operant vs. classical conditioning ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ### Stimulus control, discrimination, and generalization When a behavior's probability depends on whether a stimulus is present, the behavior is under **stimulus control**. It gets there through **discrimination training** — reinforcement in the presence of the S^D, extinction in the presence of the S^Δ — and its mirror image is **generalization**, responding to stimuli that resemble the S^D: pigeons reinforced for pecking a key lit with 550 nm light pecked less and less as the color moved away from it, tracing a smooth generalization gradient.[3] Discrimination training also reshapes that gradient (peak shift), can be arranged so the learner never makes an error (the basis of prompting and fading), and extends to the context itself, which is why an extinguished behavior can [renew](https://operantconditioning.com/extinction/) when the setting changes back. [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) ### Motivating operations Jack Michael pointed out in 1982 that two different antecedent functions were being lumped together. A stimulus can signal that a reinforcer is *available* (the discriminative function) or change how much the reinforcer is *worth* (the motivational function).[4] He called the second an **establishing operation**; the field later adopted **motivating operation** (MO) as the umbrella term, with an *establishing operation* increasing a reinforcer's value and an *abolishing operation* decreasing it.[5][6] Every MO has two effects at once: it alters the reinforcer's value *and* it alters behavior, evoking responses that have produced that reinforcer or abating them. Eight hours without food makes food a stronger reinforcer and makes you open the fridge; a large lunch does the opposite. A child who slept badly finds escape from demands more valuable, so escape behavior climbs. The vending machine says chips are available; hunger says they are worth having. MOs are the closest thing to "motivation" in the model, and unlike the folk concept they are manipulable: train the dog before dinner, not after. ### Setting events **Setting events** are more distant conditions — illness, poor sleep, an argument at breakfast — that change the strength of the immediate A-B-C relation without being part of it.[7] Most can be re-described as motivating operations, but the term survives in schools because it reminds observers to look further back than the last thirty seconds. ### Prompts A **prompt** is a supplementary antecedent — an instruction, a gesture, a demonstration, physical guidance — added when the natural S^D does not yet evoke the behavior. "What do you say?" prompts "thank you" until the gift itself does the job. Prompts must be faded so control transfers to the natural stimulus; prompting without fading produces a learner who waits to be told.[9] ## What counts as a behavior? ### Operational definitions and the dead-man test A behavior is something the organism *does*: observable, measurable, and defined so that two observers would agree whether it occurred — objective, clear, and complete, in the textbook formula.[9] The quickest screen is Ogden Lindsley's **dead-man test** from 1965: if a dead man can do it, it isn't behavior; if a dead man can't, it is.[8] "Sit still," "don't interrupt," and "stop snacking" all fail; a corpse manages every one. They name the absence of behavior, and an absence cannot be reinforced. Rewrite them as what you want to *see*: "completes the worksheet while seated," "raises a hand and waits," "eats an apple at 3 p.m." ### Response classes and the dimensions of behavior An operant is a **response class**: every response that produces the same consequence, whatever its form. Pressing the lever with the left paw, the right paw, or the nose is one operant; asking, pointing, and grabbing are one operant if they all get the cookie. Behavior is defined by function, not by which muscles moved. It also has measurable dimensions — **frequency**, **duration**, **latency** (how long after the antecedent it begins), and **intensity** — and the problem decides which you track. A tantrum is a duration problem; a morning run is a latency problem, the minutes between alarm and front door. ### Why "be more productive" is not a behavior "Be more productive," "eat healthier," and "be a better listener" are labels for outcomes. They cannot be observed at a moment in time or counted, so nothing can be made contingent on them. The translation is always the same: find a specific, countable act that would, repeated, produce the outcome. "Open the document and write one sentence before checking email" has a start and an end, either happened or didn't, and can be followed by a consequence within seconds. Start with the smallest version, because that is the one that actually gets emitted — the logic of [shaping](https://operantconditioning.com/shaping/) and of every effective [habit protocol](https://operantconditioning.com/habits/). ## What is a consequence? A consequence is a stimulus change that follows a behavior. Its effect on the future frequency of the behavior — not its appearance or the intent of whoever delivered it — determines what it is. Consequences sort into four quadrants plus one non-event: | Consequence | What happens after the behavior | Effect | Example | | --- | --- | --- | --- | | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | A stimulus is added | Increases | Dog sits → treat | | [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) | A stimulus is removed | Increases | You buckle up → chime stops | | [Positive punishment](https://operantconditioning.com/positive-punishment/) | A stimulus is added | Decreases | You touch the pan → burn | | [Negative punishment](https://operantconditioning.com/negative-punishment/) | A stimulus is removed | Decreases | Teen breaks curfew → loses car keys | | [Extinction](https://operantconditioning.com/extinction/) | The usual reinforcer no longer follows | Decreases, after a burst | Elevator button → nothing → you stop pressing | ### Immediacy and contingency Two properties decide how much of the loop closes. **Immediacy**: a reinforcer's effect falls steeply as the delay grows from seconds to minutes, which is why trainers use a click and why a monthly paycheck reinforces almost nothing in particular.[10] **Contingency**: the consequence must depend on the behavior. Pigeons fed on a timer regardless of what they did developed "superstitious" rituals from accidental pairings, and adding free reinforcers that arrive whether or not the animal responds weakens the behavior.[11][12] ### The four-term contingency Because a reinforcer only reinforces when its value is high, a complete account adds the motivating operation: MO → S^D → behavior → consequence, the **four-term contingency**.[9] It is 6 p.m. and you last ate at noon (MO); you walk into the kitchen (S^D); you open the fridge (B); you eat (C). Remove the MO with a late lunch and the same kitchen evokes nothing. Practitioners reach for all four terms because the fourth is often the easiest to change. ## Changing the antecedent is the underrated lever When most people say "behavior change" they mean changing the consequence. That works, but it acts at the hardest moment — after the old habit has already run. The antecedent acts *before* the moment of weakness. Skinner listed "changing the stimulus" among the basic techniques of self-control: remove the cue for an unwanted behavior, or put the cue for a wanted one where you cannot miss it.[1] Charge the phone in the kitchen. Put the guitar on a stand, not in a case. Two research programs outside behavior analysis reach the same conclusion. Peter Gollwitzer's **implementation intentions** are plans of the form "when situation X arises, I will do Y." Forming one links a concrete cue to a concrete response in advance, so the cue evokes the behavior with little deliberation; a meta-analysis of 94 studies found a medium-to-large effect on goal attainment.[13] An implementation intention is a discriminative stimulus installed verbally — antecedent control's cognitive cousin. Wendy Wood's habit research shows the other side. Habits are cued by context and survive on context rather than intention: students who transferred universities kept their exercise, reading, and TV habits only when the new setting resembled the old one, and habitual cinema popcorn-eaters ate stale popcorn as readily as fresh in a cinema but not in a meeting room.[14][15][16] To keep a behavior, keep its cue constant; to lose one, change the cue. > **Rule of thumb** > > If you keep failing at the moment of choice, you are trying to solve an antecedent problem with a consequence. Move the intervention earlier: remove the cue, add a cue, or change the motivating operation. ## Why notifications are weak antecedents Held against the three-term contingency, most attempts to build a habit are all B: a behavior with a checkbox, no antecedent, and a consequence that may or may not do anything. Where there is an antecedent it is usually a push notification, and notifications are weak antecedents for three reasons. Habituation: a stimulus repeated without any differential consequence loses its power to evoke a response, which is why the fortieth buzz of the day is background noise.[17] Many people switch them off. And, least appreciated, a notification is a discriminative stimulus for *picking up the phone* — the behavior that reinforcement has actually followed — not for flossing. The fix is a cue that already occurs reliably in the world: after the coffee, when you sit in the car, when the meeting ends. [Building habits with all three terms ›](https://operantconditioning.com/habits/) ## Three worked ABC examples The model's value is that it hands you three places to intervene. Each example is analyzed once, then attacked at A, at B, and at C. ### A child's tantrum in the supermarket | Aspect | Antecedent | Behavior | Consequence | | --- | --- | --- | --- | | **Analysis** | Checkout aisle lined with candy; parent busy paying; child skipped a nap (MO) | Screaming, lying on the floor | Parent hands over candy. The tantrum is positively reinforced; the giving-in is negatively reinforced by the silence. | | **Intervene** | Shop after the nap; use self-checkout; state the rule before entering ("no candy today, you choose the cereal") | Teach and prompt a replacement: "Can I have a snack at home?" — reinforced every time at first | Candy never follows screaming (expect an [extinction burst](https://operantconditioning.com/extinction/)); attention and a small privilege follow calm behavior | ### An employee who keeps missing deadlines | Aspect | Antecedent | Behavior | Consequence | | --- | --- | --- | --- | | **Analysis** | Tasks assigned verbally, without written dates; competing requests from three managers | Submits work days late | Late work is accepted with a mild remark; on-time work draws no comment. Delay is negatively reinforced (the aversive task is postponed) and never costs anything. | | **Intervene** | Every task in writing with a date and a check-in three days before; one prioritized queue | Define it as "send the draft by Thursday noon," not "be more reliable" | Same-day acknowledgment of on-time delivery; a late submission triggers an immediate re-planning conversation, not a note in the [annual review](https://operantconditioning.com/applications/#workplace) | ### Your own doom-scrolling | Aspect | Antecedent | Behavior | Consequence | | --- | --- | --- | --- | | **Analysis** | Phone on the nightstand; in bed, lights off; tired (MO: fatigue makes escape from effort more valuable) | Unlock, open the feed, scroll | Unpredictable novelty on a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/); escape from boredom and unwelcome thoughts | | **Intervene** | Charge the phone in the kitchen; delete the app; leave a book on the pillow; plan "when I get into bed, I open the book" | Replace with a behavior that serves the same function — one page of a novel, a podcast | Log out after every session so the app opens to a login screen (response cost); mark the page read and let that be the consequence you chose | ## ABC data collection in practice ### The ABC recording chart **ABC recording** means writing down, for each occurrence of a target behavior, what came immediately before and after. Formalized for field research by Bijou, Peterson, and Ault in 1968, it is now the standard first step in schools and clinics.[18] | Date / time | Antecedent | Behavior | Consequence | Possible function | | --- | --- | --- | --- | --- | | Mon 9:05 | Teacher hands out a math worksheet | Jamal tears the worksheet and shouts | Sent to the hallway for 10 minutes | Escape from the task | | Mon 10:30 | Teacher helping another student | Jamal shouts | Teacher comes over and talks with him | Attention | | Tue 9:10 | Math worksheet, harder problems | Jamal pushes the worksheet off the desk | Aide removes the worksheet | Escape from the task | Ten to twenty entries usually reveal a pattern — here, shouting reliably follows academic demands and is reliably followed by their removal. Write what you saw rather than what you inferred, and fill in the "possible function" column after the observations, not before. ### Functional behavior assessment A **functional behavior assessment** (FBA) combines indirect methods (interviews, rating scales), descriptive methods (ABC recording), and sometimes experimental methods to reach a hypothesis about what a problem behavior gets or avoids; U.S. special-education law has required one in certain disciplinary situations since 1997. The output is a behavior intervention plan that changes the antecedents, teaches a replacement behavior, and rearranges the consequences — all three terms, deliberately. ### Functional analysis and the four functions of behavior ABC recording shows correlation; a **functional analysis** tests causation. In the landmark 1982 study, Brian Iwata and colleagues exposed nine children who injured themselves to systematically arranged conditions — adult attention contingent on self-injury, escape from demands contingent on it, being alone, and a play control — and measured self-injury in each.[19] The same behavior was maintained by attention in one child, escape in another, and its own sensory consequences in a third. A later summary of 152 such analyses found escape the most common function, followed by social-positive reinforcement (attention or tangibles) and automatic reinforcement.[20] Practitioners therefore speak of four common functions — **attention**, **escape**, **access to tangibles**, and **automatic** (sensory) reinforcement — and the function dictates the treatment. Escape-maintained shouting is treated by teaching the child to ask for a break and honoring the request, not by sending him to the hallway, which is the very consequence keeping the behavior alive.[21] [How applied behavior analysis uses this ›](https://operantconditioning.com/applications/#aba) ## Common misconceptions about the ABC model - **It is not the ABC model of cognitive-behavioral therapy.** Albert Ellis's ABC, from rational emotive behavior therapy, runs Activating event → Belief → emotional Consequence: the B is a thought and the C is a feeling.[22] In the behavioral ABC, the B is an observable act and the C is an environmental event that follows it. Both date from the 1950s; they describe different things. - **The antecedent does not cause the behavior.** It evokes it, and only because of a consequence history. Change the consequences and the same antecedent stops working. - **ABC data does not tell you the function.** Descriptive records suggest hypotheses; only a functional analysis tests them. - **"Consequence" does not mean punishment.** It includes reinforcement, punishment, and the absence of either. - **A → B is not classical conditioning.** A conditioned stimulus elicits a reflex regardless of what the organism does; a discriminative stimulus sets the occasion for a behavior whose fate is decided by what follows. ## Key takeaways - The three-term contingency is antecedent, behavior, consequence. The consequence happens because the behavior happened, and the behavior is emitted in the presence of the antecedent because that is where the consequence has followed before. - Antecedents come in kinds that are changed in different ways: a discriminative stimulus signals that a reinforcer is available, a motivating operation changes how much it is worth, and a prompt is a supplementary cue that must be faded. Adding the motivating operation gives the four-term contingency. - A behavior is something observable and countable that a dead man could not do, defined by its function rather than its form. "Be more productive" is an outcome label; "write one sentence before checking email" is a behavior. - A consequence is sorted by its effect on future frequency, not by its appearance or the deliverer's intent: reinforcement increases behavior, punishment decreases it, and extinction is the usual reinforcer no longer arriving. Immediacy and contingency decide how much of the loop closes. - Changing the antecedent is the underrated lever, because it acts before the moment of weakness. ABC recording reveals patterns, but only a functional analysis tests what a behavior gets or avoids, and the function dictates the treatment. ### Check yourself **A teacher sends Jamal to the hallway every time he shouts during math. Over the month, shouting increases. Is the hallway a punishment?** No. A consequence is defined by its effect on future behavior, not by how it looks, and shouting went up, so the hallway is reinforcing it: shouting produces escape from the math demand. The fix is to teach Jamal to ask for a break and honor the request, not to keep delivering the very consequence that maintains the behavior. **Your phone buzzes and you pick it up. Does the buzz cause the behavior the way a puff of air causes a blink?** No. The buzz is a discriminative stimulus: it evokes picking up the phone only because that behavior has been reinforced after buzzes before, whereas the air puff elicits a reflex regardless of any history. Change the consequences and the same buzz stops working. **It is 6 p.m., you last ate at noon, you walk into the kitchen, and you open the fridge. Which part is the discriminative stimulus and which is the motivating operation?** The kitchen is the discriminative stimulus, because it signals that food is available; the six hours without food are the motivating operation, because they make food worth more and evoke the behavior that has produced it. Remove the motivating operation with a late lunch and the same kitchen evokes nothing. **A parent's goal for their child is "stop interrupting." Is that a behavior you can reinforce?** No. A dead man can manage not interrupting, so it fails the dead-man test: it names the absence of behavior, and an absence cannot be reinforced. Rewrite it as something you want to see, such as "raises a hand and waits," and reinforce that. **Explain it to a friend.** Explain why the antecedent does not cause the behavior, using something you did today as the example and without using the words reinforcement or punishment. ## Frequently asked questions **What is the ABC model of behavior?** A framework for analyzing any behavior in three parts: the antecedent (what came before and set the occasion), the behavior itself, and the consequence (what followed and made the behavior more or less likely in future). It is the three-term contingency that B. F. Skinner placed at the center of operant conditioning. **What is the three-term contingency?** The relation between a discriminative stimulus, a response, and a reinforcing or punishing consequence — the same thing as the ABC model. "Contingency" means the consequence depends on the behavior, and the behavior comes to depend on the antecedent because of that history. **What is a discriminative stimulus?** A stimulus in whose presence a behavior has been reinforced, so that the behavior becomes more likely when the stimulus is present. A ringing phone is a discriminative stimulus for answering; a lit "Open" sign is one for entering. Its opposite, the S-delta, is a stimulus in whose presence the behavior has not paid off. **What is the difference between a discriminative stimulus and a motivating operation?** A discriminative stimulus signals that a reinforcer is *available*; a motivating operation changes how much the reinforcer is *worth* and evokes the behavior that has produced it. A vending machine is a discriminative stimulus for inserting coins; hunger is a motivating operation that makes the snack worth buying. **What is ABC data collection?** Recording, for each instance of a target behavior, what happened immediately before (antecedent) and immediately after (consequence), usually in a table with the date and time. After ten to twenty entries a pattern typically emerges that suggests what the behavior is getting or avoiding. It is the descriptive core of a functional behavior assessment. **What is the four-term contingency?** The three-term contingency with the motivating operation added in front: MO → discriminative stimulus → behavior → consequence. It recognizes that a reinforcer only reinforces when the organism's state makes it valuable, so a full analysis includes deprivation, satiation, and similar conditions. **Is the ABC model the same as the ABC model in CBT?** No. In Albert Ellis's cognitive model, A is an activating event, B is a belief about it, and C is the emotional consequence of that belief. In the behavioral ABC model, B is an observable behavior and C is an environmental event that follows it and changes its future frequency. Same letters, different variables. **How do I use the ABC model to change my own behavior?** Write out the antecedent, behavior, and consequence for the behavior you want to change, then intervene at all three points: change the cue, define the behavior you want in small and countable terms, and arrange an immediate consequence that actually matters to you. [The full self-management protocol ›](https://operantconditioning.com/habits/) ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Guttman, N., & Kalish, H. I. (1956). Discriminability and stimulus generalization. *Journal of Experimental Psychology, 51*(1), 79–88. 4. Michael, J. (1982). Distinguishing between discriminative and motivational functions of stimuli. *Journal of the Experimental Analysis of Behavior, 37*(1), 149–155. 5. Michael, J. (1993). Establishing operations. *The Behavior Analyst, 16*(2), 191–206. 6. Laraway, S., Snycerski, S., Michael, J., & Poling, A. (2003). Motivating operations and terms to describe them: Some further refinements. *Journal of Applied Behavior Analysis, 36*(3), 407–414. 7. Wahler, R. G., & Fox, J. J. (1981). Setting events in applied behavior analysis: Toward a conceptual and methodological expansion. *Journal of Applied Behavior Analysis, 14*(3), 327–338. 8. Lindsley, O. R. (1991). From technical jargon to plain English for application. *Journal of Applied Behavior Analysis, 24*(3), 449–458. 9. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 10. Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139. 11. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 12. Hammond, L. J. (1980). The effect of contingency upon the appetitive conditioning of free-operant behavior. *Journal of the Experimental Analysis of Behavior, 34*(3), 297–304. 13. Gollwitzer, P. M. (1999). Implementation intentions: Strong effects of simple plans. *American Psychologist, 54*(7), 493–503. See also Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. *Advances in Experimental Social Psychology, 38*, 69–119. 14. Wood, W., Tam, L., & Witt, M. G. (2005). Changing circumstances, disrupting habits. *Journal of Personality and Social Psychology, 88*(6), 918–933. 15. Neal, D. T., Wood, W., Wu, M., & Kurlander, D. (2011). The pull of the past: When do habits persist despite conflict with motives? *Personality and Social Psychology Bulletin, 37*(11), 1428–1437. 16. Wood, W., & Rünger, D. (2016). Psychology of habit. *Annual Review of Psychology, 67*, 289–314. 17. Rankin, C. H., Abrams, T., Barry, R. J., et al. (2009). Habituation revisited: An updated and revised description of the behavioral characteristics of habituation. *Neurobiology of Learning and Memory, 92*(2), 135–138. 18. Bijou, S. W., Peterson, R. F., & Ault, M. H. (1968). A method to integrate descriptive and experimental field studies at the level of data and empirical concepts. *Journal of Applied Behavior Analysis, 1*(2), 175–191. 19. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1982/1994). Toward a functional analysis of self-injury. *Analysis and Intervention in Developmental Disabilities, 2*(1), 3–20. Reprinted in *Journal of Applied Behavior Analysis, 27*(2), 197–209. 20. Iwata, B. A., Pace, G. M., Dorsey, M. F., et al. (1994). The functions of self-injurious behavior: An experimental-epidemiological analysis. *Journal of Applied Behavior Analysis, 27*(2), 215–240. 21. Carr, E. G., & Durand, V. M. (1985). Reducing behavior problems through functional communication training. *Journal of Applied Behavior Analysis, 18*(2), 111–126. 22. Ellis, A. (1962). *Reason and Emotion in Psychotherapy*. Lyle Stuart. ## Related - [Build habits with operant conditioning](https://operantconditioning.com/habits/): The A-B-C protocol applied to yourself, step by step. - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The consequence that does most of the work, and how to deliver it well. - [Operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/): Why an antecedent that evokes is not a stimulus that elicits. --- # Operant vs. Classical Conditioning: The Difference, With Examples > Classical conditioning pairs stimuli to transfer a reflex; operant conditioning changes behavior through consequences. Comparison table, examples, checklist. - Source: https://operantconditioning.com/operant-vs-classical-conditioning/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Comparison* Pavlov's dogs and Skinner's rats learned two different things. Here is the difference between classical and operant conditioning, the science behind each, how they work together in real life, and a checklist for telling them apart in the wild. > **Definitions** > > **Classical conditioning** (also Pavlovian or respondent conditioning) is learning in which a neutral stimulus comes to elicit a reflexive response because it reliably predicts a stimulus that already elicits that response. The organism learns that one stimulus signals another; the key event comes *before* the response.[1] > > **Operant conditioning** (also instrumental conditioning) is learning in which a voluntary behavior becomes more or less frequent because of the consequences that follow it. The organism learns that its own behavior produces an outcome; the key event comes *after* the response.[2] **In brief** - In classical conditioning a stimulus that predicts another comes to elicit a reflex; the key event comes before the response and the organism is passive. - In operant conditioning a voluntary behavior changes because of the consequence that follows it and depends on it. - Real scenarios usually contain both: a classically conditioned emotion and an operant action, so name each part rather than forcing one label. ## Operant vs. classical conditioning: the short answer In classical conditioning, two stimuli are paired and a reflex transfers from one to the other. Pavlov's dogs already salivated at food; after a metronome repeatedly preceded food, they salivated at the metronome. Nothing the dog *did* changed what happened — food came regardless. In operant conditioning, a behavior is followed by a consequence and the behavior changes as a result. Skinner's rats pressed a lever, food arrived *because* they pressed, and pressing increased. The dog learned "metronome means food." The rat learned "pressing gets food." That difference — signal versus consequence, elicited reflex versus emitted action — is the whole distinction, and everything else in this article is a consequence of it.[3] ![Classical versus operant conditioning as sequences](https://operantconditioning.com/assets/diagrams/classical-versus-operant-conditioning-as-sequences.svg) *The two kinds of learning as sequences. In classical conditioning a stimulus comes to elicit a reflex; in operant conditioning a consequence changes how often a voluntary behavior recurs.* ## The difference between classical and operant conditioning, side by side | Aspect | Classical (Pavlovian) conditioning | Operant (instrumental) conditioning | | --- | --- | --- | | **What is learned** | A relation between two stimuli: the CS predicts the US | A relation between a behavior and its consequence, in a context | | **Type of behavior** | Reflexive, involuntary: salivation, blinking, nausea, fear, arousal, heart rate | Voluntary, "emitted": pressing, walking, speaking, studying, scrolling | | **Role of the organism** | Passive — the stimuli arrive whatever it does | Active — the consequence depends on what it does | | **Timing of the key event** | The conditioned stimulus comes *before* the response, usually by less than a few seconds (hours, in taste aversion) | The consequence comes *after* the response, ideally within seconds | | **Is the outcome contingent on behavior?** | No. Food follows the signal regardless of salivation | Yes. Food follows the press only if the press occurs | | **Key figures** | Ivan Pavlov (1890s–1927); John B. Watson (1920); Robert Rescorla (1960s–1980s) | Edward Thorndike (1898); B. F. Skinner (1937–1990) | | **Key terms** | Unconditioned stimulus and response (US, UR); conditioned stimulus and response (CS, CR); neutral stimulus | Reinforcement, punishment (positive and negative); discriminative stimulus (S^D); schedules; shaping | | **Typical laboratory preparation** | Dog in a harness with a salivary fistula; rabbit eyeblink conditioning; conditioned suppression in rats | Thorndike's puzzle box; Skinner's operant chamber with a lever or key and a cumulative recorder | | **Acquisition** | CR strengthens over repeated CS–US pairings that carry predictive information | Response rate rises as responses are reinforced; complex behavior is built by shaping | | **Extinction** | Present the CS without the US; the CR declines | Stop delivering the reinforcer after the response; the response declines, often after a burst | | **Spontaneous recovery** | Yes — the CR returns partially after a rest | Yes — the response returns partially after a rest | | **Generalization and discrimination** | CR spreads to similar stimuli; discrimination training narrows it | Behavior spreads to similar contexts; discrimination training brings it under stimulus control | | **Everyday examples** | Flinching at the dentist's drill; nausea in the chemo clinic waiting room; excitement at the sound of a treat bag | Studying for grades; buckling up to stop the chime; a dog sitting for a treat; checking a phone for likes | ## How to tell which is which: a decision checklist 1. **Name the response.** Is it something the organism does with its skeletal muscles — walks, presses, says, buys — or something that happens to it — salivates, flinches, feels sick, feels afraid, heart races? Voluntary points to operant; reflexive or emotional points to classical. 2. **Find the key event and check its timing.** Does the important stimulus come *before* the response as a signal, or *after* it as a result? Before is classical; after is operant. 3. **Test contingency.** Would the outcome have happened anyway? If food comes whether or not the dog salivates, it is classical. If food comes only if the dog sits, it is operant. 4. **State what was learned in one sentence.** "X predicts Y" is classical. "Doing X produces Y" is operant. 5. **Look for both.** Most real scenarios contain a classically conditioned emotion and an operant action. Name each part separately rather than forcing one label on the whole scene. ## Worked examples: classical or operant? | Scenario | Answer | Why | | --- | --- | --- | | You tense up when you hear a dentist's drill. | **Classical** | The drill sound (CS) preceded pain (US) in the past; tension (CR) is elicited, not chosen, and comes before anything you do. | | A teenager cleans his room and gets the car keys for the evening. | **Operant** (positive reinforcement) | A voluntary behavior is followed by an added consequence that depends on it; cleaning increases. | | A cat comes running when it hears the can opener. | **Both** | The sound is a CS that elicits excitement and salivation (classical). Running to the kitchen is an operant reinforced by food (operant). | | A chemotherapy patient feels nauseated in the clinic waiting room. | **Classical** | Clinic cues (CS) preceded the drug (US) that caused nausea (UR); now the cues elicit anticipatory nausea (CR). Nothing the patient does changes the outcome. | | A student stops raising her hand after the teacher never calls on her. | **Operant** (extinction) | A behavior that was once reinforced by being called on no longer is, and it declines. | | Your smoke alarm shrieks every time you make toast. Now you flinch when you push the lever down — and you open a window before you start. | **Both** | Flinching at the lever is a CR to a CS (classical). Opening the window is an operant, negatively reinforced by preventing the alarm (avoidance). | | You got a stomach bug hours after eating clams and now can't stand the smell of them. | **Classical** (taste aversion) | One pairing, a long delay, and a response — disgust — you cannot decide not to have. Biological preparedness at work. | | A puppy wags and drools when it sees the treat pouch, then sits when asked and gets a treat. | **Both** | The wagging and drooling are CRs to the pouch (classical). The sit is an operant reinforced by the treat (operant). The pouch is also becoming an S^D for sitting. | More practice: the [examples page](https://operantconditioning.com/examples/) has more than fifty operant scenarios sorted by quadrant, and the [quiz](https://operantconditioning.com/quiz/) mixes classical and operant items with instant explanations. ## Where people mix them up - **Deciding by whether the stimulus is pleasant.** Both kinds of conditioning use pleasant and unpleasant stimuli. Decide by timing and contingency, not by valence. - **Calling the bell the unconditioned stimulus.** The US is the stimulus that works *without* learning — the food. The bell (or metronome) is neutral, then conditioned. - **Assuming the CR is a copy of the UR.** Often it is weaker; sometimes it is the opposite (Siegel's compensatory responses), and the CR to a shock-predicting tone is freezing, not the jump the shock itself produces. - **Filing negative reinforcement under classical conditioning** because it "involves something unpleasant." Negative reinforcement is operant: a behavior removes an aversive stimulus and increases. Classical conditioning does not have reinforcement or punishment in Skinner's sense at all. - **Confusing the CS with the discriminative stimulus.** A CS elicits a reflex regardless of behavior; an S^D signals that a behavior will be reinforced. - **Treating classical conditioning as mere pairing.** Since Rescorla, the CS must *predict* the US. Pairings without predictive value produce little learning. - **Treating extinction as forgetting.** In both kinds of conditioning, extinction is new learning that inhibits the old, which is why spontaneous recovery and renewal occur. - **Forcing one label on a scene that has both.** "The dog gets excited at the leash and then sits for a treat" contains a classical part and an operant part. Full credit requires naming each. For the people, dates, and disputes behind both traditions, see the [history of operant conditioning](https://operantconditioning.com/history/). ## Going further The sections below go past the short answer: how classical conditioning really works, where the two kinds of learning meet, and the cases that blur the line. ## Classical conditioning, explained properly ### Pavlov's discovery Ivan Pavlov was a physiologist studying digestion — work that earned him the 1904 Nobel Prize — when he noticed that his dogs began salivating before food arrived: at the sight of the food dish, at the footsteps of the attendant. He called these "psychic secretions" and spent the rest of his career studying them with the rigor of a physiologist, using a surgically implanted tube to measure drops of saliva. In many of his experiments the signal was a metronome, a buzzer, a light, or a touch rather than the bell of legend. The results were published in English in 1927 as *Conditioned Reflexes*.[1] (Pavlov's own word was "conditional" — the reflex was conditional on the pairing — and "conditioned" is an early translation that stuck.) ### The four terms - **Unconditioned stimulus (US or UCS):** a stimulus that elicits a response without any learning. Food in the mouth. - **Unconditioned response (UR or UCR):** the unlearned response to it. Salivation to food. - **Conditioned stimulus (CS):** a previously neutral stimulus that, after predicting the US, elicits a response on its own. The metronome. - **Conditioned response (CR):** the learned response to the CS. Salivation to the metronome — usually similar to the UR, but often weaker and sometimes different in form. ### What Pavlov found **Acquisition** is gradual: the CR grows over pairings. **Extinction** follows when the CS is presented repeatedly without the US — but Pavlov noticed that an extinguished response reappears after a rest (**spontaneous recovery**), which told him extinction was new learning laid over the old, not erasure. Modern work confirms this: extinguished responses also return when the context changes (renewal) or when the US is encountered again (reinstatement).[4] A CR trained to one tone appears, more weakly, to similar tones (**generalization**); pairing one tone with food and another with nothing narrows the response to the first (**discrimination**). When Pavlov's laboratory made a circle-versus-ellipse discrimination progressively harder, a previously calm dog became agitated and uncooperative — the first "experimental neurosis."[1] Finally, an established CS can itself condition a new stimulus (**higher-order conditioning**), which is how a word like "dinner" ends up doing what the metronome did. ### Little Albert: the famous study and its problems In 1920 John B. Watson and Rosalie Rayner reported conditioning fear in an infant, "Albert B.," about eleven months old. Albert initially reached for a white rat without fear. Watson then struck a steel bar with a hammer behind Albert's head whenever the rat appeared. After a handful of pairings, Albert cried and turned away at the sight of the rat alone, and his distress generalized to a rabbit, a dog, a fur coat, and a Santa Claus mask.[5] The study is in every textbook, and it should be read with its problems attached. Ethically, it deliberately induced fear in an infant who could not consent, and Watson and Rayner made no attempt to remove the fear before Albert left the hospital, though they knew in advance when he would leave. Methodologically, it was a single case with no control condition; fear was rated subjectively; some of Albert's reactions were mild or inconsistent; and the responses were "freshened up" with additional pairings between tests. A 1979 review found that textbooks had for decades embellished the results, reporting deconditioning that never happened and generalization that was never tested.[6] Albert's real identity has been the subject of competing claims by historians. What survives is the modest, real finding: a fear response can be conditioned to a neutral stimulus in a human infant, and it generalizes. ### It's not what you think it is: contingency, not pairing The textbook story — "pair two stimuli enough times and the reflex transfers" — turns out to be wrong in an important way. In 1968 Robert Rescorla gave rats a tone followed by shock, but for some groups he added shocks during the silent periods too, so that the tone no longer *predicted* any change in the likelihood of shock. Those rats received exactly as many tone–shock pairings as the others, and they learned almost nothing.[7] Leon Kamin's blocking experiment made the same point from another direction: if a light already predicts shock, adding a tone alongside it teaches the animal nothing about the tone, because the shock is no longer surprising.[8] Rescorla and Allan Wagner turned these findings into a mathematical model in which learning is driven by prediction error — how much the outcome differs from what was expected.[9] Rescorla summarized the modern view in a 1988 paper whose title says it all: "Pavlovian conditioning: It's not what you think it is." Classical conditioning is not the mechanical transfer of a reflex by contiguity; it is the organism learning the predictive structure of its environment — which events signal which others — and adjusting a whole set of responses accordingly.[10] ### Taste aversion and biological preparedness The other crack in the simple story came from John Garcia. In 1966 Garcia and Robert Koelling let rats drink "bright, noisy" water — sweetened, and accompanied by a light and a click with every lick. Rats that were then made ill (with X-rays or lithium chloride) later avoided the sweet taste but drank the noisy, bright water happily. Rats that were shocked instead avoided the light and click but not the taste.[11] The animals were not equally ready to associate any stimulus with any outcome: tastes go with illness, sights and sounds go with pain. Taste aversion also broke the timing rule — it formed after a single trial with delays of an hour or more between taste and illness. Martin Seligman called this **preparedness**: evolution has made some associations easy to learn and others nearly impossible.[12] ## Operant conditioning, briefly Edward Thorndike's cats, escaping from puzzle boxes, showed that responses followed by satisfying outcomes are "stamped in" — the law of effect.[13] [B. F. Skinner](https://operantconditioning.com/bf-skinner/) made this a laboratory science: an animal in a chamber, a lever or key, a consequence delivered by the apparatus, and a cumulative record of responses over time. In a 1937 paper he formally separated the two kinds of learning, calling Pavlov's "Type S" (stimulus-elicited, respondent) and his own "Type R" (response-emitted, operant).[3] Every consequence falls into one of four quadrants — [positive reinforcement](https://operantconditioning.com/positive-reinforcement/), [negative reinforcement](https://operantconditioning.com/negative-reinforcement/), positive punishment, negative punishment — and how often it arrives is governed by [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/). Behavior that no longer pays off undergoes [extinction](https://operantconditioning.com/extinction/), complex behavior is built by [shaping](https://operantconditioning.com/shaping/), and the [antecedent](https://operantconditioning.com/abc-model/) that signals when a behavior will be reinforced brings it under stimulus control. [The complete guide to operant conditioning ›](https://operantconditioning.com/) > **The CS and the S^D are not the same thing** > > Both are "cues," which is why students confuse them. A **conditioned stimulus** *elicits* a response: the metronome makes the dog salivate whether or not it does anything. A **[discriminative stimulus](https://operantconditioning.com/glossary/#discriminative-stimulus)** *sets the occasion* for a response: the green light tells the rat that pressing will now be reinforced, but the food still depends on the press. The test is contingency. If the outcome arrives no matter what the organism does, the cue is a CS. If the outcome depends on the behavior, the cue is an S^D. ## How classical and operant conditioning work together Outside the laboratory the two processes almost never run alone, and the most useful accounts of real behavior combine them. ### Two-factor theory of avoidance Why does a rat keep jumping a barrier to avoid a shock that never comes any more? O. H. Mowrer's answer was two processes in sequence: first, a warning signal is classically conditioned to elicit fear; second, the avoidance response is operantly reinforced by escape from that fear.[14] The same logic explains why **phobias** persist: the fear may be acquired classically (a dog bite, or even a frightening story), but it is maintained operantly, because avoiding dogs is negatively reinforced by relief and the fear never gets a chance to extinguish. Exposure therapy is the deliberate blocking of the operant half so the classical half can extinguish. ### Conditioned reinforcers are classically conditioned A clicker is silent to a dog that has never heard one paired with food. Pair click and treat a few dozen times and the click acquires two properties at once. It is a **CS**: it elicits the anticipatory excitement that food elicits. And it is a **conditioned reinforcer**: delivered right after a behavior, it strengthens that behavior, bridging the seconds until the treat arrives.[15] Money, praise, grades, and the green checkmark work the same way — classically conditioned value, operantly deployed. [How clicker training uses both ›](https://operantconditioning.com/dog-training/) ### Conditioned suppression One of the cleanest laboratory demonstrations of the two processes interacting is also one of the oldest. In 1941 Estes and Skinner had rats pressing a lever for food, then sounded a tone that ended in shock. As the tone acquired fear (classical), the rats' ongoing lever pressing slowed or stopped during it (operant behavior suppressed).[16] The "conditioned emotional response" became a standard way to measure Pavlovian fear — by its effect on operant behavior. ### Addiction cues Drug taking is operant: the drug's effects reinforce the behavior that produces them. But the paraphernalia, the place, and the people become classically conditioned cues. Shepard Siegel showed that rats given morphine in a familiar environment developed conditioned *compensatory* responses — the body preparing to counteract the drug — so that tolerance was partly a learned response to the setting, and the same dose in a new place was more dangerous.[17] Craving triggered by cues, and relapse in old environments, is classical conditioning driving the person back toward the operant. ### Pavlovian-instrumental transfer A classically conditioned cue can also change the *vigor* of operant behavior without ever having been part of it. Train a rat to press a lever for food; separately, in another session, pair a tone with free food. Then play the tone while the rat is pressing: pressing speeds up, even though the tone was never a signal for pressing and pressing has never paid off during it. William Estes reported the effect in 1948, and it is now called **Pavlovian-instrumental transfer** (PIT).[18][19] It is the laboratory version of a familiar experience: the smell of the bakery does not teach you to walk in, but it makes you walk in faster. In addiction research PIT is one of the main models of how drug cues energize drug seeking. ## Behavior without reinforcement? Autoshaping and contrafreeloading Some observations look, at first, like operant behavior that no reinforcement produced, and they mark the edge of the law of effect. ### Autoshaping and sign-tracking In 1968 Brown and Jenkins lit a pigeon's response key for a few seconds and then delivered grain — whether or not the bird did anything. After a few dozen pairings the pigeons began pecking the lit key. Nobody had shaped the peck; the bird had "auto-shaped" it.[20] The following year Williams and Williams arranged that pecking the key *cancelled* the grain, so that the only way to be fed was not to peck. The pigeons kept pecking — less, but persistently — and lost food for it.[21] This **omission** result rules out reinforcement as the cause: the pecks were being punished by food loss and continued anyway. Jenkins and Moore then showed that the form of the peck matched the reinforcer — birds autoshaped with grain pecked the key as if eating it, birds autoshaped with water pecked as if drinking.[22] The behavior is directed at the signal as if it were the reward, which is why Hearst and Jenkins named it **sign-tracking**.[23] The accepted interpretation is that autoshaping is classical conditioning: the key light is a CS, the grain a US, and approaching and pecking a food-predicting stimulus is the conditioned response, as inevitable in a pigeon as salivation in a dog. The autoshaping procedure has, in fact, become one of the standard ways to *measure* Pavlovian conditioning. It also uncovered stable individual differences: some rats become sign-trackers who approach the cue, others goal-trackers who go straight to the food cup, and dopamine appears to be required for the first kind of learning but not the second.[24] The lesson for the operant–classical distinction is not that the law of effect is wrong but that a response can look operant and be Pavlovian, and the experimenter's job is to find out which contingency is actually controlling it. ### Contrafreeloading Give a rat free food in a dish and a lever that delivers the same food, and it will press the lever for a substantial share of its meals. Jensen reported the effect in 1963, and it has been found in most species tested, with cats the notable exception.[25][26] Working for food that is freely available is not what a naive reading of reinforcement predicts. It is compatible with a fuller one: the opportunity to explore, manipulate, and gather information is itself reinforcing, and a well-designed environment for a captive animal — or a person — is one that lets it work. ## Key takeaways - Classical conditioning pairs two stimuli so that a reflex transfers from one to the other; operant conditioning follows a behavior with a consequence that changes its future rate. Signal versus consequence, elicited reflex versus emitted action, is the whole distinction. - To tell them apart, name the response (reflexive or voluntary), check whether the key event comes before or after it, and test contingency. If the outcome would have happened anyway, it is classical; if it depends on the behavior, it is operant. - A conditioned stimulus elicits a response regardless of behavior; a discriminative stimulus sets the occasion for a behavior whose outcome still depends on it. Negative reinforcement is operant, not classical, because a behavior removes the aversive stimulus and increases. - Classical conditioning is prediction, not mere pairing: a stimulus that does not predict the outcome teaches almost nothing. Taste aversion shows that some associations are biologically prepared, forming in one trial across delays of an hour or more. - The two processes usually work together. A phobia is acquired classically and maintained operantly by avoidance, a clicker is a conditioned stimulus and a conditioned reinforcer at once, and autoshaping shows that a response can look operant and be Pavlovian. ### Check yourself **You buckle your seat belt to stop the chime. Because the chime is unpleasant, a classmate files this under classical conditioning. Is that right?** No. This is operant conditioning, specifically negative reinforcement: a voluntary behavior removes an aversive stimulus and becomes more frequent. Whether a stimulus is pleasant or unpleasant does not decide the type of learning; timing and contingency do, and classical conditioning has no reinforcement in Skinner's sense at all. **A pigeon pecks a lit key that is followed by grain no matter what the bird does. Since the peck looks like a lever press, is it operant behavior?** No. This is autoshaping, and the accepted interpretation is classical conditioning: the key light is a conditioned stimulus for grain, and pecking a food-predicting signal is the conditioned response. Pigeons keep pecking even when a peck cancels the grain, which rules out reinforcement as the cause. **A cat comes running when it hears the can opener. Classical or operant?** Both. The sound is a conditioned stimulus that elicits excitement and salivation, which is classical; running to the kitchen is an operant reinforced by food. Full credit means naming each part rather than forcing one label on the scene. **In Pavlov's experiment, which is the unconditioned stimulus: the metronome or the food?** The food. The unconditioned stimulus is the one that elicits the response without any learning, and food in the mouth produces salivation from the start. The metronome begins as a neutral stimulus and becomes a conditioned stimulus only after it reliably predicts the food. **Explain it to a friend.** Explain the difference between the two kinds of conditioning using a single everyday example that contains both, and say which part is which without using the words voluntary or involuntary. ## Frequently asked questions **What is the main difference between classical and operant conditioning?** Classical conditioning pairs two stimuli so that an involuntary response (salivation, fear, nausea) transfers from one to the other; the signal comes before the response and the outcome does not depend on what the organism does. Operant conditioning changes a voluntary behavior through the consequence that follows it; the consequence comes after the behavior and depends on it. **Is Pavlov's dog classical or operant conditioning?** Classical. The metronome or bell predicted food, and the dog's salivation transferred to the signal. The dog did not have to do anything for the food to arrive. If the dog had been required to press a lever to get food, that would be operant conditioning. **Can classical and operant conditioning happen at the same time?** Yes, and in real life they usually do. A clicker is a classically conditioned stimulus and an operant reinforcer at once. A phobia is typically acquired classically and maintained operantly by avoidance. Mowrer's two-factor theory of avoidance is built on exactly this combination. **Is negative reinforcement classical or operant conditioning?** Operant. Negative reinforcement means a behavior removes or prevents an aversive stimulus and becomes more frequent — buckling a seat belt to stop the chime. Classical conditioning does not involve reinforcement or punishment of behavior at all; it involves one stimulus coming to predict another. **What is an example of classical conditioning in everyday life?** Feeling your mouth water when you smell bread baking; flinching at the sound of a dentist's drill; feeling anxious when you hear the ringtone assigned to your boss; a dog getting excited at the jingle of the leash; feeling queasy at the sight of a food that once made you ill. **Which is stronger, classical or operant conditioning?** Neither — they do different jobs. Classical conditioning is the fastest way to attach an emotional or physiological response to a cue, sometimes in one trial. Operant conditioning is the only way to build a new voluntary skill or change how often someone does something. Most effective behavior change uses both: make the cue mean something, and make the behavior pay off. **What is the difference between a conditioned stimulus and a discriminative stimulus?** A conditioned stimulus elicits a reflexive response by itself, because it predicts an unconditioned stimulus; the dog salivates at the tone whatever it does. A discriminative stimulus signals that a voluntary behavior will now be reinforced; the light tells the rat that pressing will produce food, but it still has to press. **Who discovered classical and operant conditioning?** Ivan Pavlov described classical conditioning in dogs beginning in the 1890s and published *Conditioned Reflexes* in 1927. Edward Thorndike described the law of effect in 1898; B. F. Skinner named operant conditioning in 1937 and developed the experimental science of it from *The Behavior of Organisms* (1938) onward. ## References 1. Pavlov, I. P. (1927). *Conditioned Reflexes: An Investigation of the Physiological Activity of the Cerebral Cortex* (G. V. Anrep, Trans.). Oxford University Press. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279. 4. Bouton, M. E. (2004). Context and behavioral processes in extinction. *Learning & Memory, 11*(5), 485–494. 5. Watson, J. B., & Rayner, R. (1920). Conditioned emotional reactions. *Journal of Experimental Psychology, 3*(1), 1–14. 6. Harris, B. (1979). Whatever happened to Little Albert? *American Psychologist, 34*(2), 151–160. 7. Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. *Journal of Comparative and Physiological Psychology, 66*(1), 1–5. 8. Kamin, L. J. (1969). Predictability, surprise, attention, and conditioning. In B. A. Campbell & R. M. Church (Eds.), *Punishment and Aversive Behavior* (pp. 279–296). Appleton-Century-Crofts. 9. Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), *Classical Conditioning II: Current Research and Theory* (pp. 64–99). Appleton-Century-Crofts. 10. Rescorla, R. A. (1988). Pavlovian conditioning: It's not what you think it is. *American Psychologist, 43*(3), 151–160. 11. Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. *Psychonomic Science, 4*(1), 123–124. 12. Seligman, M. E. P. (1970). On the generality of the laws of learning. *Psychological Review, 77*(5), 406–418. 13. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 14. Mowrer, O. H. (1947). On the dual nature of learning — a re-interpretation of "conditioning" and "problem-solving." *Harvard Educational Review, 17*, 102–148. 15. Williams, B. A. (1994). Conditioned reinforcement: Experimental and theoretical issues. *The Behavior Analyst, 17*(2), 261–285. 16. Estes, W. K., & Skinner, B. F. (1941). Some quantitative properties of anxiety. *Journal of Experimental Psychology, 29*(5), 390–400. 17. Siegel, S. (1975). Evidence from rats that morphine tolerance is a learned response. *Journal of Comparative and Physiological Psychology, 89*(5), 498–506. 18. Estes, W. K. (1948). Discriminative conditioning II: Effects of a Pavlovian conditioned stimulus upon a subsequently established operant response. *Journal of Experimental Psychology, 38*(2), 173–177. 19. Lovibond, P. F. (1983). Facilitation of instrumental behavior by a Pavlovian appetitive conditioned stimulus. *Journal of Experimental Psychology: Animal Behavior Processes, 9*(3), 225–247. 20. Brown, P. L., & Jenkins, H. M. (1968). Auto-shaping of the pigeon's key-peck. *Journal of the Experimental Analysis of Behavior, 11*(1), 1–8. 21. Williams, D. R., & Williams, H. (1969). Auto-maintenance in the pigeon: Sustained pecking despite contingent non-reinforcement. *Journal of the Experimental Analysis of Behavior, 12*(4), 511–520. 22. Jenkins, H. M., & Moore, B. R. (1973). The form of the auto-shaped response with food or water reinforcers. *Journal of the Experimental Analysis of Behavior, 20*(2), 163–181. 23. Hearst, E., & Jenkins, H. M. (1974). *Sign-Tracking: The Stimulus-Reinforcer Relation and Directed Action*. Psychonomic Society. 24. Flagel, S. B., Clark, J. J., Robinson, T. E., Mayo, L., Czuj, A., Willuhn, I., Akers, C. A., Clinton, S. M., Phillips, P. E. M., & Akil, H. (2011). A selective role for dopamine in stimulus–reward learning. *Nature, 469*(7328), 53–57. 25. Jensen, G. D. (1963). Preference for bar pressing over "freeloading" as a function of number of rewarded presses. *Journal of Experimental Psychology, 65*(5), 451–454. 26. Inglis, I. R., Forkman, B., & Lazarus, J. (1997). Free food or earned food? A review and fuzzy model of contrafreeloading. *Animal Behaviour, 53*(6), 1171–1191. ## Related - [Operant conditioning: the complete guide](https://operantconditioning.com/): The four quadrants, schedules, extinction, shaping, and how to use it on yourself. - [History of operant conditioning](https://operantconditioning.com/history/): Thorndike, Watson, Skinner, the cognitive revolution, and reinforcement learning. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): Escape, avoidance, and the two-factor theory in depth. --- # Avoidance Learning: Escape, Avoidance, and Why It Persists > Escape and avoidance learning: signaled and Sidman avoidance, the avoidance paradox, two-factor vs. one-factor theory, and learned helplessness. - Source: https://operantconditioning.com/avoidance-learning/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Negative reinforcement · Escape and avoidance* A dog that learned to jump a barrier after a handful of shocks kept jumping for hundreds of trials without ever being shocked again. How can the absence of something reinforce behavior? The answer took thirty years and explains phobias, procrastination, and safety rituals. > **Definition** > > **Escape learning** is operant learning in which a behavior *terminates* an aversive stimulus that is already present. **Avoidance learning** is operant learning in which a behavior *prevents or postpones* an aversive stimulus that has not yet occurred. Both are forms of [negative reinforcement](https://operantconditioning.com/negative-reinforcement/): the behavior increases because something is removed or kept away.[1] > > Shielding your eyes from the sun is escape; putting on sunglasses before you go outside is avoidance. Almost every avoidance response begins its life as an escape response that got earlier. **In brief** - Escape ends an aversive stimulus that is already present; avoidance prevents or postpones one that has not yet occurred, and both are [negative reinforcement](https://operantconditioning.com/negative-reinforcement/). - In successful avoidance the consequence is that nothing happens, yet avoidance is among the most persistent behavior known: the avoidance paradox. - Avoidance persists because a successful avoider never tests the contingency, which is why blocking the response works when simple extinction does not. ## Escape: the easy case Escape is learned quickly because the organism feels the contrast directly: the shock is on, the lever is pressed, the shock is off. Thorndike's cats escaping from puzzle boxes were the first laboratory case, and the relief of a headache after aspirin is the everyday one.[2] Nothing about escape is paradoxical. The consequence is present, immediate, and obviously reinforcing. ## Discriminated (signaled) avoidance The classic procedure adds a warning signal. A light or tone comes on; a few seconds later a shock begins; a response made during the signal cancels the shock, and a response made after the shock begins ends it. Early trials are escape trials. As learning progresses the response moves earlier, into the signal, and the animal stops receiving shocks altogether: the trials have become avoidance trials. Richard Solomon and Lyman Wynne ran the definitive version in 1953 with dogs in a two-compartment shuttle box. After a few intense shocks the dogs began jumping the barrier during the signal, their latencies shortened to a second or two, and they continued to jump on trial after trial — hundreds of them — without receiving another shock. Several dogs never received more than a handful of shocks in the whole experiment.[3] When the shock generator was later switched off entirely, the jumping persisted; ordinary extinction barely touched it. What worked best was a glass barrier that physically prevented the jump, so that the dogs had to stay in the compartment and discover that no shock came — most effectively combined with a shock for jumping — and even then several dogs never fully stopped, which led Solomon and Wynne to speak of the "partial irreversibility" of the response.[4][12] ## Free-operant (Sidman) avoidance In 1953 Murray Sidman removed the warning signal altogether. Rats received brief shocks on a timer — say, every 5 seconds — unless they pressed a lever, in which case the next shock was postponed by a fixed period. Two intervals define the procedure: the **shock–shock (S–S) interval**, the time between shocks if the animal does nothing, and the **response–shock (R–S) interval**, the shock-free period each response buys. Every press restarts the R–S clock.[5][6] With nothing to warn them, rats still learned to press at a steady rate that kept shocks rare, and response rate depended in an orderly way on both intervals. Sidman avoidance mattered because it seemed to strip the procedure to its bones: no signal, so no fear of a signal, and yet learning. What exactly was being reinforced? ## The avoidance paradox Reinforcement, by definition, is a consequence that follows a response. In successful avoidance the consequence is that nothing happens. A rat that presses every few seconds never gets shocked, and an event that never occurs cannot follow a response. Worse, the better the animal performs, the less contact it has with the contingency: a perfect avoider could not tell whether the shock generator was still connected. Yet avoidance is among the most persistent behavior known. The theories below are attempts to say what the real consequence is. ## Two-factor theory O. H. Mowrer's answer, first sketched in 1939 and developed into "two-factor" theory from 1947, is that two learning processes run in sequence.[7][8] 1. **Classical conditioning of fear.** The warning signal is paired with shock and comes to elicit a conditioned emotional response — fear — with the racing heart and freezing that go with it. Neal Miller showed in 1948 that this conditioned fear works as an acquired drive: rats would learn a brand-new response, turning a wheel, simply to get out of a compartment where they had once been shocked.[9] 2. **Operant reinforcement by fear reduction.** The avoidance response terminates the signal, and with it the fear. The response is therefore reinforced — not by the absence of shock, which is nothing, but by escape from an aversive internal state, which is something. On this account the animal is not "avoiding" in the sense of anticipating the shock; it is escaping the fear the signal produces. Leon Kamin confirmed in 1956 that both parts contribute: rats learned best when a response both turned off the signal and cancelled the shock, and each element alone supported weaker learning.[10] The theory also explains the persistence: because the animal responds early, it never stays with the signal long enough to discover that shock no longer follows, so the fear never extinguishes and neither does the response. ### Problems with two-factor theory - **Well-trained avoiders show little fear.** Solomon and Wynne's dogs, after the first few trials, looked calm and businesslike. Kamin, Brimer, and Black measured fear of the signal directly, by how much it suppressed food-reinforced lever pressing, and found that fear *declined* as avoidance became well learned — while the avoidance response stayed strong.[11] If fear reduction is the reinforcer, the reinforcer seems to fade while the behavior does not. - **Sidman avoidance has no signal.** Two-factor theorists replied that the passage of time since the last response, and the animal's own proprioceptive feedback, serve as internal signals — a reasonable move, but one that makes the theory hard to test. - **Extinction is far too slow.** With the shock off, the signal should lose its fear and the response should fade. In practice avoidance can outlast measurable fear by hundreds of trials.[4][12] ## One-factor theory: the missed shock is the reinforcer Richard Herrnstein and Philip Hineline argued in 1966 that avoidance needs only one process if "consequence" is understood at the level of rates rather than single events. They built a procedure in which a response did not postpone any particular shock; it merely switched the rat from a schedule delivering shocks at a high rate to one delivering them at a lower rate, and shocks still arrived, unpredictably, after responses. Rats learned to respond anyway. The reinforcer, they concluded, is a **reduction in the overall frequency of aversive stimulation** — something organisms are demonstrably sensitive to.[13][14] James Dinsmoor pushed the point further: stimuli that reliably accompany the avoidance response — the feel of the lever, the click of a relay, or an explicit "safety signal" — are paired with shock-free time and become conditioned reinforcers in their own right. Add a brief tone after each avoidance response and rats learn faster and respond more.[15] On this view the missed shock is not nothing at all; it is a period of safety, and safety has cues. ## Expectancy and modern hybrid accounts Martin Seligman and James Johnston proposed in 1973 that avoiders learn two expectancies — "if I respond, no shock; if I don't, shock" — and that the response persists as long as the first expectancy is confirmed, which in successful avoidance it always is.[16] Cognitive language, but the same structural insight as one-factor theory: a well-trained avoider never tests the alternative. Current accounts generally combine Pavlovian fear (which clearly drives acquisition), operant reinforcement by safety and by aversive-rate reduction (which maintains the behavior), and expectancies about the contingency (which explain the immunity to extinction).[17] ## Not every response can be an avoidance response Rats learn to run or jump to avoid shock in a handful of trials and learn to press a lever to avoid it slowly and unreliably, sometimes never. Robert Bolles explained the discrepancy in 1970 with **species-specific defense reactions**: every species has a small innate repertoire for danger — freezing, fleeing, fighting — and an avoidance response is learned readily only if it is one of these or compatible with them. A rat's natural response to a threatening chamber is to freeze or flee, not to manipulate an object; requiring a lever press pits the contingency against biology.[18] It is the aversive counterpart of the Brelands' "misbehavior of organisms," and a reminder that reinforcement selects from what an animal is built to do. ## Learned helplessness: when escape is never learned In 1967 Seligman, Steven Maier, and Bruce Overmier gave dogs a series of inescapable shocks in a harness, then placed them in a shuttle box where a jump would end the shock. Dogs that had first received *escapable* shock — or no shock — learned to jump quickly. Most of the dogs that had received inescapable shock did not: they whimpered, lay down, and took the shock, even after occasionally jumping and ending it.[19][20] The decisive comparison was the "triadic design": animals in the two shocked groups received exactly the same shocks, one group's controllable and the other's not. Only uncontrollability produced the deficit. Seligman called it **learned helplessness** and proposed that the animals had learned that responding and outcomes were independent. Fifty years on, Maier and Seligman reversed the interpretation on the strength of the neuroscience. Passivity in the face of prolonged aversive stimulation turned out to be the *default*, driven by serotonergic neurons in the dorsal raphe nucleus; what animals with escapable shock actually learn is that they have control, and that learning, mediated by the ventromedial prefrontal cortex, inhibits the default and allows escape and avoidance to be learned later.[21] Helplessness is not learned. Control is. ## Why avoidance matters outside the laboratory Avoidance is the operant engine inside most anxiety problems. Fear may be acquired classically — a panic attack in a supermarket, a dog bite, a humiliating presentation — but it is *maintained* operantly, because avoiding the supermarket, the dog, or the meeting is negatively reinforced by relief, and the avoidance prevents the fear from ever being tested. Paul Salkovskis showed that even subtle "safety behaviors" — gripping the trolley, sitting near the exit, rehearsing every sentence — work the same way: they feel protective, and they keep the catastrophe uncontradicted.[22] Compulsions in obsessive–compulsive disorder are avoidance responses to an internal threat, and the treatment that works, exposure with response prevention, is Solomon and Wynne's cure applied to people: block the response so that the feared outcome can fail to arrive.[23][4] | Everyday avoidance | Aversive event kept away | Why it persists | | --- | --- | --- | | Procrastinating on a hard task | The discomfort of starting | Every postponement is reinforced immediately; the task's real cost arrives much later | | Checking the stove three times | Imagined fire | The house never burns down, which "confirms" that checking works | | Never speaking in meetings | Possible embarrassment | Silence is safe every single time; the belief is never tested | | Leaving early to beat traffic | The jam | The jam is never experienced, so the response never meets extinction | | Ordering tests a patient does not need ("defensive medicine") | A malpractice suit | The suit never comes, which looks like proof the tests work; the cost lands on someone else | | Buying insurance | Financial loss | Rational avoidance: the rate of catastrophe really is reduced | | Streak-keeping in an app | Losing the count | A designed avoidance contingency; the aversive event is manufactured | Not all avoidance is pathological — sunglasses, seat belts, and vaccines are avoidance, and adaptive. The problem cases are the ones in which the feared event would not occur, or would be tolerable, and the avoidance itself costs more than the thing avoided. The diagnostic question is always the same: has the contingency been tested lately? > **Three confusions worth clearing up** > > **Avoidance is not punishment.** In avoidance a behavior *increases* because it prevents something aversive; in punishment a behavior *decreases* because it produces something aversive. **Escape is not avoidance.** Escape ends a stimulus that is present; avoidance prevents one that is not. **Negative does not mean bad.** Both are negative reinforcement because a stimulus is subtracted, and both strengthen behavior. ## Key takeaways - Escape terminates an aversive stimulus that is present; avoidance prevents or postpones one that has not yet occurred. Both are negative reinforcement, and almost every avoidance response begins as an escape response that got earlier. - In successful avoidance nothing happens, so the reinforcer is not obvious, and a perfect avoider never contacts the change when the shock is switched off. That is why avoidance can outlast measurable fear by hundreds of trials and why response prevention works when ordinary extinction does not. - Two-factor theory says the warning signal is classically conditioned to elicit fear and the response is reinforced by escape from that fear, but well-trained avoiders show little fear and Sidman avoidance has no signal. One-factor theory names the reinforcer as a reduction in the overall rate of aversive events plus the safety signals that accompany responding; modern accounts combine both with expectancies. - Avoidance is the operant engine inside most anxiety problems: fear is acquired classically but maintained by relief, and safety behaviors keep the catastrophe uncontradicted. Exposure with response prevention blocks the response so the feared outcome can fail to arrive. - Animals given inescapable shock later fail to escape when they can, and only uncontrollability produces the deficit. Maier and Seligman later reversed the interpretation: passivity is the default, and what is learned is control. ### Check yourself **Every time a warning tone sounds, a rat jumps a barrier and the shock never comes. Since shock is aversive, is the jumping being punished?** No. Punishment decreases a behavior by producing something aversive; here jumping increases because it prevents something aversive, which is negative reinforcement in the form of avoidance. Negative does not mean bad: a stimulus is kept away, and the behavior strengthens. **Solomon and Wynne switched off the shock generator, yet the dogs kept jumping for hundreds of trials. Did extinction fail because the dogs were still terrified?** Not mainly. Well-trained avoiders show little fear; the response persists because a dog that jumps early on every trial receives exactly what it always received, no shock, so nothing signals that the contingency has changed. What worked was a barrier that prevented the jump, so the dogs had to stay put and discover that no shock came. **You take a painkiller once a headache has started. Your roommate takes one before a long day at the screen. Which is escape and which is avoidance?** Taking the painkiller once the headache is present is escape, because the behavior terminates an aversive stimulus that is already there. Taking it beforehand is avoidance, because the behavior prevents an aversive stimulus that has not yet occurred. Both are negative reinforcement. **Two groups of dogs receive exactly the same shocks in a harness, but only one group can turn them off. Which group later fails to escape in the shuttle box, and what does the modern account say the other group learned?** The group that could not control the shock: only uncontrollability produces the deficit, which is what the triadic design showed. On the modern account, passivity is the default response to prolonged aversive stimulation, and the group with escapable shock learned that it had control, which inhibits that default and allows escape and avoidance to be learned later. **Explain it to a friend.** Explain why a safety ritual that "works every time" is the hardest kind of habit to drop, using one of your own habits as the example. ## Frequently asked questions **What is the difference between escape and avoidance learning?** In escape learning the aversive stimulus is already present and the behavior terminates it (turning off a shock, taking a painkiller). In avoidance learning the behavior occurs before the aversive stimulus and prevents or postpones it (jumping during the warning signal, leaving early to miss traffic). Both are negative reinforcement. **What is the avoidance paradox?** Reinforcement is supposed to be a consequence that follows a response, but in successful avoidance the "consequence" is that the aversive event does not happen. Something that never occurs cannot follow anything. Theories of avoidance are attempts to identify the actual reinforcer: escape from conditioned fear (two-factor theory), reduction in the overall rate of aversive events and the safety signals that accompany responding (one-factor theory), or confirmed expectancies. **What is two-factor theory?** Mowrer's proposal that avoidance involves two learning processes: first the warning signal is classically conditioned to elicit fear, then the avoidance response is operantly reinforced because it terminates the signal and the fear. It explains acquisition well and persistence partly, but well-trained animals show little fear, and avoidance can be learned without any signal. **What is Sidman avoidance?** Free-operant avoidance, introduced by Murray Sidman in 1953: shocks arrive on a timer unless the animal responds, and each response postpones the next shock for a set interval. There is no warning signal. Rats learn to respond at a steady rate that keeps shocks rare, which was hard for signal-based theories to explain. **Why is avoidance so hard to extinguish?** Because a successful avoider never experiences the change. If the shock generator is switched off, an animal that responds early on every trial receives exactly what it always received — no shock — so nothing signals that responding is now unnecessary. Extinction requires contact with the new contingency, which is why response prevention (blocking the response) works when simple extinction does not. **Is learned helplessness still an accepted theory?** The phenomenon is solid — animals and people exposed to uncontrollable aversive events later fail to escape controllable ones — but the explanation has been revised by its own authors. Maier and Seligman (2016) concluded that passivity is the brain's default response to prolonged aversive stimulation, and that what is learned is control, which inhibits the default. The practical lesson is unchanged: experiences of control are protective. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 3. Solomon, R. L., & Wynne, L. C. (1953). Traumatic avoidance learning: Acquisition in normal dogs. *Psychological Monographs, 67*(4), 1–19. 4. Solomon, R. L., Kamin, L. J., & Wynne, L. C. (1953). Traumatic avoidance learning: The outcomes of several extinction procedures with dogs. *Journal of Abnormal and Social Psychology, 48*(2), 291–302. 5. Sidman, M. (1953). Avoidance conditioning with brief shock and no exteroceptive warning signal. *Science, 118*(3058), 157–158. 6. Sidman, M. (1953). Two temporal parameters of the maintenance of avoidance behavior by the white rat. *Journal of Comparative and Physiological Psychology, 46*(4), 253–261. 7. Mowrer, O. H. (1947). On the dual nature of learning — a re-interpretation of "conditioning" and "problem-solving." *Harvard Educational Review, 17*, 102–148. 8. Mowrer, O. H. (1960). *Learning Theory and Behavior*. Wiley. 9. Miller, N. E. (1948). Studies of fear as an acquirable drive: I. Fear as motivation and fear-reduction as reinforcement in the learning of new responses. *Journal of Experimental Psychology, 38*(1), 89–101. 10. Kamin, L. J. (1956). The effects of termination of the CS and avoidance of the US on avoidance learning. *Journal of Comparative and Physiological Psychology, 49*(4), 420–424. 11. Kamin, L. J., Brimer, C. J., & Black, A. H. (1963). Conditioned suppression as a monitor of fear of the CS in the course of avoidance training. *Journal of Comparative and Physiological Psychology, 56*(3), 497–501. 12. Solomon, R. L., & Wynne, L. C. (1954). Traumatic avoidance learning: The principles of anxiety conservation and partial irreversibility. *Psychological Review, 61*(5), 353–385. 13. Herrnstein, R. J., & Hineline, P. N. (1966). Negative reinforcement as shock-frequency reduction. *Journal of the Experimental Analysis of Behavior, 9*(4), 421–430. 14. Herrnstein, R. J. (1969). Method and theory in the study of avoidance. *Psychological Review, 76*(1), 49–69. 15. Dinsmoor, J. A. (2001). Stimuli inevitably generated by behavior that avoids electric shock are inherently reinforcing. *Journal of the Experimental Analysis of Behavior, 75*(3), 311–333. 16. Seligman, M. E. P., & Johnston, J. C. (1973). A cognitive theory of avoidance learning. In F. J. McGuigan & D. B. Lumsden (Eds.), *Contemporary Approaches to Conditioning and Learning* (pp. 69–110). Winston-Wiley. 17. Krypotos, A.-M., Effting, M., Kindt, M., & Beckers, T. (2015). Avoidance learning: A review on theoretical and experimental approaches. *Frontiers in Behavioral Neuroscience, 9*, 189. 18. Bolles, R. C. (1970). Species-specific defense reactions and avoidance learning. *Psychological Review, 77*(1), 32–48. 19. Seligman, M. E. P., & Maier, S. F. (1967). Failure to escape traumatic shock. *Journal of Experimental Psychology, 74*(1), 1–9. 20. Overmier, J. B., & Seligman, M. E. P. (1967). Effects of inescapable shock upon subsequent escape and avoidance responding. *Journal of Comparative and Physiological Psychology, 63*(1), 28–33. 21. Maier, S. F., & Seligman, M. E. P. (2016). Learned helplessness at fifty: Insights from neuroscience. *Psychological Review, 123*(4), 349–367. 22. Salkovskis, P. M. (1991). The importance of behaviour in the maintenance of anxiety and panic: A cognitive account. *Behavioural Psychotherapy, 19*(1), 6–19. 23. Meyer, V. (1966). Modification of expectations in cases with obsessional rituals. *Behaviour Research and Therapy, 4*(4), 273–280. ## Related - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): The quadrant avoidance belongs to, and why it is not punishment. - [Operant vs. classical](https://operantconditioning.com/operant-vs-classical-conditioning/): Two-factor theory is where the two kinds of learning meet. - [Extinction](https://operantconditioning.com/extinction/): Why avoidance resists it, and what response prevention does. --- # The Matching Law: Herrnstein's Equation for Choice, Explained > The matching law: behavior is allocated in proportion to reinforcement. Herrnstein's experiment, the equations, sports and classroom evidence, self-control. - Source: https://operantconditioning.com/matching-law/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Choice · How time gets divided* Every behavior is a choice among alternatives, and organisms divide their behavior among alternatives in proportion to what each one pays. That single regularity, found in pigeons in 1961, now predicts basketball shot selection, classroom disruption, and why we buy things we do not need. > **Definition** > > The **matching law** states that when two or more responses are available, the *relative rate* of each response equals — matches — the *relative rate of reinforcement* it produces. For two alternatives: > > (B1)/(B1 + B2) = (R1)/(R1 + R2) > > where B is the rate of each behavior and R the rate of reinforcement it earns. An option that delivers 70% of the available reinforcement gets about 70% of the behavior.[1] > > **A worked example.** A pigeon can peck either of two keys. Key A pays off about 30 times an hour, key B about 10 times an hour, so A delivers 30 ÷ (30 + 10) = 75% of the reinforcement. The matching law predicts that the pigeon will make about 75% of its pecks on A and 25% on B — not all of them on A, even though A is plainly better. That the pigeon does not simply pick the better key every time, and that the poorer option keeps a share proportional to what it pays, is the surprise of the law: on schedules like these, behavior is spread in proportion to payoff rather than piled onto the best option. **In brief** - Organisms allocate behavior in proportion to reinforcement: an option that delivers 70% of the available reinforcement gets about 70% of the behavior. - A behavior's strength depends not only on what it earns but on what everything else earns, so enriching the alternatives reduces it without punishment. - Because a reinforcer's value falls hyperbolically with delay, preference flips toward a small, soon reward as it approaches; commitment means choosing before the flip. ## Herrnstein's experiment Richard Herrnstein put pigeons in a chamber with two response keys, each paying grain on its own [variable-interval schedule](https://operantconditioning.com/schedules-of-reinforcement/) — a **concurrent VI VI** schedule. A bird could switch between keys whenever it liked. He varied how the reinforcement was divided between the keys while holding the total constant (VI 3-minute against VI 3-minute, VI 2.25 against VI 4.5, and so on), and to stop the birds simply alternating, he added a **changeover delay**: a brief period after each switch during which no reinforcement could be collected.[1] The result was startlingly orderly. Plot the proportion of pecks on the left key against the proportion of reinforcers earned on the left key and the points fall on the diagonal. Birds did not go exclusively to the richer key, and they did not split their pecks evenly; they matched. ![The matching relation](https://operantconditioning.com/assets/diagrams/the-matching-relation.svg) *Schematic of the matching relation. Real data cluster around the diagonal; systematic departures from it are captured by the generalized matching law below.* ## From matching to a new law of effect In 1970 Herrnstein drew the larger conclusion. If behavior is allocated in proportion to reinforcement, then even a single response in a Skinner box is a choice — between pressing the lever and everything else the rat could do (grooming, exploring, resting), all of which produce some reinforcement of their own. Writing the "everything else" reinforcement as Re, the matching law for one measured response becomes a hyperbola: B = (kR)/(R + Re) Response rate rises with reinforcement rate but with diminishing returns, leveling off at a maximum k. This fit decades of single-schedule data that the original law of effect had described only qualitatively, and Herrnstein proposed it as the quantitative law of effect.[2] It also carried a practical implication: a behavior's strength depends not just on what it earns but on what everything else earns. Enrich the alternatives and the behavior declines without any punishment at all. ## The generalized matching law Real organisms do not match perfectly. William Baum showed in 1974 that the departures are systematic and can be captured by two parameters. Taking logarithms of the ratio form: log ⁡ ( (B1)/(B2) ) = a ⁢ log ⁡ ( (R1)/(R2) ) + log ⁡ b The slope **a** is sensitivity. When a = 1, matching is perfect; the usual finding is **undermatching**, a slope around 0.8, meaning organisms are somewhat less extreme in their preference than the reinforcement ratio warrants. The intercept **b** is bias: a constant preference for one alternative unrelated to reinforcement — a key that is easier to reach, a side the animal favors.[3][4] The generalized form has been fitted to hundreds of data sets across species, and its parameters turn out to be sensitive to procedural details in interpretable ways: shorter changeover delays produce more undermatching, for instance, because switching itself is reinforced. Reinforcement rate is not the only thing organisms match to. Reinforcer magnitude, delay, and quality all enter the equation, which is what makes the framework a general account of choice rather than a fact about grain.[4] ## Matching beyond the pigeon ### Sports A basketball player choosing between a two-point and a three-point attempt is on a concurrent schedule. Vollmer and Bourret analyzed a season of college basketball and found that the proportion of three-point shots taken by teams and by individual players matched the proportion of points those shots produced.[5] Reed, Critchfield, and Martens applied the generalized matching law to NFL play-calling and found that the ratio of passing to rushing plays tracked the ratio of yards each type gained, with the undermatching and bias the generalized law allows for.[6] No coach was computing logarithms; the law describes what allocation looks like when consequences are doing the selecting. ### Classrooms and problem behavior Martens and Houk observed a student whose disruptive and on-task behavior each drew teacher attention at different rates, and found the two behaviors allocated in proportion to the attention each earned.[7] This is the theoretical spine of [differential reinforcement of alternative behavior](https://operantconditioning.com/glossary/#differential-reinforcement-of-alternative-behavior): to reduce a problem behavior maintained by attention, you do not need to punish it. You need the alternative to pay better — more attention, more reliably, sooner. A large applied literature has since evaluated problem behavior as choice, with concurrent-schedule arrangements that make the appropriate response the richer option.[8] ### Everyday human behavior McDowell argued in 1988 that Herrnstein's hyperbola predicts a common frustration: adding a little reinforcement for a desired behavior has a large effect in an environment that is otherwise barren and almost none in an environment that is already rich. The same praise that transforms a child's behavior in a bleak classroom does nothing in one full of competing reinforcers.[9] Humans, it must be said, match less cleanly than pigeons; people given instructions or forming their own rules about a schedule often follow the rule rather than the contingency, a theme that runs through all human operant research. ## Melioration and the mechanism of matching The matching law is a description, not a mechanism. Two candidate mechanisms competed. **Maximizing** accounts, borrowed from economics, hold that organisms distribute behavior so as to obtain the most total reinforcement, and on concurrent VI VI schedules matching happens to be nearly optimal.[10] Herrnstein and Vaughan's **melioration** holds instead that organisms shift behavior toward whichever alternative currently has the higher *local* rate of return until the local rates are equal — a myopic rule that produces matching without any computation of totals, and that predicts the systematically suboptimal choices people make when a locally better option worsens the long-run outcome.[11] Melioration is one reason the matching law connects so naturally to impulsiveness. ## Self-control: when the alternatives differ in time The most consequential extension of matching is to choices between a smaller reinforcer available sooner and a larger one available later. Rachlin and Green showed in 1972 that pigeons facing that choice directly took the small immediate grain, but that if the choice was made well in advance, the same pigeons committed themselves to the larger, later reward — the first laboratory demonstration of a commitment device.[12] George Ainslie explained why in 1975: if the value of a reinforcer falls with delay along a *hyperbola* rather than an exponential curve, the curves for a small-soon and a large-late reward cross, so preference reverses as the small reward approaches.[13] James Mazur's adjusting-delay procedure confirmed the hyperbolic shape precisely.[14] Steep [delay discounting](https://operantconditioning.com/glossary/#delay-discounting) — a fast drop in value with delay — has since been documented in people with substance-use disorders, in problem gamblers, and in smokers, and is studied as a process that cuts across many conditions.[15] The everyday translation: the environment that makes you impulsive is one in which the small reward is near and the large reward is far, and the fix is to move the choice point earlier, when the curves have not yet crossed. Organisms are not uniformly impulsive, though. Cole found that rats on a schedule in which retrieving food pellets from the tray started a one-minute period without further pellets learned to let pellets accumulate and collect them in batches — **operant hoarding**, a form of self-control the impulsivity findings would not have predicted, and a reminder that the details of the contingency matter.[16] ## Behavioral economics: demand, price, and elasticity Once behavior is allocation, the tools of economics apply. Steven Hursh proposed in 1980 that a schedule requirement is a **price** (responses per reinforcer), that consumption plotted against price gives a **demand curve**, and that the slope of that curve — **elasticity** — measures how essential a reinforcer is.[17] Food in a closed economy, where the animal earns all of its food in the chamber, is inelastic: raise the price and the animal works harder to keep consumption up. Sweetened water in an open economy is elastic: raise the price and consumption collapses. Whether reinforcers are **substitutes** (one replaces another) or **complements** (consumed together) can be measured the same way.[18] Kagel, Battalio, and Green showed, in a research program running from the 1970s onward, that rats and pigeons obey demand theory in detail, including some of its odder predictions, such as Giffen goods.[19] The approach has direct policy uses. Demand curves for cigarettes, alcohol, and drugs measured in the laboratory predict how consumption responds to taxation, and "essential value" derived from demand analysis compares the reinforcing efficacy of drugs on a common scale.[20] It also explains a stubborn feature of behavior change: a reinforcer you offer competes in a market, and if the problem behavior is a cheap, inelastic, non-substitutable good, small incentives for the alternative will not move it. ## Limits and criticisms - **Matching is descriptive.** It tells you the outcome of allocation, not how the organism gets there; melioration, maximizing, and momentary-maximizing accounts all reproduce it under various conditions and are hard to separate. - **Humans often follow rules instead.** Verbal instructions and self-generated rules can override contingencies, so human matching is weaker and more variable than animal matching unless the schedule is hard to describe.[21] - **Ratio schedules break the pattern.** On concurrent ratio schedules, exclusive preference for the better option is the optimal strategy and is what animals do, so matching in its simple form applies mainly to interval schedules, where spreading behavior across options pays. - **Parameters need estimating.** The generalized law fits almost anything with a free slope and intercept; its value lies in the parameters being stable and interpretable, which they generally are, not in the fit alone. Within those limits, the matching law is the closest thing behavior analysis has to a physical law. It made choice measurable, connected the laboratory to economics, and gave clinicians a simple instruction that holds up: to change what someone does, change what the alternatives pay. ## Key takeaways - On concurrent variable-interval schedules, organisms do not pile behavior onto the best option; they spread it in proportion to payoff. Herrnstein's pigeons matched the proportion of pecks on each key to the proportion of reinforcers it delivered. - Even a single response is a choice against everything else the organism could do. Herrnstein's hyperbola makes response rate rise with reinforcement at diminishing returns, which is why the same praise transforms behavior in a barren environment and does almost nothing in a rich one. - Real organisms deviate systematically. The generalized matching law adds sensitivity (usually undermatching, a slope around 0.8) and bias (a constant preference unrelated to reinforcement), and reinforcer magnitude, delay, and quality enter the equation alongside rate. - Matching is a description, not a mechanism; melioration, which shifts behavior toward the locally richer option, is one candidate. Humans match less cleanly because rules can override contingencies, and on concurrent ratio schedules exclusive preference for the better option is optimal and is what animals do. - The practical instruction: to change what someone does, change what the alternatives pay. Problem behavior maintained by attention is reduced by making the appropriate alternative pay better, and impulsive choices are avoided by moving the choice point earlier, before the value curves cross. ### Check yourself **A pigeon can peck key A, which pays about 30 times an hour, or key B, which pays about 10. A classmate predicts the bird will peck A almost exclusively, since A is plainly better. What does the matching law predict?** About 75% of pecks on A and 25% on B, because A delivers 30 out of every 40 reinforcers. On concurrent variable-interval schedules behavior is spread in proportion to payoff rather than piled onto the best option; the poorer key keeps a share proportional to what it pays. **A praise program that transformed a student's behavior in one classroom does nothing in another. Was the praise too weak?** Not necessarily. In Herrnstein's hyperbola a behavior's strength depends on its reinforcement relative to the reinforcement for everything else, so adding a little reinforcement has a large effect in a barren environment and almost none in one already full of competing reinforcers. The praise is the same; the alternatives are not. **A student's disruptive behavior earns teacher attention more reliably than on-task behavior does. Without punishing anything, how does the matching law say to reduce the disruption?** Make the alternative pay better: attend to on-task behavior more, more reliably, and sooner, so that it earns the larger share of attention and therefore draws the larger share of behavior. This is differential reinforcement of alternative behavior, and the matching law is its theoretical spine. **Pigeons choosing between a small immediate grain and a larger delayed one take the small one, yet when the same choice is offered well in advance they commit to the larger one. Why does preference reverse?** Because value falls with delay along a hyperbola rather than an exponential curve, the value curves for the small-soon and large-late rewards cross. Far from both rewards the larger one is worth more; as the small one becomes imminent it overtakes, so choosing early, before the curves cross, is a commitment device. **Explain it to a friend.** Explain the matching law to someone who dislikes math, using either the basketball or the classroom example and no numbers at all. ## Frequently asked questions **What is the matching law in simple terms?** Organisms spread their behavior across options in proportion to how much reinforcement each option provides. If one option delivers twice as much reinforcement as another, it gets about twice as much behavior. Richard Herrnstein discovered it in pigeons in 1961, and it holds, with some systematic deviations, across species and settings. **What is the generalized matching law?** Baum's 1974 extension, which adds two parameters: sensitivity (how strongly behavior tracks reinforcement; usually a little less than 1, called undermatching) and bias (a constant preference for one option unrelated to reinforcement). It is written as a straight line in logarithmic ratios and fits most choice data. **How does the matching law explain problem behavior?** Problem behavior and appropriate behavior are alternatives on a concurrent schedule. If misbehavior earns attention more reliably than good behavior, the matching law predicts a lot of misbehavior. The treatment is to make the appropriate alternative pay more — differential reinforcement of alternative behavior — rather than to punish the problem. **What is the difference between the matching law and the law of effect?** Thorndike's law of effect says responses followed by satisfying consequences are strengthened. Herrnstein's matching law quantifies it: response strength is relative, depending on the reinforcement for a behavior compared with the reinforcement for everything else. Herrnstein proposed the hyperbolic form of matching as the quantitative law of effect. **Does the matching law apply to humans?** Yes, though less cleanly. Sports play-calling, classroom behavior, and conversation have been shown to match. Human deviations mostly come from rules and instructions: people who can describe a schedule often follow their description rather than the contingency. **What does the matching law have to do with self-control?** Choices between a smaller, sooner reward and a larger, later one are matching choices in which delay reduces value. Because value falls hyperbolically with delay, preference flips toward the small reward as it becomes imminent. Making the choice early, before the flip, is what commitment devices do. ## References 1. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 2. Herrnstein, R. J. (1970). On the law of effect. *Journal of the Experimental Analysis of Behavior, 13*(2), 243–266. 3. Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. *Journal of the Experimental Analysis of Behavior, 22*(1), 231–242. 4. Baum, W. M. (1979). Matching, undermatching, and overmatching in studies of choice. *Journal of the Experimental Analysis of Behavior, 32*(2), 269–281. 5. Vollmer, T. R., & Bourret, J. (2000). An application of the matching law to evaluate the allocation of two- and three-point shots by college basketball players. *Journal of Applied Behavior Analysis, 33*(2), 137–150. 6. Reed, D. D., Critchfield, T. S., & Martens, B. K. (2006). The generalized matching law in elite sport competition: Football play calling as operant choice. *Journal of Applied Behavior Analysis, 39*(3), 281–297. 7. Martens, B. K., & Houk, J. L. (1989). The application of Herrnstein's law of effect to disruptive and on-task behavior of a retarded adolescent girl. *Journal of the Experimental Analysis of Behavior, 51*(1), 17–27. 8. Fisher, W. W., & Mazur, J. E. (1997). Basic and applied research on choice responding. *Journal of Applied Behavior Analysis, 30*(3), 387–410. 9. McDowell, J. J. (1988). Matching theory in natural human environments. *The Behavior Analyst, 11*(2), 95–109. 10. Rachlin, H., Green, L., Kagel, J. H., & Battalio, R. C. (1976). Economic demand theory and psychological studies of choice. In G. H. Bower (Ed.), *The Psychology of Learning and Motivation* (Vol. 10). Academic Press. 11. Herrnstein, R. J., & Vaughan, W. (1980). Melioration and behavioral allocation. In J. E. R. Staddon (Ed.), *Limits to Action: The Allocation of Individual Behavior*. Academic Press. 12. Rachlin, H., & Green, L. (1972). Commitment, choice and self-control. *Journal of the Experimental Analysis of Behavior, 17*(1), 15–22. 13. Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. *Psychological Bulletin, 82*(4), 463–496. 14. Mazur, J. E. (1987). An adjusting procedure for studying delayed reinforcement. In M. L. Commons, J. E. Mazur, J. A. Nevin, & H. Rachlin (Eds.), *Quantitative Analyses of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value* (pp. 55–73). Erlbaum. 15. Bickel, W. K., & Marsch, L. A. (2001). Toward a behavioral economic understanding of drug dependence: Delay discounting processes. *Addiction, 96*(1), 73–86. 16. Cole, M. R. (1990). Operant hoarding: A new paradigm for the study of self-control. *Journal of the Experimental Analysis of Behavior, 53*(2), 247–262. 17. Hursh, S. R. (1980). Economic concepts for the analysis of behavior. *Journal of the Experimental Analysis of Behavior, 34*(2), 219–238. 18. Hursh, S. R. (1984). Behavioral economics. *Journal of the Experimental Analysis of Behavior, 42*(3), 435–452. 19. Kagel, J. H., Battalio, R. C., & Green, L. (1995). *Economic Choice Theory: An Experimental Analysis of Animal Behavior*. Cambridge University Press. 20. Hursh, S. R., & Silberberg, A. (2008). Economic demand and essential value. *Psychological Review, 115*(1), 186–198. 21. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley. ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): The concurrent schedules matching was discovered on. - [Premack principle](https://operantconditioning.com/premack-principle/): The other great relativity of reinforcement. - [Building habits](https://operantconditioning.com/habits/): Move the choice point before the curves cross. --- # The Premack Principle: Definition, Examples, and the Response Deprivation Hypothesis > The Premack principle: a more probable behavior can reinforce a less probable one. Premack's experiments, Grandma's rule, response deprivation, and mistakes. - Source: https://operantconditioning.com/premack-principle/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Grandma’s rule* "Eat your vegetables, then you can have dessert." Every grandmother knows the rule; David Premack showed why it works, discovered that it runs backward under the right conditions, and changed what a reinforcer is. > **Definition** > > The **Premack principle** states that the opportunity to engage in a *more probable* behavior will reinforce a *less probable* behavior when access to the first is made contingent on performing the second. Reinforcers, on this view, are not stimuli but **behaviors**: it is not the dessert that reinforces vegetable-eating but the eating of the dessert.[1] > > Informally it is called **Grandma's rule**, or in classrooms the **first–then** rule: first the work, then the play. **In brief** - The opportunity to perform a more probable behavior reinforces a less probable one when access to the first is made contingent on the second. - Reinforcers are behaviors, not objects, and the relation reverses with circumstances: water-deprived rats ran to drink, running-deprived rats drank to run. - The response deprivation hypothesis is the better-supported statement: a contingency reinforces whenever it restricts the contingent behavior below its free baseline. ## Premack's experiments David Premack's 1959 study gave first-grade children free access to a candy dispenser and a pinball machine and measured which each child used more. He then made the two contingent on each other. For children who preferred pinball, making the machine available only after they ate a candy increased candy eating; making candy available only after a game did nothing much. For children who preferred candy, the result reversed: candy eating reinforced pinball, and pinball did not reinforce candy eating.[1] Whether an activity worked as a reinforcer depended entirely on whether it was the more probable of the pair for that child. The decisive test came in 1962 with rats, running wheels, and water. Rats deprived of water would run in a wheel to earn access to drinking — the standard result, drinking reinforcing running. But rats given free water and *deprived of running* learned to drink in order to unlock the wheel: running reinforced drinking.[2] The same two behaviors reversed roles with the animal's circumstances. Drinking was not a reinforcer in itself; it was a reinforcer only when drinking was the more probable behavior. ![The Premack principle as a reversal of baseline probabilities](https://operantconditioning.com/assets/diagrams/the-premack-principle-as-a-reversal-of-baseline-probabilities.svg) *Premack's reversal. Whichever behavior is more probable at baseline can reinforce the other; deprivation changes which one that is.* > **Why this was radical** > > Thorndike and Skinner had treated reinforcers as things — food, water, a pellet — that were reinforcing more or less everywhere. Premack argued that reinforcement is a *relation between behaviors*: any behavior can reinforce any less probable one and be reinforced by any more probable one. The list of reinforcers is not a list of objects; it is a ranking of what the organism would do right now if it could.[3] ## The punishment side Premack later extended the principle to punishment. If being made to perform a *less* probable behavior is made contingent on a more probable one, the more probable one declines. Forcing a rat that would rather drink to run a wheel after each drink suppresses drinking.[4] Reinforcement and punishment become two sides of a single relation: the direction of the probability difference determines which you get. ## The response deprivation hypothesis The principle has a known failure. Eisenberger, Karpman, and Trattner found in 1967 that a *less* probable behavior can reinforce a more probable one, provided the schedule restricts the less probable behavior below the level the organism would normally choose.[5] William Timberlake and James Allison built that into a more general rule in 1974, the **response deprivation hypothesis**. Measure how much of each behavior the organism performs when both are freely available — the baseline. A contingency will make one behavior reinforce another if, under the contingency, performing the instrumental behavior at its baseline level would leave the organism with *less* of the contingent behavior than its baseline. Formally, with I the required instrumental responses, C the contingent responses earned, and O the baseline levels: (I)/(C) > (OI)/(OC) When the inequality holds, the organism is "response deprived," and it will increase the instrumental behavior to recover access to the contingent one — regardless of which was more probable to begin with.[6] Premack's principle turns out to be the special case in which the contingent behavior is the more probable one, which almost guarantees deprivation. Timberlake and Allison's version explains the reversals, and it explains why a small requirement often fails: "read one page, then play for an hour" deprives the child of nothing. Applied tests followed. Konarski and colleagues showed in classrooms that schedules meeting the response-deprivation condition increased children's academic work even when the contingent activity was the *less* preferred one, exactly as the hypothesis and not the original principle predicted.[7] A later review concluded that the response deprivation hypothesis is the better-supported statement, and that both are best understood through the lens of [motivating operations](https://operantconditioning.com/glossary/#motivating-operation): restricting a behavior below baseline is an establishing operation for it.[8] ## How to use the Premack principle 1. **Observe before you arrange.** Watch what the person (or animal, or you) actually does when free to choose. The most probable behaviors are the reinforcers available to you, whatever anyone says they enjoy. 2. **Make the probable behavior contingent on the improbable one.** First homework, then screen; first the walk, then the coffee; first three sits, then the game of tug. Access to the reinforcing activity is granted *only* after the target behavior. 3. **Keep the ratio deprivational.** The requirement has to leave the person with less of the preferred activity than they would otherwise have taken, or there is no reason to work for it. Short requirements, delivered often, are usually better than a huge requirement for a huge reward. 4. **Deliver promptly.** The preferred activity should begin within seconds of the target behavior ending, or be bridged with a conditioned reinforcer such as a token or a check-off. 5. **Watch for satiation.** Probabilities change. After an hour of screen time, screen time is no longer the most probable behavior, and the contingency stops working until it is restored. ## Examples | Setting | Less probable behavior (first) | More probable behavior (then) | | --- | --- | --- | | Parenting | Clearing the table | Playing outside | | Classroom | Ten minutes of math problems | Five minutes of free choice | | Special education | Completing a task strip | Time with a preferred toy, shown on a "first–then" board | | Dog training | Sitting at the door | Going through it for the walk | | Self-management | Writing 200 words | Checking messages | | Exercise | The workout | The podcast you only listen to at the gym | | Workplace | Filing the expense report | Starting the interesting design task | The dog example is the one most people already use without a name. A dog that wants to go outside will sit, wait, and make eye contact for the privilege, and the door opening is a more reliable reinforcer than any treat because, at that moment, going through the door is the dog's most probable behavior.[9] The classroom "first–then" board, standard in [applied behavior analysis](https://operantconditioning.com/applications/#aba), is Premack made visible. ## Common mistakes - **Reversing the order.** "You can play now if you promise to do homework after" delivers the reinforcer before the behavior. It reinforces promising. - **Guessing the probabilities.** Parents and managers routinely pick a "reward" nobody would choose. The only test is observation: what does the person do when free? - **Requirements that deprive nobody.** If the person could get as much of the preferred activity as they want anyway, the contingency has no bite. Access must actually be restricted. - **Making the target behavior aversive.** A punishing requirement — an hour of tedium for five minutes of play — teaches avoidance of the whole arrangement. Small, frequent contingencies work better. - **Forgetting that the reinforcer is an activity.** Ending the preferred activity abruptly to start the next requirement can function as negative punishment and provoke resistance. Signal transitions in advance. ## Where it fits in the theory The Premack principle and its successor sit alongside the [matching law](https://operantconditioning.com/matching-law/) as the two great "relativity" results of operant research. Matching says a behavior's strength depends on what the alternatives pay; Premack says whether something reinforces at all depends on what the organism would otherwise be doing. Both replaced the picture of reinforcers as fixed objects with a picture of organisms distributing their time among activities, and both feed directly into [motivating operations](https://operantconditioning.com/glossary/#motivating-operation), the modern term for the deprivation and satiation that make an activity more or less probable.[8] ## Key takeaways - A more probable behavior reinforces a less probable one when access to it is granted only after the less probable one. Grandma's rule and the classroom first–then board are the same idea. - Reinforcement is a relation between behaviors, not a property of objects. Which behavior reinforces which depends on what the organism would otherwise be doing, deprivation can reverse the roles, and running the relation the other way produces punishment. - The response deprivation hypothesis is the more accurate statement: a contingency reinforces the instrumental behavior whenever it leaves the organism with less of the contingent behavior than its free baseline, regardless of which behavior was more probable to begin with. - To use it, observe what the person actually does when free, make the probable behavior contingent on the improbable one, keep the requirement deprivational but small and frequent, deliver promptly, and watch for satiation. - The common mistakes are delivering the preferred activity first, which reinforces promising; guessing the probabilities instead of observing them; requirements that deprive nobody; and requirements so large that the whole arrangement becomes aversive. ### Check yourself **A manager announces a team lunch as a reward for finishing reports on time. Reports get no faster. Why might the lunch have failed?** A reinforcer is not whatever the manager thinks people enjoy; it is what people would actually be doing if free to choose, and the only test is observation. If the lunch is not a more probable activity than what it displaces, or if people would get it anyway, the contingency has no bite. **A child spends more free time on math worksheets than on coloring. The teacher makes coloring available only after math and restricts it below the amount the child would normally do. Can coloring reinforce math, the more probable behavior?** Yes. The original Premack principle says only a more probable behavior can reinforce, but the response deprivation hypothesis says a contingency reinforces whenever it restricts the contingent behavior below its baseline, whichever behavior was more probable. Konarski's classroom studies found exactly this, which is why the response deprivation hypothesis is the better-supported statement. **A parent says, "You can play video games now if you promise to do your homework afterward." Is this the Premack principle?** No. The preferred activity is delivered before the target behavior, so the arrangement reinforces promising, not homework. Access to the probable behavior has to come only after the less probable one. **"Read one page, then you can play for an hour" produces almost no reading. What went wrong?** The requirement deprives the child of nothing: one page for an hour of play leaves the child with as much play as ever, so there is no reason to work for it. The ratio has to be deprivational, and short requirements delivered often work better than a huge requirement for a huge reward. **Explain it to a friend.** Explain why "first the work, then the play" works and when it stops working, in three sentences that use two activities from your own routine. ## Frequently asked questions **What is the Premack principle in simple terms?** A behavior you are likely to do can be used to reinforce a behavior you are unlikely to do, if the likely one is only allowed after the unlikely one. Vegetables before dessert; homework before video games; the walk before the coffee. It is often called Grandma's rule. **What is an example of the Premack principle?** A child who would rather play than tidy up is told that play begins once the toys are put away. Tidying increases. In the laboratory, water-deprived rats ran in a wheel to earn drinking, while rats deprived of running drank to earn access to the wheel — the same two behaviors, reversed by circumstances. **What is the response deprivation hypothesis?** Timberlake and Allison's 1974 refinement: a contingency reinforces the instrumental behavior whenever it restricts the contingent behavior below the amount the organism would freely perform. It explains cases where a less-preferred activity reinforces a more-preferred one, which the original Premack principle cannot, and it is generally considered the more accurate statement. **Is the Premack principle positive reinforcement?** Yes. Access to the preferred activity is added after the target behavior, and the target behavior increases. The novelty is in what counts as the reinforcer — an opportunity to behave, rather than a stimulus — not in the quadrant. **Does the Premack principle work on yourself?** Yes, with the caveat that you are both the person setting the rule and the person tempted to break it. It works best when the preferred activity is physically gated — the podcast that only exists at the gym, the café you only visit after the writing — so that the contingency does not depend on willpower. **Who was David Premack?** An American psychologist (1925–2015) who, besides the principle that bears his name, was a pioneer of primate cognition research: he taught the chimpanzee Sarah to communicate with plastic symbols and, with Guy Woodruff, coined the term "theory of mind" in 1978. ## References 1. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 2. Premack, D. (1962). Reversibility of the reinforcement relation. *Science, 136*(3512), 255–257. 3. Premack, D. (1965). Reinforcement theory. In D. Levine (Ed.), *Nebraska Symposium on Motivation* (Vol. 13, pp. 123–180). University of Nebraska Press. 4. Premack, D. (1971). Catching up with common sense or two sides of a generalization: Reinforcement and punishment. In R. Glaser (Ed.), *The Nature of Reinforcement* (pp. 121–150). Academic Press. 5. Eisenberger, R., Karpman, M., & Trattner, J. (1967). What is the necessary and sufficient condition for reinforcement in the contingency situation? *Journal of Experimental Psychology, 74*(3), 342–350. 6. Timberlake, W., & Allison, J. (1974). Response deprivation: An empirical approach to instrumental performance. *Psychological Review, 81*(2), 146–164. 7. Konarski, E. A., Johnson, M. R., Crowell, C. R., & Whitman, T. L. (1980). Response deprivation and reinforcement in applied settings: A preliminary analysis. *Journal of Applied Behavior Analysis, 13*(4), 595–609. 8. Klatt, K. P., & Morris, E. K. (2001). The Premack principle, response deprivation, and establishing operations. *The Behavior Analyst, 24*(2), 173–180. 9. Pryor, K. (1984). *Don't Shoot the Dog! The New Art of Teaching and Training*. Simon & Schuster. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): Types of reinforcers, including activity reinforcers. - [The matching law](https://operantconditioning.com/matching-law/): The other relativity: strength depends on the alternatives. - [Building habits](https://operantconditioning.com/habits/): Put the principle to work on yourself. --- # The Neuroscience of Operant Conditioning: Dopamine, Reward Prediction Error, and Habit > What happens in the brain during operant conditioning: Olds and Milner, dopamine as a prediction-error signal, wanting vs. liking, how actions become habits. - Source: https://operantconditioning.com/neuroscience/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *The brain · Dopamine and prediction error* Skinner deliberately treated the organism as a black box. Seventy years of neuroscience have opened it, and what is inside looks remarkably like the law of effect: a broadcast teaching signal, a window of a few seconds, and two separate systems for wanting and liking. > **In one paragraph** > > Reinforcement has a physical address. When a consequence is better than the brain predicted, midbrain **dopamine** neurons fire a brief burst that is broadcast to the striatum and frontal cortex, strengthening whichever synapses were active in the preceding seconds. When the consequence is exactly as predicted, they stay quiet; when an expected reinforcer fails to arrive, they dip below baseline. This **reward prediction error** is the teaching signal of operant conditioning, and it is why immediacy, contingency, and unpredictability matter so much.[4][5] **In brief** - Midbrain dopamine neurons signal reward prediction error: a burst when a consequence is better than predicted, silence when expected, a dip when worse. - Dopamine drives wanting, not liking: rats without dopamine still enjoy sugar but stop seeking it, and addiction is sensitized wanting. - Dopamine strengthens only synapses active in the preceding seconds, which is why immediacy matters, and extended training turns goal-directed actions into [habits](https://operantconditioning.com/habits/). ## Reinforcement has an address: Olds and Milner In 1953 James Olds and Peter Milner, working at McGill, implanted an electrode into the septal area of a rat's brain and arranged for a lever press to deliver a brief electrical pulse. The rat pressed, and kept pressing; they published the finding the following year. Olds went on to map the sites that supported **intracranial self-stimulation**: rats would respond thousands of times an hour and cross electrified grids to reach the lever. Routtenberg and Lindy later gave hungry rats a daily hour with both a food lever and a stimulation lever; some spent the hour on stimulation and lost weight.[1][2][3] For the first time, a reinforcer had been produced by acting directly on the nervous system, and the anatomy of the effective sites — the medial forebrain bundle and the pathways it carries — pointed toward a particular chemical system. ## Dopamine is a prediction-error signal, not a pleasure signal The pathways Olds had stimulated carry axons of dopamine neurons from the midbrain (the ventral tegmental area and substantia nigra) to the striatum and prefrontal cortex. Through the 1980s the working assumption was that dopamine *was* pleasure. Wolfram Schultz's recordings from monkeys overturned that. Schultz trained monkeys on a simple task in which a light or sound was followed, a second or two later, by a squirt of juice, while recording individual dopamine neurons. Early in training the neurons fired when the juice arrived. Once the cue reliably predicted juice, the burst moved to the *cue*, and the fully predicted juice produced no response at all. And when the cue was followed by no juice, the neurons paused — activity dropped below baseline at the moment the juice should have come.[4][5] That is exactly the profile of a **prediction error**: positive for better-than-expected, zero for as-expected, negative for worse-than-expected. In 1996 Montague, Dayan, and Sejnowski had shown that this is the quantity a temporal-difference learning algorithm needs, and the 1997 paper by Schultz, Dayan, and Montague joined the two: the brain appeared to be running the same computation that computer scientists had derived from the law of effect.[6][4][7] Correlation became causation with optogenetics. Activating dopamine neurons with light at the moment a reward is delivered is enough to make rats learn about a cue that would otherwise be blocked, and phasic stimulation of these neurons is sufficient to produce a conditioned preference for a place.[8][9] The dopamine burst does not merely accompany learning; it drives it. > **Why this explains schedules of reinforcement** > > A reinforcer that is fully predictable generates no prediction error and, eventually, no dopamine response: a continuous schedule teaches fast, then goes flat. A reinforcer that arrives after an unpredictable number of responses cannot be fully predicted, so every one produces a burst. Schultz's group later found that dopamine neurons also show a slow ramp of activity that is largest when reward probability is 50% — maximal uncertainty.[10] Whether that ramp is what makes gambling compelling is still argued, but the [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/) has, at minimum, a plausible neural signature. ## Wanting is not liking If dopamine is not pleasure, what is it for? Kent Berridge and Terry Robinson answered with a dissociation that has held up for three decades. Rats whose dopamine neurons were destroyed with the toxin 6-hydroxydopamine stopped eating — they would starve unless tube-fed — yet when sugar was placed on their tongues they showed exactly the lip-licking, paw-licking facial reactions that intact rats show. They still *liked* sugar. What they had lost was *wanting*: the motivation to work for it, approach it, and treat cues for it as attractive.[11][12] Berridge and Robinson called this attribution of attractiveness to reinforcers and their cues **incentive salience**. Liking, meanwhile, turned out to depend on tiny opioid "hedonic hotspots" in the nucleus accumbens and elsewhere, anatomically separate from the dopamine system.[13] Reinforcement, at the level of neurons, is therefore mostly about wanting. Dopamine also sets how much effort an animal will spend: with dopamine reduced, rats still choose food, but they shift from a lever that pays well and requires many presses to freely available, less preferred chow.[14] The everyday consequence is familiar to anyone who has kept doing something they no longer enjoy. ### Addiction as sensitized wanting The dissociation explains the central puzzle of addiction: people keep wanting drugs that have long since stopped delivering much pleasure. Robinson and Berridge's incentive-sensitization theory proposes that repeated drug exposure sensitizes the dopamine system's response to the drug and to its cues, so that wanting grows even as liking shrinks. The syringe, the bar, the friend, the time of day become cues with enormous incentive salience, which is why relapse is so often triggered by the environment rather than by withdrawal.[15][16] In operant terms: drug taking is positively reinforced by the drug and negatively reinforced by relief from withdrawal, and the cues that precede it become discriminative stimuli and conditioned reinforcers with a sensitized grip. ## Learning from the stick: punishment in the brain Worse-than-expected outcomes produce a dopamine dip, and the brain appears to learn from dips through a different route than from bursts. Michael Frank and colleagues gave a probabilistic learning task to people with Parkinson's disease, a condition of dopamine loss. Off medication, patients were better at learning from negative feedback — which choices to avoid — than from positive feedback. On dopamine-replacing medication, the pattern reversed: they learned from positive feedback and became worse at learning from negative.[17] Frank's model attributes the two to separate striatal pathways, one facilitating action when dopamine is high and one suppressing it when dopamine is low. Separately, neurons in the lateral habenula fire when a reward is omitted or a punishment predicted, and inhibit dopamine neurons — a candidate source of the dip.[18][36] Reinforcement and punishment are not mirror images at the neural level any more than they are at the behavioral one. ## The window of a few seconds: why immediacy matters Every practical guide to reinforcement says the consequence must be immediate. The neural reason is now visible. A dopamine burst cannot strengthen every synapse in the striatum; it strengthens the ones that were recently active, which carry a short-lived molecular "eligibility trace." Yagishita and colleagues, using glutamate uncaging on single dendritic spines, found that dopamine enlarged a spine only if it arrived within roughly 0.3 to 2 seconds after the spine had been stimulated. Earlier or later, nothing happened.[19] That window is the cellular basis of [contiguity](https://operantconditioning.com/glossary/#contiguity): a reinforcer that comes thirty seconds after a behavior finds the trace gone and strengthens whatever happened in the last two seconds instead. Dopamine is not the only teaching signal. Neurons of the nucleus basalis release acetylcholine across the cortex and respond to reinforcers and to the stimuli that predict them; pairing a tone with electrical stimulation of the nucleus basalis is enough to expand the tone's representation in the auditory cortex, without any behavior at all.[20][21] Reinforcement, in other words, reshapes perception as well as action. ## From action to habit: two systems in the striatum Press a lever a few hundred times for food and then make the food unappealing — pair it with a mild poison, or feed the animal to satiety. A rat with moderate training stops pressing; it "knows" what the lever produces and no longer wants it. A rat with extensive training keeps pressing anyway.[22][23] Anthony Dickinson used this reinforcer-devaluation test to distinguish **goal-directed actions**, which are sensitive to the current value of their outcome, from **habits**, which are triggered by antecedents and run off regardless. The two have different homes. Lesions of the dorsomedial striatum leave animals unable to act on outcome value, while lesions of the dorsolateral striatum prevent habits from forming, so that over-trained animals stay sensitive to devaluation.[24][25][26] Training on interval schedules produces habits faster than training on ratio schedules, presumably because on an interval schedule the connection between how much you respond and how much you get is loose.[27] Human imaging finds the same division of labor: prediction errors in the ventral striatum during learning, with the dorsal striatum engaged when the learning must guide action.[28] Everitt and Robbins argued that addiction is this transition run to its end — from action to habit to compulsion, with control migrating from ventral to dorsal striatum as the behavior becomes cue-driven and insensitive to consequences.[29] This is the neuroscience behind a piece of practical advice on [the habits page](https://operantconditioning.com/habits/): a well-formed habit survives the loss of motivation, for good and ill. The cue keeps producing the behavior after the outcome has lost its appeal, which is why habits are hard to break by deciding to and easier to break by changing the antecedent. ## Operant conditioning in a single neuron The principle scales down remarkably far. In 1969 Eberhard Fetz reinforced monkeys with food pellets whenever a single recorded neuron in motor cortex fired faster; within minutes the monkeys raised that neuron's firing rate, with no instruction about what they were doing.[30] In the sea slug *Aplysia*, Brembs and colleagues reinforced a feeding movement by stimulating a dopaminergic nerve immediately after it, and then reproduced the learning in a single identified neuron in a dish: contingent dopamine applied to neuron B51 changed its excitability the way training changed the whole animal's behavior.[31] Operant conditioning is not a trick of large brains; it is a property of neurons. ## The brain as a reinforcement learner Reinforcement learning, the branch of artificial intelligence in which an agent learns from reward signals, was built on Thorndike and Skinner and on temporal-difference learning, and the dopamine findings turned it into a theory of the brain.[7][32] The current picture has two learners running in parallel: a "model-free" system that caches the value of actions from prediction errors — the habit system — and a "model-based" system that plans using a map of how actions lead to outcomes — the goal-directed system — with control shifting between them according to which is more reliable.[33] The same framework connects to classical conditioning through the Rescorla–Wagner model of 1972, which was itself a prediction-error rule.[34] ## What this means in practice - **Immediacy is not a rule of thumb; it is a molecular window.** If the real reinforcer must be delayed, deliver a conditioned reinforcer — a click, a word, a checkmark — inside the window and let it bridge the gap. - **Predictable reinforcers stop teaching.** Once a behavior is learned, thinning to an intermittent schedule keeps prediction errors, and dopamine, alive. The same mechanism is what makes slot machines and feeds hard to leave. - **Wanting and liking come apart.** A behavior can be maintained by cues long after its outcome stopped being enjoyable; treat the cues, not just the outcome. - **Habits outlive motivation.** An over-trained behavior is insensitive to devaluation. To change it, change the antecedent or make the response impossible rather than relying on wanting it less. - **Reinforcement and punishment use different circuitry.** Which one a person learns from best can depend on the state of their dopamine system — one reason blanket claims that "punishment doesn't work" or "rewards don't work" are both too simple. ## What is still unsettled The prediction-error account is the best-supported theory of dopamine, not the whole story. Dopamine also ramps up as animals approach rewards, participates in movement and vigor, and is released in patterns that a single scalar error signal does not obviously explain; some researchers argue it broadcasts several different messages on different timescales.[35][14] Most of the causal work is in rodents, and human evidence rests on imaging and on patient groups. And none of it changes the functional definitions: a reinforcer is still whatever increases the behavior it follows. The neuroscience explains why the law of effect holds; it does not replace it. ## Key takeaways - Dopamine is a prediction-error signal, not a pleasure signal: midbrain dopamine neurons burst when a consequence is better than predicted, stay quiet when it is as predicted, and dip when it is worse. Because a fully predictable reinforcer produces no error, continuous reinforcement teaches fast and then goes flat, while intermittent schedules keep prediction errors alive. - Reinforcement and punishment are not mirror images in the brain. Learning from dips runs through a different striatal pathway than learning from bursts, and which one a person learns from best can depend on the state of their dopamine system. - Wanting and liking are separate systems: dopamine drives wanting (incentive salience) and effort, while liking depends on opioid hotspots. Addiction is sensitized wanting, which is why cues trigger relapse long after the drug stopped being enjoyable. - A dopamine burst strengthens only synapses that were active within roughly 0.3 to 2 seconds before it, the cellular basis of contiguity. A delayed reinforcer strengthens whatever happened just before it arrived, so bridge any delay with a conditioned reinforcer. - With extended training, control shifts from a goal-directed system in the dorsomedial striatum, sensitive to outcome value, to a habit system in the dorsolateral striatum, triggered by cues regardless of value. Habits outlive motivation, so they are easier to break by changing the antecedent than by wanting the outcome less. ### Check yourself **A monkey has learned that a light predicts juice. When the juice arrives exactly as predicted, do its dopamine neurons fire?** No. Once the cue reliably predicts juice, the burst moves to the cue and the fully predicted juice produces no response, because dopamine signals prediction error rather than pleasure. If the juice is then omitted, the neurons dip below baseline at the moment it should have come. **A rat whose dopamine neurons have been destroyed stops eating and would starve unless tube-fed. Has it lost the ability to enjoy food?** No. When sugar is placed on its tongue it shows the same lip-licking and paw-licking reactions as an intact rat, so liking is intact. What it has lost is wanting: the motivation to seek food, work for it, and treat its cues as attractive, which depends on dopamine while liking depends on separate opioid hotspots. **A trainer gives a treat about thirty seconds after a good sit, once she has found the treat bag. What does the dopamine burst strengthen?** Whatever the dog did in the last couple of seconds before the treat arrived, not the sit. Dopamine enlarges a synapse only if it arrives within roughly 0.3 to 2 seconds of the synapse's activity, and the sit's eligibility trace is long gone. The fix is a conditioned reinforcer, such as a click, delivered inside the window to bridge the gap. **Two rats were trained to press a lever for food, one moderately and one extensively. The food is then made unappealing. Which rat keeps pressing, and why?** The extensively trained rat. Its pressing has become a habit, supported by the dorsolateral striatum and triggered by antecedents regardless of the outcome's current value. The moderately trained rat's pressing is still goal-directed, supported by the dorsomedial striatum, so it stops once the outcome is devalued. **Explain it to a friend.** Explain why a reinforcer that arrives every single time eventually stops teaching anything, using the word "surprise" and without using the word "dopamine." ## Frequently asked questions **Is dopamine the "pleasure chemical"?** No. Dopamine neurons signal reward prediction error — how much better or worse an outcome was than expected — and drive wanting (motivation and cue attraction). Pleasure, or "liking," depends on separate opioid systems. Animals with dopamine removed still show every sign of enjoying sugar; they simply stop seeking it. **What is a reward prediction error?** The difference between the reward that arrives and the reward that was predicted. A positive error (better than expected) produces a burst of dopamine and strengthens the preceding behavior; zero error (as expected) produces nothing; a negative error (worse than expected, including an omitted reward) produces a dip. It is the neural version of the law of effect and the core of reinforcement-learning algorithms. **Which part of the brain is responsible for operant conditioning?** No single part. Dopamine neurons in the midbrain provide the teaching signal; the ventral striatum learns predictions; the dorsomedial striatum supports goal-directed action; the dorsolateral striatum supports habits; the prefrontal cortex supports planning and rule use; and the amygdala and lateral habenula handle aversive outcomes. Operant learning also occurs in invertebrates with far simpler nervous systems, and even in single neurons. **Why must reinforcement be immediate?** Because the synapses that were active during a behavior stay "eligible" for strengthening only briefly. In mouse striatal neurons, dopamine strengthened a synapse only if it arrived within about 0.3–2 seconds of the synapse's activity. A delayed reinforcer strengthens whatever happened just before it arrived, not the behavior you intended. **Does the brain treat punishment as the opposite of reinforcement?** Not exactly. Omitted rewards and predicted punishments produce dopamine dips, partly driven by the lateral habenula, and learning from them appears to run through a different striatal pathway than learning from rewards. In people with Parkinson's disease, dopamine medication improves learning from positive feedback and worsens learning from negative feedback. **How does this relate to habits?** With extended practice, control of a behavior shifts from a goal-directed system (dorsomedial striatum, sensitive to whether the outcome is still valuable) to a habit system (dorsolateral striatum, triggered by cues regardless of outcome value). That is why long-standing habits continue after their rewards have lost appeal, and why changing cues works better than willpower. ## References 1. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427. 2. Olds, J. (1958). Self-stimulation of the brain. *Science, 127*(3294), 315–324. 3. Routtenberg, A., & Lindy, J. (1965). Effects of the availability of rewarding septal and hypothalamic stimulation on bar pressing for food under conditions of deprivation. *Journal of Comparative and Physiological Psychology, 60*(2), 158–161. 4. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 5. Schultz, W. (1998). Predictive reward signal of dopamine neurons. *Journal of Neurophysiology, 80*(1), 1–27. 6. Montague, P. R., Dayan, P., & Sejnowski, T. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. *Journal of Neuroscience, 16*(5), 1936–1947. 7. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. 8. Steinberg, E. E., Keiflin, R., Boivin, J. R., Witten, I. B., Deisseroth, K., & Janak, P. H. (2013). A causal link between prediction errors, dopamine neurons and learning. *Nature Neuroscience, 16*(7), 966–973. 9. Tsai, H.-C., Zhang, F., Adamantidis, A., Stuber, G. D., Bonci, A., de Lecea, L., & Deisseroth, K. (2009). Phasic firing in dopaminergic neurons is sufficient for behavioral conditioning. *Science, 324*(5930), 1080–1084. 10. Fiorillo, C. D., Tobler, P. N., & Schultz, W. (2003). Discrete coding of reward probability and uncertainty by dopamine neurons. *Science, 299*(5614), 1898–1902. 11. Berridge, K. C., Venier, I. L., & Robinson, T. E. (1989). Taste reactivity analysis of 6-hydroxydopamine-induced aphagia: Implications for arousal and anhedonia hypotheses of dopamine function. *Behavioral Neuroscience, 103*(1), 36–45. 12. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? *Brain Research Reviews, 28*(3), 309–369. 13. Peciña, S., & Berridge, K. C. (2005). Hedonic hot spot in nucleus accumbens shell: Where do μ-opioids cause increased hedonic impact of sweetness? *Journal of Neuroscience, 25*(50), 11777–11786. 14. Salamone, J. D., & Correa, M. (2012). The mysterious motivational functions of mesolimbic dopamine. *Neuron, 76*(3), 470–485. 15. Robinson, T. E., & Berridge, K. C. (1993). The neural basis of drug craving: An incentive-sensitization theory of addiction. *Brain Research Reviews, 18*(3), 247–291. 16. Volkow, N. D., Koob, G. F., & McLellan, A. T. (2016). Neurobiologic advances from the brain disease model of addiction. *New England Journal of Medicine, 374*(4), 363–371. 17. Frank, M. J., Seeberger, L. C., & O'Reilly, R. C. (2004). By carrot or by stick: Cognitive reinforcement learning in parkinsonism. *Science, 306*(5703), 1940–1943. 18. Matsumoto, M., & Hikosaka, O. (2007). Lateral habenula as a source of negative reward signals in dopamine neurons. *Nature, 447*(7148), 1111–1115. 19. Yagishita, S., Hayashi-Takagi, A., Ellis-Davies, G. C. R., Urakubo, H., Ishii, S., & Kasai, H. (2014). A critical time window for dopamine actions on the structural plasticity of dendritic spines. *Science, 345*(6204), 1616–1620. 20. Richardson, R. T., & DeLong, M. R. (1990). Context-dependent responses of primate nucleus basalis neurons in a go/no-go task. *Journal of Neuroscience, 10*(8), 2528–2540. 21. Kilgard, M. P., & Merzenich, M. M. (1998). Cortical map reorganization enabled by nucleus basalis activity. *Science, 279*(5357), 1714–1718. 22. Adams, C. D., & Dickinson, A. (1981). Instrumental responding following reinforcer devaluation. *Quarterly Journal of Experimental Psychology B, 33*(2), 109–121. 23. Dickinson, A. (1985). Actions and habits: The development of behavioural autonomy. *Philosophical Transactions of the Royal Society B, 308*(1135), 67–78. 24. Yin, H. H., Knowlton, B. J., & Balleine, B. W. (2004). Lesions of dorsolateral striatum preserve outcome expectancy but disrupt habit formation in instrumental learning. *European Journal of Neuroscience, 19*(1), 181–189. 25. Yin, H. H., Ostlund, S. B., Knowlton, B. J., & Balleine, B. W. (2005). The role of the dorsomedial striatum in instrumental conditioning. *European Journal of Neuroscience, 22*(2), 513–523. 26. Yin, H. H., & Knowlton, B. J. (2006). The role of the basal ganglia in habit formation. *Nature Reviews Neuroscience, 7*(6), 464–476. 27. Dickinson, A., Nicholas, D. J., & Adams, C. D. (1983). The effect of the instrumental training contingency on susceptibility to reinforcer devaluation. *Quarterly Journal of Experimental Psychology B, 35*(1), 35–51. 28. O'Doherty, J., Dayan, P., Schultz, J., Deichmann, R., Friston, K., & Dolan, R. J. (2004). Dissociable roles of ventral and dorsal striatum in instrumental conditioning. *Science, 304*(5669), 452–454. 29. Everitt, B. J., & Robbins, T. W. (2005). Neural systems of reinforcement for drug addiction: From actions to habits to compulsion. *Nature Neuroscience, 8*(11), 1481–1489. 30. Fetz, E. E. (1969). Operant conditioning of cortical unit activity. *Science, 163*(3870), 955–958. 31. Brembs, B., Lorenzetti, F. D., Reyes, F. D., Baxter, D. A., & Byrne, J. H. (2002). Operant reward learning in *Aplysia*: Neuronal correlates and mechanisms. *Science, 296*(5573), 1706–1709. 32. Dayan, P., & Niv, Y. (2008). Reinforcement learning: The good, the bad and the ugly. *Current Opinion in Neurobiology, 18*(2), 185–196. 33. Daw, N. D., Niv, Y., & Dayan, P. (2005). Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control. *Nature Neuroscience, 8*(12), 1704–1711. 34. Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), *Classical Conditioning II: Current Research and Theory* (pp. 64–99). Appleton-Century-Crofts. 35. Berke, J. D. (2018). What does dopamine mean? *Nature Neuroscience, 21*(6), 787–793. 36. Matsumoto, M., & Hikosaka, O. (2009). Representation of negative motivational value in the primate lateral habenula. *Nature Neuroscience, 12*(1), 77–84. ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Why unpredictable reinforcers hold behavior — now with a neural reason. - [Building habits](https://operantconditioning.com/habits/): Use the action-to-habit transition on purpose. - [History](https://operantconditioning.com/history/): From Thorndike to reinforcement learning. --- # The Skinner Box (Operant Conditioning Chamber): What It Is, How It Works, and Why It Mattered > A Skinner box (operant conditioning chamber) is the apparatus Skinner built to measure how consequences change behavior: parts, a typical session, and myths. - Source: https://operantconditioning.com/skinner-box/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Apparatus · Methods* A small chamber, a lever, a food tray, and a pen that stepped up once per press. Here is what the operant conditioning chamber is, part by part, what it measured that nothing before it could, how to read the records it drew — and what the "baby in a box" story gets wrong. > **Definition** > > A **Skinner box** — formally an **operant conditioning chamber** — is an enclosed apparatus in which an animal can make a simple, repeatable response, such as pressing a lever or pecking a lit disk, that automatic equipment records and can follow with a programmed consequence, usually food. B. F. Skinner built the first versions at Harvard in the early 1930s to measure the *rate* of a freely emitted behavior over time. > > Its purpose is isolation, not confinement: one response, one consequence, and nothing else changing, so that the relation between them can be measured for hours without an experimenter in the room.[1] **In brief** - A Skinner box, formally an operant conditioning chamber, lets an animal make one simple response that equipment records and can follow with a consequence. - Unlike puzzle boxes and mazes, it records the rate of a free operant continuously; on the [cumulative record](https://operantconditioning.com/glossary/#cumulative-record), slope is rate. - Skinner never called it a Skinner box, and the "baby in a box" was a climate-controlled crib, not an experiment. ## What is a Skinner box? Most students meet the Skinner box as a picture: a rat in a metal cage pressing a bar. The picture is accurate but misses the point. The chamber was never the discovery; it was the instrument that made a kind of discovery possible. It let the animal respond whenever and as often as it liked and recorded every response automatically. That made *how often* an animal did something a measurable quantity, and almost everything in [operant conditioning](https://operantconditioning.com/) — reinforcement, extinction, [schedules](https://operantconditioning.com/schedules-of-reinforcement/), [shaping](https://operantconditioning.com/shaping/), [stimulus control](https://operantconditioning.com/stimulus-control/) — was first seen as a change in that quantity. By Skinner's account the box evolved by simplification: a runway along which a rat ran for food, whose timing proved surprisingly orderly, was shortened step by step until the rat needed only to press a horizontal bar to work a food dispenser.[2] *The Behavior of Organisms* (1938), which laid out the resulting science, is built almost entirely on records from it.[1] ### Why Skinner never called it a Skinner box The name was not his. It came out of Clark Hull's laboratory at Yale, which adopted a version of the apparatus in the 1930s; Skinner never used it, and the field settled on *operant conditioning chamber*.[3] The eponym misleads in two ways. It makes a method sound like a gadget, when the box was only a way of getting at the method — measuring the rate of a free operant. And within a few years the name attached itself to something else entirely: the enclosed crib Skinner built for his daughter, the subject of the most durable myth about him (below). ## The parts of a Skinner box ![Schematic of an operant conditioning chamber for a rat](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber-for-a-rat.svg) *An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a pellet on any schedule; the cue light signals when pressing will pay off; the recorder draws responses against time.* | Part | Rat version | Pigeon version | What it is for | | --- | --- | --- | --- | | **Operandum** | A small lever that closes a switch when pressed | A translucent disk, the "key," at head height, lit from behind; a peck closes the switch | Defines the response that counts. Anything that closes the switch — left paw, right paw, nose — is the same operant. | | **Food magazine** | A dispenser drops a 45-milligram pellet into a tray | A hopper of grain is raised into an opening, lit, for a few seconds | Delivers the reinforcer. Its click precedes every pellet and becomes a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer) that bridges press and food. | | **Stimulus lights** | A house light and cue lights above the lever | Colors or patterns projected onto the key | Signal when responding will pay: the discriminative stimulus and its opposite, the S-delta. | | **Speaker** | Tones, clicks, or white noise; the cubicle's fan masks outside sounds | Auditory stimuli, and a chamber that sounds the same every session. | | | **Grid floor** | Parallel metal rods over a droppings tray | Easy to clean; in punishment and [avoidance](https://operantconditioning.com/avoidance-learning/) experiments the rods can carry a brief shock. | | | **Programming and recording** | Originally relays and timers; now a computer | Runs the [schedule](https://operantconditioning.com/schedules-of-reinforcement/) and timestamps every response, so the experimenter can leave the room. | | ## How a Skinner box experiment runs A chamber with an untrained rat in it produces nothing: the rat sniffs the corners, grooms, and occasionally leans on the lever by accident. A working session is built in stages, each depending on the one before. 1. **Deprivation.** The animal is kept mildly hungry — typically around 80 to 85 percent of its free-feeding weight — so that food functions as a reinforcer. Without this [motivating operation](https://operantconditioning.com/abc-model/), nothing else works. 2. **Magazine training.** Pellets arrive on a timer, no response required. Within a few dozen deliveries the rat goes to the tray the instant it hears the click, now a conditioned reinforcer that can be delivered from across the chamber. 3. **Shaping.** The experimenter delivers a pellet by hand switch whenever the rat does something closer to a lever press: facing it, approaching, rearing, touching. Each approximation is reinforced until frequent, then dropped for the next. [How shaping works ›](https://operantconditioning.com/shaping/) 4. **Continuous reinforcement, then schedules.** Every press pays at first. Then the requirement is thinned — every fifth press, an unpredictable number, the first press after a fixed time — and the recorder shows each schedule's signature. 5. **Discrimination training.** Presses pay only while a light is on. Pressing comes under stimulus control — high in the light, near zero in the dark — the laboratory version of the antecedent in the [A-B-C model](https://operantconditioning.com/abc-model/). [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) 6. **Extinction.** Food stops. The record shows a burst of rapid, variable pressing, then a decline to almost nothing, with some recovery at the next session. [What extinction looks like ›](https://operantconditioning.com/extinction/) Later experiments are built on this base: punishment (a brief shock for a press that is still reinforced), choice (two keys on different schedules, the procedure behind the [matching law](https://operantconditioning.com/matching-law/)), drugs (a dose before the session, with the change in rate as the measure). You can run the magazine-training, shaping, schedule, and extinction stages in the virtual chamber below. ## Run a Skinner box yourself This lab puts an untrained, mildly hungry rat in a chamber and hands you the pellet button. Magazine-train it, shape a lever press one approximation at a time, put the press on a schedule, then stop the food and watch extinction. The cumulative record under the chamber draws every press. The rat is a stylized model — its tendencies shift with what you reinforce, drift back when you don't, and follow the schedule patterns Ferster and Skinner documented — not a replay of real data. Real shaping takes longer and real rats are more surprising. ## What a Skinner box measures, and why that mattered Edward Thorndike's cats, in 1898, were placed one at a time in a latched "puzzle box" with food outside; Thorndike timed each escape and plotted the times across trials.[4] The mazes and runways that dominated animal psychology for the next forty years worked the same way. These are **discrete-trial** methods: the experimenter decides when the animal may behave, each trial yields one number, and learning appears as a curve across trials. In the operant chamber the lever is always there. The animal decides when to respond and how often — a **[free operant](https://operantconditioning.com/glossary/#free-operant)** — and the natural measure becomes the **rate of response**, presses per minute, moment by moment. Skinner argued that rate was the datum psychology had been missing: it changes continuously rather than in trial-sized lumps, it can be read at a glance from a cumulative record, and it proved sensitive to almost every variable an experimenter could manipulate.[1][5] | Aspect | Thorndike's puzzle box | Maze or runway | Operant chamber | | --- | --- | --- | --- | | **Who starts a trial** | Experimenter places the cat | Experimenter places the rat at the start | The animal, whenever it likes | | **What is measured** | Seconds to escape | Time to the goal; wrong turns | Responses per unit time, continuously | | **What a session yields** | One number per trial | One or two numbers per trial | A record of every response and reinforcer | | **Best suited to** | Showing that consequences "stamp in" behavior | Spatial learning and motivation | Schedules, extinction, stimulus control, choice, drug effects | Some of the field's landmarks were accidents of the apparatus, and Skinner said so. The cumulative record began as a lucky by-product of the way his early food magazine was built. The first extinction curve he ever saw appeared when the pellet dispenser jammed and the rat went on pressing. And the first intermittent schedule was born when, running short of the pellets he made by hand, he reinforced only one press a minute to save them — and found the record perfectly orderly.[2] In the same kind of chamber, feeding pigeons every fifteen seconds regardless of what they did produced "superstitious" rituals in six of eight birds, the classic demonstration that accidental contingencies shape behavior.[6] The papers describing the method and these discoveries are collected in *Cumulative Record*.[7] ## Pigeons vs. rats *The Behavior of Organisms* is a book about rats. Skinner switched to pigeons during the Second World War, when Project Pigeon trained them to guide a missile by pecking at a target image, and largely stayed with them.[8] The bird pecks quickly and can sustain thousands of responses an hour, has sharp color vision (ideal for work on stimulus control), is cheap and hardy, and lives for a decade or more, so one subject can be studied for years. *Schedules of Reinforcement*, the 1957 catalog of schedule effects, is built mostly on pigeon records.[9] Rats never went away. They are nocturnal and nearly color-blind by comparison, but their physiology, genetics, and brain are far better mapped, which is why the rat (and now the mouse) chamber is standard in pharmacology and neuroscience. The logic is the same for any species — a defined response, an automated consequence, a continuous record of rate — and chambers have been built for monkeys, fish, and humans. ## How to read a cumulative record Skinner's second invention made the data visible. In the **cumulative recorder**, a strip of paper moves under a pen at constant speed; each response steps the pen a small fixed distance upward, and each reinforcer is marked with a short diagonal tick, or "pip." When the pen reaches the top of the paper it drops back to the bottom and continues. The [cumulative record](https://operantconditioning.com/glossary/#cumulative-record) is the field's native graph, and every schedule leaves a signature on it.[9] ![An annotated cumulative record](https://operantconditioning.com/assets/diagrams/an-annotated-cumulative-record.svg) *Reading a cumulative record: slope is rate, ticks are reinforcers, flat stretches are pauses, and the curved "scallop" is the fixed-interval pattern. The line only ever goes up; a behavior that has stopped draws a horizontal line.* - **Slope is rate.** Steep means fast responding; flat means the animal has stopped. A change in slope is a change in behavior, visible the moment it happens. - **Pauses.** A flat stretch right after a tick is the [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause), which lengthens as ratios and intervals grow. The fixed-ratio record is "break and run": pause, then a burst at full speed. - **Scallops.** On a fixed-interval schedule the line curves — nothing just after a reinforcer, then accelerating responding as the interval runs out. Variable schedules erase both pause and curve and draw a steady, nearly straight line. - **Extinction.** When reinforcement stops, the record rises steeply for a while (the extinction burst) and then bends toward horizontal. Ferster and Skinner's book contains hundreds of such records; the reader is expected to see the effects, not compute them. [Draw your own records in the schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) ## The "baby in a box": the Skinner box myth In 1944, before his second daughter Deborah was born, Skinner built what he called a **baby tender** and later an **air crib**: an enclosed crib with a large safety-glass front and a stretched canvas floor, supplied with filtered, warmed air, so that a baby could sleep and play in a diaper alone — no blankets, layers of clothing, or bars — at a constant temperature. It was a bed with climate control. Deborah was taken out to be fed, changed, held, and played with like any other infant. *Ladies' Home Journal* published Skinner's article about it in October 1945 under a headline he did not write, "Baby in a Box," and a number of families later used commercial versions.[10] > **What is false** > > The headline, and the fact that "Skinner box" was already a phrase, produced a rumor that circulated for decades: that Skinner raised his daughter in a Skinner box as an experiment, and that she became psychotic, sued him, or died by suicide. None of it is true. No experiment was ever run in the air crib; it had no lever, no dispenser, and no contingencies. Deborah Skinner Buzan grew up to be an artist living in Britain, has described a happy childhood with an affectionate father, and answered the rumor in 2004 in an article titled "I was not a lab rat," after a popular book repeated it.[11] The myth survives because it fits a picture of Skinner as a cold manipulator. The facts of the box cut against it: it was a measuring instrument, its reinforcers were overwhelmingly food rather than shock, and Skinner spent fifty years arguing against punishment. [B. F. Skinner: life, work, critics, and myths ›](https://operantconditioning.com/bf-skinner/) ## Modern descendants of the Skinner box The chamber remains standard laboratory equipment, though it rarely looks like Skinner's. - **Touchscreen and home-cage chambers.** The lever becomes a screen a rat or mouse touches with its nose, so tasks used to test memory and attention in human patients can be run in animals. In home-cage systems the operandum lives inside the animals' enclosure; they identify themselves electronically and work when they choose, around the clock, without handling stress. - **Drug self-administration.** Since James Weeks fitted rats with an intravenous line in 1962 so that a lever press delivered a dose of morphine, the operant chamber has been the principal animal model of addiction: the drugs people abuse are the ones animals will work for.[12] - **Brain stimulation and optogenetics.** In 1954 James Olds and Peter Milner showed that a rat would press a lever for a brief pulse of current to a region of its own brain — the first direct evidence that reinforcement has a physical address.[13] In the current version, the lever fires a laser that switches on a genetically targeted set of neurons, so the reinforcing effect of specific dopamine cells can be tested press by press. [The neuroscience of reinforcement ›](https://operantconditioning.com/neuroscience/) ### The human "Skinner box" The phrase now describes any product engineered to keep people responding: slot machines, infinite feeds, loot boxes. The analogy has substance. Skinner himself pointed to the gambling device as a variable-ratio schedule, the arrangement that produces the highest and most persistent rates in the laboratory, and the same schedule runs the pull-to-refresh feed.[14] > **A caution about the metaphor** > > A chamber isolates one response and one consequence under an experimenter's complete control. An app competes with every other reinforcer in your life, and you can walk away in a way a rat cannot. The metaphor is useful for naming the schedule and misleading when it suggests people are helpless: the same analysis says what to do — change the cue, the schedule, or the cost of responding. [Operant conditioning in technology and product design ›](https://operantconditioning.com/applications/#technology) ## Criticisms and limitations of the Skinner box **Artificiality.** The chamber studies a hungry animal, a bare environment, and a single arbitrary response, and critics have argued from the start that this tells us about rats in boxes rather than organisms in the world. Two findings from inside the operant tradition gave the objection teeth. Keller and Marian Breland, who left Skinner's laboratory to train animals commercially, reported that trained behavior drifted toward species-typical patterns — raccoons "washing" the coins they were supposed to deposit — whatever the contingencies said.[15] And the pigeon's key peck, the chamber's signature response, turned out to be partly Pavlovian: a pigeon will begin pecking a lit key that merely precedes food, even though pecking is not required and changes nothing.[16] The box does not create behavior from nothing; it selects from what the species brings. **Generalization to humans.** Noam Chomsky's charge was that terms precise in the pigeon laboratory become loose metaphors when stretched over human language and thought.[17] Skinner's later work conceded part of the point: humans follow rules and instructions, and rule-governed behavior can look quite different from behavior shaped directly by contingencies.[18] Human volunteers on laboratory schedules, who arrive with instructions and hypotheses of their own, often fail to show the animal patterns for that reason. **The reply.** The chamber is a method for isolating variables, as a physicist's frictionless plane is, and its findings have to be tested outside it. They were: reinforcement, extinction, schedule effects, shaping, and stimulus control replicated across species and then across classrooms, clinics, zoos, and workplaces, where [applied behavior analysis](https://operantconditioning.com/applications/) measures them against real behavior. Whether that vindicates the philosophy Skinner built on the box is a different, and still open, question. [The history of operant conditioning, including the cognitive critique ›](https://operantconditioning.com/history/) ## Key takeaways - The chamber was never the discovery; it was the instrument. It isolates one response and one consequence, with nothing else changing, so their relation can be measured for hours without an experimenter in the room. - Its measure is the rate of a free operant. Unlike Thorndike's puzzle box or a maze, where the experimenter starts each trial and gets one number, the animal decides when and how often to respond, and every response is recorded. - On a cumulative record, slope is rate, ticks are reinforcers, and flat stretches are pauses. Each schedule leaves a signature: the fixed-ratio break and run, the fixed-interval scallop, the steady line of variable schedules, and the burst-then-flatten of extinction. - A session is built in stages that each depend on the one before: deprivation, magazine training, shaping, continuous reinforcement then schedules, discrimination training, and extinction. Punishment, choice, and drug experiments are built on that base. - The "baby in a box" was an air crib, a climate-controlled bed with no lever, dispenser, or contingencies; no experiment was ever run in it. The chamber's real limitation is artificiality: it selects from what the species brings rather than creating behavior from nothing. ### Check yourself **On a cumulative record, the line goes flat for a stretch right after every reinforcer tick, then shoots up steeply. A student concludes the pellet must be punishing the rat, since pressing stops each time one arrives. What is wrong with that reading?** The flat stretch is the post-reinforcement pause, the first half of the fixed-ratio "break and run" pattern, and the steep run that follows shows pressing is as strong as ever. A behavior that has actually stopped draws a line that stays horizontal; a pause followed by a burst at full speed is the schedule's signature, not evidence of punishment. **Thorndike's cat and Skinner's rat both learn to work a mechanism to get food. Why did Skinner treat his chamber as a different kind of measurement?** The puzzle box is a discrete-trial method: the experimenter starts each trial and gets one number, the time to escape. In the chamber the lever is always there, so the animal decides when and how often to respond, and the measure is the rate of a free operant, recorded continuously and read from the slope of the record. **A pigeon begins pecking a lit key that comes on just before free food, even though pecking is not required and changes nothing. Does this show operant conditioning at work?** Not on its own. Because pecking has no consequence, the light merely preceding food is enough to produce it, which is why the key peck is described as partly Pavlovian. It is one of the findings behind the criticism that the box does not create behavior from nothing but selects from what the species brings. **An untrained rat is placed in a chamber with the lever already wired to the pellet dispenser. Why does nothing happen, and which two stages have to come first?** The rat has no reason to press and presses only by accident. It must first be mildly hungry, so that food functions as a reinforcer, and then magazine-trained, so that the dispenser's click becomes a conditioned reinforcer that can be delivered from across the chamber. Only then can shaping reinforce closer and closer approximations to a press. **Explain it to a friend.** Explain what a Skinner box measures that Thorndike's puzzle box could not, without using the words "rate" or "trial." ## Frequently asked questions **What is a Skinner box?** An enclosed apparatus, formally called an operant conditioning chamber, in which an animal can make a simple response — pressing a lever or pecking a lit key — that equipment records automatically and can follow with a consequence such as food. B. F. Skinner built it in the early 1930s to measure the rate of freely emitted behavior and study how its consequences change it. **What did Skinner's box experiment show?** There was no single experiment but a research program. Its central findings: behavior followed by a reinforcer becomes more frequent and fades when the reinforcer stops (extinction); how often reinforcement is delivered — the schedule — controls the rate, pattern, and persistence of behavior; and behavior comes under the control of stimuli that signal when reinforcement is available. The 1948 "superstition" study showed that accidental reinforcement is enough to build a ritual. **Is the Skinner box cruel?** In the typical experiment the animal is kept mildly hungry, works for food in a daily session of an hour or so, and is otherwise housed and fed normally; the great majority of operant research uses positive reinforcement. Some experiments on punishment and avoidance use brief shock through the grid floor, and those are the ones most people object to; today all such work is reviewed by institutional animal-care committees. Given free food in the chamber, many animals still work the lever — a phenomenon called [contrafreeloading](https://operantconditioning.com/operant-vs-classical-conditioning/#contrafreeloading). **Are Skinner boxes still used?** Yes. Modern chambers with touchscreens, home-cage versions, and computer control are standard in neuroscience, pharmacology, and behavioral genetics, and drug self-administration in an operant chamber is the main animal model of addiction. The principles first found in the box also underlie applied behavior analysis and reinforcement-based animal training. **Did Skinner put his daughter in a Skinner box?** No. He built a climate-controlled crib, the air crib, in which his daughter Deborah slept as an infant. It was a bed, not an experiment, and had no lever, dispenser, or contingencies. The rumor that she was harmed by it, sued him, or died by suicide is false; she said so herself in a 2004 newspaper article. **What is the difference between a Skinner box and a puzzle box?** Thorndike's puzzle box (1898) was a discrete-trial method: the cat was put in, escaped by working a latch, and was put back for the next trial, with escape time as the measure. A Skinner box is a free-operant method: the animal stays in, responds whenever it likes, and the rate of responding is recorded continuously. The puzzle box showed that consequences strengthen behavior; the Skinner box made it possible to measure exactly how. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Skinner, B. F. (1956). A case history in scientific method. *American Psychologist, 11*(5), 221–233. 3. Skinner, B. F. (1979). *The Shaping of a Behaviorist*. Knopf. 4. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 5. Skinner, B. F. (1950). Are theories of learning necessary? *Psychological Review, 57*(4), 193–216. 6. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 7. Skinner, B. F. (1959). *Cumulative Record*. Appleton-Century-Crofts. 8. Skinner, B. F. (1960). Pigeons in a pelican. *American Psychologist, 15*(1), 28–37. 9. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 10. Skinner, B. F. (1945, October). Baby in a box. *Ladies' Home Journal*. 11. Buzan, D. S. (2004, March 12). I was not a lab rat. *The Guardian*. 12. Weeks, J. R. (1962). Experimental morphine addiction: Method for automatic intravenous injections in unrestrained rats. *Science, 138*(3537), 143–144. 13. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427. 14. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 15. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 16. Brown, P. L., & Jenkins, H. M. (1968). Auto-shaping of the pigeon's key-peck. *Journal of the Experimental Analysis of Behavior, 11*(1), 1–8. 17. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58. 18. Skinner, B. F. (1969). *Contingencies of Reinforcement: A Theoretical Analysis*. Appleton-Century-Crofts. ## Related - [B. F. Skinner](https://operantconditioning.com/bf-skinner/): The life, the work, the critics, and the myths of the man who built the box. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): The patterns the cumulative recorder revealed — with a live simulator. - [Shaping](https://operantconditioning.com/shaping/): How a rat that has never seen a lever learns to press it, one approximation at a time. --- # B. F. Skinner: Biography, the Skinner Box, and Operant Conditioning > B. F. Skinner (1904–1990) founded radical behaviorism and the science of operant conditioning. His life, the Skinner box, major works, critics, and the myths. - Source: https://operantconditioning.com/bf-skinner/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *People* The psychologist who turned the law of effect into a laboratory science, built the box that carries his name, and spent fifty years arguing that behavior is selected by its consequences. His life, his work, his critics — and the myths that still follow him. > **Who he was** > > **B. F. Skinner** (Burrhus Frederic Skinner, March 20, 1904 – August 18, 1990) was an American psychologist at Harvard University who founded **radical behaviorism** and the **experimental analysis of behavior** — the science of [operant conditioning](https://operantconditioning.com/). He coined the term "operant," invented the operant conditioning chamber (the "Skinner box") and the cumulative recorder, and argued that behavior is selected by its consequences in the way species are selected by their environments. > > Born in Susquehanna, Pennsylvania, he died of leukemia in Cambridge, Massachusetts, eight days after receiving the American Psychological Association's first citation for an outstanding lifetime contribution to psychology. A 2002 survey ranked him the most eminent psychologist of the twentieth century.[1] **In brief** - B. F. Skinner (1904–1990) founded radical behaviorism and the experimental analysis of behavior, arguing that behavior is selected by its consequences. - He invented the [operant conditioning chamber](https://operantconditioning.com/skinner-box/) and the cumulative recorder, which made the rate of freely emitted behavior a measurable quantity. - Radical behaviorism does not deny thoughts and feelings; it treats them as private behavior to be explained rather than as ultimate causes. ## Who was B. F. Skinner? Skinner grew up in a small railroad town in northeastern Pennsylvania, the son of a lawyer, and spent his boyhood building things — wagons, rafts, a perpetual-motion machine that did not work. He went to Hamilton College intending to become a writer, and after graduating with a degree in English in 1926 he tried. Robert Frost had read three of his short stories and sent an encouraging letter. A year in his parents' house produced almost nothing, and Skinner later concluded that he had failed as a writer because he had nothing important to say. He decided that literature described behavior and that science might explain it.[2] What pushed him toward psychology was reading: John B. Watson's *Behaviorism*, Bertrand Russell's essays on Watson, and Ivan Pavlov's *Conditioned Reflexes*, newly translated into English in 1927. He entered Harvard's graduate program in 1928 with no formal training in the subject, worked largely in the physiology laboratory of William Crozier, and set himself a schedule of study so strict that he later described it with some embarrassment. By 1931 he had a PhD; by 1938 he had a book that defined a new field.[3] The thread running through everything he did afterward is a single ambition: to make psychology a natural science of behavior, with the same standing as physics or biology, by finding orderly relationships between what an organism does and the conditions under which it does it — without appealing to a mind inside to do the explaining. ## B. F. Skinner's life and career: a timeline - **1904** — Born in Susquehanna, Pennsylvania March 20. His father, William, was a lawyer; his mother, Grace, gave him her maiden name, Burrhus. Friends and family called him Fred. - **1926** — Hamilton College, then the "dark year" Graduates with a BA in English. Spends a year at home in Scranton trying to write fiction, then a few months in Greenwich Village. Reads Watson, Pavlov, and Russell, and decides to become a psychologist.[2] - **1928–1931** — Harvard graduate school Builds his own apparatus — a series of increasingly automated boxes for studying rats' eating and lever-pressing — and, almost by accident, the cumulative recorder. Completes his PhD in 1931 with a thesis on the concept of the reflex.[3] - **1933–1936** — Harvard Society of Fellows One of the first Junior Fellows of the newly founded Society, which gives him three years to do research with no teaching duties. In 1935 he publishes the paper distinguishing two types of conditioned reflex that will lead to the word "operant."[4] - **1936** — University of Minnesota Takes his first faculty post. Marries Yvonne (Eve) Blue the same year. Daughter Julie is born in 1938; Deborah in 1944. - **1937** — "Operant" and "respondent" In a reply to the Polish physiologists Konorski and Miller, Skinner introduces the terms that separate behavior *emitted* and controlled by its consequences from reflexive behavior *elicited* by a prior stimulus.[5] - **1938** — *The Behavior of Organisms* His first book lays out the experimental analysis of behavior: rate of response as the basic datum, reinforcement, extinction, discrimination, and the first schedule effects, all from rats in his apparatus.[6] - **1940–1944** — Project Pigeon Trains pigeons to guide a missile by pecking at a target image projected inside its nose cone. The birds worked; the military, committed to radar and electronic guidance, canceled the project in 1944. Along the way, in 1943, Skinner and his collaborators discover [shaping](https://operantconditioning.com/shaping/) while teaching a pigeon to "bowl."[7][8] - **1945** — Indiana, the air crib, and radical behaviorism Becomes chair of psychology at Indiana University. Publishes "Baby in a Box" in *Ladies' Home Journal* describing the climate-controlled crib he built for Deborah, and "The Operational Analysis of Psychological Terms," the founding statement of radical behaviorism.[9][10] - **1948** — Back to Harvard; *Walden Two*; superstition Returns to Harvard, where he stays for the rest of his career. Publishes the utopian novel *Walden Two*, written in seven weeks in 1945, and the "superstition in the pigeon" experiment.[11] - **1953** — *Science and Human Behavior* and the teaching machine Extends operant analysis to thinking, self-control, government, religion, and education. In November, a visit to his daughter Deborah's fourth-grade arithmetic class convinces him that classrooms violate everything known about reinforcement, and he builds his first teaching machine within days.[12][13] - **1957** — Two big books *Schedules of Reinforcement*, with Charles Ferster, catalogs thousands of hours of cumulative records. *Verbal Behavior*, developed from his 1947 William James Lectures, treats language as operant behavior.[14][15] - **1958–1959** — A journal, a chair, and a famous review The *Journal of the Experimental Analysis of Behavior* is founded. Skinner is named Edgar Pierce Professor of Psychology and receives the APA's Distinguished Scientific Contribution Award. In 1959 Noam Chomsky publishes his review of *Verbal Behavior*.[16] - **1968–1971** — National Medal of Science; *Beyond Freedom and Dignity* Receives the National Medal of Science in 1968. In 1971 *Beyond Freedom and Dignity* becomes a bestseller and puts him on the cover of *Time*, arguing that "autonomous man" is a fiction and that a culture must design its own contingencies.[17] - **1974** — *About Behaviorism*; retirement Retires from Harvard as professor emeritus and publishes his clearest statement of what radical behaviorism does and does not claim, written to correct twenty common misreadings.[18] - **1976–1983** — The autobiography Publishes a three-volume life — *Particulars of My Life*, *The Shaping of a Behaviorist*, and *A Matter of Consequences* — written in the same plain, external style he used for everything else. - **1990** — Final address and death On August 10, at the APA convention in Boston, he accepts the first Citation for Outstanding Lifetime Contribution to Psychology and delivers a last talk, "Can Psychology Be a Science of Mind?" He dies of leukemia at home in Cambridge, Massachusetts, on August 18.[19] ## What is a Skinner box? A **Skinner box** — Skinner's own term was **operant conditioning chamber** — is an enclosed, sound-attenuated space in which an animal can make a simple, repeatable response that the apparatus records and, under programmed conditions, rewards. Its purpose is not to trap the animal but to isolate one behavior and one consequence so that the relationship between them can be measured precisely, for hours at a time, without an experimenter in the room. [A full tour of the box, part by part ›](https://operantconditioning.com/skinner-box/) ### How the operant conditioning chamber works - **The operandum.** For rats, a small lever that closes a switch when pressed. For pigeons, a translucent plastic disk — the "key" — mounted at head height and lit from behind; a light peck closes a switch and registers a response. - **The food magazine.** A dispenser that drops a 45-milligram pellet into a tray for rats, or raises a hopper of grain for a few seconds for pigeons. The sound of the magazine operating becomes a conditioned reinforcer, bridging the gap between the response and the food. - **Stimuli.** Lights, tones, and the color of the key light let the experimenter signal when responding will pay off — the basis of [stimulus control](https://operantconditioning.com/abc-model/) and the discriminative stimulus. - **Programming.** Originally electromechanical relays and timers, later computers, deliver reinforcement on whatever [schedule](https://operantconditioning.com/schedules-of-reinforcement/) the experiment requires and record every response with its time. The crucial feature is that the animal is free to respond at any time and at any rate. This **free-operant** method, as opposed to discrete trials in mazes and runways, is what made *rate of response* available as a measure — and rate turned out to be exquisitely sensitive to the conditions of reinforcement.[6] ### The cumulative recorder Skinner's second invention was the instrument that made the data visible. A strip of paper moves at a constant speed under a pen; each response steps the pen a small fixed distance upward, and a diagonal tick marks each reinforcement. The result is a **cumulative record** whose slope is the rate of responding: steep means fast, flat means the animal has stopped, and the shape of the curve between reinforcers reveals the pattern a schedule produces. The famous "scallop" of the fixed-interval schedule and the "break-and-run" of the fixed-ratio schedule were first seen as pen strokes on this paper.[14] ### Why Skinner disliked the name The nickname "Skinner box" came from Clark Hull's laboratory at Yale, not from Skinner, who never used it and objected to it — partly because it suggested a gadget rather than a method, and partly because the term became attached in the popular mind to something quite different, which is the next story.[3] ### The "baby in a box": the air crib and a persistent rumor In 1944, before his second daughter Deborah was born, Skinner built what he called a **baby tender** and later an **air crib**: an enclosed crib with a large safety-glass front, a stretched canvas floor, and filtered, warmed air, so that the baby could sleep and play in a diaper alone, without blankets or layers of clothing, at a controlled temperature. It was a crib, used for sleeping and unattended play; Deborah was taken out to be fed, held, and played with like any other infant. *Ladies' Home Journal* published his article about it in October 1945 under a headline he did not choose, "Baby in a Box," and a number of families later used commercial versions.[9] > **What is false** > > A rumor circulated for decades that Skinner raised his daughter in a "Skinner box" as an experiment, and that she became psychotic, sued her father, or died by suicide. None of it is true. Deborah Skinner Buzan is an artist living in Britain, has said she was a happy child with a loving father, and publicly answered the rumor in 2004 in an article titled "I was not a lab rat" after a popular book repeated it.[20] The air crib was a bed, not an operant chamber, and no experiments were run in it. ## Skinner's famous experiments Four experiments come up again and again, partly because they are vivid and partly because each one makes a point that still matters. ### "Superstition" in the pigeon (1948) Skinner put hungry pigeons in a chamber in which the food hopper appeared for a few seconds at regular intervals — every fifteen seconds — no matter what the bird was doing. Six of eight birds developed a distinctive ritual: one turned counterclockwise around the cage between feedings, another thrust its head repeatedly into an upper corner, a third made a "tossing" motion as if lifting an invisible bar with its head, two swung their heads and bodies like a pendulum, and one made brushing movements toward the floor. Skinner's explanation was **adventitious reinforcement**: whatever the bird happened to be doing when food arrived was strengthened, which made it more likely to be under way at the next delivery, which strengthened it again. Contingency in the world is not required; contiguity is enough.[11] The experiment is also a lesson in how science corrects itself. When Staddon and Simmelhag repeated it in 1971 and recorded behavior throughout each interval, they found that the "terminal" behaviors just before food were much the same across birds — mostly pecking near the hopper — and looked less like accidentally reinforced rituals than like responses induced by the periodic arrival of food itself. The rituals were real; the explanation was more complicated than Skinner's.[30] The glossary entry on [superstitious behavior](https://operantconditioning.com/glossary/#superstitious-behavior) gives the everyday version. ### Project Pigeon (1940–1944) During the Second World War Skinner trained pigeons to guide a missile. A lens in the nose cone projected the image of the target onto a screen; the pigeon, harnessed inside, pecked at the image, and the position of its pecks steered the missile. Three pigeons voted for reliability. With support first from General Mills and then, in 1943, from the government's Office of Scientific Research and Development, the birds learned to track a target through distortion, noise, and cold, and kept pecking steadily under conditions that unsettled human observers. The project was cancelled in 1944 — the committee could not take a pigeon-guided bomb seriously against electronic guidance — and was briefly revived by the Navy after the war as Project ORCON.[7] Its lasting product was accidental: in 1943, while teaching a pigeon to "bowl" a ball down a miniature alley, Skinner and his colleagues discovered how quickly behavior could be built by reinforcing successive approximations by hand — [shaping](https://operantconditioning.com/shaping/).[8] ### The ping-pong pigeons In a demonstration still shown in classrooms, two pigeons stand at either end of a small table and bat a ping-pong ball back and forth with their beaks. Each bird was first shaped to peck the ball, then to peck it toward the other end; once both could rally, food was delivered to the bird that got the ball past its opponent. Skinner presented it as a "synthetic social relation": competition assembled from individual contingencies, with no need to assume the birds understood the game.[31] The companion demonstration in the same paper built cooperation the same way — two pigeons reinforced only when they pecked matching keys at nearly the same moment. ### Teaching machines (1954–1958) On a visit to his daughter Deborah's fourth-grade arithmetic class in November 1953, Skinner watched some children sit idle after finishing a problem sheet while others struggled, all of them waiting a day or more to learn whether their answers were right. By his own account he built a prototype teaching machine within days. The machine presented material in small steps, required the student to compose an answer rather than pick one, showed the correct answer immediately, and let each student move at his or her own pace. He described the principles at a 1954 conference and set out the case for "teaching machines" in *Science* in 1958, crediting Sidney Pressey's self-scoring devices of the 1920s as a precursor.[13][22] The machines themselves went out of fashion within a decade; the principles — small steps, active responding, immediate feedback, self-pacing — reappear in programmed instruction and in most learning software written since. ## Skinner's key contributions to psychology | Year | Contribution | Why it matters | | --- | --- | --- | | 1930–1938 | Operant chamber, cumulative recorder, rate as the datum | Made moment-to-moment behavior measurable and reproducible; the method the entire field still uses.[6] | | 1935–1937 | Operant vs. respondent behavior | Separated behavior controlled by consequences from reflexes elicited by stimuli; "operant conditioning" gets its name.[4][5] | | 1938 | *The Behavior of Organisms* | First systematic account of reinforcement, extinction, discrimination, and schedule effects in the free-operant method.[6] | | 1940–1944 | Project Pigeon | Demonstrated fine stimulus control in pigeons; produced the discovery of shaping; a famous example of practical operant technology ahead of its time.[7] | | 1943; 1951 | Shaping by successive approximation | Showed how new behavior can be built by reinforcing closer and closer approximations; explained to the public in "How to Teach Animals."[8][21] | | 1945 | Radical behaviorism | Brought private events — thinking, feeling — inside the analysis as behavior, rather than excluding them as Watson had.[10] | | 1948 | "Superstition" in the pigeon | Showed that accidental reinforcement produces and maintains behavior; contingency, not the animal's understanding, does the work.[11] | | 1948 | *Walden Two* | A novel imagining a community designed on behavioral principles; inspired real intentional communities and decades of argument. | | 1953 | *Science and Human Behavior* | The textbook that applied operant analysis to self-control, thinking, social behavior, and institutions.[12] | | 1954; 1958 | Teaching machines and programmed instruction | Small steps, active responding, immediate feedback, self-pacing — principles that survive in modern learning software.[13][22] | | 1957 | *Schedules of Reinforcement* (with Ferster) | Established that how often behavior is reinforced controls its rate, pattern, and persistence.[14] | | 1957 | *Verbal Behavior* | A functional analysis of language (mands, tacts, intraverbals) that underlies much of modern language intervention in applied behavior analysis.[15] | | 1971 | *Beyond Freedom and Dignity* | Argued that behavior is always controlled and that cultures should design their contingencies deliberately; his most controversial book.[17] | | 1974 | *About Behaviorism* | His definitive answer to critics and misreadings of the philosophy behind the science.[18] | | 1981 | "Selection by Consequences" | Placed operant conditioning alongside natural selection and cultural evolution as a third kind of selection.[23] | ## What is radical behaviorism? **Radical behaviorism** is Skinner's philosophy of the science of behavior. "Radical" means thoroughgoing, not extreme: where Watson's *methodological* behaviorism ruled private experience out of psychology because it could not be observed by a second person, Skinner ruled it *in*. Thinking, feeling, imagining, and seeing with your eyes closed are, in his account, behavior — covert, but subject to the same variables as any other behavior, and open to study through the verbal reports a community teaches each of us to make about them.[10][18] What radical behaviorism rejects is not the *existence* of mental life but its use as an *explanation*. To say a student studies "because she is motivated" or a rat presses "because it expects food" is, for Skinner, to stop the analysis one step too soon: the motivation and the expectation must themselves be explained, and when they are, the explanation turns out to be a history of consequences and a present set of conditions. He called inner causes that merely restate the behavior they are supposed to explain "explanatory fictions."[12] Three commitments follow: - **Behavior is a subject matter in its own right**, not a symptom of something happening elsewhere (the mind, the brain). Neuroscience is welcome; it fills in mechanism, but it does not replace the functional relations. - **The causes of behavior lie in the organism's genetic endowment, its history of reinforcement, and its current environment.** Feelings are real but are collateral products of those same contingencies — we feel the effects of the causes, not the causes themselves. - **Selection by consequences.** Just as natural selection explains the appearance of design in organisms without a designer, reinforcement explains the appearance of purpose in behavior without purpose as a prior cause.[23] > **The most common misreading** > > Skinner did not deny that people think and feel, and he did not say the organism is "empty." The first chapter of *About Behaviorism* opens with a list of twenty things commonly said about behaviorism — that it ignores consciousness, that it treats people as robots, that it cannot account for creativity — and calls every one of them wrong.[18] The dispute is about where the causes are, not about what exists. ## Controversies and critiques of Skinner ### Chomsky's review of *Verbal Behavior* The most consequential criticism Skinner ever received was Noam Chomsky's 1959 review of *Verbal Behavior*. Chomsky argued that terms like "stimulus," "response," and "reinforcement," precise in the pigeon laboratory, became vacuous metaphors when stretched over human language, and that children acquire grammar far too fast, from far too little input, for reinforcement to be the mechanism.[16] The review is often credited with helping launch the cognitive revolution. Skinner never published a reply, and later said he had read only part of it. The detailed answer came from Kenneth MacCorquodale in 1970, who argued that Chomsky had reviewed Hull-style stimulus–response psychology rather than Skinner's book, had misdescribed reinforcement as a theory of drive reduction, and had criticized the book for failing to be a theory of grammar when it set out to be a functional analysis of why speakers say what they say.[24] Both essays remain worth reading; on the specific question of syntax, most linguists sided with Chomsky, while *Verbal Behavior*'s functional categories went on to become the basis of a large applied literature on teaching language to children with autism. ### Free will and *Beyond Freedom and Dignity* Skinner's 1971 book argued that the idea of an autonomous inner agent who freely chooses is a pre-scientific holdover, and that the real question is not whether behavior will be controlled — it always is — but by what and for whose benefit.[17] The reaction was intense. Chomsky reviewed it, too, under the title "The Case Against B. F. Skinner," and philosophers objected that a science of behavior cannot settle a metaphysical question by fiat. Skinner's answer was that the traditional view had not produced a technology for solving problems like overpopulation, pollution, or war, and that dignity, properly understood, is the credit we give people when we cannot see what controls them. > Give me the specifications, and I'll give you the man! > > *— Frazier, the fictional founder of the community in Skinner's *Walden Two*, 1948* ### The ethics of behavioral control Lines like Frazier's are why critics heard totalitarianism in *Walden Two*. Skinner's position was that a science of behavior is a tool, that the alternative to designed contingencies is accidental ones, and that the safeguard against abuse is **countercontrol**: the controlled must be able to control the controller. Whether that is enough is a live question, and it is why the ethics codes of applied behavior analysis put consent, least-restrictive procedures, and the client's own goals at the center. [How ABA is practiced today ›](https://operantconditioning.com/applications/#aba) ### Skinner on punishment Skinner was, throughout his career, an opponent of punishment as a method of control. His view rested on early evidence — his own 1938 experiments and W. K. Estes's 1944 dissertation — that punishment only temporarily suppressed responding while the reinforced behavior remained intact underneath, together with its side effects: fear, aggression, escape, and avoidance.[6][25] Later work by Azrin and Holz showed that sufficiently immediate and intense punishment *can* produce lasting suppression, so Skinner's empirical claim was too strong.[26] His practical conclusion — build behavior with [reinforcement](https://operantconditioning.com/positive-reinforcement/), and treat [punishment](https://operantconditioning.com/positive-punishment/) as a last resort — is nonetheless where the applied field ended up. ## Skinner's legacy Few psychologists have left as many working descendants: - **Applied behavior analysis (ABA).** A licensed profession in most U.S. states, built directly on Skinner's principles and his students' work, used in autism intervention, developmental disabilities, brain-injury rehabilitation, and behavioral medicine. - **Education.** Programmed instruction, mastery learning, precision teaching, and school-wide Positive Behavioral Interventions and Supports (PBIS) all trace to his 1950s work on teaching. [Operant conditioning in education ›](https://operantconditioning.com/applications/#education) - **Animal training.** Shaping, conditioned reinforcers, and the clicker were carried from his laboratory into the training world by Keller and Marian Breland and later Karen Pryor. [Reinforcement-based dog training ›](https://operantconditioning.com/dog-training/) - **Behavioral economics.** The matching law, delay discounting, and the quantitative analysis of choice grew out of the operant laboratory, largely through his student Richard Herrnstein. - **Reinforcement learning in AI.** The branch of machine learning behind game-playing programs, robotics, and the fine-tuning of large language models is a mathematical formalization of learning from consequences, and its founders cite the operant tradition explicitly.[27] - **Self-management.** Skinner's own late work on managing one's own behavior in old age, and the whole tradition of [habit formation through antecedents and consequences](https://operantconditioning.com/habits/), applies the three-term contingency to oneself. His works are kept in print by the B. F. Skinner Foundation, led by his daughter Julie S. Vargas. [The full history of operant conditioning, from Thorndike to reinforcement learning ›](https://operantconditioning.com/history/) ## Myths about B. F. Skinner | Myth | What is actually true | | --- | --- | | He raised his daughter in a Skinner box. | He built a climate-controlled crib. She slept and played in it as an infant, was raised normally, and is alive; she has publicly refuted the story.[20] | | He denied that thoughts and feelings exist. | Radical behaviorism treats thoughts and feelings as private behavior, real and worth studying. What he denied is that they are the ultimate *causes* of what we do.[18] | | He believed people are blank slates shaped entirely by environment. | He wrote explicitly that behavior is the joint product of genetic endowment, individual history, and current setting, and devoted a 1966 paper to the role of evolution.[28] | | He said "give me a dozen healthy infants" and he could make them into anything. | That is John B. Watson, writing in 1924, decades before Skinner's career — and Watson himself added that he was going beyond the evidence.[29] | | His main method was shocking animals. | His research was overwhelmingly about food reinforcement in mildly hungry rats and pigeons. He argued against punishment for fifty years; the definitive punishment research was done by others.[26] | | Chomsky "destroyed" behaviorism in 1959. | The review was influential on the study of language, but the experimental analysis of behavior grew steadily afterward, founding its applied journal in 1968 and a profession in the decades since. | ## Selected bibliography - *The Behavior of Organisms: An Experimental Analysis* (1938) - *Walden Two* (1948) - *Science and Human Behavior* (1953) - *Schedules of Reinforcement*, with C. B. Ferster (1957) - *Verbal Behavior* (1957) - *Cumulative Record* (1959; expanded editions 1961, 1972) — collected papers - *The Technology of Teaching* (1968) - *Contingencies of Reinforcement: A Theoretical Analysis* (1969) - *Beyond Freedom and Dignity* (1971) - *About Behaviorism* (1974) - *Particulars of My Life* (1976), *The Shaping of a Behaviorist* (1979), *A Matter of Consequences* (1983) — autobiography - *Enjoy Old Age: A Program of Self-Management*, with M. E. Vaughan (1983) None of Skinner's books is in the public domain, so none can be republished here. The books he built on — Thorndike, Morgan, James, Yerkes, Pfungst, Darwin — are, and they are in the [library](https://operantconditioning.com/library/) in full text. ## Key takeaways - Skinner's single ambition was a natural science of behavior: orderly relations between what an organism does and the conditions under which it does it, without appealing to a mind inside to do the explaining. - The operant chamber and the cumulative recorder made rate of response the basic datum. Because the animal is free to respond at any time, rate became available as a measure, and it proved sensitive to the conditions of reinforcement. - Radical behaviorism rules private experience in, not out: thinking and feeling are covert behavior, subject to the same variables as any other behavior. What it rejects is using mental states as explanations, which Skinner called explanatory fictions. - Chomsky's 1959 review charged that laboratory terms become metaphors when stretched over language, and most linguists sided with him on syntax; Azrin and Holz showed that immediate, intense punishment can produce lasting suppression, so Skinner's empirical claim against punishment was too strong. His practical conclusion, build behavior with reinforcement and treat punishment as a last resort, is where the applied field ended up. - The "baby in a box" was a climate-controlled crib, not an experiment. Skinner did not deny that thoughts and feelings exist, did not claim people are blank slates, and the "dozen healthy infants" line belongs to Watson. ### Check yourself **A classmate says Skinner, like Watson, held that psychology should ignore thoughts and feelings because a second person cannot observe them. Is that right?** No. That is Watson's methodological behaviorism, which ruled private experience out. Skinner's radical behaviorism ruled it in: thinking and feeling are covert behavior, real and worth studying; what he rejected was treating them as the ultimate causes of what a person does. **In the 1948 experiment, pigeons were fed every fifteen seconds no matter what they did, yet six of eight developed rituals. A student concludes the birds had figured out that the ritual caused the food. What was Skinner's explanation, and what did later work add?** Skinner's explanation was adventitious reinforcement: whatever the bird happened to be doing when food arrived was strengthened, making it more likely to be under way at the next delivery. No understanding is required; contiguity is enough. Staddon and Simmelhag later found that the behaviors just before food were much the same across birds and looked induced by the periodic arrival of food itself, so the rituals were real but the explanation was more complicated. **Skinner argued for fifty years that punishment only temporarily suppresses behavior. Did the evidence support that claim?** Only partly. His view rested on his own 1938 experiments and Estes's 1944 dissertation, but Azrin and Holz later showed that sufficiently immediate and intense punishment can produce lasting suppression, so the empirical claim was too strong. His practical conclusion still stands in applied work: build behavior with reinforcement and treat punishment as a last resort. **Why did Skinner treat the free-operant method as a breakthrough rather than just a new gadget?** In mazes and runways the experimenter runs discrete trials. In the chamber the animal is free to respond at any time and at any rate, which made rate of response available as a measure, and rate turned out to be exquisitely sensitive to the conditions of reinforcement. The cumulative recorder then made that rate visible as the slope of a line. **Explain it to a friend.** Explain the difference between Watson's behaviorism and Skinner's radical behaviorism in two sentences a twelve-year-old would follow. ## Frequently asked questions **What is B. F. Skinner best known for?** Operant conditioning — the science of how consequences change behavior. He named it, invented the apparatus (the operant conditioning chamber, or Skinner box) and the cumulative recorder used to study it, discovered shaping and schedules of reinforcement, and founded radical behaviorism, the philosophy behind the science. **What does B. F. stand for?** Burrhus Frederic. Burrhus was his mother's maiden name. Family and colleagues called him Fred. **What is a Skinner box and what was it used for?** A Skinner box, or operant conditioning chamber, is an enclosed apparatus in which an animal can make a simple response — pressing a lever, pecking a lit key — that the equipment records and can follow with a consequence such as food. It was used to study how reinforcement, extinction, schedules, and discriminative stimuli control the rate of behavior, with the results drawn on a cumulative recorder. **Did B. F. Skinner really raise his daughter in a box?** No. He built an enclosed, temperature-controlled crib called the air crib in which his daughter Deborah slept as an infant. It was a bed, not an experiment. The rumor that she was damaged by it, sued him, or died by suicide is false; she has said so publicly and is alive and well. **What is Skinner's theory of operant conditioning?** That behavior is selected by its consequences. Responses followed by reinforcement become more frequent; those followed by punishment or by nothing at all become less frequent. The unit of analysis is the three-term contingency — antecedent, behavior, consequence — and the basic measure is the rate of responding. [Read the complete guide to operant conditioning](https://operantconditioning.com/). **What is the difference between Skinner and Pavlov?** Pavlov studied respondent (classical) conditioning, in which a reflex such as salivation comes to be elicited by a stimulus that precedes it. Skinner studied operant conditioning, in which voluntary behavior is strengthened or weakened by the consequences that follow it. Skinner coined "respondent" and "operant" in 1937 to mark the distinction. [Full comparison](https://operantconditioning.com/operant-vs-classical-conditioning/). **What is radical behaviorism in simple terms?** The view that thoughts and feelings are real but are themselves behavior to be explained, not the ultimate explanation of what we do. The causes of behavior lie in genetics, learning history, and the present environment. It differs from Watson's behaviorism, which excluded private experience from science altogether. **When and how did B. F. Skinner die?** He died of leukemia on August 18, 1990, at his home in Cambridge, Massachusetts, at age 86. Eight days earlier he had received the American Psychological Association's first Citation for Outstanding Lifetime Contribution to Psychology and given his final public talk. ## References 1. Haggbloom, S. J., Warnick, R., Warnick, J. E., Jones, V. K., Yarbrough, G. L., Russell, T. M., Borecky, C. M., McGahhey, R., Powell, J. L., Beavers, J., & Monte, E. (2002). The 100 most eminent psychologists of the 20th century. *Review of General Psychology, 6*(2), 139–152. 2. Skinner, B. F. (1976). *Particulars of My Life*. Knopf. 3. Skinner, B. F. (1979). *The Shaping of a Behaviorist*. Knopf. See also Skinner, B. F. (1956). A case history in scientific method. *American Psychologist, 11*(5), 221–233. 4. Skinner, B. F. (1935). Two types of conditioned reflex and a pseudo type. *Journal of General Psychology, 12*, 66–77. 5. Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279. 6. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 7. Skinner, B. F. (1960). Pigeons in a pelican. *American Psychologist, 15*(1), 28–37. 8. Peterson, G. B. (2004). A day of great illumination: B. F. Skinner's discovery of shaping. *Journal of the Experimental Analysis of Behavior, 82*(3), 317–328. 9. Skinner, B. F. (1945, October). Baby in a box. *Ladies' Home Journal*. 10. Skinner, B. F. (1945). The operational analysis of psychological terms. *Psychological Review, 52*, 270–277. 11. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 12. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 13. Skinner, B. F. (1954). The science of learning and the art of teaching. *Harvard Educational Review, 24*(2), 86–97. 14. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 15. Skinner, B. F. (1957). *Verbal Behavior*. Appleton-Century-Crofts. 16. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58. 17. Skinner, B. F. (1971). *Beyond Freedom and Dignity*. Knopf. 18. Skinner, B. F. (1974). *About Behaviorism*. Knopf. 19. Skinner, B. F. (1990). Can psychology be a science of mind? *American Psychologist, 45*(11), 1206–1210. 20. Buzan, D. S. (2004). I was not a lab rat. *The Guardian*. 21. Skinner, B. F. (1951). How to teach animals. *Scientific American, 185*(6), 26–29. 22. Skinner, B. F. (1958). Teaching machines. *Science, 128*(3330), 969–977. 23. Skinner, B. F. (1981). Selection by consequences. *Science, 213*(4507), 501–504. 24. MacCorquodale, K. (1970). On Chomsky's review of Skinner's *Verbal Behavior*. *Journal of the Experimental Analysis of Behavior, 13*(1), 83–99. 25. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 26. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 27. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. 28. Skinner, B. F. (1966). The phylogeny and ontogeny of behavior. *Science, 153*, 1205–1213. 29. Watson, J. B. (1930). *Behaviorism* (rev. ed.). Norton. (First edition 1924.) 30. Staddon, J. E. R., & Simmelhag, V. L. (1971). The "superstition" experiment: A reexamination of its implications for the principles of adaptive behavior. *Psychological Review, 78*(1), 3–43. 31. Skinner, B. F. (1962). Two "synthetic social relations." *Journal of the Experimental Analysis of Behavior, 5*(4), 531–533. ## Related - [History of operant conditioning](https://operantconditioning.com/history/): From Thorndike's cats to dopamine neurons and reinforcement learning. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): The cumulative records Skinner and Ferster drew — in a live simulator. - [Shaping](https://operantconditioning.com/shaping/): The technique Skinner discovered while teaching a pigeon to bowl. --- # The History of Operant Conditioning: From Thorndike's Cats to Reinforcement Learning > Who discovered operant conditioning? Thorndike found the law of effect in 1898; Skinner named and systematized it in 1937–38. The timeline, Aristotle to AI. - Source: https://operantconditioning.com/history/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *History* Who discovered operant conditioning, where behaviorism came from, what the cognitive revolution did and did not overturn, and how a law about cats in boxes became a profession, a finding about dopamine, and the algorithm behind modern AI. > **Who discovered operant conditioning?** > > The principle was discovered by **Edward L. Thorndike**, whose puzzle-box experiments with cats (1898) produced the **law of effect**: responses followed by a satisfying consequence are strengthened, and those followed by discomfort are weakened. **B. F. Skinner** named it *operant conditioning* in 1937, invented the methods for studying it, and built the science around it beginning with *The Behavior of Organisms* (1938).[4][11] > > So the honest answer is two names: Thorndike found the law; Skinner turned it into a field. Behind both stand a century of earlier thinking about how animals learn from the results of their actions. **In brief** - Thorndike found the law of effect with his 1898 puzzle-box cats; [Skinner](https://operantconditioning.com/bf-skinner/) named operant conditioning in 1937 and built the science around it. - The cognitive revolution bounded the law of effect rather than overturning it: reinforcement still changes behavior, within limits set by evolution. - The tradition continues as applied behavior analysis, in the dopamine prediction-error signal, and in reinforcement learning, a formalization of learning from consequences. ## Who discovered operant conditioning: Thorndike or Skinner? Thorndike was the first to show experimentally, with learning curves rather than anecdotes, that the consequences of a response change its future probability. Skinner was the first to treat that fact as the foundation of a science: he defined the **operant** as a class of behavior controlled by its consequences, distinguished it from Pavlov's reflexes, built the free-operant chamber and cumulative recorder, and discovered the phenomena — [schedules](https://operantconditioning.com/schedules-of-reinforcement/), [shaping](https://operantconditioning.com/shaping/), [extinction](https://operantconditioning.com/extinction/), stimulus control — that make up the subject today. Thorndike's own term, *instrumental learning*, survives as "instrumental conditioning," a synonym still used in the discrete-trial tradition. ## Timeline: the history of operant conditioning - **c. 350 BCE** — Aristotle on habit In the *Nicomachean Ethics* Aristotle argues that character is built by repetition — “we become just by doing just acts” — and that the pleasure or pain attached to an act is what decides whether it is repeated. It is not the law of effect: there is no experiment, no measurement of response frequency, and no functional definition of a reinforcer. But the intuition that consequences shape conduct is about as old as written philosophy, and Thorndike's contribution was to make it testable rather than to think of it.[48] - **1855–1882** — Bain and Romanes The Scottish philosopher Alexander Bain describes how spontaneous movements that happen to produce pleasure are retained and those producing pain dropped — a law of effect before the name. George Romanes collects anecdotes of animal cleverness in *Animal Intelligence* (1882), a method soon discredited.[1][2] - **1894** — Morgan's canon C. Lloyd Morgan rules that no animal behavior should be explained by a higher mental faculty if a lower process will do. Comparative psychology adopts the rule, and the search for simple learning processes begins.[3] - **1898** — Thorndike's puzzle boxes For his Columbia dissertation, Edward Thorndike puts hungry cats in latched boxes with food outside and times their escapes. Escape times fall gradually, with no sudden insight; he proposes that satisfying consequences "stamp in" the connection between situation and response.[4] - **1911–1932** — The law of effect, stated and then truncated *Animal Intelligence* (1911) names the law of effect. Two decades later, on the basis of human learning experiments, Thorndike concludes that "annoyers" do not weaken connections the way "satisfiers" strengthen them, and drops the negative half of the law — the first evidence that punishment and reinforcement are not mirror images.[5][6] - **1913** — Watson's behaviorist manifesto John B. Watson declares that psychology's subject matter is behavior, not consciousness, and its goal prediction and control. Behaviorism becomes a movement — though its early learning theory rests on Pavlov's reflex rather than Thorndike's law.[7] - **1927** — Pavlov in English Ivan Pavlov's *Conditioned Reflexes* is translated, giving American psychology the terms "reinforcement," "extinction," "generalization," and "discrimination" — all of which Skinner will borrow for a different kind of learning.[8] - **1928** — Konorski and Miller's "type II" reflex Two Polish physiologists, Jerzy Konorski and Stefan Miller, passively flex a dog's leg and then feed it; the dog begins flexing on its own. They call this a second type of conditioned reflex, distinct from Pavlov's, and later argue the point with Skinner in print.[9][10] - **1930–1938** — Skinner: the chamber, the operant, the book At Harvard, [B. F. Skinner](https://operantconditioning.com/bf-skinner/) builds the [operant chamber](https://operantconditioning.com/skinner-box/) and cumulative recorder and defines behavior by its function rather than its form (1935); at Minnesota, where he takes his first faculty post in 1936, he coins "operant" and "respondent" (1937) and publishes *The Behavior of Organisms* (1938).[10][11] - **1943–1944** — Hull's system; Estes on punishment Clark Hull's *Principles of Behavior* offers a rival, mathematical behaviorism in which reinforcement works by reducing a biological drive. Skinner's student W. K. Estes shows that punishment suppresses lever-pressing only temporarily, evidence Skinner will cite against punishment for the rest of his life.[12][13] - **1948** — Cognitive maps and superstitious pigeons Edward Tolman argues from latent-learning and maze studies that rats form "cognitive maps," not chains of responses. The same year Skinner shows that pigeons fed on a timer develop idiosyncratic rituals — accidental contingencies shape behavior with no understanding required.[14][15] - **1950** — A textbook and a manifesto Fred Keller and William Schoenfeld's *Principles of Psychology* teaches introductory psychology entirely from reinforcement principles. Skinner's "Are Theories of Learning Necessary?" argues for describing functional relations rather than inventing internal mechanisms.[16][17] - **1957–1958** — Schedules; a journal Ferster and Skinner publish *Schedules of Reinforcement*. The *Journal of the Experimental Analysis of Behavior* is founded in 1958, giving the field its own outlet and its own methods: single organisms, steady states, within-subject replication.[18][19] - **1959–1961** — Premack, Herrnstein, the Brelands, Chomsky David Premack shows that a more probable behavior can reinforce a less probable one. Richard Herrnstein finds that pigeons match response ratios to reinforcement ratios — the [matching law](https://operantconditioning.com/matching-law/). The Brelands report trained animals drifting toward instinct — "instinctive drift." Noam Chomsky's review of *Verbal Behavior* opens the cognitive revolution.[20][21][22][23] - **1968** — Applied behavior analysis is born Teodoro Ayllon and Nathan Azrin publish *The Token Economy*, from their work on a psychiatric ward. The *Journal of Applied Behavior Analysis* launches, and Baer, Wolf, and Risley's paper in its first issue defines what makes an analysis "applied."[24][25] - **1972** — The Rescorla–Wagner model A model of classical conditioning in which learning is driven by *surprise* — the gap between what was predicted and what occurred. Though about Pavlovian learning, it reshapes all of learning theory and, decades later, turns out to describe what dopamine neurons compute.[26] - **1982–1987** — Functional analysis, momentum, and Lovaas Brian Iwata and colleagues show that self-injury can be experimentally traced to its reinforcer, making assessment of *function* the first step in treatment. John Nevin's behavioral momentum research shows that persistence depends on the rate of reinforcement in a context. Ivar Lovaas reports that intensive early intervention brought nine of nineteen autistic children to normal educational functioning.[27][28][29] - **1997–1998** — Dopamine, reinforcement learning, and a credential Schultz, Dayan, and Montague show that midbrain dopamine neurons signal a reward prediction error. Sutton and Barto publish *Reinforcement Learning: An Introduction*. The Behavior Analyst Certification Board is founded, and behavior analysis becomes a credentialed profession.[30][31] - **2015–2025** — Deep reinforcement learning Reinforcement learning combined with deep neural networks masters Atari games from pixels and then Go; the same family of methods is used to fine-tune large language models from human feedback. In 2025 Sutton and Barto are awarded the ACM Turing Award (for 2024) for founding the field.[32][33] > **Read the sources.** The books this timeline begins with are in the public domain and republished in full in the [library](https://operantconditioning.com/library/): Thorndike's [*Animal Intelligence*](https://operantconditioning.com/library/thorndike-animal-intelligence/) (with the 1898 monograph), Morgan's [*Animal Behaviour*](https://operantconditioning.com/library/morgan-animal-behaviour/), James's [*Principles of Psychology*](https://operantconditioning.com/library/james-principles-of-psychology-vol-1/), Yerkes's [*Dancing Mouse*](https://operantconditioning.com/library/yerkes-the-dancing-mouse/), Pfungst's [*Clever Hans*](https://operantconditioning.com/library/pfungst-clever-hans/), Romanes and Darwin. Skinner's own books are still under copyright. ## From the law of effect to the operant: what Skinner changed It is tempting to read Skinner as Thorndike with better equipment. The differences are deeper, and they explain why the field is called the *experimental analysis of behavior* rather than the study of trial-and-error learning. | Aspect | Thorndike (1898–1932) | Skinner (1938 onward) | | --- | --- | --- | | **Basic measure** | Time to escape on each trial | **Rate of response** — continuous, sensitive, and the measure that made schedule effects visible[11] | | **Method** | Discrete trials; experimenter resets the box | **Free operant**; the animal responds whenever it likes and the cumulative recorder draws rate as a slope | | **Unit of behavior** | A specific movement connected to a situation | The **operant**: a class of responses defined by their common consequence, whatever their form (left paw, right paw, nose)[34] | | **Why reinforcement works** | "Satisfiers" stamp in connections (Hull later: drive reduction) | No theory required: a reinforcer is whatever strengthens the behavior it follows; deprivation is an operation you perform and measure[17] | | **Research design** | Learning curves across trials, later group comparisons | **Single organisms** studied to steady state, with effects shown by reversal in the same animal — codified by Sidman in 1960[35] | | **Scope** | Animal learning and education | All behavior, including thinking, language, and culture — via the three-term contingency of [antecedent, behavior, consequence](https://operantconditioning.com/abc-model/) | Add the distinction from Pavlovian conditioning and you have the framework every page on this site uses. [Operant vs. classical conditioning ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## The cognitive revolution: what it overturned and what it didn't Between roughly 1956 and 1970 psychology's center of gravity moved from behavior to mind. George Miller's paper on the limits of short-term memory (1956), Chomsky's review of *Verbal Behavior* (1959), and Ulric Neisser's *Cognitive Psychology* (1967) are the usual markers.[23][36] Within learning research itself, a series of findings showed that consequences were not the whole story: - **Learning without performance.** Tolman's rats explored mazes without reward and then, once food appeared, ran them almost immediately — they had learned the layout all along.[14] - **Learning by observation.** Albert Bandura's children imitated an adult's aggression toward an inflatable doll without ever being reinforced for it.[37] - **Biological constraints.** The Brelands' raccoons "washed" the coins they were supposed to deposit; John Garcia's rats associated taste with nausea over hours but could not associate a light with nausea at all. Organisms come prepared to learn some things and not others.[22][38] - **Contingency, not contiguity.** Robert Rescorla showed that pairing alone does not produce conditioning; what matters is whether one event *predicts* another. The Rescorla–Wagner model made this quantitative.[26] > **What survived** > > None of these findings overturned the law of effect; they bounded it. Reinforcement still changes behavior, schedules still produce their signature patterns, extinction bursts still occur, and shaping and stimulus control still work — in rats, pigeons, dolphins, children, and adults. What changed is that operant conditioning is now understood as one powerful learning process among several, operating within limits set by evolution, and describable in the language of prediction error. Behavior analysts kept publishing in their own journals throughout; the "death of behaviorism" was mostly a change in who wrote the introductory textbook. ## Behavior analysis today The experimental tradition became a profession in the decades after 1968. **Applied behavior analysis** now has a certifying body (the BACB, founded in 1998), a credential (the Board Certified Behavior Analyst), licensure in most U.S. states, insurance coverage for autism services in every state, and a research literature spanning developmental disabilities, education, organizational management, addiction treatment, brain-injury rehabilitation, and animal welfare.[39] Its methods — [reinforcement](https://operantconditioning.com/positive-reinforcement/)-based teaching, functional assessment before treatment, single-subject designs with continuous measurement — descend directly from the laboratory. [How ABA is practiced ›](https://operantconditioning.com/applications/#aba) ### The neurodiversity critique An honest history has to note that ABA is also contested. Many autistic adults and advocates object to its early history (Lovaas's original 1960s program used aversives, including electric shock, a practice long since abandoned), to goals that aimed at making autistic children "indistinguishable" from peers, to suppressing harmless self-stimulatory behavior, and to intensive programs delivered without the child's assent.[40] Practitioners have responded that modern ABA is reinforcement-based, individualized, and increasingly assent-driven, and the field's own journals now publish on these concerns.[40] The debate is about goals, consent, and history, not about whether reinforcement changes behavior — which everyone in it agrees it does. ### Beyond the clinic Outside ABA, the operant tradition underlies reinforcement-based [animal training](https://operantconditioning.com/dog-training/), contingency management in addiction medicine, school-wide positive behavior supports, organizational behavior management, and the analysis of choice that became behavioral economics. Herrnstein's matching law and George Ainslie's hyperbolic discounting — the finding that we overvalue immediate consequences along a predictable curve — gave economists a behavioral account of impulsiveness years before it was fashionable.[21][41] ## The neuroscience of reinforcement The last chapter of the history is being written by neuroscience. In 1953 James Olds and Peter Milner found that a rat would press a lever for hours to deliver a pulse of current to its own brain — the first demonstration that reinforcement had an identifiable neural substrate.[42] In 1997 Wolfram Schultz, Peter Dayan, and Read Montague showed that midbrain **dopamine** neurons report a *prediction error* — firing to unexpected reward, falling silent to predicted reward, dipping when a predicted reward fails — precisely the quantity in the Rescorla–Wagner model and in temporal-difference learning.[30][43] Kent Berridge and Terry Robinson then separated "wanting" from "liking": animals without dopamine still enjoy sugar but no longer work for it.[44] The operant laboratory's strangest findings — the grip of the variable-ratio schedule, the fading of a fully predictable reinforcer — suddenly had a mechanism. [What happens in the brain, in depth ›](https://operantconditioning.com/neuroscience/) ## Operant conditioning in artificial intelligence **Reinforcement learning** is the branch of machine learning in which an agent learns by acting in an environment and receiving a numerical reward. Its founders were explicit about the lineage: Richard Sutton and Andrew Barto's textbook opens with Thorndike's law of effect, and their early work with Charles Anderson on "neuronlike adaptive elements" was an attempt to build a learning rule that behaved like an animal under reinforcement.[31][45] Sutton's 1988 **temporal-difference learning** algorithm updates predictions from the difference between successive predictions — the computational cousin of Rescorla–Wagner — and it was TD error that Montague, Dayan, and Sejnowski proposed in 1996 as the thing dopamine neurons encode, a year before Schultz's data confirmed it.[43][46] The vocabulary crossed over intact. AI researchers speak of *reward*, *policy*, *exploration and exploitation*, and *reward shaping* — the practice of adding intermediate rewards to guide an agent toward a hard-to-reach goal, named for Skinner's procedure.[47] Deep reinforcement learning, which pairs these algorithms with neural networks, learned Atari games from raw pixels in 2015 and beat the world's best Go players in 2016; *reinforcement learning from human feedback* is now a standard stage in training large language models, with human preferences standing in for the food hopper.[32][33] The differences matter too. An RL agent's reward is written by an engineer, not discovered functionally; the agent can explore millions of episodes an animal never could; and the algorithms include explicit value estimates that Skinner would have called explanatory fictions. But the central claim — that a system with no built-in knowledge of a task can acquire complex, purposeful behavior purely from the consequences of its actions — is Thorndike's and Skinner's, vindicated at a scale neither imagined. ## Key takeaways - Two names answer the question of discovery. Thorndike showed experimentally, with learning curves, that the consequences of a response change its future probability; Skinner treated that fact as the foundation of a science, defining the operant and inventing the free-operant methods that revealed schedules, shaping, extinction, and stimulus control. - Skinner was not Thorndike with better equipment. He replaced time to escape with rate of response, discrete trials with the free operant, a specific movement with an operant defined by its consequence, theories of why reinforcement works with a functional definition, and group comparisons with single organisms studied to steady state. - Thorndike himself dropped the negative half of the law of effect in the 1930s, the first evidence that punishment and reinforcement are not mirror images. Estes's 1944 finding that punishment suppresses responding only temporarily became Skinner's lifelong argument against it. - The cognitive revolution bounded the law of effect rather than overturning it. Latent learning, observational learning, biological constraints, and contingency over contiguity showed that reinforcement is one powerful process among several; its findings still replicate, and behavior analysts kept publishing throughout. - The tradition became a credentialed profession, applied behavior analysis, whose current debate is about goals, consent, and history rather than whether reinforcement works. It found a mechanism in the dopamine prediction-error signal and was formalized as reinforcement learning, whose founders cite the law of effect directly. ### Check yourself **Thorndike or Skinner: who discovered operant conditioning? Make the case for each.** Thorndike discovered the principle: his 1898 puzzle-box experiments showed, with learning curves rather than anecdotes, that consequences change the future probability of a response, and he named it the law of effect. Skinner named the process in 1937, defined the operant as a class of behavior controlled by its consequences, built the chamber and cumulative recorder, and discovered schedules, shaping, extinction, and stimulus control. Thorndike found the law; Skinner turned it into a field. **Behaviorism began with Watson's 1913 manifesto, so it seems safe to say early behaviorist learning theory was built on Thorndike's law of effect. Is that right?** No. Watson's early learning theory rested on Pavlov's reflex, not on Thorndike's law. It was Skinner, in the 1930s, who borrowed Pavlov's vocabulary of reinforcement, extinction, generalization, and discrimination and applied it to a different kind of learning, behavior controlled by its consequences rather than elicited by a prior stimulus. **A textbook states that Chomsky's 1959 review and the cognitive revolution killed behaviorism. What does the historical record show?** Mainstream psychology's attention did shift to mental processes, and findings such as latent learning, observational learning, and biological constraints showed that consequences are not the whole story. But those findings bounded the law of effect rather than overturning it: reinforcement, schedule effects, extinction bursts, shaping, and stimulus control still replicate, and behavior analysts kept publishing in their own journals, founding an applied journal in 1968 and a credentialing body in 1998. The "death of behaviorism" was mostly a change in who wrote the introductory textbook. **Sutton and Barto open their reinforcement learning textbook with Thorndike. What does a reinforcement learning agent share with a cat in a puzzle box, and where does the comparison break down?** Both acquire complex, purposeful behavior with no built-in knowledge of the task, purely from the consequences of their actions, and the vocabulary of reward, exploration, and reward shaping crossed over intact. The differences: an agent's reward is written by an engineer rather than discovered functionally, the agent can explore millions of episodes an animal never could, and its algorithms carry explicit value estimates that Skinner would have called explanatory fictions. **Explain it to a friend.** Explain why the cognitive revolution did not overturn the law of effect, without using the words "behaviorism" or "cognitive." ## Frequently asked questions **Who discovered operant conditioning?** Edward Thorndike discovered the underlying principle, the law of effect, in his 1898 puzzle-box experiments with cats. B. F. Skinner named the process "operant conditioning" in 1937, invented the experimental methods to study it, and founded the science with *The Behavior of Organisms* (1938). Both are correctly credited. **When was operant conditioning discovered?** The law of effect dates to Thorndike's 1898 monograph and was named in his 1911 book. The term "operant conditioning" dates to 1937, and the systematic science to 1938. **Who is the father of operant conditioning?** B. F. Skinner is usually called the father of operant conditioning because he named it, built the apparatus and methods, and developed the concepts — reinforcement schedules, shaping, stimulus control, the operant itself. Thorndike is the grandfather: he found the law of effect on which everything rests. **What is the difference between Thorndike and Skinner?** Thorndike studied discrete trials, measured time to respond, and explained learning as connections "stamped in" by satisfaction. Skinner let animals respond freely, measured rate of response, defined reinforcement purely by its effect, and rejected explanations in terms of satisfaction, drive, or connections. **What are the origins of behaviorism?** Behaviorism as a movement began with John B. Watson's 1913 paper "Psychology as the Behaviorist Views It," which argued that psychology should study observable behavior rather than consciousness. Its roots include Pavlov's conditioned reflexes, Thorndike's animal learning, and Morgan's canon of parsimony. Skinner's radical behaviorism, from 1945, is a later and different philosophy. **Did the cognitive revolution disprove operant conditioning?** No. It showed that reinforcement is not the only way organisms learn — observation, latent learning, and biological preparedness matter too — and it shifted mainstream psychology's attention to mental processes. The findings of operant conditioning still replicate, and the field continued as behavior analysis, with its own journals, profession, and applications. **How is operant conditioning related to artificial intelligence?** Reinforcement learning, a major branch of machine learning, is a mathematical formalization of learning from consequences. Its founders, Sutton and Barto, cite Thorndike's law of effect directly; temporal-difference learning parallels the Rescorla–Wagner model; and "reward shaping" takes its name from Skinner's procedure. The same prediction-error signal appears in the firing of dopamine neurons. **Is behaviorism still used today?** Yes, as applied behavior analysis (a licensed profession), reinforcement-based animal training, contingency management for addiction, positive behavior supports in schools, organizational behavior management, and the self-management methods behind habit-building apps. Its concepts also run through behavioral economics, neuroscience, and AI. ## References 1. Bain, A. (1855). *The Senses and the Intellect*. John W. Parker. See also Bain, A. (1859). *The Emotions and the Will*. John W. Parker. 2. Romanes, G. J. (1882). *Animal Intelligence*. Kegan Paul, Trench. 3. Morgan, C. L. (1894). *An Introduction to Comparative Psychology*. Walter Scott. 4. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 5. Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. 6. Thorndike, E. L. (1932). *The Fundamentals of Learning*. Teachers College, Columbia University. 7. Watson, J. B. (1913). Psychology as the behaviorist views it. *Psychological Review, 20*(2), 158–177. 8. Pavlov, I. P. (1927). *Conditioned Reflexes* (G. V. Anrep, Trans.). Oxford University Press. 9. Miller, S., & Konorski, J. (1928). Sur une forme particulière des réflexes conditionnels. *Comptes Rendus des Séances de la Société de Biologie, 99*, 1155–1157. English translation: Miller, S., & Konorski, J. (1969). On a particular form of conditioned reflex. *Journal of the Experimental Analysis of Behavior, 12*(1), 187–189. 10. Konorski, J., & Miller, S. (1937). On two types of conditioned reflex. *Journal of General Psychology, 16*, 264–272; Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279. 11. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 12. Hull, C. L. (1943). *Principles of Behavior: An Introduction to Behavior Theory*. Appleton-Century. 13. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 14. Tolman, E. C. (1948). Cognitive maps in rats and men. *Psychological Review, 55*(4), 189–208. See also Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. *University of California Publications in Psychology, 4*, 257–275. 15. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 16. Keller, F. S., & Schoenfeld, W. N. (1950). *Principles of Psychology: A Systematic Text in the Science of Behavior*. Appleton-Century-Crofts. 17. Skinner, B. F. (1950). Are theories of learning necessary? *Psychological Review, 57*(4), 193–216. 18. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 19. Skinner, B. F. (1956). A case history in scientific method. *American Psychologist, 11*(5), 221–233. 20. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 21. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 22. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 23. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58. 24. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 25. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97. 26. Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), *Classical Conditioning II: Current Research and Theory* (pp. 64–99). Appleton-Century-Crofts. See also Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. *Journal of Comparative and Physiological Psychology, 66*(1), 1–5. 27. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 28. Nevin, J. A., Mandell, C., & Atak, J. R. (1983). The analysis of behavioral momentum. *Journal of the Experimental Analysis of Behavior, 39*(1), 49–59. See also Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 29. Lovaas, O. I. (1987). Behavioral treatment and normal educational and intellectual functioning in young autistic children. *Journal of Consulting and Clinical Psychology, 55*(1), 3–9. 30. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 31. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. (First edition 1998.) 32. Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2015). Human-level control through deep reinforcement learning. *Nature, 518*(7540), 529–533. 33. Silver, D., Huang, A., Maddison, C. J., et al. (2016). Mastering the game of Go with deep neural networks and tree search. *Nature, 529*(7587), 484–489. See also Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. *Advances in Neural Information Processing Systems, 30*. 34. Skinner, B. F. (1935). The generic nature of the concepts of stimulus and response. *Journal of General Psychology, 12*, 40–65. 35. Sidman, M. (1960). *Tactics of Scientific Research: Evaluating Experimental Data in Psychology*. Basic Books. 36. Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. *Psychological Review, 63*(2), 81–97; Neisser, U. (1967). *Cognitive Psychology*. Appleton-Century-Crofts. 37. Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. *Journal of Abnormal and Social Psychology, 63*(3), 575–582. 38. Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. *Psychonomic Science, 4*(1), 123–124. See also Seligman, M. E. P. (1970). On the generality of the laws of learning. *Psychological Review, 77*(5), 406–418. 39. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 40. Leaf, J. B., Cihon, J. H., Leaf, R., McEachin, J., Liu, N., Russell, N., Unumb, L., Shapiro, S., & Khosrowshahi, D. (2022). Concerns about ABA-based intervention: An evaluation and recommendations. *Journal of Autism and Developmental Disorders, 52*(6), 2838–2853. On the early use of aversives, see Lovaas, O. I., Schaeffer, B., & Simmons, J. Q. (1965). Building social behavior in autistic children by use of electric shock. *Journal of Experimental Research in Personality, 1*, 99–109. 41. Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. *Psychological Bulletin, 82*(4), 463–496. 42. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427. 43. Montague, P. R., Dayan, P., & Sejnowski, T. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. *Journal of Neuroscience, 16*(5), 1936–1947. 44. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? *Brain Research Reviews, 28*(3), 309–369. 45. Barto, A. G., Sutton, R. S., & Anderson, C. W. (1983). Neuronlike adaptive elements that can solve difficult learning control problems. *IEEE Transactions on Systems, Man, and Cybernetics, 13*(5), 834–846. 46. Sutton, R. S. (1988). Learning to predict by the methods of temporal differences. *Machine Learning, 3*(1), 9–44. 47. Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. In *Proceedings of the Sixteenth International Conference on Machine Learning* (pp. 278–287). Morgan Kaufmann. 48. Aristotle. (c. 350 BCE). *Nicomachean Ethics*, Book II (W. D. Ross, Trans.). See especially 1103a–1103b on habituation and 1104b on pleasure and pain as signs of character. ## Related - [B. F. Skinner](https://operantconditioning.com/bf-skinner/): Biography, the Skinner box, radical behaviorism, and the myths. - [Operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/): Pavlov's reflexes and Skinner's operants, side by side. - [Applications](https://operantconditioning.com/applications/): Where the science is used today: ABA, education, parenting, animals, work, health, technology. --- # Applications of Operant Conditioning: Therapy, Education, Parenting, Work, Health, Technology, and More > Operant conditioning in the classroom, parenting, the workplace, therapy, animal training, technology, and AI — and how strong the evidence is for each. - Source: https://operantconditioning.com/applications/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications* The same three-term contingency runs a therapy session, a classroom, a family dinner, a factory floor, a dolphin show, and a slot machine. Here is where operant conditioning is used, how, and how strong the evidence really is in each field. > **Definition** > > The **applications of operant conditioning** are the deliberate arrangements of antecedents and consequences to change behavior that matters — in clinics, schools, homes, workplaces, zoos, software, and in the design of one's own life. > > The discipline built to do this systematically is **applied behavior analysis**, founded as a field in 1968, but the principles are used far beyond it, in some places rigorously and in others loosely.[1] The table below rates each field's evidence as candidly as the literature allows. **In brief** - Applications of operant conditioning are deliberate arrangements of antecedents and consequences to change behavior in clinics, schools, homes, workplaces, zoos, and software. - Applied behavior analysis, founded in 1968, is the discipline built to do this systematically; its core tool is the functional analysis, and treatment follows function. - The evidence varies by field: strong for classroom programs, parent training, contingency management, and behavioral activation; weaker for early intensive autism intervention and gamification. | Field | Core techniques | Evidence strength | Learn more | | --- | --- | --- | --- | | **Applied behavior analysis** | Functional analysis, reinforcement, differential reinforcement, shaping, prompting and fading | Strong for targeted behavior change in single-case research; low-certainty and contested for early intensive intervention in autism | ABA › | | **Education** | Token economies, Good Behavior Game, PBIS, precision teaching, direct instruction, praise ratios | Strong: randomized trials with follow-ups into adulthood; large meta-analyses for direct instruction | Classroom › | | **Parenting** | Specific praise, planned ignoring, time-out, response cost, consistency | Strong for behavioral parent-training programs; strong evidence *against* corporal punishment | Parenting › | | **Self-management** | Antecedent control, tiny behaviors, self-monitoring, immediate self-administered consequences | Moderate: self-monitoring and implementation intentions well supported; "self-reinforcement" mechanism debated | Self-management › · [Habits guide ›](https://operantconditioning.com/habits/) | | **Health and clinical** | Contingency management, behavioral activation, exposure with response prevention, habit reversal | Strong: contingency management and behavioral activation are among the best-supported psychosocial treatments in their areas | Health › | | **Animal training** | Marker signals, shaping, reinforcement-based husbandry training | Strong for reinforcement-based methods; aversive methods linked to welfare costs | Animals › · [Dog training ›](https://operantconditioning.com/dog-training/) | | **Workplace** | Pinpointing, performance feedback, behavior-based safety, incentives | Moderate to strong within-organization studies; feedback reliably works; gamification mixed | Workplace › | | **Technology and product design** | Variable-ratio feeds, notifications, streaks, loot boxes | Strong that the mechanisms drive engagement; contested whether population-level harms are large | Technology › | | **Behavioral economics and AI** | Matching law, delay discounting, reinforcement learning | Strong: quantitative laws replicated across species; reinforcement learning is foundational to modern AI | Economics & AI › | > **How to read the evidence column** > > "Strong" means randomized trials or many replicated single-case experiments with durable effects. "Moderate" means consistent results from weaker designs, or strong results for some components and not others. Where a field's reputation outruns its data, the table says so. ## Applied behavior analysis (ABA) **Applied behavior analysis** is the science in which procedures derived from the principles of behavior are applied to improve socially significant behavior, and experimentation is used to show that the procedures were responsible for the change.[2] Baer, Wolf, and Risley's founding paper set out seven dimensions that still define the field: work should be applied, behavioral, analytic, technological, conceptually systematic, effective, and general.[1] Its core tool is the **functional analysis**: experimentally testing what consequence maintains a problem behavior before trying to change it, a method established by Iwata and colleagues' 1982 study of self-injury.[3] Treatment then follows function. If a child hits to get attention, the analyst teaches a replacement — tapping a shoulder and saying "excuse me" — and reinforces it while the hitting no longer produces attention. That is **differential reinforcement of alternative behavior** (DRA), the most-used procedure in the field; its cousins reinforce the *absence* of the behavior for a set interval (DRO) or a behavior physically *incompatible* with it (DRI).[4] [How ABC data and functional analysis work ›](https://operantconditioning.com/abc-model/) The evidence deserves a careful statement. For targeted behavior change — reducing self-injury, teaching communication, building daily-living skills — hundreds of single-case experiments show large, reliable effects. The claim that early intensive behavioral intervention (EIBI) transforms outcomes in autism rests on a smaller base. Lovaas reported in 1987 that 47% of children receiving roughly 40 hours a week of one-to-one therapy reached normal-range IQ and unsupported first-grade placement, versus 2% of a comparison group.[5] Later reviews were more modest: a Cochrane review found only low-certainty evidence, from a handful of mostly non-randomized studies, that EIBI improves adaptive behavior and IQ, and a 2020 meta-analysis found support for behavioral approaches weakened substantially when limited to randomized trials with outcomes not reported by caregivers.[6][7] ABA also faces a principled critique from autistic adults and the neurodiversity movement: that its historical goal of making autistic children "indistinguishable" from peers, its early use of aversives (Lovaas's 1960s studies used contingent electric shock), and its emphasis on compliance were harmful whatever the outcome measures said.[8] Contemporary practice has moved toward assent, client-chosen goals, and reinforcement-only procedures. Practitioners are credentialed by the Behavior Analyst Certification Board (the BCBA credential dates from 1998), and the profession is licensed in most U.S. states. ## Operant conditioning in the classroom Schools were the first institutions to use operant technology at scale, and they still produce some of its best evidence. The token economy — points earned for target behaviors and exchanged later for backup reinforcers — moved from a psychiatric ward into classrooms almost as soon as Ayllon and Azrin described it, and it works because tokens are generalized conditioned reinforcers that can be delivered the instant a behavior occurs.[9] The Good Behavior Game, a team-based contingency introduced in a fourth-grade classroom in 1969, is among the most thoroughly studied classroom procedures of all: in randomized trials that followed first-graders into their twenties, the boys who had played it showed lower rates of substance-use and antisocial-personality disorders than controls.[10][11] School-wide positive behavioral supports scale the same logic across a building, and Skinner's own contribution — programmed instruction, with its small steps, active responding, and immediate feedback — survives in adaptive software and Direct Instruction.[12][13][14] Praise, the cheapest reinforcer in the room, works only when it is contingent, specific, and credible.[56] [The classroom in depth: praise, token economies, the Good Behavior Game, PBIS, and what the evidence says ›](https://operantconditioning.com/classroom/) ## Operant conditioning in parenting A parent's most powerful reinforcer is attention, and the most common parenting error is spending it on misbehavior. "Catching them being good" reverses the allocation: notice the behavior you want and name it specifically and immediately. Gerald Patterson's observations of families showed how the opposite pattern escalates — the child's tantrum is negatively reinforced when the parent gives in, and the parent's giving in is negatively reinforced when the tantrum stops, a *coercive family process* that trains both sides.[15] Parent management training, the best-supported treatment for childhood conduct problems, teaches the reverse contingencies: specific praise, planned ignoring, effective commands, brief time-out from reinforcement, and point systems.[16][17] On the punishment side, the largest meta-analysis of spanking — 75 studies and more than 160,000 children — found it associated with worse outcomes and no benefit.[18] [Parenting in depth: tantrums, time-out, sticker charts, bedtime, and the evidence ›](https://operantconditioning.com/parenting/) ## Self-management The three-term contingency turned inward. Skinner argued that a person controls their own behavior with the same tools used on anyone else's — changing the stimulus, restraining themselves physically, arranging deprivation and satiation, and delivering their own consequences.[19] The best-supported components are antecedent control (implementation intentions have a medium-to-large meta-analytic effect on goal attainment) and self-monitoring (recording your own behavior reliably changes it).[20][21] Whether a self-delivered reward is "really" reinforcement is debated; what is not debated is that a cue you cannot miss, a behavior small enough to emit, and a consequence that arrives immediately outperform resolve. [The complete habit-building protocol ›](https://operantconditioning.com/habits/) ## Health and clinical applications **Contingency management** for substance use is the clearest case of operant principles applied to a hard medical problem. In 1991 Stephen Higgins and colleagues offered cocaine-dependent outpatients vouchers exchangeable for retail goods, contingent on cocaine-negative urine samples, with the value escalating for consecutive clean samples and resetting after a positive one; retention and abstinence far exceeded standard counseling.[22] Two 2006 meta-analyses confirmed moderate, reliable effects across drugs, strongest for stimulants and opioids, and a 2008 meta-analytic review found contingency management produced the largest effects of any psychosocial treatment examined.[23][24][25] Its limits are what the theory predicts: effects fade after incentives stop unless natural reinforcers have taken over, and uptake has been slowed more by regulation and squeamishness about "paying people to stay clean" than by the data. Chronic pain was one of the first medical problems given an operant analysis. Wilbert Fordyce observed that "pain behaviors" — guarding, limping, grimacing, resting, taking medication, talking about pain — are behaviors, and that they are reinforced: by attention and sympathy, by relief from unwanted duties, and by medication given in response to complaints. His inpatient program at the University of Washington, reported in 1973, reversed the contingencies. Staff attended to activity rather than complaints, exercise quotas were raised gradually with rest as the reinforcer for meeting them rather than for pain, and medication was given on a time schedule instead of on demand; patients' activity rose and their medication use and reported pain fell.[57][58] Operant-behavioral treatment remains a component of modern pain programs; in a randomized trial for fibromyalgia, operant-behavioral and cognitive-behavioral treatments each outperformed an attention-placebo condition, with the operant program's largest gains in physical functioning.[59] Nobody in this literature claims the pain is imaginary; the claim is that what a person does about pain is learned, and can be re-learned. **Behavioral activation** treats depression as a collapse in response-contingent positive reinforcement, an analysis Charles Ferster offered in 1973.[26] The treatment schedules activities, grades difficult tasks, and dismantles avoidance so the person contacts reinforcement again. A 1996 dismantling study found the activation component alone matched the full cognitive therapy package; a 2006 randomized trial found it comparable to antidepressant medication and superior to cognitive therapy among more severely depressed patients; and a 2016 trial found it non-inferior to CBT when delivered by junior mental-health workers at lower cost.[27][28][29] **Exposure therapy** has an operant engine inside its classical shell. In Mowrer's two-factor account, fear is classically conditioned, but *avoidance* is negatively reinforced by the relief it brings, and avoidance is what prevents the fear from extinguishing.[30] Exposure with response prevention blocks the escape so extinction can proceed. **Habit reversal training** — awareness training, a competing response, and social support — was introduced by Azrin and Nunn in 1973 for tics and nervous habits and is now the core of the leading behavioral treatment for Tourette syndrome.[31][32] **Biofeedback** is operant conditioning of physiological responses; Neal Miller's 1969 claims of visceral learning in rats proved hard to replicate, but clinical biofeedback has a real evidence base for specific problems such as incontinence and tension headache.[33] And financial incentives for **medication adherence** consistently improve adherence while they are in place.[34] ## Animal training Modern animal training is operant conditioning with the lecture removed. A **marker signal** — a clicker or a word — is a conditioned reinforcer that bridges the gap between the behavior and the treat; [shaping](https://operantconditioning.com/shaping/) builds complex behavior from approximations; and reinforcement-based methods have displaced the dominance-and-correction approaches of the mid-twentieth century. [Dog training with operant conditioning ›](https://operantconditioning.com/dog-training/) The commercial lineage runs through Keller and Marian Breland, Skinner's former students, who founded Animal Behavior Enterprises in 1943 and trained thousands of animals for advertising, exhibitions, and their "IQ Zoo." Their 1961 paper "The misbehavior of organisms," reporting raccoons that "washed" coins instead of depositing them, remains the classic demonstration that reinforcement works with, not against, an animal's evolved tendencies.[35] Marine-mammal training grew from the same roots; in a famous 1969 study, dolphins reinforced only for behaviors they had never shown before began producing novel actions.[36] The quietest revolution is in zoos and laboratories. **Husbandry training** teaches animals to present a limb for a blood draw, open their mouths for dental checks, and enter a crate voluntarily, replacing restraint and anesthesia with cooperation and reducing measurable stress.[37] Reviews of aversive dog-training methods, meanwhile, link them to stress and problem behaviors without any advantage in effectiveness.[38] ## Operant conditioning in the workplace **Organizational behavior management** (OBM) applies the same analysis to employees: pinpoint a behavior, measure it, give feedback, and reinforce it. The field has had its own journal since 1977. The signature demonstration is **behavior-based safety**: in a 1978 study at a food-manufacturing plant, researchers defined specific safe behaviors, posted graphed feedback on how often they occurred, and watched safe performance rise sharply — then fall when feedback was withdrawn and rise again when it returned.[39] A review of performance-feedback studies from 1985 to 1998 found feedback effective in most applications, most consistently when it was graphic, frequent, and combined with goals and reinforcement; a meta-analysis of behavior-modification programs across organizations reported an average performance improvement of about 17%.[40][41] Aubrey Daniels, the field's best-known popularizer, summarizes the timing problem with a three-letter test: consequences that are positive, immediate, and certain control behavior; consequences that are negative, delayed, or uncertain barely register.[42] The annual performance review fails on every count. It arrives months after the behavior it addresses, it is delivered once, and for most people it is aversive, so it mainly evokes escape — the polished self-assessment, the defensive meeting — rather than changing what anyone does on Tuesday. Daily feedback from a supervisor who knows what to look for costs less and does more. Two cautions. **Gamification** — points, badges, leaderboards — reliably produces short-term engagement and unreliably produces lasting change; a review of the empirical studies found positive but context-dependent effects that often faded with novelty.[43] And reinforcing a proxy reinforces the proxy. When sales targets and incentives at Wells Fargo rewarded accounts opened, employees opened millions of accounts customers had not asked for. The contingency worked as designed; the design was the problem. ## Technology and product design Consumer software is the largest deployment of operant conditioning in history, and much of it is aimed at the user rather than for them. A social feed is a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/): pull to refresh, and the reinforcer — a like, a message, something novel — arrives after an unpredictable number of pulls, the schedule that produces the highest response rates and the most persistence when reinforcement stops. Notifications are discriminative stimuli for opening the app. Streaks convert a habit into an avoidance contingency — the reinforcer becomes not losing the count — and borrow loss aversion to make a missed day feel like a fine. Gambling is the original engineered variable-ratio schedule, and slot machines add a refinement the laboratory did not anticipate: the **near miss**. Two cherries and a lemon is a loss, but it looks like almost winning, and Reid argued in 1986 that near misses encourage continued play as if they were partial reinforcement.[60] Brain-imaging work later found that near misses recruit some of the same reward circuitry as wins and increase the urge to keep playing, more so in people who gamble more heavily.[61] Game designers borrowed the whole toolkit early; a 2001 article by the psychologist John Hopson laid out how to apply schedules of reinforcement to game design so that players keep playing; the industry's later name for the resulting anticipation–activity–reward cycle is the **compulsion loop**, and most games since have one.[62] **Loot boxes** in video games are the starkest case. Researchers who examined popular games found that a large share of loot-box systems met the psychological criteria for gambling, and spending on them correlates with problem-gambling severity; Belgium's gaming regulator ruled paid loot boxes illegal gambling in 2018.[44][45] BJ Fogg's **behavior model** — behavior occurs when motivation, ability, and a prompt converge — is a design-oriented cousin of the four-term contingency: the prompt is the antecedent, ability stands in for response effort, and motivation for the motivating operation.[46] When those tools trick users into choices they would not otherwise make, the designs are called **dark patterns**, and they are increasingly the target of regulation. Honesty requires a caveat: whether the resulting engagement causes large population-level harm is contested. Large-sample analyses of adolescent well-being find the association with digital technology use to be tiny.[47] That the mechanisms work is not in doubt; how much damage they do is. The same science runs the other way: environment design against the feeds — the phone in the kitchen, the app logged out — and tools that make the user set the antecedent and consequence for behavior they have chosen. [Using the loop for your own goals ›](https://operantconditioning.com/habits/) ## Behavioral economics and artificial intelligence Operant research turned quantitative in 1961, when Richard Herrnstein showed that pigeons offered two keys, each paying on its own variable-interval schedule, distributed their pecks in proportion to the reinforcement each key delivered — the **[matching law](https://operantconditioning.com/matching-law/)**.[48] Matching turned choice into something that could be modeled with the tools of economics, and by the 1990s laboratory animals had been shown to obey demand curves, substitute between goods, and respond to price much as human consumers do.[49] The most consequential offshoot is **delay discounting**. Using an adjusting procedure, James Mazur showed that the value of a delayed reinforcer falls along a hyperbola rather than the exponential curve standard economics assumed.[50] George Ainslie had already worked out the implication: hyperbolic curves cross, so a person who prefers the larger, later reward from a distance will flip to the smaller, sooner one as it approaches — a mathematical account of impulsiveness and of why commitment devices are needed.[51] Steeper discounting has since been documented in smokers, in people with substance-use disorders, and in problem gamblers, and it is now studied as a process cutting across many conditions.[52] Public policy has borrowed from both halves of the contingency. **Nudges** — Thaler and Sunstein's term for changes in "choice architecture" that steer behavior without changing incentives, such as making retirement saving the default — are antecedent interventions: they alter the stimulus conditions under which a choice is made, not its consequences, and they are cheap precisely because no reinforcer has to be delivered.[63] Incentive programs — conditional cash transfers, deposit contracts, sin taxes — work on the consequence side. The behavioral analysis predicts what the evaluations find: nudges produce modest, durable effects when the behavior is a one-off choice, and incentives produce larger effects that fade when the incentive stops unless a natural reinforcer takes over. The law of effect also became an algorithm. **Reinforcement learning**, the branch of machine learning in which an agent learns by acting and receiving reward signals, descends directly from Thorndike and Skinner by way of temporal-difference learning, and its standard textbook says so in its opening pages.[53] In 1997 Schultz, Dayan, and Montague showed that midbrain dopamine neurons fire in the pattern of a temporal-difference prediction error — the brain running the same computation.[54] [The neuroscience of operant conditioning ›](https://operantconditioning.com/neuroscience/) Reinforcement learning drove the systems that mastered Go, and reinforcement learning from human feedback is one of the methods used to train large language models to behave as people prefer.[55] [The history from Thorndike to reinforcement learning ›](https://operantconditioning.com/history/) ## Military training The most frequently cited military application is also the most disputed. After the Second World War, the U.S. Army historian S. L. A. Marshall reported, on the basis of group interviews, that only about 15–20% of riflemen had fired their weapons at the enemy in combat.[64] Training changed in response: bullseye targets were replaced by man-shaped silhouettes that pop up briefly and fall when hit — immediate feedback on a realistic discriminative stimulus — and the conditioned response was practiced hundreds of times. Dave Grossman argued in *On Killing* that this was operant conditioning in all but name, and credited it with raising reported firing rates to about 55% in Korea and over 90% in Vietnam.[65] The account should be read with care: Marshall's figures have been challenged by historians who found no record of the systematic interviews he described, and the later rates are not measured the same way.[66] What is not in dispute is the design of the training itself, which is a textbook contingency, or Skinner's own wartime contribution: Project Pigeon, in which he shaped pigeons to peck at an image of a target so that their pecks could steer a glide bomb — a working system that the military declined to deploy.[67] ## Coercion: the dark side of the contingency The same principles that build a token economy can hold a person in a harmful relationship or a toxic workplace, and behavior analysts have said so plainly. Murray Sidman's *Coercion and Its Fallout* catalogued what aversive control does to the people subjected to it: it produces escape and avoidance, countercontrol, aggression, and a generalized suppression of behavior that reaches far beyond the punished response.[68] **Traumatic bonding** is the clearest case. Donald Dutton and Susan Painter proposed that two features of abusive relationships — a power imbalance and the *intermittency* of abuse, with cycles of cruelty and reconciliation — produce unusually strong attachment to the abuser, on the same logic by which [intermittent reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) produces the most persistent behavior in the laboratory. Their follow-up study of women who had left abusive partners found that intermittency of abuse and dominance predicted continued attachment months later.[69][70] Popular accounts of psychological manipulation list the same operations — intermittent reward, negative reinforcement by the removal of hostility, punishment, and one-trial traumatic learning — and they are recognizable to anyone who has read this far. Organizations run the same contingencies at scale. Blake Ashforth's analysis of "petty tyranny" in management described supervisors who rule through arbitrary punishment, belittling, and the withholding of consideration, and traced the consequences the laboratory predicts: helplessness, low initiative, and the disappearance of any behavior not strictly required.[71] A culture of fear is a culture of avoidance, and avoidance, as [its own page](https://operantconditioning.com/avoidance-learning/) explains, is the behavior least sensitive to whether the threat is real. The ethical guidance of behavior analysis — reinforcement first, the least restrictive procedure, consent, the client's own goals — exists because the field knows exactly how effective the alternative is. ## Key takeaways - The same three-term contingency runs every application. Applied behavior analysis is the discipline built to apply it systematically, and its core tool is the functional analysis: find what consequence maintains a behavior, then teach and reinforce an alternative while the problem behavior no longer pays. - Evidence strength varies by field, and a field's reputation can outrun its data. The Good Behavior Game, behavioral parent training, contingency management, and behavioral activation have randomized trials with durable effects; early intensive behavioral intervention in autism rests on low-certainty evidence and faces a principled critique from autistic adults. - Consequences control behavior when they are positive, immediate, and certain; delayed, infrequent, and aversive consequences such as the annual review mainly evoke escape. Reinforcing a proxy reinforces the proxy. - Consumer software is the largest deployment of operant conditioning in history: feeds run variable-ratio schedules, notifications are discriminative stimuli, streaks are avoidance contingencies, and loot boxes reproduce the structure of gambling. That the mechanisms work is not in doubt; how much population-level harm they do is. - The principles cut both ways. Intermittent abuse produces traumatic bonding on the same logic that makes intermittent reinforcement persistent, and coercive management produces avoidance and helplessness. The field's ethics, reinforcement first, the least restrictive procedure, and consent, exist because it knows how effective the alternative is. ### Check yourself **A parent gives in to a tantrum and the tantrum stops. A classmate says the child has been rewarded and the parent has been punished by losing the standoff. What is actually happening, in operant terms?** Both behaviors are being strengthened, not one. The child's tantrum is negatively reinforced when the parent gives in, and the parent's giving in is negatively reinforced when the tantrum stops. This is Patterson's coercive family process: each side trains the other, which is why parent management training teaches the reverse contingencies. **A clinic offers vouchers for cocaine-negative urine samples. A critic objects that paying people to stay clean cannot work because motivation has to come from within. What does the evidence say, and what is the program's real limitation?** The evidence says it works: meta-analyses find moderate, reliable effects across drugs, strongest for stimulants and opioids, and a review of psychosocial treatments found contingency management had the largest effects of any approach examined. Its real limitation is what the theory predicts: effects fade after the incentives stop unless natural reinforcers have taken over. **A bank pays incentives on the number of accounts opened, and the number of accounts opened soars. Did the contingency fail?** No. It worked exactly as designed, and the design was the problem. Reinforcing a proxy reinforces the proxy, so employees opened millions of accounts customers had not asked for. Organizational behavior management starts by pinpointing the behavior that matters, not a stand-in for it. **One agency wants more people to enroll in a retirement plan. Another wants people to keep exercising for a year. Which tool suits each, a nudge or an incentive, and why?** Enrollment is a one-off choice, so a nudge fits: making saving the default changes the stimulus conditions under which the choice is made, delivers no reinforcer, and produces a modest, durable effect. Exercising for a year is ongoing behavior, so an incentive works on the consequence side and produces a larger effect, but that effect will fade when the incentive stops unless a natural reinforcer has taken over. **Explain it to a friend.** Explain what a classroom token economy and a slot machine have in common, and what separates them, without using the word "reinforcement." ## Frequently asked questions **What are the main applications of operant conditioning?** Applied behavior analysis (including autism and developmental-disability services), classroom management and instruction, parenting and parent training, animal training, organizational behavior management and workplace safety, clinical treatments such as contingency management and behavioral activation, product and game design, self-management and habit formation, and — in its mathematical form — behavioral economics and reinforcement learning in AI. **How is operant conditioning used in the classroom?** Through token economies, the Good Behavior Game, school-wide PBIS, high praise-to-reprimand ratios, and instructional methods built on immediate feedback such as programmed instruction, precision teaching, and Direct Instruction. The Good Behavior Game has randomized trials with follow-ups showing benefits into early adulthood. **How is operant conditioning used in parenting?** By reinforcing wanted behavior with specific, immediate attention; using planned ignoring for attention-maintained misbehavior; using brief, calm time-out and response cost for serious misbehavior; and being consistent rather than severe. Behavioral parent-training programs such as PMT and PCIT package these skills and are among the best-supported treatments for childhood behavior problems. **How is operant conditioning used in the workplace?** Organizational behavior management pinpoints specific behaviors, measures them, and delivers frequent feedback and reinforcement. Behavior-based safety programs and graphic performance feedback have decades of supporting studies. Annual reviews fail as consequences because they are delayed, infrequent, and aversive. **Is ABA therapy the same as operant conditioning?** No. Operant conditioning is the basic learning process; applied behavior analysis is the professional discipline that applies its principles (and others) to socially important behavior, with its own methods, credentials, and ethics code. ABA is one application of operant conditioning among many. **Is contingency management effective for addiction?** Yes. Meta-analyses find it produces reliable reductions in drug use, with the strongest effects for stimulants and opioids, and a broad review of psychosocial treatments found it had the largest effects of any approach. Its main limitation is that gains can fade after incentives end unless natural reinforcers have taken over. **How do apps and games use operant conditioning?** Feeds deliver unpredictable rewards on a variable-ratio schedule, notifications act as cues to open the app, streaks turn use into loss-avoidance, and loot boxes reproduce the structure of gambling. The same principles can be used deliberately for your own goals by controlling the cues in your environment and attaching immediate consequences to behaviors you choose. **Does operant conditioning apply to artificial intelligence?** Yes. Reinforcement learning — the AI method in which an agent learns from reward signals — is a direct mathematical descendant of the law of effect, and dopamine neurons in the brain have been shown to compute the same prediction-error signal its algorithms use. Reinforcement learning underlies game-playing systems and is used in training large language models. ## References 1. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97. 2. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 3. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1982/1994). Toward a functional analysis of self-injury. *Analysis and Intervention in Developmental Disabilities, 2*(1), 3–20. Reprinted in *Journal of Applied Behavior Analysis, 27*(2), 197–209. 4. Carr, E. G., & Durand, V. M. (1985). Reducing behavior problems through functional communication training. *Journal of Applied Behavior Analysis, 18*(2), 111–126. 5. Lovaas, O. I. (1987). Behavioral treatment and normal educational and intellectual functioning in young autistic children. *Journal of Consulting and Clinical Psychology, 55*(1), 3–9. 6. Reichow, B., Hume, K., Barton, E. E., & Boyd, B. A. (2018). Early intensive behavioral intervention (EIBI) for young children with autism spectrum disorders (ASD). *Cochrane Database of Systematic Reviews*, Issue 5, CD009260. 7. Sandbank, M., Bottema-Beutel, K., Crowley, S., et al. (2020). Project AIM: Autism intervention meta-analysis for studies of young children. *Psychological Bulletin, 146*(1), 1–29. 8. Lovaas, O. I., Schaeffer, B., & Simmons, J. Q. (1965). Building social behavior in autistic children by use of electric shock. *Journal of Experimental Research in Personality, 1*, 99–109. 9. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 10. Barrish, H. H., Saunders, M., & Wolf, M. M. (1969). Good behavior game: Effects of individual contingencies for group consequences on disruptive behavior in a classroom. *Journal of Applied Behavior Analysis, 2*(2), 119–124. 11. Kellam, S. G., Brown, C. H., Poduska, J. M., et al. (2008). Effects of a universal classroom behavior management program in first and second grades on young adult behavioral, psychiatric, and social outcomes. *Drug and Alcohol Dependence, 95*(Suppl. 1), S5–S28. 12. Bradshaw, C. P., Mitchell, M. M., & Leaf, P. J. (2010). Examining the effects of schoolwide positive behavioral interventions and supports on student outcomes: Results from a randomized controlled effectiveness trial in elementary schools. *Journal of Positive Behavior Interventions, 12*(3), 133–148. 13. Skinner, B. F. (1954). The science of learning and the art of teaching. *Harvard Educational Review, 24*(2), 86–97. See also Skinner, B. F. (1958). Teaching machines. *Science, 128*(3330), 969–977. 14. Stockard, J., Wood, T. W., Coughlin, C., & Rasplica Khoury, C. (2018). The effectiveness of Direct Instruction curricula: A meta-analysis of a half century of research. *Review of Educational Research, 88*(4), 479–507. 15. Wolf, M. M., Risley, T. R., & Mees, H. (1964). Application of operant conditioning procedures to the behaviour problems of an autistic child. *Behaviour Research and Therapy, 1*, 305–312. 16. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 17. Patterson, G. R. (1982). *Coercive Family Process*. Castalia. 18. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 19. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 20. Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. *Advances in Experimental Social Psychology, 38*, 69–119. 21. Harkin, B., Webb, T. L., Chang, B. P. I., et al. (2016). Does monitoring goal progress promote goal attainment? A meta-analysis of the experimental evidence. *Psychological Bulletin, 142*(2), 198–229. 22. Higgins, S. T., Delaney, D. D., Budney, A. J., Bickel, W. K., Hughes, J. R., Foerg, F., & Fenwick, J. W. (1991). A behavioral approach to achieving initial cocaine abstinence. *American Journal of Psychiatry, 148*(9), 1218–1224. 23. Prendergast, M., Podus, D., Finney, J., Greenwell, L., & Roll, J. (2006). Contingency management for treatment of substance use disorders: A meta-analysis. *Addiction, 101*(11), 1546–1560. 24. Lussier, J. P., Heil, S. H., Mongeon, J. A., Badger, G. J., & Higgins, S. T. (2006). A meta-analysis of voucher-based reinforcement therapy for substance use disorders. *Addiction, 101*(2), 192–203. 25. Dutra, L., Stathopoulou, G., Basden, S. L., Leyro, T. M., Powers, M. B., & Otto, M. W. (2008). A meta-analytic review of psychosocial interventions for substance use disorders. *American Journal of Psychiatry, 165*(2), 179–187. 26. Ferster, C. B. (1973). A functional analysis of depression. *American Psychologist, 28*(10), 857–870. 27. Jacobson, N. S., Dobson, K. S., Truax, P. A., et al. (1996). A component analysis of cognitive-behavioral treatment for depression. *Journal of Consulting and Clinical Psychology, 64*(2), 295–304. 28. Dimidjian, S., Hollon, S. D., Dobson, K. S., et al. (2006). Randomized trial of behavioral activation, cognitive therapy, and antidepressant medication in the acute treatment of adults with major depression. *Journal of Consulting and Clinical Psychology, 74*(4), 658–670. 29. Richards, D. A., Ekers, D., McMillan, D., et al. (2016). Cost and outcome of behavioural activation versus cognitive behavioural therapy for depression (COBRA): A randomised, controlled, non-inferiority trial. *The Lancet, 388*(10047), 871–880. 30. Mowrer, O. H. (1960). *Learning Theory and Behavior*. Wiley. 31. Azrin, N. H., & Nunn, R. G. (1973). Habit-reversal: A method of eliminating nervous habits and tics. *Behaviour Research and Therapy, 11*(4), 619–628. 32. Piacentini, J., Woods, D. W., Scahill, L., et al. (2010). Behavior therapy for children with Tourette disorder: A randomized controlled trial. *JAMA, 303*(19), 1929–1937. 33. Miller, N. E. (1969). Learning of visceral and glandular responses. *Science, 163*(3866), 434–445. 34. DeFulio, A., & Silverman, K. (2012). The use of incentives to reinforce medication adherence. *Preventive Medicine, 55*(Suppl.), S86–S94. 35. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 36. Pryor, K. W., Haag, R., & O'Reilly, J. (1969). The creative porpoise: Training for novel behavior. *Journal of the Experimental Analysis of Behavior, 12*(4), 653–661. 37. Laule, G. E., Bloomsmith, M. A., & Schapiro, S. J. (2003). The use of positive reinforcement training techniques to enhance the care, management, and welfare of primates in the laboratory. *Journal of Applied Animal Welfare Science, 6*(3), 163–173. 38. Ziv, G. (2017). The effects of using aversive training methods in dogs — A review. *Journal of Veterinary Behavior, 19*, 50–60. 39. Komaki, J., Barwick, K. D., & Scott, L. R. (1978). A behavioral approach to occupational safety: Pinpointing and reinforcing safe performance in a food manufacturing plant. *Journal of Applied Psychology, 63*(4), 434–445. 40. Alvero, A. M., Bucklin, B. R., & Austin, J. (2001). An objective review of the effectiveness and essential characteristics of performance feedback in organizational settings (1985–1998). *Journal of Organizational Behavior Management, 21*(1), 3–29. 41. Stajkovic, A. D., & Luthans, F. (1997). A meta-analysis of the effects of organizational behavior modification on task performance, 1975–95. *Academy of Management Journal, 40*(5), 1122–1149. 42. Daniels, A. C. (2000). *Bringing Out the Best in People: How to Apply the Astonishing Power of Positive Reinforcement* (2nd ed.). McGraw-Hill. 43. Hamari, J., Koivisto, J., & Sarsa, H. (2014). Does gamification work? — A literature review of empirical studies on gamification. *Proceedings of the 47th Hawaii International Conference on System Sciences*, 3025–3034. 44. Drummond, A., & Sauer, J. D. (2018). Video game loot boxes are psychologically akin to gambling. *Nature Human Behaviour, 2*(8), 530–532. 45. Zendle, D., & Cairns, P. (2018). Video game loot boxes are linked to problem gambling: Results of a large-scale survey. *PLoS ONE, 13*(11), e0206767. 46. Fogg, B. J. (2009). A behavior model for persuasive design. *Proceedings of the 4th International Conference on Persuasive Technology*, Article 40. 47. Orben, A., & Przybylski, A. K. (2019). The association between adolescent well-being and digital technology use. *Nature Human Behaviour, 3*(2), 173–182. 48. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 49. Kagel, J. H., Battalio, R. C., & Green, L. (1995). *Economic Choice Theory: An Experimental Analysis of Animal Behavior*. Cambridge University Press. 50. Mazur, J. E. (1987). An adjusting procedure for studying delayed reinforcement. In M. L. Commons, J. E. Mazur, J. A. Nevin, & H. Rachlin (Eds.), *Quantitative Analyses of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value* (pp. 55–73). Erlbaum. 51. Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. *Psychological Bulletin, 82*(4), 463–496. 52. Bickel, W. K., Odum, A. L., & Madden, G. J. (1999). Impulsivity and cigarette smoking: Delay discounting in current, never, and ex-smokers. *Psychopharmacology, 146*(4), 447–454. 53. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. 54. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 55. Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. *Advances in Neural Information Processing Systems, 30*. 56. Brophy, J. (1981). Teacher praise: A functional analysis. *Review of Educational Research, 51*(1), 5–32. 57. Fordyce, W. E., Fowler, R. S., Lehmann, J. F., DeLateur, B. J., Sand, P. L., & Trieschmann, R. B. (1973). Operant conditioning in the treatment of chronic pain. *Archives of Physical Medicine and Rehabilitation, 54*(9), 399–408. 58. Fordyce, W. E. (1976). *Behavioral Methods for Chronic Pain and Illness*. Mosby. 59. Thieme, K., Flor, H., & Turk, D. C. (2006). Psychological pain treatment in fibromyalgia syndrome: Efficacy of operant behavioural and cognitive behavioural treatments. *Arthritis Research & Therapy, 8*(4), R121. 60. Reid, R. L. (1986). The psychology of the near miss. *Journal of Gambling Behavior, 2*(1), 32–39. 61. Clark, L., Lawrence, A. J., Astley-Jones, F., & Gray, N. (2009). Gambling near-misses enhance motivation to gamble and recruit win-related brain circuitry. *Neuron, 61*(3), 481–490. 62. Hopson, J. (2001, April 27). Behavioral game design. *Gamasutra*. 63. Thaler, R. H., & Sunstein, C. R. (2008). *Nudge: Improving Decisions About Health, Wealth, and Happiness*. Yale University Press. 64. Marshall, S. L. A. (1947). *Men Against Fire: The Problem of Battle Command in Future War*. William Morrow. 65. Grossman, D. (1995). *On Killing: The Psychological Cost of Learning to Kill in War and Society*. Little, Brown. 66. Spiller, R. J. (1988). S.L.A. Marshall and the ratio of fire. *RUSI Journal, 133*(4), 63–71. 67. Skinner, B. F. (1960). Pigeons in a pelican. *American Psychologist, 15*(1), 28–37. 68. Sidman, M. (1989). *Coercion and Its Fallout*. Authors Cooperative. 69. Dutton, D. G., & Painter, S. L. (1981). Traumatic bonding: The development of emotional attachments in battered women and other relationships of intermittent abuse. *Victimology, 6*(1–4), 139–155. 70. Dutton, D. G., & Painter, S. (1993). Emotional attachments in abusive relationships: A test of traumatic bonding theory. *Violence and Victims, 8*(2), 105–120. 71. Ashforth, B. (1994). Petty tyranny in organizations. *Human Relations, 47*(7), 755–778. ## Related - [Build habits with operant conditioning](https://operantconditioning.com/habits/): The self-management protocol, step by step. - [Dog training](https://operantconditioning.com/dog-training/): Markers, shaping, and the evidence on aversive methods. - [50+ examples](https://operantconditioning.com/examples/): Every quadrant, every setting. --- # Operant Conditioning in the Classroom: Praise, Token Economies, the Good Behavior Game, and What the Evidence Says > Operant conditioning in the classroom: what the evidence says about praise, token economies, the Good Behavior Game, PBIS, and Skinner's teaching machines. - Source: https://operantconditioning.com/classroom/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications · Education* Schools were the first institutions to use operant conditioning at scale, and they still produce some of its best evidence. What works — praise, tokens, the Good Behavior Game, school-wide systems, feedback-driven teaching — and what quietly backfires. > **Definition** > > **Operant conditioning in the classroom** is the deliberate arrangement of antecedents (rules, cues, task design) and consequences (attention, praise, tokens, feedback, lost privileges) so that learning-related behavior increases and disruptive behavior decreases. > > Every classroom already runs on consequences; the question is whether they are arranged on purpose. The first controlled demonstration that a teacher's attention could switch a child's studying on and off appeared in the first issue of the *Journal of Applied Behavior Analysis*.[1] **In brief** - Every classroom already runs on consequences; operant conditioning in the classroom arranges antecedents and consequences on purpose so learning behavior rises and disruption falls. - Contingent, specific praise is the cheapest reinforcer in the building, and the Good Behavior Game is one of the most thoroughly studied classroom management procedures there is. - A consequence is whatever the behavior says it is: a reprimand can be attention, and removal from class can reinforce misbehavior that escapes work. ## Why classrooms adopted operant conditioning first Skinner's interest in education began on a school visit. In 1953 he watched his daughter's fourth-grade class work identical arithmetic problems at the same pace, with results returned a day later — feedback so delayed that, by everything he knew from the laboratory, it could hardly teach. His 1954 lecture "The science of learning and the art of teaching" diagnosed education's failures as failures of reinforcement: too few reinforced responses, too long a delay, no program of small steps.[2] Psychiatric wards ran the earliest behavior-change programs, but classrooms were where the technology first spread widely; the early issues of the *Journal of Applied Behavior Analysis*, launched in 1968, were full of schools.[1][3][4] The reasons are structural: one adult controlling most of the consequences, behavior that is visible and countable, a fixed schedule, twenty-five learners at once — the operant chamber with the roles reversed. And discipline was a problem everyone agreed on: measurable, urgent, and public. ## The four quadrants in a classroom | Procedure | What happens | Classroom example | Effect | | --- | --- | --- | --- | | Positive reinforcement | Something is added | A student raises her hand instead of calling out; the teacher calls on her: "Thanks for raising your hand." | Hand-raising increases | | Negative reinforcement | Something is removed | Students who get the first ten problems right are excused from the last ten. | Careful work increases | | Positive punishment | Something aversive is added | A student shouts across the room; the teacher gives a quiet, immediate reprimand. | Shouting decreases — unless, for this student, a reprimand is attention | | Negative punishment | A reinforcer is removed | A student pushes in line and loses two minutes of recess (response cost). | Pushing decreases | | Extinction | The maintaining reinforcer stops | A student blurts out answers for the teacher's reaction; the teacher stops reacting and calls only on raised hands. | Blurting rises briefly, then declines | Two rows deserve a second look. For a student who gets little adult attention, a reprimand *is* attention, and the shouting will go up. And extinction works only on the reinforcer actually maintaining the behavior; if the class is laughing, the teacher's silence changes nothing. A consequence is whatever the behavior says it is. [More examples for every quadrant ›](https://operantconditioning.com/examples/) ## Praise that works and praise that doesn't Teacher attention is the cheapest reinforcer in the building. When teachers attended to six elementary pupils only while they studied, studying rose; when attention went back to dawdling, it fell; when contingent attention returned, it rose again.[1] In a companion study, posting rules did little, adding ignoring did little, and disruption fell only when the teachers added praise for appropriate behavior — probably, the authors concluded, the key to classroom management.[3] The reverse holds too: withdraw approval from a well-behaved class and it turns disruptive; add disapproval and it gets worse.[5] Yet most classroom praise is not reinforcement. Jere Brophy's review found that teachers praise infrequently, often without regard to what the student has just done, and often as a management reflex; vague, non-contingent praise changes nothing.[6] A later review added that praise supports intrinsic motivation when it is sincere, names effort or strategy rather than ability, and sets a reachable standard — and undermines it when it compares students with one another or gushes over easy work.[7] | Praise that reinforces | Praise that doesn't | | --- | --- | | Contingent and specific: "You showed every step of the working" | Global or vague: "good class today," "nice job" | | Credible: matched to the difficulty, in a normal voice | Effusive praise for trivial work, which signals low expectations | | About effort, strategy, and improvement on the student's own past work | About ability ("you're so smart") or rank against classmates | | In a form the student can accept — privately, for many adolescents | Public praise that embarrasses, which functions as punishment | Quantity matters too. A three-year observational study of elementary classrooms found that the higher a teacher's praise-to-reprimand ratio, the more time students spent on task, with no ceiling.[8] Many programs recommend at least four praise statements per reprimand; most classrooms run far below that, because misbehavior is more salient than quiet work, so the ratio has to be engineered — a tally on the desk, a timer that prompts a scan of the room. ## Token economies: how to build one that works A [token economy](https://operantconditioning.com/glossary/#token-economy) delivers points, stars, or chips contingent on target behaviors and lets students exchange them later for backup reinforcers. Ayllon and Azrin developed it on a psychiatric ward in the early 1960s,[9] and classrooms adopted it almost immediately: in a 1967 program, a public-school class of seventeen children with serious behavior problems earned ratings exchangeable for small prizes, and disruption dropped sharply and stayed down.[4] Tokens work because they are [generalized conditioned reinforcers](https://operantconditioning.com/positive-reinforcement/#generalized-reinforcers): instant to deliver, slow to satiate, exchangeable for whatever each student values. Alan Kazdin's reviews, from 1972 to 1982, judged the evidence strong across schools, wards, prisons, and homes, and named the recurring weaknesses: gains faded when tokens were withdrawn, did not transfer to settings without tokens, and depended on staff who often stopped delivering them.[10][11] A well-built system anticipates all three. 1. **Define three to five target behaviors as things to do.** "Start work within a minute of the bell," not "don't waste time." 2. **Deliver the token in a second, with specific praise every time.** Praise inherits the token's power and still works when the tokens are gone. 3. **Build the menu from what students do when free to choose,** start cheap, and exchange the same day. Long saving periods are a schedule most children have not yet learned to work on. 4. **Keep fines rare and small.** [Response cost](https://operantconditioning.com/negative-punishment/) works only while students have tokens to lose. 5. **Plan the fade from day one.** Raise prices, space the exchanges, shift from points to praise to natural consequences — and count the target behavior weekly. ## The Good Behavior Game The **Good Behavior Game** is one of the most thoroughly studied classroom management procedures there is, and it costs nothing. In the 1969 original, a fourth-grade class was split into two teams during math and reading; a mark went against a team whenever any member left their seat or talked out; and teams finishing under a set number of marks — both could win — earned privileges such as end-of-day free time and lining up first. Out-of-seat and talking-out behavior fell dramatically and returned when the game was withdrawn.[12] It is an *interdependent group contingency*: each student's behavior affects the team's outcome, so peers prompt and reinforce one another. The follow-up is what makes it remarkable. In the mid-1980s Sheppard Kellam's group randomly assigned first-grade classrooms in Baltimore public schools to the game, to a curriculum intervention, or to standard practice, and followed the children into adulthood. At ages 19 to 21, the men who had been the most aggressive first-graders and had played the game showed lower rates of drug and alcohol disorders, regular smoking, and antisocial personality disorder than their counterparts from control classrooms.[13] > **How to run the Good Behavior Game** > > 1. **Pick two or three rules** stated as visible behaviors: in your seat, hand up to talk, hands to yourself. > 2. **Split the class into teams** of roughly equal difficulty, and announce when the game is on — one short period a day at first, in the lesson that suffers most. > 3. **Mark violations calmly and visibly,** with no lecture, and keep teaching. > 4. **Let every team under the criterion win,** immediately at first — five minutes of a preferred activity — then at the end of the day, then the week. > 5. **Expand slowly,** stop announcing the start, and give a deliberate saboteur a team of one. ## School-wide PBIS **Positive Behavioral Interventions and Supports** (PBIS) applies the same logic to a whole building: three to five expectations, taught explicitly in each setting; tickets or praise for meeting them; a predictable response to violations; monthly review of office-referral data; and tiered support, from the universal system up to individual plans built on a functional behavioral assessment. The evidence is respectable rather than spectacular: in a randomized trial of 37 Maryland elementary schools, the 21 trained in PBIS showed reductions in office discipline referrals and suspensions relative to the 16 that were not.[14] PBIS is not a product. It is the adults in a building agreeing on what they will reinforce, and then doing it. ## Skinner's teaching machines, precision teaching, and Direct Instruction Skinner's own contribution was to instruction. His **teaching machine** presented a frame of material, required a composed answer, revealed the correct one immediately, and moved on through steps so small that the student was nearly always right.[2][15] He credited Sidney Pressey's devices of the 1920s for scoring answers automatically; what was new was the program. **Programmed instruction** named the principles — small steps, active responding, immediate feedback, self-pacing — so that the reinforcement of being right arrives thousands of times a semester rather than a few dozen. The machines are gone; the principles survive in mastery learning and adaptive tutoring software. **Precision teaching**, developed by Skinner's student Ogden Lindsley, kept the laboratory's measure: rate. Students do brief timed practice and chart correct and incorrect responses per minute, and the teacher changes the teaching when the curve flattens. The target is *fluency* — accuracy plus speed — and the rule is that the child knows best: if the child is not learning, the program is wrong.[16] **Direct Instruction**, Siegfried Engelmann's scripted, fast-paced lessons with choral responding, immediate correction, and mastery before advancement, is the most evaluated instructional program of the past half-century; a 2018 meta-analysis of more than 300 studies found consistently positive effects across subjects and across five decades of research.[17] All three share what Skinner saw in 1953: many responses, immediate consequences, a program that keeps the learner succeeding. ## The negative-reinforcement trap: misbehavior that escapes work The commonest classroom mistake is not a failure of punishment but a misreading of function. A student who finds the work aversive — too hard, too long, humiliating in front of peers — discovers that acting out makes it go away: the lesson stops, the worksheet is put aside, the student is sent to the hallway or the office. Each of those is [negative reinforcement](https://operantconditioning.com/negative-reinforcement/), and the behavior that produced it will happen more. The intended punishment is functioning as a reward. In functional analyses of problem behavior, escape from demands is the single most common maintaining consequence identified — 38% of 152 cases of self-injury in the largest series,[18] and time-out is the classic casualty: in a pair of experiments it reduced problem behavior when the environment the child left was rich, and *increased* it when leaving offered an escape.[19] Since 1997, U.S. special-education law has required a functional behavioral assessment when a student with a disability is removed from school for behavior that turns out to be a manifestation of the disability, for exactly this reason. The escape-maintained student needs work pitched where success is likely, an acceptable way to ask for a break that is honored every time, and — where safe — the demand kept in place so that acting out no longer ends it. [Escape extinction vs. ignoring ›](https://operantconditioning.com/extinction/#extinction-is-not-ignoring) ## Extinction and planned ignoring: what they can and can't do **Planned ignoring** — withholding attention from a behavior that attention has been maintaining — is [extinction](https://operantconditioning.com/extinction/), and it fits exactly one class of behavior: attention-seeking aimed at the teacher. It is the wrong tool for behavior maintained by peers' laughter, for escape, and for anything dangerous. Even where it fits, it works slowly, produces an [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) first, and on its own did little in the classic experiment; combined with praise for the behavior wanted instead, it worked well.[3] The professional version is differential reinforcement: the blurt is ignored and the raised hand is called on within seconds, every time. > **Ignoring is a decision, not a default** > > Ignore a behavior only after answering three questions: is my attention what maintains it, can I withhold it every single time, and can I outlast the burst? Ignoring for four minutes and reacting on the fifth puts the behavior on an intermittent schedule and makes it stronger. ## Do rewards kill intrinsic motivation? The overjustification caveat The objection every teacher hears is that rewards turn learning into work. The evidence is real but narrow. In the classic study, preschoolers who liked drawing were promised a certificate for drawing with markers; a week or two later they drew less in free play than children given the same certificate unexpectedly, or none at all.[20] A meta-analysis of 128 experiments found that *expected, tangible* rewards for an already interesting activity reduced later free-choice engagement, while praise did not — it raised it in college students and had no reliable effect in children; a competing meta-analysis found the effect small and confined to narrow conditions.[21][22] So: use tokens and tangible rewards for behavior students do not already do, fade them as the behavior meets its natural reinforcers, use praise and progress feedback freely, and never pay students for what they already do for pleasure. [The intrinsic-motivation evidence in detail ›](https://operantconditioning.com/positive-reinforcement/#does-positive-reinforcement-undermine-intrinsic-motivation) ## Before you reach for consequences: an antecedent checklist Consequences are the last third of the [ABC model](https://operantconditioning.com/abc-model/). Most persistent classroom problems have an antecedent solution that costs less than any reward or penalty. - **Taught, not just posted?** Rules alone change little.[3] Model, practice, and reinforce each one where it applies. - **Doable?** Work that is too hard is the antecedent for escape; work that is too easy, for everything else. - **Transitions signaled?** A two-minute warning and a visible timer are discriminative stimuli for stopping. - **Enough chances to respond?** Choral responses, whiteboards, and turn-and-talk multiply the opportunities for reinforcement; a lecture offers almost none. - **Motivating operations?** A hungry, tired, or anxious student values escape more and praise less. - **A cue for the behavior you want?** A voice-level chart, a first–then card, a checklist on the desk: [stimulus control](https://operantconditioning.com/stimulus-control/) is cheaper than consequence control. ## Common mistakes with operant conditioning in the classroom - **Reprimanding the attention-starved.** For some students a scolding is the richest attention of the day. If the behavior rises, the reprimand is reinforcing it. - **Sending the escape-maintained student out.** The hallway, the office, and in-school suspension all end the lesson. Check function before removing anyone. - **Praising the class instead of the behavior.** "Good job, everyone" reinforces nothing in particular. - **Announcing consequences you will not deliver.** Unenforced threats teach that the rule is inactive. Small consequences every time beat large ones occasionally. - **Never fading.** Tokens are scaffolding. If the system is the same in June as in September, the behavior belongs to the system, not the student. The through-line is the same as everywhere else: identify the behavior, find what actually reinforces it, make the right behavior easy, reinforce it immediately and often, and let the data decide. [All applications of operant conditioning ›](https://operantconditioning.com/applications/) · [Operant conditioning in parenting ›](https://operantconditioning.com/parenting/) ## Key takeaways - Every classroom already runs on consequences; the question is whether they are arranged on purpose. Skinner diagnosed education's failures as failures of reinforcement: too few reinforced responses, too long a delay, no program of small steps. - Teacher attention is the cheapest reinforcer in the building, but most classroom praise is not reinforcement. Praise works when it is contingent, specific, credible, and about effort or strategy, and the praise-to-reprimand ratio has to be engineered because misbehavior is more salient than quiet work. - Tokens are generalized conditioned reinforcers and the evidence for token economies is strong, but gains fade when tokens are withdrawn unless the fade is planned from day one. The Good Behavior Game, an interdependent group contingency, costs nothing and has randomized trials with benefits into adulthood. - A consequence is whatever the behavior says it is. A reprimand can be attention, and sending an escape-maintained student out of the room is negative reinforcement, so function has to be checked before anything is removed or ignored. - Expected, tangible rewards for an already interesting activity can reduce later free-choice engagement; praise and informational feedback do not. Use tangible rewards for behavior students do not already do, fade them, and fix antecedents first: rules taught, work doable, transitions signaled. ### Check yourself **A teacher gives a quiet reprimand every time a student shouts across the room, and over the month the shouting increases. What is going on?** For a student who gets little adult attention, a reprimand is attention, so the intended punishment is functioning as positive reinforcement. The behavior is the evidence: if it rises, the reprimand is reinforcing it. A consequence is whatever the behavior says it is, not what the teacher meant it to be. **A student acts out during long worksheets and is sent to the office each time, yet the acting out gets worse. The teacher concludes the office is not a harsh enough punishment. What is the better analysis?** The behavior is escape-maintained. Being sent out ends work the student finds aversive, so each trip to the office is negative reinforcement, and the behavior that produced it will happen more. The fix is antecedent and functional: work pitched where success is likely, an acceptable way to ask for a break that is honored every time, and, where safe, the demand kept in place so that acting out no longer ends it. **A teacher decides to ignore a student's blurting, holds out for four minutes, then reacts on the fifth. What has happened to the blurting?** It has been put on an intermittent schedule, which makes it stronger and more persistent. Planned ignoring is extinction, and it works only when the teacher's attention is the maintaining reinforcer, only when it is withheld every single time, and only when the teacher can outlast the extinction burst, ideally while calling on raised hands within seconds instead. **Two classes run token economies. In one the system is identical in June and September; in the other, tokens have been faded toward praise and natural consequences. Which class's behavior is more likely to survive the end of the year, and why?** The faded one. Kazdin's reviews found that gains faded when tokens were withdrawn and did not transfer to settings without tokens, so a system that never fades leaves the behavior belonging to the system, not the student. Pairing every token with specific praise lets praise inherit the token's power and keep working when the tokens are gone. **Explain it to a friend.** Explain why sending a disruptive student to the hallway can make the disruption worse, without using the words "positive" or "negative." ## Frequently asked questions **Does positive reinforcement work in the classroom?** Yes; it is the best-supported classroom-management tool there is. Contingent teacher attention and specific praise reliably increase on-task behavior, and token systems and the Good Behavior Game have decades of controlled studies behind them. It fails when praise is vague or non-contingent, when the "reward" does not actually reinforce that student, or when it arrives too late. **What is a token economy in the classroom?** A system in which students earn tokens — points, stamps, chips — immediately after target behaviors and exchange them later for privileges, activities, or small prizes. Good systems start with cheap, same-day exchanges, pair every token with specific praise, keep fines rare, and fade the tokens as praise and natural consequences take over. **What is the Good Behavior Game?** A classroom procedure, first published in 1969, in which the class is split into teams, a mark is recorded against a team whenever a member breaks a posted rule, and every team finishing under a set number of marks wins a small privilege. Randomized trials in Baltimore found that first-graders who played it had lower rates of substance-use disorders and antisocial personality disorder as young adults. **Do rewards kill intrinsic motivation?** Only under specific conditions. Expected, tangible rewards for an activity students already find interesting can reduce their later free-choice engagement with it; praise, informational feedback, and rewards for behavior students would not otherwise do generally do not. Use tangible rewards to establish behavior that is not happening, fade them, and rely on praise and feedback for the rest. **What is planned ignoring?** Deliberately withholding attention from a behavior that attention has been maintaining, so that it extinguishes. It works only when the teacher's attention is the reinforcer, only when applied every time, and only alongside reinforcement of the behavior you want instead. It is the wrong tool for behavior maintained by peers, for behavior that escapes work, and for anything dangerous. **How do I use operant conditioning in my classroom?** Start with antecedents: teach the expectations, make the work doable, signal transitions. Raise your ratio of specific praise to reprimands well above four to one, ignore minor attention-seeking while reinforcing the alternative, and add the Good Behavior Game or a simple token economy for the periods that need it. Before punishing persistent misbehavior, ask what it is getting the student; if the answer is escape from work, removal will make it worse. ## References 1. Hall, R. V., Lund, D., & Jackson, D. (1968). Effects of teacher attention on study behavior. *Journal of Applied Behavior Analysis, 1*(1), 1–12. 2. Skinner, B. F. (1954). The science of learning and the art of teaching. *Harvard Educational Review, 24*(2), 86–97. 3. Madsen, C. H., Jr., Becker, W. C., & Thomas, D. R. (1968). Rules, praise, and ignoring: Elements of elementary classroom control. *Journal of Applied Behavior Analysis, 1*(2), 139–150. 4. O'Leary, K. D., & Becker, W. C. (1967). Behavior modification of an adjustment class: A token reinforcement program. *Exceptional Children, 33*(9), 637–642. 5. Thomas, D. R., Becker, W. C., & Armstrong, M. (1968). Production and elimination of disruptive classroom behavior by systematically varying teacher's behavior. *Journal of Applied Behavior Analysis, 1*(1), 35–45. 6. Brophy, J. (1981). Teacher praise: A functional analysis. *Review of Educational Research, 51*(1), 5–32. 7. Henderlong, J., & Lepper, M. R. (2002). The effects of praise on children's intrinsic motivation: A review and synthesis. *Psychological Bulletin, 128*(5), 774–795. 8. Caldarella, P., Larsen, R. A. A., Williams, L., Downs, K. R., Wills, H. P., & Wehby, J. H. (2020). Effects of teachers' praise-to-reprimand ratios on elementary students' on-task behaviour. *Educational Psychology, 40*(10), 1306–1322. 9. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 10. Kazdin, A. E., & Bootzin, R. R. (1972). The token economy: An evaluative review. *Journal of Applied Behavior Analysis, 5*(3), 343–372. See also Kazdin, A. E. (1977). *The Token Economy: A Review and Evaluation*. Plenum Press. 11. Kazdin, A. E. (1982). The token economy: A decade later. *Journal of Applied Behavior Analysis, 15*(3), 431–445. 12. Barrish, H. H., Saunders, M., & Wolf, M. M. (1969). Good behavior game: Effects of individual contingencies for group consequences on disruptive behavior in a classroom. *Journal of Applied Behavior Analysis, 2*(2), 119–124. 13. Kellam, S. G., Brown, C. H., Poduska, J. M., et al. (2008). Effects of a universal classroom behavior management program in first and second grades on young adult behavioral, psychiatric, and social outcomes. *Drug and Alcohol Dependence, 95*(Suppl. 1), S5–S28. 14. Bradshaw, C. P., Mitchell, M. M., & Leaf, P. J. (2010). Examining the effects of schoolwide positive behavioral interventions and supports on student outcomes: Results from a randomized controlled effectiveness trial in elementary schools. *Journal of Positive Behavior Interventions, 12*(3), 133–148. 15. Skinner, B. F. (1958). Teaching machines. *Science, 128*(3330), 969–977. 16. Lindsley, O. R. (1992). Precision teaching: Discoveries and effects. *Journal of Applied Behavior Analysis, 25*(1), 51–57. 17. Stockard, J., Wood, T. W., Coughlin, C., & Rasplica Khoury, C. (2018). The effectiveness of Direct Instruction curricula: A meta-analysis of a half century of research. *Review of Educational Research, 88*(4), 479–507. 18. Iwata, B. A., Pace, G. M., Dorsey, M. F., Zarcone, J. R., Vollmer, T. R., Smith, R. G., Rodgers, T. A., Lerman, D. C., Shore, B. A., Mazaleski, J. L., Goh, H.-L., Cowdery, G. E., Kalsher, M. J., McCosh, K. C., & Willis, K. D. (1994). The functions of self-injurious behavior: An experimental-epidemiological analysis. *Journal of Applied Behavior Analysis, 27*(2), 215–240. The method itself: Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 19. Solnick, J. V., Rincover, A., & Peterson, C. R. (1977). Some determinants of the reinforcing and punishing effects of timeout. *Journal of Applied Behavior Analysis, 10*(3), 415–424. 20. Lepper, M. R., Greene, D., & Nisbett, R. E. (1973). Undermining children's intrinsic interest with extrinsic reward: A test of the "overjustification" hypothesis. *Journal of Personality and Social Psychology, 28*(1), 129–137. 21. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 22. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423. ## Related - [All applications](https://operantconditioning.com/applications/): ABA, health, work, animals, technology, and more — with the evidence rated. - [Operant conditioning in parenting](https://operantconditioning.com/parenting/): Tantrums, time-out, sticker charts, and what the evidence says. - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): What a reinforcer is, what makes it work, and why so many attempts fail. --- # Operant Conditioning in Parenting: Praise, Tantrums, Time-Out, Sticker Charts, Bedtime, and What the Evidence Says > Operant conditioning in parenting: praise, the coercion trap behind tantrums, time-out, sticker charts, bedtime extinction, and the evidence on spanking. - Source: https://operantconditioning.com/parenting/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-09 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications · Parenting* Every household runs on consequences, whether or not anyone planned them. What the best-tested parenting programs teach, why tantrums train parents as surely as parents train children, and what the evidence says about time-out, sticker charts, bedtime, and spanking. > **Definition** > > **Operant conditioning in parenting** is the use of consequences — attention, praise, privileges, and their removal — and of antecedents such as routines, clear instructions, and cues to make wanted behavior more frequent and unwanted behavior less frequent. > > Parents cannot opt out; the only choice is whether the contingencies are deliberate. The deliberate version, behavioral parent training, has been refined in controlled trials for half a century and is among the best-supported treatments for childhood behavior problems.[1][2] **In brief** - Parents cannot opt out of consequences; the only choice is whether the contingencies are deliberate, and behavioral parent training is the deliberate version. - A parent's most powerful reinforcer is attention, and households usually pay it to misbehavior; catching them being good reverses the payroll. - When a parent gives in to a tantrum, both the tantrum and the giving in are [negatively reinforced](https://operantconditioning.com/negative-reinforcement/#the-negative-reinforcement-trap-how-it-maintains-problem-behavior): Patterson's coercive family process. ## The four quadrants at home | Procedure | What the parent does | Example | Where it goes wrong | | --- | --- | --- | --- | | Positive reinforcement | Adds something after the behavior | "You put your plate in the sink — thank you." Plates go in the sink more often. | Delivered late, vaguely, or for behavior that was not happening | | Negative reinforcement | Removes something after the behavior | The nagging stops the moment the coat goes on. The coat goes on faster. | Usually runs the other way: the whining stops when the parent gives in | | Positive punishment | Adds something aversive | A sharp "No" as a hand reaches for the stove. Reaching decreases. | Scolding is attention, and attention reinforces | | Negative punishment | Removes a reinforcer | Hitting a sibling ends the game: two minutes of time-out. Hitting decreases. | Time-out from a boring room is an escape, not a loss | | Extinction | Stops delivering the reinforcer | Whining for dessert never again produces dessert. Whining declines, after a burst. | Giving in on the fifth whine teaches five-whine persistence | The rows are not equal. Positive reinforcement builds behavior, and parent-training programs spend most of their time on it. Negative punishment is for the short list of behaviors that must stop. Positive punishment has the weakest evidence and the worst side effects. Extinction is the tool everyone uses accidentally, in reverse. [Examples for every quadrant ›](https://operantconditioning.com/examples/) ## Catch them being good: the attention economy of a household A parent's most powerful reinforcer is attention, and a household has a fixed budget of it. Think of the budget as a payroll. Misbehavior is paid immediately, reliably, and in full: the shout, the lecture, the parent crossing the room. Good behavior — playing quietly, sharing unasked, shoes by the door — mostly goes unpaid, because a quiet child is a chance to do something else. Over a thousand repetitions the child learns exactly which behaviors the household pays for. **Catching them being good** reverses the payroll. Notice the behavior you want and name it, specifically and at once: "You waited until I was off the phone — that was patient," not a global "good girl" an hour later. Parent-training programs treat this **labeled praise** as a skill to be modeled and rehearsed in session; Parent–Child Interaction Therapy, for instance, coaches parents until they deliver ten labeled praises in five minutes of play, because that is what it takes to overturn the existing allocation.[1][19] > **A scolding is still attention** > > If a behavior keeps happening after it has been scolded a hundred times, the scolding is not punishing it. It is probably paying for it. Watch what the behavior does, not what you meant. ## The reinforcement trap: Patterson's coercive family process Gerald Patterson's observations of families of aggressive children in their homes found a pattern that is the most important idea on this page. The parent makes a request; the child whines, argues, or screams; the parent, worn down, withdraws the request or hands over what the child wanted. Two things were just learned. The child's aversive behavior was *negatively reinforced*: it made the demand disappear. The parent's giving in was also negatively reinforced: it made the screaming stop. Each has trained the other, and neither meant to.[3] The cycle escalates because it also involves [shaping](https://operantconditioning.com/shaping/). A parent who holds out through the whine and gives in at the scream has reinforced screaming; next time the child starts closer to the scream. A parent who gives in only sometimes has put the tantrum on an intermittent schedule, the schedule that produces the most persistent behavior of all. Patterson called the result the **coercive family process**, and his group's longitudinal work traced where it leads: coercive exchanges in early childhood predict conduct problems, rejection by ordinary peers, school failure, drift toward deviant friends, and delinquency in adolescence.[4] [More on negative-reinforcement traps ›](https://operantconditioning.com/negative-reinforcement/#the-negative-reinforcement-trap-how-it-maintains-problem-behavior) > **Both of you are being trained** > > The question after any standoff is not "who won?" but "what did each of us just get?" If the child got out of the task and you got out of the noise, both behaviors will be back tomorrow, slightly stronger. ## What parent management training teaches The remedy Patterson's group developed became **Parent Management Training** (PMT), now a family of programs with a shared core.[5] Alan Kazdin's version, refined over decades of trials with children referred for oppositional and aggressive behavior, teaches parents a small set of skills through modeling, role-play, and rehearsal — not lectures — and sends them home with practice assignments.[1] Reviews applying the strictest criteria for evidence-based treatment consistently put parent training at the top of the list for childhood disruptive behavior; the Oregon model was among the first to meet the "well-established" standard.[2] The order matters: praise and attending come first and take most of the program's time, commands next, time-out and point systems last. | Skill | What it looks like | Operant mechanism | | --- | --- | --- | | **Specific, immediate praise** | "You started your homework the first time I asked." | Positive reinforcement; pairing praise with attention makes praise itself a reinforcer | | **Planned ignoring** | No eye contact, comment, or reaction to whining; full attention the moment it stops | Extinction of attention-maintained behavior, plus reinforcement of its absence | | **Effective commands** | One instruction at a time, stated as a direction ("Put the blocks in the box"), up close, followed by a short wait | A clear discriminative stimulus; vague, chained, or question-form commands evoke less compliance[6] | | **Time-out** | Brief, calm, immediate, for a short list of serious behaviors | Negative punishment: removal of access to reinforcement | | **Point charts** | Points for two or three target behaviors, exchanged daily for privileges | A token economy; conditioned reinforcers bridge the delay | | **Monitoring** (older children) | Knowing where the adolescent is and with whom | Keeps consequences contingent; unmonitored behavior meets no parental consequence | ## Time-out done correctly **Time-out** is short for time-out from positive reinforcement, and the full name is the instruction. It works only if the environment the child leaves — the **time-in** — is warm, engaged, and reinforcing; a child removed from a dull room to a bedroom full of toys has been rewarded, not punished. Done properly it is immediate, brief (a few minutes; roughly a minute per year of age for young children), calm, free of lectures, and ended when the child has been quiet for a moment rather than while she is still protesting. Time-out has been studied with parents for more than fifty years and is part of every major evidence-based parenting program.[7] The claim that it damages attachment has been examined directly: a 2019 review concluded that, implemented as designed, the evidence does not support the concern, and that steering parents away from time-out risks pushing them toward tools with far worse evidence.[8] The American Academy of Pediatrics, which advises against spanking and recommends positive reinforcement and limit-setting in its place, teaches time-out in its guidance for parents.[9] [The step-by-step time-out procedure ›](https://operantconditioning.com/negative-punishment/#how-to-do-time-out-correctly) ## Sticker charts and token systems that work A sticker chart is a token economy, and it obeys the same rules as the ones on hospital wards and in classrooms.[10] Most home charts fail on one of four details. 1. **Immediacy.** The sticker goes on within seconds of the behavior, with praise — not at bedtime for "being good today." For a young child the sticker is the consequence; the prize it buys is a bonus. 2. **Small steps.** "A clean room for a week" is a goal, not a behavior. Start with "clothes in the hamper before dinner" and shape upward once that is reliable. 3. **Cheap, quick exchanges.** Three stickers should buy something today — choosing dinner, ten extra minutes before bed, a game with a parent. 4. **A planned fade.** Once the behavior is steady, raise the price, space the exchanges, and let praise and natural consequences take over. A chart still running unchanged six months later has become the reason for the behavior. Two cautions. Do not fine stickers off the chart for misbehavior; a child at zero has nothing to work for. And do not pay for things the child already loves: tangible rewards for an activity that was already interesting can reduce interest in it once the rewards stop, whereas praise, and rewards for behavior the child would not otherwise do, generally do not.[11] Charts are for building behavior that is not happening, not for decorating behavior that is. ## Bedtime and the extinction burst Bedtime is where most parents meet extinction for the first time, usually by accident. A child who cries when the parent leaves, and whose parent returns, has been reinforced for crying; stop returning, and the crying goes through an [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) — louder and longer on the first nights — before it declines. The [classic 1959 case](https://operantconditioning.com/extinction/#examples-of-extinction) followed exactly this course, including a relapse when a relative went back in.[12] Sleep researchers have tested both the unmodified version and **graduated extinction**, in which the parent checks in briefly at lengthening intervals rather than not at all. A review of 52 treatment studies found that the large majority reported clinically significant improvements, with unmodified extinction and preventive parent education best supported and graduated extinction close behind.[13] A randomized trial that followed infants for a year found graduated extinction reduced the time to fall asleep and the number of night wakings, with no rise in stress hormones — infant cortisol was, if anything, slightly lower — and no effect on attachment or later emotional and behavioral problems;[14] a five-year follow-up of a larger trial found neither lasting harms nor lasting benefits.[15] > **What honesty about bedtime extinction sounds like** > > It works for many families, it is not the only option, and the burst is not a sign that it is failing — it is the procedure working. Decide in advance whether you can hold the line for a week, get every caregiver to agree, and expect a smaller recovery after any night the crying is answered. A gentler method you will actually follow beats a faster one you will abandon on night three, because abandoning it on night three is intermittent reinforcement of the loudest crying yet. ## Spanking: what the evidence says Spanking is [positive punishment](https://operantconditioning.com/positive-punishment/), and it has been studied more than any other parenting consequence. The most careful meta-analysis, restricted to ordinary open-handed spanking rather than abuse, pooled 75 studies covering 160,927 children and found spanking significantly associated with 13 of the 17 outcomes examined — more aggression, more antisocial behavior, more mental-health problems, worse parent–child relationships — every one in the harmful direction and none in the beneficial direction.[16] The data are mostly correlational, and difficult children are spanked more; but the associations held in the longitudinal studies the review examined, and no comparable evidence finds benefits. In 2018 the American Academy of Pediatrics advised parents not to spank, hit, or otherwise physically punish children, not to use words that shame or humiliate, and to rely instead on positive reinforcement, limit-setting, and brief time-out.[9] The operant analysis predicts the result. Spanking as practiced is delayed, inconsistent, and escalating; it delivers a burst of parental attention; it models the aggression it is meant to stop; and it makes the parent an aversive stimulus, which produces avoidance and lying. [The corporal-punishment evidence in detail ›](https://operantconditioning.com/positive-punishment/#corporal-punishment-what-the-evidence-shows) ## Consistency beats severity Parents whose consequences are not working tend to make them bigger. The laboratory says to make them more reliable instead. In Azrin and Holz's classic review, punishment delivered intermittently was far less effective than punishment delivered every time, and punishment introduced mildly and then escalated produced adaptation: the organism learned to tolerate each new level.[17] In practice consistency means three things. **Immediacy**: the consequence follows within seconds or minutes, which is why "wait until your father gets home" changes nothing except the child's feelings about the front door. **Contingency**: it follows the behavior and nothing else — not withheld because the parent is busy, not delivered because the parent is tired. **Agreement**: both parents, the grandparents, and the babysitter enforce the same short list of rules, because a rule enforced by one adult in three is a rule on a variable schedule. A rule you cannot enforce every time is better dropped than announced. ## The Premack principle at home "First homework, then screen" is the [Premack principle](https://operantconditioning.com/premack-principle/): a more probable behavior can reinforce a less probable one.[18] Grandmothers knew it as "first your vegetables, then dessert." Its power at home is that it needs no prizes: the reinforcers are the things the child already does when free — playing outside, watching a show, being read to — made contingent on the things he does not. The preferred activity comes *after*, promptly, and in an amount small enough that the child could not have had it anyway. "You can play now if you promise to do homework later" reverses the order and reinforces promising. ## Screens and the variable-ratio problem Screens complicate every contingency in the house for a specific reason: the device delivers its own reinforcement on a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/#why-variable-ratio-is-so-powerful-and-so-dangerous), the arrangement that produces the highest rates of behavior and the greatest resistance to extinction. A game or a feed pays out unpredictably, so "one more minute" is not defiance; it is the burst a pigeon shows when the key stops paying. Ending screen time is therefore a loss, the transition off a screen is the most reliable trigger of protest in the modern household, and a parent who hands the tablet back to end the protest has reinforced it. - Put screens on the "then" side of every first–then, never the "first." - Signal the end with a timer the child can see, so the timer rather than the parent is the cue for stopping. - Never hand over a screen to stop a tantrum; that is the supermarket candy with a battery. - Set the daily amount when everyone is calm; a limit renegotiated during a protest has just been placed on an intermittent schedule. [How apps and games use operant conditioning ›](https://operantconditioning.com/applications/#technology) ## What not to do - **Bribing before the behavior.** "If you stop screaming I'll buy you the toy" delivers the reinforcer for screaming. Reinforcement follows the behavior you want; a bribe precedes it and is triggered by the one you don't. - **Delayed consequences.** A privilege lost next weekend for something done on Tuesday is felt as arbitrary, and by Saturday it is punishing Saturday's behavior. - **Punishing after the fact.** The mess discovered an hour later cannot be connected to the act, for a puppy or a toddler. A consequence discovered late is better skipped than delivered. - **Rewards that satiate.** Candy after every good act stops working by mid-afternoon. Attention, activities, and choices satiate far more slowly, and they are free. - **Threats you will not carry out.** An unenforced warning teaches that warnings are noise, so the next real one is ignored too. ## A worked example: the supermarket tantrum in ABC The [ABC model page analyzes a supermarket tantrum](https://operantconditioning.com/abc-model/#a-child-s-tantrum-in-the-supermarket) and shows how to intervene at the antecedent, the behavior, and the consequence. Here is what that analysis leaves out: how the tantrum was built, and why there are two contingencies in the aisle rather than one. | Whose behavior | Antecedent | Behavior | Consequence | What was learned | | --- | --- | --- | --- | --- | | **The child's** | Checkout aisle, candy at eye level, parent busy, no snack since lunch | Asks, then whines, then screams and drops to the floor | Candy appears; so does the parent's full attention | Screaming is positively reinforced by candy and attention | | **The parent's** | Screaming child, a queue watching, a card machine waiting | Hands over the candy | The screaming stops; the stares stop | Giving in is negatively reinforced by the end of the noise and the embarrassment | Play it forward. On the first Saturday the parent gives in at the whine. On the second, resolved to be firmer, she holds out through the whine and gives in at the scream — which shapes the scream. On the third she holds out through the scream and gives in when the child hits the floor. By the fourth Saturday the child starts on the floor, because that is the response that has been paid, and the parent gives in at once, because that is the response that ends it fastest. Nobody in this story is weak or bad. Two ordinary learners have shaped each other on an intermittent schedule, and the aisle now controls both of them. The way out is the one the ABC page lays out: change the antecedent, teach and reinforce a replacement, and make sure candy never again follows a scream. The honest addition is that the next scream will be the loudest, that giving in to it will make the one after worse, and that the parent who has decided in advance what she will do — with a snack and a job for the child already in her bag — rarely reaches the floor at all. [All applications of operant conditioning ›](https://operantconditioning.com/applications/) · [Operant conditioning in the classroom ›](https://operantconditioning.com/classroom/) ## Key takeaways - Every household runs on consequences whether or not anyone planned them. Positive reinforcement builds behavior and takes most of a parent-training program's time; negative punishment is for the short list of behaviors that must stop; positive punishment has the weakest evidence and the worst side effects. - Attention is the household's most powerful reinforcer, and misbehavior is usually paid first and in full. Labeled praise, specific and immediate, reverses the allocation; a scolding is still attention, so a behavior that survives a hundred scoldings is probably being paid by them. - In the coercive family process, the child's whining is negatively reinforced when the demand disappears and the parent's giving in is negatively reinforced when the noise stops. Holding out and then giving in shapes a louder version and puts it on an intermittent schedule, the most persistent of all. - Time-out works only when time-in is warm and reinforcing, and done as designed it does not damage attachment. A sticker chart is a token economy: immediate delivery, small steps, cheap same-day exchanges, and a planned fade. - Spanking is associated with worse outcomes on 13 of 17 measures and better outcomes on none, and the operant analysis predicts why. Consistency beats severity: immediacy, contingency, and agreement among every adult, and a rule that cannot be enforced every time is better dropped than announced. ### Check yourself **A four-year-old keeps drawing on the wall even though he is scolded every time. The parent concludes the punishment is not strong enough. What does the behavior say?** A scolding is attention, and attention reinforces. If the behavior keeps happening after it has been scolded a hundred times, the scolding is not punishing it; it is probably paying for it. Watch what the behavior does, not what the consequence was meant to do, and move the attention to the behavior wanted instead. **During a dull afternoon a child hits her brother and is sent to her toy-filled bedroom for two minutes. Is this time-out?** Not in function. Time-out is short for time-out from positive reinforcement, so it works only if the time-in the child leaves is warm, engaged, and reinforcing. Leaving a boring room for a room full of toys is an escape and a reward, not a loss, and hitting is likely to go up. **A parent who used to give in at the first whine resolves to be firmer. She now holds out through the whine and gives in when the child screams. Has she made progress?** No. She has shaped screaming: the louder response is the one that was paid, so next time the child starts closer to the scream. Because she gives in only sometimes, the tantrum is now on an intermittent schedule, the schedule that produces the most persistent behavior of all. The way out is decided in advance, not during the standoff. **A parent stops returning to a crying child at bedtime. On the first two nights the crying is louder and longer than ever. Is the procedure failing?** No. That is the extinction burst, and it is the procedure working: crying that returning had reinforced gets louder before it declines. Giving in on night three would be intermittent reinforcement of the loudest crying yet, which is why the decision to hold the line, and every caregiver's agreement, has to come before night one. **Explain it to a friend.** Explain how a tantrum trains a parent, using a standoff from your own household or childhood as the example. ## Frequently asked questions **Does positive reinforcement work on children?** Yes; it is the core of every evidence-based parenting program. Specific praise and attention delivered immediately after a behavior reliably increase it. It fails when the praise is vague or late, when the "reward" is not actually reinforcing for that child, or when misbehavior is still being paid more reliably in attention. **Is time-out harmful?** Not when done as designed: brief, calm, immediate, for a short list of serious behaviors, from a home rich in positive attention. A 2019 review of the claim that time-out damages attachment found the evidence does not support it, and the American Academy of Pediatrics, which advises against physical punishment, includes time-out in its guidance for parents. It fails — and can backfire — when time-in is not rewarding or the child is glad to leave. **Do sticker charts work?** They work when they follow the rules of a token economy: the sticker arrives within seconds of a small, specific behavior; stickers buy something the same day at first; nothing is taken away for misbehavior; and the chart is faded once the behavior is steady. They fail when the goal is too large, the prize too distant, the chart abandoned after a week, or the reward attached to something the child already enjoys. **Why does my child ignore consequences?** Usually because the consequence is not functioning as one. It may be too delayed, delivered only sometimes, or not something the child values losing; it may even be reinforcing — a scolding is attention, and being sent to a room full of toys is a reward. Ask what the behavior is getting the child (attention, escape, an item) and whether your consequence removes that or supplies it. **Is spanking effective?** Not by the evidence. The largest meta-analysis of spanking alone, covering 75 studies and more than 160,000 children, found it associated with worse outcomes on 13 of 17 measures — including more aggression and antisocial behavior — and better outcomes on none. The American Academy of Pediatrics advises against it and recommends positive reinforcement, limit-setting, redirection, and clear expectations instead. **How do I stop giving in to tantrums?** Decide before the tantrum, not during it: giving in at the loudest point reinforces the loudest version and puts the tantrum on an intermittent schedule. Change the antecedents (a snack, a job, the rule stated before entering the store), teach and reinforce a replacement way of asking, ignore the tantrum if attention maintains it while keeping the child safe, and expect it to get louder for a few episodes before it fades. ## References 1. Kazdin, A. E. (2005). *Parent Management Training: Treatment for Oppositional, Aggressive, and Antisocial Behavior in Children and Adolescents*. Oxford University Press. 2. Eyberg, S. M., Nelson, M. M., & Boggs, S. R. (2008). Evidence-based psychosocial treatments for children and adolescents with disruptive behavior. *Journal of Clinical Child & Adolescent Psychology, 37*(1), 215–237. 3. Patterson, G. R. (1982). *Coercive Family Process*. Castalia. 4. Patterson, G. R., DeBaryshe, B. D., & Ramsey, E. (1989). A developmental perspective on antisocial behavior. *American Psychologist, 44*(2), 329–335. 5. Forgatch, M. S., & Patterson, G. R. (2010). Parent Management Training — Oregon Model: An intervention for antisocial behavior in children and adolescents. In J. R. Weisz & A. E. Kazdin (Eds.), *Evidence-Based Psychotherapies for Children and Adolescents* (2nd ed.). Guilford Press. 6. Forehand, R. L., & McMahon, R. J. (1981). *Helping the Noncompliant Child: A Clinician's Guide to Parent Training*. Guilford Press. 7. Everett, G. E., Hupp, S. D. A., & Olmi, D. J. (2010). Time-out with parents: A descriptive analysis of 30 years of research. *Education and Treatment of Children, 33*(2), 235–259. 8. Dadds, M. R., & Tully, L. A. (2019). What is it to discipline a child: What should it be? A reanalysis of time-out from the perspective of child mental health, attachment, and trauma. *American Psychologist, 74*(7), 794–808. 9. Sege, R. D., Siegel, B. S., Council on Child Abuse and Neglect, & Committee on Psychosocial Aspects of Child and Family Health. (2018). Effective discipline to raise healthy children. *Pediatrics, 142*(6), e20183112. 10. Kazdin, A. E. (1977). *The Token Economy: A Review and Evaluation*. Plenum Press. 11. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 12. Williams, C. D. (1959). The elimination of tantrum behavior by extinction procedures. *Journal of Abnormal and Social Psychology, 59*(2), 269. 13. Mindell, J. A., Kuhn, B., Lewin, D. S., Meltzer, L. J., & Sadeh, A. (2006). Behavioral treatment of bedtime problems and night wakings in infants and young children. *Sleep, 29*(10), 1263–1276. 14. Gradisar, M., Jackson, K., Spurrier, N. J., et al. (2016). Behavioral interventions for infant sleep problems: A randomized controlled trial. *Pediatrics, 137*(6), e20151486. 15. Price, A. M. H., Wake, M., Ukoumunne, O. C., & Hiscock, H. (2012). Five-year follow-up of harms and benefits of behavioral infant sleep intervention: Randomized trial. *Pediatrics, 130*(4), 643–651. 16. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469. 17. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts. 18. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 19. McNeil, C. B., & Hembree-Kigin, T. L. (2010). *Parent–Child Interaction Therapy* (2nd ed.). Springer. ## Related - [All applications](https://operantconditioning.com/applications/): ABA, health, work, animals, technology, and more — with the evidence rated. - [Operant conditioning in the classroom](https://operantconditioning.com/classroom/): Praise, token economies, the Good Behavior Game, and what the evidence says. - [Negative punishment](https://operantconditioning.com/negative-punishment/): Time-out and response cost, step by step. --- # Operant Conditioning in Dog Training: The Four Quadrants, the Evidence, and How to Train > Operant conditioning in dog training: the four quadrants, the evidence on reward-based vs. aversive methods, clicker training, timing, and a recall protocol. - Source: https://operantconditioning.com/dog-training/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications · Animal training* Every dog trainer uses operant conditioning, whether they know the vocabulary or not. Here are the four quadrants with dog examples, what the research actually says about reward-based and aversive methods, how markers, shaping, timing, and schedules work, and a protocol for a recall you can trust. > **Definition** > > **Operant conditioning in dog training** is the deliberate use of consequences to change what a dog does: behaviors that are followed by reinforcement (a treat, play, release to sniff, or relief from pressure) happen more often, and behaviors followed by punishment (an added aversive or a lost reward) happen less often. Trainers arrange the [antecedent](https://operantconditioning.com/abc-model/) — the cue and the setting — and the consequence, and the dog's behavior changes in between. > > A "reward" only counts as a reinforcer if the behavior actually increases, and a "correction" only counts as a punisher if the behavior actually decreases. The dog, not the trainer, decides what works.[1] **In brief** - The [four quadrants](https://operantconditioning.com/#the-four-quadrants-reinforcement-and-punishment) are not equal choices: R+ and P− need nothing unpleasant present, while R− and P+ depend on an aversive the trainer introduces. - No controlled study has found aversive methods more effective than reward-based ones, and aversive methods are associated with stress and aggressive responses. - A reward counts as a reinforcer only if the behavior actually increases; the dog, not the trainer, decides what works. ## The four quadrants in dog training Every consequence a dog experiences can be sorted by two questions: was something *added* or *removed*, and did the behavior go *up* or *down*? Trainers usually abbreviate the results as R+, R−, P+, and P−. | Quadrant | What happens | Dog training example | Effect and notes | | --- | --- | --- | --- | | R+ Positive reinforcement | Something the dog wants is added after the behavior | Dog sits → treat, or a thrown ball, or the door opens | Sitting increases. The foundation of modern, reward-based training. | | R− Negative reinforcement | Something the dog dislikes is removed after the behavior | Steady leash pressure → dog steps toward the handler → pressure released | Moving toward the handler increases. Requires an aversive to be present first; used in traditional and some "balanced" training. | | P+ Positive punishment | Something the dog dislikes is added after the behavior | Dog pulls → leash pop; dog barks → spray bottle | Pulling or barking decreases, if it works at all. Associated with stress and aggressive responses (see the evidence below). | | P− Negative punishment | Something the dog wants is removed after the behavior | Dog jumps up → person turns away; dog mouths hand → play ends for 30 seconds | Jumping or mouthing decreases. The mildest way to reduce behavior, and it pairs naturally with R+ for an alternative. | Two things follow from the table. First, "positive" and "negative" describe adding and removing, not kind and cruel: the spray bottle is *positive* punishment. Second, the quadrants are not four equal choices. R+ and P− require nothing unpleasant to be present, while R− and P+ both depend on an aversive stimulus the trainer has to introduce, which brings side effects with it. That asymmetry is why most evidence-based trainers work almost entirely in the top-left cell of the [quadrant grid](https://operantconditioning.com/#the-four-quadrants-reinforcement-and-punishment) (R+) and reach for P− when a behavior needs to shrink. ## What the evidence says about reward-based vs. aversive dog training Trainers argue about methods constantly; the research is smaller than the argument, but it points consistently in one direction. ### Owner surveys Hiby, Rooney, and Bradshaw surveyed 364 dog owners about the methods they used for common tasks and the behavior of their dogs. Owners who relied on reward-based methods reported higher obedience, and the use of punishment-based methods was associated with a higher number of problem behaviors.[2] Herron, Shofer, and Reisner surveyed 140 owners whose dogs had been referred to a veterinary behavior clinic. For several confrontational techniques — hitting or kicking the dog, growling at it, the "alpha roll," staring it down — roughly a quarter or more of owners reported that the dog had responded with aggression. Reward-based techniques rarely produced aggressive responses.[3] ### Reviews and welfare studies Ziv's 2017 review of the literature concluded that aversive methods carry welfare risks and that there is no evidence they are more effective than reward-based methods.[4] Vieira de Castro and colleagues then went beyond questionnaires: they filmed dogs from training schools that used reward-based methods and schools that used aversive methods, sampled salivary cortisol, and ran a cognitive bias test. Dogs from aversive-based schools showed more stress-related behaviors and body postures during training, higher post-training cortisol, and in the cognitive bias test they approached an ambiguous bowl more slowly — the "pessimistic" pattern seen in animals in poorer welfare states.[5] ### Experimental comparisons The strongest single piece of efficacy evidence is an experiment rather than a survey. China, Mills, and Cooper assigned pet dogs with known off-lead problems to training with remote electronic collars by industry-approved trainers, to the same trainers without collars, or to reward-based trainers. The reward-based group responded to "sit" and "come" more reliably and more quickly; the e-collar added no measurable benefit.[6] > **The honest caveats** > > Most of this evidence is correlational. Owners who choose punishment may already have more difficult dogs, and owners who choose reward-based methods may differ in other ways too. Surveys rely on self-report; clinic samples are not typical dogs; and school-based comparisons cannot fully separate the method from the trainer. What can be said is this: no controlled study has found aversive methods to be *more* effective, several have found reward-based methods equally or more effective, and the welfare and aggression findings point the same way across designs. That is enough for professional bodies to recommend reward-based training as the default.[7] [More on aversives and their side effects ›](https://operantconditioning.com/positive-punishment/#aversives-in-dog-training-what-the-evidence-shows) ## The dominance myth A great deal of popular dog training rests on the idea that dogs are trying to become the "alpha" of the household and must be shown their place. The idea traces to studies of unrelated wolves confined together in captivity in the mid-twentieth century, which showed constant fighting for rank. L. David Mech's 1970 book helped spread the "alpha wolf" concept; his own later fieldwork on wild wolves showed that a pack is simply a family — a breeding pair and their offspring — and that the parents lead the way parents do, without ritualized dominance contests. Mech published the correction in 1999, has spent years asking people to drop the term, and has said he asked his publisher to stop printing the 1970 book.[8][9] Domestic dogs are not wolves in any case, and studies of free-ranging dogs and of dog–human interactions find little support for the notion that problem behavior is a bid for status.[10] The dog that pulls on the leash is not staging a coup; pulling has simply been reinforced by getting where it wants to go. The American Veterinary Society of Animal Behavior's position statement on dominance theory recommends against confrontational "dominance" techniques, and its 2021 statement on humane dog training recommends reward-based methods and advises against aversive tools.[7][11] ## Marker and clicker training A treat that arrives three seconds after a sit reinforces whatever the dog was doing three seconds after it sat — usually standing up and sniffing your hand. A **marker** solves the timing problem. A clicker (or a short word like "yes") is paired with food until the sound itself becomes a **conditioned reinforcer**: through classical conditioning it comes to predict the treat, and through operant conditioning it can then strengthen whatever it follows.[12] Trainers call it a **bridge** because it spans the gap between the behavior and the primary reinforcer. [How classical and operant conditioning combine ›](https://operantconditioning.com/operant-vs-classical-conditioning/) The technique is older than most people think. Keller and Marian Breland, two of Skinner's early students, left academia in the 1940s to train animals commercially and used a hand-held clicker as a conditioned reinforcer across dozens of species.[13] Skinner himself described the method for a general audience in 1951, explaining how to train a dog with a conditioned reinforcer and successive approximations.[14] Marine-mammal trainers adopted a whistle as their bridge, and Karen Pryor, a former dolphin trainer, brought the approach to dog owners through *Don't Shoot the Dog* and the clicker-training movement that followed.[15] ### Charging the clicker Click, then treat; click, then treat — twenty or thirty times over a couple of short sessions, with the click always *preceding* the treat by about a second. When the dog's head whips toward you at the sound, the marker is charged. From then on the rule is simple: every click earns a treat, and the click marks the exact instant the behavior you want occurs. The treat can follow a moment later; the click has already done the teaching. ## Shaping, luring, and capturing There are three ways to get a behavior to happen so you can reinforce it. - **Capturing.** Wait for the dog to do the behavior on its own — lie down, make eye contact — and mark it. Slow for rare behaviors, excellent for common ones, and it produces behavior the dog "owns." - **Luring.** Use a treat in the hand to guide the dog into position: raise it over the nose and the rear drops into a sit. Fast, but the lure must be faded within a few repetitions or the dog learns that the behavior happens only when food is visible. - **Shaping.** Reinforce successive approximations — first a glance at the mat, then a step toward it, then a paw on it, then lying on it. Shaping builds behaviors that can't be lured and teaches the dog to experiment.[1] [How shaping works, step by step ›](https://operantconditioning.com/shaping/) ## What makes reinforcement work ### Timing: the one-second window Laboratory work on delayed reinforcement shows that a consequence loses much of its power within a few seconds of the behavior, and a delayed reinforcer tends to strengthen whatever happened just before *it* rather than the behavior you intended.[16] In practice: mark within about a second, and deliver the treat where you want the dog to be (feed a "down" on the floor between the paws, feed heel position at your left knee). ### Rate of reinforcement Early in learning, aim for many reinforced repetitions per minute — ten or fifteen is not unusual in a good session. A high rate keeps the dog engaged and out-competes distractions. ### Reinforcer value Reinforcers are not interchangeable. Most dogs rank kibble below cheese, cheese below roast chicken, and, for some, a tug toy above all of it. Match the reinforcer to the difficulty: kibble for a sit in the kitchen, chicken for a recall past a squirrel. And remember motivating operations: a dog trained right after dinner is a dog for whom food has stopped being a reinforcer. ### Life rewards and the Premack principle Anything the dog wants to do can reinforce something you want it to do — David Premack's [principle](https://operantconditioning.com/premack-principle/) that a higher-probability behavior reinforces a lower-probability one.[17] Sit, and the door opens. Look at me, and you are released to sniff that fascinating hydrant. Come when called, and you get sent back to play. ## Schedules of reinforcement in dog training While the dog is learning, reinforce *every* correct response — a continuous schedule, which produces the fastest acquisition. Once the behavior is reliable, thin to a **variable-ratio** schedule: reinforce most responses, then some, in an unpredictable pattern, while keeping the best responses on a "jackpot."[18] This is the step most owners skip, and skipping it in either direction causes trouble. Never thinning produces a dog that works only when food is visible. Thinning too fast, or reinforcing so rarely that the behavior stops paying, produces a dog that stops offering the behavior. > **Why occasional treats make behavior stronger, not weaker** > > Owners worry that "not rewarding every time" will erode the behavior. The opposite is true. Behavior maintained on an intermittent schedule is *more* resistant to extinction than behavior reinforced every time — the partial-reinforcement extinction effect that Ferster and Skinner documented in thousands of hours of cumulative records. A dog whose sit is reinforced unpredictably keeps sitting through long stretches without a treat, because long stretches without a treat are exactly what it has learned to expect.[18] The same principle works against you when counter-surfing pays off once a month. [Run the schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) ## Stimulus control, cues, and generalization ### Add the cue after the behavior is reliable A cue is a [discriminative stimulus](https://operantconditioning.com/glossary/#discriminative-stimulus): a signal that the behavior will now be reinforced. Saying "sit" to a dog that does not yet sit reliably teaches it nothing except that "sit" is background noise. The cleaner sequence is: get the behavior (by capturing, luring, or shaping), reinforce it until the dog offers it readily, and then say the cue *just as* the dog starts to do it. Within a few dozen pairings the word predicts the behavior and the reinforcement, and the dog starts responding to the word alone. A behavior is under stimulus control when it happens promptly on cue, does not happen without the cue, and does not happen to other cues.[1] ### Dogs don't generalize well — so train everywhere A dog that sits perfectly in the kitchen has learned "sit, in the kitchen, facing my owner, with the treat pouch on." Take it to the park and the behavior can vanish, not out of stubbornness but because none of the antecedent conditions match. Behavior analysts learned long ago that generalization has to be programmed rather than hoped for: train in many places, with many people, at many distances and levels of distraction, and reinforce across all of them.[19] Trainers summarize this as the three D's — distance, duration, distraction — and the rule that you raise only one at a time. ## Extinction and extinction bursts When a behavior that used to be reinforced stops working, it does not vanish quietly. It gets worse first. The dog whose jumping has always earned a hand, a voice, or a push-off will, when the attention stops, jump higher, faster, and with more mouth — an **extinction burst** — before the behavior declines.[20] The same happens with demand barking: ignore it and the dog will bark louder for a while. Two rules make extinction workable. Be completely consistent, because reinforcing the louder bark teaches the dog that louder is the new price. And always reinforce an alternative — four paws on the floor, a quiet sit — so the dog has a behavior that *does* pay. Pure extinction without a replacement is slow and unkind. [Extinction, bursts, and spontaneous recovery in depth ›](https://operantconditioning.com/extinction/) ## Management: control the antecedent before you train Every time a dog practices a behavior and it pays off, the behavior gets stronger. So before any training plan, arrange the environment so the unwanted behavior can't be rehearsed: baby gates at the front door, food off the counters, a leash on the dog when guests arrive, blinds down for the window barker, a long line for the dog whose recall isn't ready. This is antecedent control, the "A" of the [A-B-C model](https://operantconditioning.com/abc-model/), and it is what lets the reinforcement-based plan win: the old behavior stops being reinforced in the background while you teach the new one. ## How to teach a reliable recall, step by step Recall is the behavior that keeps dogs alive, and it is the one most often ruined by good intentions. The protocol below uses nothing but positive reinforcement, management, and schedules. 1. **Choose a fresh cue.** If "come" has been shouted at a dog that ignored it, or used to end walks, it is contaminated. Pick a new word — "here," a whistle — and protect it: never use it when you can't make it pay. 2. **Charge the cue.** Indoors, with the dog beside you, say the cue once and immediately feed three to five tiny pieces of something excellent, one after another. Ten repetitions, twice a day, for three or four days. The cue now predicts a jackpot before the dog has done anything. 3. **Add the behavior at short range.** Say the cue when the dog is a few feet away and likely to come anyway. When it reaches you, take the collar gently, *then* feed. The collar touch becomes part of the reinforced chain, so a hand reaching for the collar never becomes a signal to dodge. 4. **Build distance and distraction separately.** Increase distance indoors, then move to a quiet yard at short distance, then a long line in a park. Raise one variable at a time. The long line is management: it guarantees the cue is never disobeyed successfully. 5. **Send the dog back to the fun.** Most recalls in practice should end with "go play." If coming when called usually ends the walk, the recall is being punished by the loss of freedom — negative punishment of the very behavior you want. Use the Premack principle so that coming is the price of more freedom, not the end of it. 6. **Thin the schedule, keep the jackpots.** Once the recall is reliable, reinforce with food unpredictably, keep the surprise jackpots for fast responses through distractions, and lean on life rewards. Never let the schedule thin to zero. 7. **Never punish a recall.** If the dog comes slowly, reinforce anyway — a slow recall is still a recall, and scolding it teaches the dog that arriving is dangerous. If the dog doesn't come, go and get it calmly, then make the next repetition easier. 8. **Maintain for life.** A few surprise recalls on every walk, always paid, keep the behavior strong. ## Common mistakes in operant dog training - **Repeating the cue.** "Sit. Sit! SIT!" teaches the dog that the cue is "sit-sit-SIT." Say it once; if nothing happens, make the situation easier and try again. - **Poisoning the cue.** A cue that is sometimes followed by reinforcement and sometimes by a correction becomes ambiguous — Karen Pryor's term is a *poisoned cue* — and dogs respond to it slowly and with signs of stress. Keep each cue attached to one kind of consequence. - **Punishing the recall by ending the fun.** Calling the dog only to leave the park, get a bath, or be crated trains the dog to keep its distance. - **Treat dependence from never thinning.** If every sit for two years has produced a visible treat, the treat has become part of the cue. Fade the lure early and thin the schedule once the behavior is reliable. - **Bribing instead of reinforcing.** Showing the treat before the behavior is a lure or a bribe. Reinforcement comes *after*. - **Reinforcing the wrong moment.** A treat handed to a dog that sat and then stood reinforces standing. Mark the sit; feed in the sit. - **Training when the dog is over threshold.** A dog that is frantic about another dog across the street cannot learn. Add distance until it can eat and think, then train. ## Common problem behaviors: what maintains them, and a reinforcement-based plan The first question is never "how do I stop this?" but "what is this behavior getting?" Once you can name the reinforcer, the plan almost writes itself: manage so the old reinforcer stops arriving, and reinforce a behavior that can replace it. | Behavior | Likely maintaining reinforcer | Reinforcement-based plan | | --- | --- | --- | | Jumping on people | Attention: eye contact, voices, hands — even a push-off is contact R+ | Manage with a leash or gate at the door. All attention stops the instant paws leave the floor P−; four-on-the-floor or a sit earns the greeting, generously. Recruit guests, and expect a burst. | | Pulling on the leash | Forward progress toward smells and dogs R+ | A tight leash stops all forward motion (pulling no longer works); a loose leash makes the walk go on, plus frequent treats at your side. A front-clip harness for management while the new behavior builds. | | Barking at passersby from the window | The passerby always leaves R−, plus the arousal itself | Block the view (film on the glass, closed blinds). Teach and heavily reinforce "go to your mat" when someone passes; reinforce quiet glances at the window. | | Demand barking | Food, play, the door, or attention delivered to stop the noise R+ | Barking pays nothing, ever. Teach a quiet alternative request (a sit, a nose-touch) and pay it fast and often. Consistency from everyone in the house, and expect the burst. | | Counter-surfing | Food, on an intermittent schedule — the most durable kind VR | Management is non-negotiable: clear counters, closed kitchen. Reinforce lying on a mat in the kitchen while you cook. One sandwich a month will maintain the behavior indefinitely. | | Begging at the table | Scraps, from at least one family member, occasionally VR | Nobody feeds from the table, no exceptions. Give the dog a stuffed food toy on its bed during meals so lying there becomes the behavior that pays. | | Lunging and barking at dogs on leash | Distance: the other dog goes away, or the owner retreats R−, usually driven by fear or frustration | Work at a distance where the dog can eat and think. Pair the appearance of other dogs with excellent food (counterconditioning), and reinforce looking at the dog and back at you. Do this with a qualified reward-based trainer. | > **Aggression, fear, and resource guarding need a professional** > > Growling, snapping, biting, guarding food or objects, and severe fear are not obedience problems, and punishing them tends to suppress the warning while leaving the emotion intact — the dog that no longer growls may go straight to biting. Seek a certified reward-based trainer or, for aggression and anxiety, a veterinary behaviorist, who can also rule out pain and medical causes. ## Key takeaways - "Positive" and "negative" mean added and removed, not kind and cruel. R+ and P− require nothing unpleasant to be present; R− and P+ both depend on an aversive the trainer introduces, which is why evidence-based trainers work almost entirely in R+ and reach for P− when a behavior needs to shrink. - The research is smaller than the argument but points one way: no controlled study has found aversive methods more effective, reward-based methods were equally or more effective, and aversive methods are associated with stress, higher cortisol, and aggressive responses. The dominance idea rests on captive wolves and was retracted by the researcher who spread it. - A marker is a conditioned reinforcer that bridges the gap between the behavior and the treat; without it, a treat three seconds late reinforces whatever the dog was doing three seconds later. Mark within about a second and feed where you want the dog to be. - Reinforce every correct response while the dog is learning, then thin to a variable-ratio schedule and keep the jackpots. Intermittent reinforcement makes behavior more resistant to extinction, which is why a sit survives long stretches without treats and why one sandwich a month maintains counter-surfing. - Cues are discriminative stimuli added after the behavior is reliable, and dogs generalize poorly, so train everywhere and raise distance, duration, and distraction one at a time. Before training, manage the antecedent so the old behavior stops being rehearsed, and ask what the behavior is getting before asking how to stop it. ### Check yourself **An owner calls her dog at the park, and when it arrives she clips on the leash and goes home. She never scolds it, yet over the weeks the recall gets slower. Why is the recall weakening?** Coming when called reliably ends the dog's freedom, so the recall is being punished by the loss of play: negative punishment of the very behavior she wants. The fix is the Premack principle in reverse of what she has been doing: most recalls should end with "go play," so that coming is the price of more freedom rather than the end of it. **A trainer applies steady leash pressure and releases it the instant the dog steps toward her. A student calls this punishment, since leash pressure is unpleasant. Which quadrant is it, and what is the real concern?** It is negative reinforcement: something the dog dislikes is removed after the behavior, and stepping toward the handler increases. "Negative" means removed, not bad. The real concern is that R− requires the aversive to be present first, which is where the welfare cost lies and why it sits on the same side of the asymmetry as P+. **A dog jumps on every guest, and every guest pushes it off with a firm "No." Months later it still jumps. What is maintaining the behavior?** Attention: eye contact, voices, and hands, and even a push-off is contact. The intended correction is functioning as positive reinforcement, which the behavior proves by continuing. The plan is management at the door, all attention stopping the instant paws leave the floor, a generous greeting for four-on-the-floor or a sit, and an extinction burst to be expected first. **An owner worries that reinforcing sits only some of the time will erode the behavior, so she keeps treating every one. Is her worry justified?** No; the opposite is true. Behavior maintained on an intermittent, variable-ratio schedule is more resistant to extinction than behavior reinforced every time, the partial-reinforcement extinction effect. Reinforcing every sit for years makes the visible treat part of the cue. The errors to avoid are thinning too fast, or letting the schedule thin to zero. **Explain it to a friend.** Explain why a clicker works, without using the words "reinforcer," "conditioned," or "bridge." ## Frequently asked questions **What are the four quadrants of dog training?** Positive reinforcement (add something the dog wants — a treat for a sit), negative reinforcement (remove something unpleasant — leash pressure released when the dog moves), positive punishment (add something unpleasant — a leash pop), and negative punishment (remove something the dog wants — turning away when it jumps). "Positive" and "negative" mean added and removed, not good and bad. **Is positive reinforcement dog training effective?** Yes. Owner surveys associate reward-based methods with higher obedience and fewer problem behaviors, a controlled comparison found reward-based training more effective than electronic-collar training for recall and sit, and no controlled study has found aversive methods to be more effective. Reward-based methods also avoid the stress and aggression associated with confrontational techniques. **Do I have to give my dog treats forever?** No. Reinforce every correct response while the dog is learning, then shift to reinforcing unpredictably — a variable-ratio schedule — while replacing many treats with life rewards such as play, sniffing, and going through doors. Intermittent reinforcement makes behavior more durable, not less. What you should never do is let reinforcement stop entirely. **Is negative reinforcement bad for dogs?** Negative reinforcement is not punishment — it strengthens behavior by removing something unpleasant. But it requires the unpleasant thing to be present first, which is where the welfare cost lies. Mild, brief pressure released the instant the dog responds is used by many trainers; methods built on shock, prong collars, or sustained discomfort are associated with stress indicators and are discouraged by veterinary behavior organizations. **Is the alpha or dominance theory of dog training true?** No. The "alpha wolf" idea came from unrelated captive wolves forced to live together. Wild wolf packs are families led by the parents, as L. David Mech, who helped popularize the term, later showed and retracted. Dogs are not wolves, and studies of dog behavior do not support the idea that misbehavior is a bid for rank. Problem behavior is almost always a reinforcement problem, not a status problem. **What is clicker training and how does it work?** Clicker training uses a distinct sound that has been paired with food until it becomes a conditioned reinforcer. The click marks the precise instant the dog does the right thing and bridges the gap until the treat arrives. It solves the timing problem — you can click within a fraction of a second even if the treat takes longer — and it lets you shape complex behaviors in small steps. **Why does my dog only listen at home?** Because that is where the behavior was trained. Dogs learn cues together with the context — the room, the person, the pouch, the absence of distractions — and do not generalize well on their own. Retrain each behavior in new places, with new people, at greater distances and with more distraction, raising one variable at a time and reinforcing generously as you go. **Should I punish my dog for growling?** No. A growl is information — the dog is telling you it is uncomfortable. Punishing it can suppress the warning without changing the feeling, producing a dog that bites without growling first. Move the dog away from what is bothering it, and consult a certified reward-based trainer or veterinary behaviorist to address the underlying fear or guarding. ## References 1. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 2. Hiby, E. F., Rooney, N. J., & Bradshaw, J. W. S. (2004). Dog training methods: Their use, effectiveness and interaction with behaviour and welfare. *Animal Welfare, 13*(1), 63–69. 3. Herron, M. E., Shofer, F. S., & Reisner, I. R. (2009). Survey of the use and outcome of confrontational and non-confrontational training methods in client-owned dogs showing undesired behaviors. *Applied Animal Behaviour Science, 117*(1–2), 47–54. 4. Ziv, G. (2017). The effects of using aversive training methods in dogs — A review. *Journal of Veterinary Behavior, 19*, 50–60. 5. Vieira de Castro, A. C., Fuchs, D., Morello, G. M., Pastur, S., de Sousa, L., & Olsson, I. A. S. (2020). Does training method matter? Evidence for the negative impact of aversive-based methods on companion dog welfare. *PLoS ONE, 15*(12), e0225023. 6. China, L., Mills, D. S., & Cooper, J. J. (2020). Efficacy of dog training with and without remote electronic collars vs. a focus on positive reinforcement. *Frontiers in Veterinary Science, 7*, 508. 7. American Veterinary Society of Animal Behavior. (2021). *Position Statement on Humane Dog Training*. AVSAB. 8. Mech, L. D. (1999). Alpha status, dominance, and division of labor in wolf packs. *Canadian Journal of Zoology, 77*(8), 1196–1203. 9. Mech, L. D. (2008). Whatever happened to the term alpha wolf? *International Wolf, 18*(4), 4–8. 10. Bradshaw, J. W. S., Blackwell, E. J., & Casey, R. A. (2009). Dominance in domestic dogs — useful construct or bad habit? *Journal of Veterinary Behavior, 4*(3), 135–144. 11. American Veterinary Society of Animal Behavior. (2008). *Position Statement on the Use of Dominance Theory in Behavior Modification of Animals*. AVSAB. 12. Williams, B. A. (1994). Conditioned reinforcement: Experimental and theoretical issues. *The Behavior Analyst, 17*(2), 261–285. 13. Breland, K., & Breland, M. (1951). A field of applied animal psychology. *American Psychologist, 6*(6), 202–204. 14. Skinner, B. F. (1951). How to teach animals. *Scientific American, 185*(6), 26–29. 15. Pryor, K. (1984). *Don't Shoot the Dog! The New Art of Teaching and Training*. Simon & Schuster. 16. Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139. 17. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 18. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 19. Stokes, T. F., & Baer, D. M. (1977). An implicit technology of generalization. *Journal of Applied Behavior Analysis, 10*(2), 349–367. 20. Lerman, D. C., & Iwata, B. A. (1995). Prevalence of the extinction burst and its attenuation during treatment. *Journal of Applied Behavior Analysis, 28*(1), 93–94. ## Related - [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The quadrant modern dog training is built on — timing, contingency, and reinforcer types. - [Shaping](https://operantconditioning.com/shaping/): Successive approximations and chaining — how complex behaviors are built one step at a time. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): When to thin the treats, and why intermittent reinforcement makes behavior last. --- # How to Build Habits With Operant Conditioning: A Science-Based, Seven-Step Protocol > How to build habits with operant conditioning: what habit formation psychology shows (66 days, context cues), a 7-step protocol, and how to break bad habits. - Source: https://operantconditioning.com/habits/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Apply it · Self-management* Seventy years of behavioral science and a decade of habit research agree on the mechanics. Here is the evidence, a seven-step protocol built on the antecedent–behavior–consequence loop, and what to do when it breaks. > **Definition** > > A **habit** is a behavior that has come under strong stimulus control: it is cued automatically by a context, performed with little deliberation, and relatively insensitive to what you happen to want in the moment.[1] > > In operant terms a habit is a well-worn [three-term contingency](https://operantconditioning.com/abc-model/) — a cue (antecedent), a response (behavior), and the history of reinforcement that built and maintains it. Building a habit with operant conditioning means arranging all three terms on purpose, which is what most habit advice — and most habit apps — leave to chance. **In brief** - A habit is a behavior under strong stimulus control: cued automatically by context, performed with little deliberation, and insensitive to momentary wants. - Motivation is not a term in the [contingency](https://operantconditioning.com/abc-model/); build on an unmissable cue, a tiny behavior, and a consequence that arrives within seconds. - Automaticity took a median of 66 days (range 18–254) in the best real-world study, and one missed day made no material difference. ## What habit formation psychology actually shows The most-cited study of real-world habit formation followed 96 volunteers who each chose a new eating, drinking, or activity behavior and tied it to a once-daily cue — for example, a piece of fruit with lunch or a run before dinner. Self-reported automaticity rose along an asymptotic curve: fast gains early, then a plateau.[2] 66 median days to reach peak automaticity 18–254 range across individuals and behaviors 1 missed day: no material effect on the process Three findings matter more than the headline number. The range was enormous — drinking a glass of water plateaued fast, exercise slowly. Missing a single opportunity did not measurably disrupt the curve. And the "21 days" figure that circulates everywhere appears nowhere in the data.[2] Wendy Wood's research adds the mechanism: habits are cued by **context**, and they persist on context rather than on intention. Students who transferred universities kept their exercise, reading, and TV habits only when the new environment resembled the old one.[3] Habitual cinema popcorn-eaters ate stale, week-old popcorn as readily as fresh — in a cinema. In a meeting room, taste took over.[4] So a stable cue is not optional, and habits change most easily when context is already disrupted — a move, a new job, a new term. Temporal landmarks work the same way. Katy Milkman and colleagues found that searches for "diet," gym visits, and goal commitments all spike at the start of a week, a month, a year, and after birthdays — the **fresh start effect**.[5] A landmark separates the old self from a new one and changes the antecedent conditions under which the behavior is attempted. Use one, but do not wait for one. Finally, habit researchers measure habit as **automaticity** — behavior that is efficient, unintentional, and hard to control — rather than as frequency, because a behavior you do daily through gritted teeth is not yet a habit.[6] The test is not "did I do it?" but "did I have to decide to?" ## Why willpower and motivation are the wrong frame Notice what is missing from the three-term contingency: motivation. It is not a variable in the loop. The nearest thing to it — the [motivating operation](https://operantconditioning.com/abc-model/) — is deprivation or satiation, which changes how much a reinforcer is worth and which fluctuates hour to hour. Building a routine on how much you want it today is building on the one term guaranteed to be different tomorrow. Willpower fares no better. The influential "ego depletion" studies of the late 1990s reported that self-control is a limited resource that runs down with use; a preregistered replication across 23 laboratories in 2016 found an effect indistinguishable from zero.[7][8] Whatever willpower is, its footing is contested enough that no protocol should depend on it. The more robust finding cuts the other way: in experience-sampling studies, people high in self-control do not report resisting more temptations — they report *fewer* — and their advantage in life outcomes is carried largely by beneficial habits, because they have arranged their environments and routines so that the desired behavior runs automatically.[9] Self-control, in practice, is antecedent control. > **The reframe** > > You are not trying to want it more. You are trying to arrange a cue that is unmissable, a behavior that is small enough to be emitted, and a consequence that arrives fast enough to count. Motivation is what you feel while the arrangement is bad. ## The operant protocol for building a habit The seven steps below are the three-term contingency applied in order, with the schedule and failure-planning that the laboratory and the habit literature both insist on. 1. **Pick the behavior and make it tiny.** "Exercise" is an outcome; "put on running shoes and step outside" is a behavior. BJ Fogg's Tiny Habits method starts with a version that takes under a minute — two push-ups, one sentence, flossing one tooth — which is [shaping](https://operantconditioning.com/shaping/)'s first approximation by another name.[10] A tiny behavior gets emitted, and only emitted behavior can be reinforced. Once it is automatic, grow it. 2. **Choose a stable antecedent.** Anchor the new behavior to something that already happens reliably every day: after I pour the coffee, when I sit down at the desk. State it as an implementation intention — "when X happens, I will do Y" — which links cue to response in advance; a meta-analysis of 94 studies found a medium-to-large effect on goal attainment.[11] Then make the cue physically present: shoes by the door, book on the pillow, phone charging in the kitchen. Do not use a notification as the cue. It habituates, many people disable it, and it cues picking up the phone rather than the behavior you want. 3. **Arrange an immediate consequence.** The natural consequences of good habits arrive in weeks (fitness) or years (health), which is precisely why they fail to control behavior. So arrange one yourself, within seconds of finishing. Marking the behavior done works as a [conditioned reinforcer](https://operantconditioning.com/positive-reinforcement/) once it is paired with something that matters — visible progress, a small pleasure you allow only afterward, a message to someone watching. **Temptation bundling** pairs the behavior with something you already crave: gym-goers given audiobooks they could only hear at the gym went more often.[12] That is the **[Premack principle](https://operantconditioning.com/premack-principle/)** — a higher-probability behavior reinforces a lower-probability one — in a pair of headphones.[13] 4. **Reinforce every time at first.** During acquisition use continuous reinforcement: the consequence follows every occurrence. It is the fastest way to build a behavior, and it is where self-managers are too stingy, saving the reward for a "real" workout and never reinforcing the tiny one that had to come first.[14] 5. **Thin to an intermittent schedule once the behavior is stable.** Continuous reinforcement builds fragile behavior: stop the consequence and it extinguishes quickly. After a few weeks of reliable performance, shift to reinforcing unpredictably — a variable-ratio schedule — for the steadiest responding and the greatest resistance to [extinction](https://operantconditioning.com/extinction/). [How schedules work, with a simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) 6. **Track the behavior, not the outcome.** You cannot reinforce "lose ten pounds" on a Tuesday; you can reinforce "walked after lunch." Recording is itself an intervention: **self-monitoring is reactive** — observing and writing down your own behavior changes it, usually in the desired direction.[15] A meta-analysis of 138 experiments found that prompting people to monitor progress reliably improved goal attainment, more so when progress was physically recorded and reported to someone else.[16] 7. **Plan for extinction bursts, lapses, and resurgence.** When a new behavior stops being reinforced — you get sick, you travel — the old one it replaced tends to return; behavior analysts call this [resurgence](https://operantconditioning.com/extinction/), and it is the mechanism of most relapse. A lapse is not a failure; the 66-day data showed a missed day barely registered.[2] What turns a lapse into a collapse is the **abstinence violation effect**: the "I've blown it" judgment that follows one slip and licenses the next.[17] Decide the recovery response in advance — "if I miss a day, I do the tiny version tomorrow, no catching up" — and reinforce that. ### Writing the contingency down Whatever you use to track a habit — a notebook, a spreadsheet, an app — write all three terms for each habit: an antecedent (a real-world cue such as a time, a place, or a preceding routine), a behavior scoped as small as it needs to be, and a consequence you will actually deliver the moment the behavior is done. A record of only the middle term is a to-do list, not a contingency, and the value of writing the other two down is that it refuses to let you skip them. > **The app: Operant runs this loop for you.** A habit app for iPhone and Apple Watch from the publisher of this site. Each habit is set up as an antecedent, a behavior and a consequence — the three terms, not just the middle one. Free to download and try; a subscription unlocks the full app. [About the app](https://operantconditioning.com/app/) · [Download on the App Store](https://apps.apple.com/us/app/operant-behavior-change-app/id6802081776) ## Self-management in behavior analysis: what Skinner actually said Skinner devoted a chapter of *Science and Human Behavior* to self-control, and his position was simple: a person controls their own behavior the same way they control anyone else's — by manipulating the variables of which it is a function. One response (the *controlling* response) alters the conditions under which another (the *controlled* response) occurs.[18] His catalogue reads like a habit book written in 1953: - **Physical restraint and physical aid.** Walk out of the room; put the phone in a drawer; lay out the equipment the behavior needs. - **Changing the stimulus.** Remove the cues for the unwanted behavior and add cues for the wanted one — the whole of "environment design." - **Deprivation and satiation.** Eat before the party; build up an appetite for the reinforcer you plan to use. - **Manipulating emotional conditions and using aversive stimulation.** Set an alarm; make a public commitment you would be embarrassed to break. - **Operant conditioning and punishment of one's own behavior.** Self-administered reinforcers and penalties. - **Doing something else.** Emit an incompatible behavior — the seed of differential reinforcement. The last two items raised a debate that is still open: can you really reinforce yourself? Charles Catania argued in 1975 that a reinforcer you can take at any moment is not contingent on anything, so "self-reinforcement" is a misnomer for what is really rule-following.[19] A 1985 experiment sharpened the point: the benefit of a self-reward procedure disappeared when participants set their goals privately rather than publicly, suggesting the active ingredient was the social contingency.[20] The practical lesson survives whichever side wins: self-administered consequences work best when they are reliably withheld until the behavior occurs, externalized in something you cannot quietly waive — an app, a partner, a deposit — and backed by a social or financial contingency. Richard Malott's framing is the most useful: the natural consequences of most habits are too small, too delayed, or too improbable to control behavior, so the job of self-management is to add consequences that are sizable, immediate, and probable.[21] ## How to break a bad habit with operant conditioning A bad habit is a good contingency working for the wrong behavior. The steps mirror building one, in reverse. 1. **Identify the maintaining reinforcer.** Keep an [ABC record](https://operantconditioning.com/abc-model/) for three days. What does the behavior get (stimulation, food, attention) or escape (boredom, anxiety, an unpleasant task)? A habit maintained by escape needs a different replacement than one maintained by novelty. 2. **Make the cue unavailable.** The cheapest intervention is on the antecedent. No phone in the bedroom; no snacks on the counter; a route home that does not pass the bakery. Wood's popcorn study made the point dramatically: disrupting the habitual motor pattern — eating with the non-dominant hand — was enough to bring the behavior back under the control of taste.[4] 3. **Add friction and response cost.** Log out after every session so the feed opens to a password screen. Delete the app and reinstall it when you actually want it. Small increases in effort produce large decreases in a behavior that runs on automaticity, because automatic behavior stops at the first obstacle that requires a decision. 4. **Reinforce an alternative that serves the same function.** This is **differential reinforcement of alternative behavior**, and it is the step people skip. If scrolling was escape from boredom, queue a podcast; if the evening drink marked the boundary between work and rest, replace it with a walk that marks the same boundary. Remove a behavior without replacing it and the function goes unmet, and the old behavior resurges. 5. **Expect the burst and the recovery.** Withholding a reinforcer produces a temporary spike — the [extinction burst](https://operantconditioning.com/extinction/) — and an extinguished behavior can reappear after time away. Neither means the method failed. Giving in during the burst, however, reinforces a stronger version of the habit on an intermittent schedule, the worst of all outcomes. > **Do not rely on self-punishment** > > A penalty you administer to yourself is a penalty you can waive, and after the first waiver it is a penalty in name only. If you want an aversive consequence in the loop, hand it to a third party — a deposit contract, a friend who collects — so that the contingency is real. Even then, use it alongside reinforcement of the replacement behavior, not instead of it. [Why punishment is a poor first choice ›](https://operantconditioning.com/positive-punishment/) ## Common goals translated into tiny behaviors, antecedents, and consequences | Goal (outcome) | Tiny behavior (B) | Antecedent (A) | Immediate consequence (C) | | --- | --- | --- | --- | | Get fit | Put on running shoes and step outside | After the morning coffee is poured; shoes by the door | Mark it done; the podcast you only play while moving starts | | Read more | Read one page | When you get into bed; book on the pillow, phone in the kitchen | Mark it done; a moved bookmark is visible progress | | Meditate | Sit and take three slow breaths | After brushing teeth; cushion visible from the sink | Mark it done; the first sip of coffee follows the breaths | | Floss | Floss one tooth (the rest usually follows) | After putting the toothbrush down; floss pick on the brush | Mark it done; say "done" out loud — Fogg's celebration | | Journal | Write one sentence | After closing the laptop for the day; notebook open on the desk | Mark it done; the tea you make only after the sentence | | Use the phone less | Dock the phone at the kitchen charger | When you start the dishwasher after dinner | Mark it done; the evening show starts only once the phone is docked | | Study | Open the notes and answer one flashcard | When you sit down after the 4 p.m. class; deck open on the desk | Mark it done; a text to a study partner who replies | | Drink more water | Drink one glass | When the kettle is switched on; glass kept beside it | Mark it done; the coffee comes after the water | Notice the pattern in the consequence column. Almost every entry uses the Premack principle: a thing you were going to do anyway (the coffee, the show, the podcast) is made contingent on the tiny behavior. That costs nothing, it is immediate, and it is a reinforcer you already know works because you already do it. ## Commitment devices and behavioral economics A **commitment device** is an arrangement your present self makes to constrain your future self: a deadline you cannot move, money you forfeit if you fail, a membership that charges you whether or not you go. In operant terms it converts a consequence that is small, delayed, and cumulative into one that is large, certain, and near — Malott's prescription, implemented through a third party. The classic demonstration is Dan Ariely and Klaus Wertenbroch's deadline experiment. Students allowed to set their own binding deadlines for three papers set them earlier than they had to and performed better than students with a single end-of-term deadline — but worse than students given evenly spaced deadlines by the instructor. People know they procrastinate and will precommit to fight it; they just do it imperfectly.[22] Deposit contracts exploit **loss aversion**, the finding from prospect theory that a loss looms larger than an equivalent gain.[23] In a 16-week randomized trial, obese adults who put their own money at risk — refunded, with a match, only if they hit monthly weight targets — lost roughly three times as much weight as a control group given the same goals and weigh-ins; a lottery-incentive group did about as well. Much of the weight returned after the incentives ended, which is not an argument against the contingency so much as a demonstration of it: consequences control behavior while they are in force — the same pattern seen in [contingency management](https://operantconditioning.com/applications/#health) for addiction.[24] Plan the maintenance schedule before the acquisition schedule runs out. Every device on this list puts a consequence out of reach of your own leniency. That is the honest reason a partner, a coach, or an app that records the miss can outperform sheer resolve: none of them can be talked out of it at 10 p.m. ## Key takeaways - A habit is a well-worn three-term contingency: a cue, a response, and the reinforcement history that built it. Habits persist on context rather than intention, so a stable cue is not optional, and they change most easily when context is already disrupted. - Motivation and willpower are the wrong frame. Motivation is not a variable in the loop, ego depletion failed a 23-laboratory replication, and people high in self-control report fewer temptations rather than more resistance; self-control in practice is antecedent control. - The protocol runs the contingency in order: make the behavior tiny, anchor it to a stable cue stated as an implementation intention, arrange an immediate consequence (the Premack principle costs nothing), reinforce every time at first, then thin to an intermittent schedule, and track the behavior rather than the outcome. - Plan for lapses. A missed day barely registers; what turns a lapse into a collapse is the abstinence violation effect, so decide the recovery response in advance. When a new behavior stops being reinforced, the old one it replaced tends to resurge. - Breaking a habit mirrors building one: find the maintaining reinforcer, make the cue unavailable, add friction, and reinforce an alternative that serves the same function. Self-administered consequences work only when they cannot be quietly waived, which is what commitment devices and deposit contracts are for. ### Check yourself **Someone decides her morning coffee will be the reward for her morning run. On days she skips the run, she drinks the coffee anyway. Is the coffee reinforcing the run?** No. A reinforcer you can take at any moment is not contingent on anything, which was Catania's objection to "self-reinforcement." The coffee only works as a Premack consequence if it comes after the run and not otherwise; self-administered consequences work when they are withheld until the behavior occurs and externalized in something that cannot be quietly waived. **A person has reinforced a new habit every single time for six weeks and plans to keep doing so, reasoning that continuous reinforcement builds the strongest behavior. What does the schedule research say?** Continuous reinforcement is the fastest way to build a behavior, but it builds fragile behavior: stop the consequence and it extinguishes quickly. Once performance is stable, the protocol thins to an intermittent, variable-ratio schedule, which produces the steadiest responding and the greatest resistance to extinction. **Someone has done a behavior every day for two months but still has to force herself each time. Is it a habit yet?** Not by the researchers' definition. Habit is measured as automaticity, behavior that is efficient, unintentional, and hard to control, not as frequency; a behavior done daily through gritted teeth is not yet a habit. The test is not "did I do it?" but "did I have to decide to?" **A person who stopped late-night scrolling by charging the phone in the kitchen goes on a work trip, and in the hotel the scrolling returns. Did the method fail?** No. Habits are cued by context, and the antecedent arrangement that had been doing the work stayed at home; when a new behavior stops being reinforced, the old one it replaced resurges, which is the mechanism of most relapse. A lapse is not a collapse unless the "I've blown it" judgment makes it one, so the pre-decided recovery response, the tiny version the next day, is the behavior to reinforce. **Explain it to a friend.** Explain why a habit that keeps failing is usually an arrangement problem, using one habit you have tried and dropped, and without using the words "motivation" or "willpower." ## Frequently asked questions **How long does it take to form a habit?** In the best real-world study, the median was 66 days to reach peak automaticity, with a range of 18 to 254 days depending on the person and the behavior. Simple behaviors such as drinking a glass of water became automatic fastest; exercise took longest. The popular "21 days" figure has no basis in that data. **Does missing one day ruin a habit?** No. Missing a single opportunity had no material effect on habit formation in the 66-day study. What damages a habit is the "I've blown it" reaction that turns one lapse into a week of them. Decide in advance that after a miss you do the tiny version the next day, and treat that recovery as the behavior to reinforce. **What is the best reinforcer for building a habit?** One that is immediate, that you actually care about, and that you can reliably withhold until the behavior happens. The Premack principle is the easiest source: make something you already do daily (the coffee, the show, the podcast) contingent on the tiny behavior. A check-off works as a conditioned reinforcer once it is paired with visible progress. **Why don't habit-app notifications work as cues?** Three reasons. Repeated notifications habituate, so they stop grabbing attention. Many people disable them. And a notification is a cue for picking up the phone, not for the behavior you want, so it puts phone-checking under stimulus control instead. A stable real-world cue — a time, a place, a routine you already have — is a far stronger antecedent. **Can you really reinforce yourself?** Behavior analysts have argued about this since the 1970s. A reward you can take any time is not truly contingent, and experiments suggest public goal-setting often does the real work. In practice, self-administered consequences work when they are withheld until the behavior occurs, externalized in something you cannot quietly waive (an app, a partner, a deposit), and backed by a social or financial contingency. **How do you break a bad habit with operant conditioning?** Find what is reinforcing it, remove or change its cue, add friction so it can no longer run automatically, and reinforce an alternative behavior that serves the same function. Expect a temporary increase (the extinction burst) and occasional reappearances; neither means the method is failing. **What is habit stacking, and does it work?** Habit stacking (BJ Fogg calls it anchoring) attaches a new behavior to an existing routine: "after I pour my coffee, I will write one sentence." It works because the existing routine is a stable, reliable antecedent, and because stating the plan as an if-then implementation intention links cue to response in advance. It is antecedent control in plain language. **Are streaks a good idea?** A streak turns a habit into an avoidance contingency: the reinforcer becomes not losing the count. That can help while the streak is intact and hurt badly when it breaks, because the loss often triggers the abstinence violation effect. If you use streaks, decide in advance that a miss resets nothing but the number, and reinforce the return rather than mourning the run. ## References 1. Wood, W., & Rünger, D. (2016). Psychology of habit. *Annual Review of Psychology, 67*, 289–314. 2. Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. *European Journal of Social Psychology, 40*(6), 998–1009. 3. Wood, W., Tam, L., & Witt, M. G. (2005). Changing circumstances, disrupting habits. *Journal of Personality and Social Psychology, 88*(6), 918–933. 4. Neal, D. T., Wood, W., Wu, M., & Kurlander, D. (2011). The pull of the past: When do habits persist despite conflict with motives? *Personality and Social Psychology Bulletin, 37*(11), 1428–1437. 5. Dai, H., Milkman, K. L., & Riis, J. (2014). The fresh start effect: Temporal landmarks motivate aspirational behavior. *Management Science, 60*(10), 2563–2582. 6. Gardner, B. (2015). A review and analysis of the use of 'habit' in understanding, predicting and influencing health-related behaviour. *Health Psychology Review, 9*(3), 277–295. 7. Baumeister, R. F., Bratslavsky, E., Muraven, M., & Tice, D. M. (1998). Ego depletion: Is the active self a limited resource? *Journal of Personality and Social Psychology, 74*(5), 1252–1265. 8. Hagger, M. S., Chatzisarantis, N. L. D., Alberts, H., et al. (2016). A multilab preregistered replication of the ego-depletion effect. *Perspectives on Psychological Science, 11*(4), 546–573. 9. Hofmann, W., Baumeister, R. F., Förster, G., & Vohs, K. D. (2012). Everyday temptations: An experience sampling study of desire, conflict, and self-control. *Journal of Personality and Social Psychology, 102*(6), 1318–1335. See also Galla, B. M., & Duckworth, A. L. (2015). More than resisting temptation: Beneficial habits mediate the relationship between self-control and positive life outcomes. *Journal of Personality and Social Psychology, 109*(3), 508–525. 10. Fogg, B. J. (2020). *Tiny Habits: The Small Changes That Change Everything*. Houghton Mifflin Harcourt. 11. Gollwitzer, P. M. (1999). Implementation intentions: Strong effects of simple plans. *American Psychologist, 54*(7), 493–503. See also Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. *Advances in Experimental Social Psychology, 38*, 69–119. 12. Milkman, K. L., Minson, J. A., & Volpp, K. G. M. (2014). Holding the Hunger Games hostage at the gym: An evaluation of temptation bundling. *Management Science, 60*(2), 283–299. 13. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 14. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 15. Nelson, R. O., & Hayes, S. C. (1981). Theoretical explanations for reactivity in self-monitoring. *Behavior Modification, 5*(1), 3–14. 16. Harkin, B., Webb, T. L., Chang, B. P. I., et al. (2016). Does monitoring goal progress promote goal attainment? A meta-analysis of the experimental evidence. *Psychological Bulletin, 142*(2), 198–229. 17. Marlatt, G. A., & Gordon, J. R. (Eds.). (1985). *Relapse Prevention: Maintenance Strategies in the Treatment of Addictive Behaviors*. Guilford Press. 18. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. (Chapter 15, "Self-control.") 19. Catania, A. C. (1975). The myth of self-reinforcement. *Behaviorism, 3*(2), 192–199. 20. Hayes, S. C., Rosenfarb, I., Wulfert, E., Munt, E. D., Korn, Z., & Zettle, R. D. (1985). Self-reinforcement effects: An artifact of social standard setting? *Journal of Applied Behavior Analysis, 18*(3), 201–214. 21. Malott, R. W. (1989). The achievement of evasive goals: Control by rules describing contingencies that are not direct acting. In S. C. Hayes (Ed.), *Rule-Governed Behavior: Cognition, Contingencies, and Instructional Control* (pp. 269–322). Plenum. 22. Ariely, D., & Wertenbroch, K. (2002). Procrastination, deadlines, and performance: Self-control by precommitment. *Psychological Science, 13*(3), 219–224. 23. Kahneman, D., & Tversky, A. (1979). Prospect theory: An analysis of decision under risk. *Econometrica, 47*(2), 263–291. 24. Volpp, K. G., John, L. K., Troxel, A. B., Norton, L., Fassbender, J., & Loewenstein, G. (2008). Financial incentive-based approaches for weight loss: A randomized trial. *JAMA, 300*(22), 2631–2637. ## Related - [The ABC model](https://operantconditioning.com/abc-model/): Antecedent, behavior, consequence — the loop every habit runs on. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): When to reinforce every time and when to thin — with a live simulator. - [Extinction](https://operantconditioning.com/extinction/): Bursts, resurgence, and why relapse is predictable. --- # 50+ Operant Conditioning Examples: Everyday Life, Classroom, Work, and Animals > 50+ operant conditioning examples from everyday life, the classroom, work, and animal training — each labeled by quadrant, plus schedules and tricky cases. - Source: https://operantconditioning.com/examples/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Examples* More than fifty examples of operant conditioning in everyday life, at home, in the classroom, at work, in animal training, in apps, and in nature — each one labeled by quadrant and explained, plus the schedules and the tricky cases that fool people. > **The four quadrants, in two sentences** > > **Operant conditioning** is learning from consequences: a behavior that is followed by [positive reinforcement](https://operantconditioning.com/positive-reinforcement/) (something added) or [negative reinforcement](https://operantconditioning.com/negative-reinforcement/) (something removed) becomes more frequent, and a behavior followed by [positive punishment](https://operantconditioning.com/positive-punishment/) (something added) or [negative punishment](https://operantconditioning.com/negative-punishment/) (something removed) becomes less frequent. "Positive" and "negative" mean added and removed, not good and bad, and a consequence counts as reinforcement or punishment only by its actual effect on the behavior.[1] ## How to read an example of operant conditioning Every example on this page has the same three parts, the [A-B-C](https://operantconditioning.com/abc-model/) of behavior analysis: an **antecedent** (the situation), a **behavior** (what the person or animal does), and a **consequence** (what happens next). To classify the consequence, ask two questions in order: 1. **Did the behavior become more or less likely afterward?** More likely means reinforcement. Less likely means punishment. 2. **Was a stimulus added or removed?** Added means "positive." Removed means "negative." ![The four quadrants of operant conditioning](https://operantconditioning.com/assets/diagrams/the-four-quadrants-of-operant-conditioning.svg) *The four quadrants. The column answers "added or removed?"; the row answers "more or less likely?" Positive and negative are arithmetic signs, not judgments.* Two more labels cover what the quadrants leave out. Extinction means a behavior that used to be reinforced no longer is, and it fades.[11] Schedule marks examples where the interesting part is not *what* the consequence is but *how often* it arrives.[2] > **The label depends on the effect, not the intention** > > Every row below states the effect on behavior ("barking increases," "swearing decreases") because the label is only correct if that effect actually happens. A scolding that makes a child act out *more* is reinforcement, whatever the parent meant by it. When you analyze your own examples, watch the behavior over time before you name the quadrant.[1] ## Examples of operant conditioning in everyday life | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | You hold the door for a stranger | They smile and thank you; you hold doors more | Positive reinforcement | Social approval is added; the behavior increases. | | You try a new restaurant | The meal is excellent; you go back | Positive reinforcement | A good meal is added after the choice; returning increases. | | You tell a joke at dinner | Everyone laughs; you tell more jokes | Positive reinforcement | Laughter is added; joke-telling increases. | | You put on noise-cancelling headphones on the train | The noise disappears; you reach for them every ride | Negative reinforcement | An aversive stimulus (noise) is removed; the behavior increases. This is escape. | | You feed the parking meter | No ticket; feeding the meter becomes automatic | Negative reinforcement | The behavior prevents an aversive event. This is avoidance — the ticket never has to happen for the habit to hold. | | You mute a group chat that buzzes constantly | The buzzing stops; you mute chats faster in future | Negative reinforcement | An aversive stimulus is removed contingent on the behavior; muting increases. | | You skip sunscreen at the beach | Painful sunburn; you skip it less often | Positive punishment | An aversive stimulus is added; the behavior decreases. A natural punisher — no one had to deliver it. | | You text while walking | You walk into a lamppost; you text-and-walk less | Positive punishment | Pain is added; the behavior decreases. | | You leave your bike unlocked | The bike is stolen; you never leave one unlocked again | Negative punishment | A valued item is removed; the behavior decreases. One trial was enough. | | You show up late to a friend's dinners, repeatedly | The invitations stop; you become punctual with other friends | Negative punishment | A reinforcer (invitations) is withdrawn contingent on lateness; lateness decreases. | | You wave at a neighbor who never waves back | Nothing; after a few weeks you stop waving | Extinction | The behavior used to produce a wave back. With the reinforcer gone, it fades. | | You buy a lottery ticket every week | Occasional small wins; you keep buying | Schedule · VR | Reinforcement after an unpredictable number of responses — variable ratio, the most persistent pattern of all. | Notice how many everyday examples are natural consequences that nobody arranged. Sunburn, a stolen bike, and a good meal shape behavior exactly the way a food pellet shapes a rat's lever press. [The complete guide to operant conditioning ›](https://operantconditioning.com/) ## Examples at home and in parenting | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Child shares a toy with a sibling | Parent notices and says exactly what was good; sharing increases | Positive reinforcement | Specific attention is added right after the behavior. "Catch them being good." | | Child brushes teeth without a fuss | Then comes the bedtime story; brushing gets easier | Positive reinforcement | A preferred activity follows a less-preferred one — the Premack principle.[3] | | Teenager texts "home safe" | Parent's stream of check-in calls stops; texting increases | Negative reinforcement | The aversive stimulus (repeated calls) is removed; the behavior increases. | | Child eats the agreed three bites of broccoli | Excused from the table; bites happen faster | Negative reinforcement | Sitting at the table is aversive; the behavior ends it. Escape. | | Child whines for a snack before dinner | Parent gives in "just this once" — every few days; whining intensifies | Positive reinforcement | The snack is added, on a variable-ratio schedule. Giving in occasionally builds more persistent whining than giving in every time.[2] | | Child rides a bike without a helmet | The bike is put away for the rest of the day; helmet-less riding decreases | Negative punishment | A reinforcer (the bike) is removed; the behavior decreases. | | Child swears at dinner | A brief, sharp reprimand; swearing at dinner decreases | Positive punishment | An aversive stimulus is added; the behavior decreases. If swearing had *increased*, the reprimand would be attention — see the tricky cases below. | | Child calls out from bed for a fifth glass of water | Parent answers the first request only, then stops responding; after a noisy week the calling fades | Extinction | A behavior maintained by attention no longer produces it. Expect a burst before the decline.[4] | Families are full of *reciprocal* contingencies: the child's behavior is shaped by the parent's response, and the parent's response is shaped by what stops the child. Gerald Patterson called the escalating version of this the coercive family process.[5] [Operant conditioning in parenting ›](https://operantconditioning.com/applications/#parenting) ## Operant conditioning examples in the classroom | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Student hands in homework on time | Written feedback returned the next day; on-time work increases | Positive reinforcement | Feedback is added soon after the behavior. Prompt feedback reinforces; feedback three weeks later mostly doesn't. | | Class lines up quietly after recess | Teacher adds a marble to the jar; a full jar earns a class party | Positive reinforcement | A token — a conditioned, generalized reinforcer — is added. The basis of every token economy.[6] | | Student finishes every problem in class | Tonight's homework is waived; finishing in class increases | Negative reinforcement | An aversive task is removed contingent on the behavior. "Homework passes" work this way. | | Student who dreads reading aloud acts silly when called on | Teacher sighs and moves to the next student; acting silly increases | Negative reinforcement | The demand is removed. This is escape-maintained problem behavior, and it is extremely common.[7] | | Student turns in sloppy, illegible work | Has to redo it during free time; sloppy work decreases | Positive punishment | Extra effort is added contingent on the behavior; the behavior decreases. | | Student uses a phone during a lesson | Phone is kept at the desk until the bell; phone use decreases | Negative punishment | A reinforcer is removed for a set time; the behavior decreases. | | Student makes wisecracks that used to get laughs | Classmates stop laughing; the wisecracks dry up | Extinction | The reinforcer (peer attention) is no longer delivered. Peer attention is often the reinforcer teachers can't control. | | Students stay on task during independent work | Teacher circulates and gives praise at unpredictable moments | Schedule · VI | The first on-task moment after an unpredictable interval is reinforced — variable interval, which produces steady, sustained work. | Classroom studies in the 1960s established that teacher attention is a powerful reinforcer, that praise for on-task behavior plus planned ignoring of mild disruption reduces disruption, and that reprimands can backfire when the attention they carry is what the student is working for.[8] [Operant conditioning in education ›](https://operantconditioning.com/applications/#education) ## Examples at work | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Employee proposes a process fix | Manager credits her by name in the team meeting; proposals increase | Positive reinforcement | Public recognition is added. For many people it outperforms a small bonus. | | Employee completes the mandatory compliance module | The nagging pop-up at every login disappears; completion happens sooner | Negative reinforcement | An aversive stimulus is removed. Most corporate "reminders" are negative-reinforcement systems. | | New hire offers ideas in meetings | Nobody responds; after a month the ideas stop | Extinction | No reinforcer follows the behavior, so it fades. Silence trains people as surely as criticism does. | | Employee interrupts colleagues | A colleague calls it out in front of the group; interruptions drop | Positive punishment | An aversive stimulus is added; the behavior decreases — with the usual side effect of resentment toward the punisher. | | Employee expenses a personal dinner | Loses bonus eligibility for the quarter; padding expenses stops | Negative punishment | A reinforcer is removed contingent on the behavior. This is response cost. | | Worker assembles units on a production line | Paid a fixed amount per unit | Schedule · FR | Piece-rate pay is a fixed-ratio schedule: high rates of work with a pause after each payout. | [How organizational behavior management applies these ideas ›](https://operantconditioning.com/applications/#workplace) ## Animal training examples | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Dolphin touches a target pole | Whistle, then a fish; target-touching increases | Positive reinforcement | The whistle is a conditioned reinforcer that bridges the gap to the fish. Same logic as a clicker. | | Horse steps forward | Rider's leg pressure is released; the horse moves off pressure more readily | Negative reinforcement | Pressure is removed contingent on the behavior. Most traditional horsemanship is negative reinforcement. | | Cat jumps onto the kitchen counter | A motion-triggered puff of air; counter-jumping decreases | Positive punishment | An aversive stimulus is added. Because the device delivers it, the cat doesn't learn to avoid the owner. | | Parrot screams for attention | Owner leaves the room; screaming decreases | Negative punishment | The reinforcer (the owner's presence) is removed contingent on the behavior. | | Rat presses a lever that used to deliver food | Food no longer comes; pressing surges, then fades | Extinction | The laboratory original. The surge is the extinction burst.[4] | | Pigeon pecks a lighted key | Food after an average of 50 pecks | Schedule · VR | Variable ratio: pigeons on this schedule peck at very high, steady rates for long stretches.[2] | Modern dog training leans almost entirely on the first quadrant, for reasons the evidence supports. [Operant conditioning in dog training ›](https://operantconditioning.com/dog-training/) ## Examples in technology and apps | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | You complete a lesson in a language app | Chime, confetti, points; lesson completion increases | Positive reinforcement | Conditioned reinforcers are added immediately — the timing is the whole trick. | | You open a loot box in a game | Occasionally a rare item; opening boxes increases | Schedule · VR | Variable ratio, same as a slot machine. This is why regulators treat loot boxes as gambling-adjacent. | | You pay a bill in the banking app | The red "overdue" badge disappears; you pay sooner next month | Negative reinforcement | An aversive stimulus is removed contingent on the behavior. | | You hit "reply all" on a company-wide email | Two hundred "please remove me" replies; you never reply-all again | Positive punishment | An aversive flood is added; the behavior decreases sharply. | | You post spam in a forum | Post deleted and account suspended for 24 hours; spamming decreases | Negative punishment | Access (a reinforcer) is removed for a period. Time-out, in software. | | You open an app's notifications | They never contain anything useful; you start swiping them away unread | Extinction | Opening was reinforced by relevant content. Without it, the behavior fades — which is why irrelevant notifications destroy an app's reach. | [Operant design in technology ›](https://operantconditioning.com/applications/#technology) ## Examples in health and therapy | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Patient in addiction treatment submits a drug-negative urine sample | Voucher whose value rises with each consecutive negative sample; abstinence increases | Positive reinforcement | Contingency management — one of the best-supported treatments for stimulant use disorders.[9] | | Person afraid of flying drives twelve hours instead | The dread evaporates; driving instead of flying becomes the rule | Negative reinforcement | Avoidance removes fear and is reinforced by that relief. The fear itself never gets tested, so it never fades. | | Client in exposure therapy stays in the feared situation | Nothing bad happens; fear declines across sessions | Extinction | The avoidance response is blocked, and the conditioned fear extinguishes. [More on extinction ›](https://operantconditioning.com/extinction/) | | Patient takes a new medication | Nausea within the hour; doses get skipped | Positive punishment | An aversive stimulus is added after the behavior; taking the pill decreases. A major, under-recognized cause of non-adherence. | | Patient on a hospital ward shouts at staff | Loses tokens that buy privileges; shouting decreases | Negative punishment | Response cost within a token economy.[6] | [Operant conditioning in health and clinical practice ›](https://operantconditioning.com/applications/#health) ## Examples in nature and evolution | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | Crow drops a walnut onto a road | A passing car cracks it open; dropping nuts on roads increases | Positive reinforcement | Food access is added after the behavior. No trainer required. | | Toad snaps at a bumblebee | Stung on the tongue; snapping at striped insects decreases | Positive punishment | Pain is added; the behavior decreases, often after a single trial. | | Gull begs at a picnic table where nobody feeds it | Nothing; it stops coming to that table | Extinction | Begging that was reinforced at other tables is not reinforced here, so it fades — and comes under stimulus control of the tables that do pay. | Skinner argued that operant conditioning is a second kind of selection: natural selection picks traits across generations, and reinforcement picks behaviors within a lifetime.[10] Both work by consequences, and neither needs a plan. ## Examples for yourself: self-management | Behavior | Consequence | Quadrant | Why | | --- | --- | --- | --- | | You write 200 words | You cross off today's box on a paper tracker; writing sessions increase | Positive reinforcement | A conditioned reinforcer, delivered immediately by you. Small and instant beats big and delayed. | | You wash the dishes right after dinner | No crusted pile waiting in the morning; the habit sticks | Negative reinforcement | The behavior prevents an aversive state — avoidance, working in your favor. | | You eat a huge lunch | A 3 p.m. slump; big lunches decrease, slowly | Positive punishment | An aversive state is added — but two hours late, which is why natural punishers for eating are so weak. | | You miss a goal in a commitment contract | Money you pledged goes to a cause you dislike; missing decreases | Negative punishment | Response cost, arranged in advance by you against your future self. | | You get a generic "time to work out!" notification | Nothing happens whether you obey it or not; within two weeks you swipe it away unread | Extinction | A cue with no contingent consequence loses control over behavior. Most habit-app reminders die this way. | [The full guide to building habits with operant conditioning ›](https://operantconditioning.com/habits/) ## Examples of schedules of reinforcement in daily life Ferster and Skinner spent a decade cataloguing what happens when reinforcement arrives only some of the time.[2] Four schedules cover most of daily life. | Schedule | Reinforcer arrives… | Three everyday examples | Characteristic pattern | | --- | --- | --- | --- | | **Fixed ratio (FR)** | after a set number of responses | A coffee card stamped every purchase, free on the tenth · Piece-rate pay for each unit sewn · "After every 25 flashcards, a five-minute break" (self-set) | Fast, steady work with a pause right after each payout — the "just got my free coffee, no hurry" lull. | | **Variable ratio (VR)** | after an unpredictable number of responses | A slot machine · Cold-call sales, where roughly one call in thirty closes · Loot boxes and gacha games | Very high, very steady rates and extreme resistance to extinction. The hardest schedule to walk away from. | | **Fixed interval (FI)** | for the first response after a set time | Studying that ramps up the night before a weekly Friday quiz · Peering down the street for a bus that runs every 15 minutes · Checking the washing machine as its 45-minute cycle nears the end | The "scallop": almost nothing right after reinforcement, accelerating as the interval runs out. | | **Variable interval (VI)** | for the first response after an unpredictable time | Checking email · Redialing a busy customer-service line · A surfer paddling for waves that arrive at irregular intervals | Moderate, remarkably steady rates. Checking never quite stops because the next one might be the one. | The practical rule in every setting: reinforce continuously while a behavior is being learned, then thin to an intermittent schedule so it lasts. [Run the interactive schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) ## Tricky examples that fool people Textbook examples are clean. Real ones usually contain two contingencies at once, or a consequence whose effect is the opposite of its intent. Five that trip up students, parents, and trainers: ### 1. The candy that ends the tantrum A child screams for candy in the checkout line; the parent, mortified, hands it over; the screaming stops. Two things were learned. The child's tantrum was followed by candy — **positive reinforcement of tantrums**. The parent's giving-in was followed by the end of the screaming — **negative reinforcement of giving in**. Each person trained the other, and each will do it faster next time.[5] The fix is not a bigger punishment for the tantrum; it is making sure tantrums stop working (extinction) while *asking politely* starts working (reinforcement). ### 2. The scolding that is really attention A teacher tells a student to sit down every time he wanders; the wandering gets worse. Students almost always label this "positive punishment that failed." It isn't. The behavior increased, so the consequence was a reinforcer: the reprimand delivered attention, and attention was what the wandering was for. Classroom studies from the 1960s onward found exactly this — reprimands can increase the behavior they target.[8] In applied behavior analysis, this is called attention-maintained behavior, and it is identified with a functional analysis rather than guessed.[7] ### 3. The snooze button "The alarm is positive punishment for sleeping in." No: the alarm is an antecedent, not a consequence, and the behavior that matters is *pressing snooze*. Pressing snooze removes the alarm instantly, every single time — **negative reinforcement on a continuous schedule with zero delay**, which is about as strong as a contingency gets. That is why the only alarm that reliably works is the one across the room: it makes standing up, not pressing a button, the escape response. ### 4. The grounding that removes nothing A teenager is "grounded for a week" for a bad grade. She spends the week in her room, texting her friends, exactly as she would have anyway. Nothing that functions as a reinforcer was removed, so this is not **negative punishment** no matter what it is called — and if being grounded also excuses her from family dinners she dislikes, it may be negative *reinforcement* of whatever produced the grade. The test for negative punishment is that a reinforcer the person actually contacts is withdrawn and the behavior decreases. [How negative punishment actually works ›](https://operantconditioning.com/negative-punishment/) ### 5. The seat-belt chime People often call the chime "punishment for not wearing a seat belt." But "not buckling" is not a behavior that can be punished; it is the absence of one. The chime is present until you buckle; buckling ends it — **negative reinforcement of buckling** (escape). After a few weeks most drivers buckle before the chime ever starts, which is **avoidance**: the behavior now prevents the aversive stimulus rather than ending it. Engineers who design these systems are doing behavior analysis whether they call it that or not. > **A fast test for the hard cases** > > Name the behavior first — the specific thing the person or animal *did*. Then ask what changed in the environment right after it, and whether the behavior went up or down over the following days. If you find two behaviors (the child's and the parent's, the dog's and the owner's), analyze each one separately. Most "tricky" examples are just two ordinary examples stacked on top of each other. ## Classify it yourself Take any behavior from your own day and run it through the two questions. If you want a check, the [interactive quadrant finder on the home page](https://operantconditioning.com/#which-quadrant-is-it-an-interactive-check) asks the same two questions and names the quadrant for you, and the [20-question quiz](https://operantconditioning.com/quiz/) tests you on scenarios like the ones above with instant explanations. For the vocabulary, see the [glossary](https://operantconditioning.com/glossary/). ## Frequently asked questions **What are some examples of operant conditioning in everyday life?** Holding a door and getting a thank-you (positive reinforcement), putting on headphones to cut out noise (negative reinforcement), getting sunburned after skipping sunscreen (positive punishment), and having a bike stolen after leaving it unlocked (negative punishment). Lottery tickets and slot machines are variable-ratio schedules, and you stop waving at a neighbor who never waves back through extinction. **What is an example of operant conditioning in the classroom?** A teacher adds a marble to a jar when the class lines up quietly, and a full jar earns a party — positive reinforcement with tokens. A student who finishes all the problems in class has homework waived — negative reinforcement. A phone used during a lesson is kept until the bell — negative punishment. A student who acts silly to escape reading aloud, and gets skipped, is being negatively reinforced for acting silly. **What are the four types of operant conditioning with examples?** Positive reinforcement: a dolphin touches a target and gets a fish, so touching increases. Negative reinforcement: a horse steps forward and leg pressure is released, so stepping forward increases. Positive punishment: a cat jumps on the counter and gets a puff of air, so jumping decreases. Negative punishment: a parrot screams and the owner leaves the room, so screaming decreases. **Is a speeding ticket positive or negative punishment?** It depends which part you analyze, and textbooks disagree. Being pulled over and handed a citation adds an aversive event (positive punishment). Paying the fine and losing license points removes money and privileges (negative punishment). If speeding decreases, both descriptions are correct; if it doesn't, neither is, because punishment is defined by its effect. **What is an example of negative reinforcement that is not punishment?** Feeding a parking meter. Nothing unpleasant happens, and no behavior decreases; the behavior of feeding the meter increases because it prevents a ticket. Negative reinforcement always strengthens a behavior by removing or preventing something aversive. Punishment weakens a behavior. **What is an example of a variable-ratio schedule in everyday life?** A slot machine pays out after an unpredictable number of pulls; a salesperson closes roughly one cold call in thirty; a loot box occasionally contains a rare item. In each case reinforcement depends on the number of responses but the number varies, which produces high, steady rates of behavior that are very hard to extinguish. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 3. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 4. Lerman, D. C., & Iwata, B. A. (1995). Prevalence of the extinction burst and its attenuation during treatment. *Journal of Applied Behavior Analysis, 28*(1), 93–94. 5. Patterson, G. R. (1982). *Coercive Family Process*. Castalia. 6. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 7. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 8. Madsen, C. H., Becker, W. C., & Thomas, D. R. (1968). Rules, praise, and ignoring: Elements of elementary classroom control. *Journal of Applied Behavior Analysis, 1*(2), 139–150. 9. Higgins, S. T., Budney, A. J., Bickel, W. K., Foerg, F. E., Donham, R., & Badger, G. J. (1994). Incentives improve outcome in outpatient behavioral treatment of cocaine dependence. *Archives of General Psychiatry, 51*(7), 568–576. 10. Skinner, B. F. (1981). Selection by consequences. *Science, 213*(4507), 501–504. 11. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. ## Related - [Can you spot it?](https://operantconditioning.com/quiz/): Twenty scenarios, instant explanations. Can you name the quadrant every time? - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Fixed, variable, ratio, interval — with a live cumulative-record simulator. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): The quadrant behind the snooze button, the seat-belt chime, and most of the tricky cases. --- # Operant Conditioning Glossary: 128 Behavior Analysis Terms Defined > Operant conditioning glossary: 128 behavior analysis terms, from abolishing operation to variable ratio, defined precisely and linked to full explanations. - Source: https://operantconditioning.com/glossary/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-09 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reference* Every term you are likely to run into — in a paper, a therapy session, a dog-training class, or an argument on the internet — defined precisely, with links to the full explanations. - **ABC model**: The antecedent–behavior–consequence framework for reading any operant: what set the occasion (A), what the organism did (B), and what followed (C). It is another name for the three-term contingency and the basis of functional behavior assessment. [The ABC model in depth ›](https://operantconditioning.com/abc-model/) - **Abolishing operation (AO)**: A motivating operation that temporarily *decreases* the effectiveness of a reinforcer (or punisher) and decreases the current frequency of behavior that has produced it. A large meal is an abolishing operation for food: food reinforces less, and food-seeking drops. The opposite of an establishing operation. [Motivating operations ›](https://operantconditioning.com/abc-model/) - **Acquisition**: The phase of learning in which a new behavior is being established and its rate rises from baseline as reinforcement takes hold. Continuous reinforcement produces the fastest acquisition; behavior is then usually shifted to an intermittent schedule so that it persists. [Schedules of reinforcement ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Antecedent**: Any stimulus or condition present before a behavior that influences whether it occurs: a cue, an instruction, a place, a time of day, a bodily state. The "A" in A-B-C. The two most important kinds are discriminative stimuli, which signal that reinforcement is available, and motivating operations, which change how much the reinforcer is worth. [Antecedents explained ›](https://operantconditioning.com/abc-model/) - **Applied behavior analysis (ABA)**: The discipline that applies the principles of operant conditioning to socially significant behavior and measures whether the change is real. Baer, Wolf, and Risley defined the field in 1968 by seven dimensions: applied, behavioral, analytic, technological, conceptually systematic, effective, and showing generality. Used in autism intervention, education, organizational management, and behavioral medicine. [ABA in practice ›](https://operantconditioning.com/applications/#aba) - **Automatic reinforcement**: Reinforcement produced directly by the behavior itself rather than delivered by another person — scratching an itch, humming, rocking, twirling hair. Behavior that persists when the person is alone is often automatically reinforced, which is why "alone" is one of the conditions in a functional analysis. - **Autoshaping (sign-tracking)**: The emergence of a response — a pigeon pecking a lit key, a rat nosing a lever — when a stimulus is repeatedly followed by a reinforcer *whether or not the animal responds*. Discovered by Brown and Jenkins in 1968, it persists even when responding cancels the reinforcer, so it is not maintained by reinforcement; it is classical conditioning producing behavior directed at the signal. [Autoshaping explained ›](https://operantconditioning.com/operant-vs-classical-conditioning/#behavior-without-reinforcement-autoshaping-and-contrafreeloading) - **Aversive stimulus**: A stimulus an organism will work to escape or avoid. Functionally, one whose removal reinforces behavior (negative reinforcement) and whose presentation punishes it (positive punishment). Like reinforcers, aversive stimuli are identified by their effect on behavior, not by how unpleasant they seem to an observer. - **Avoidance**: Behavior that prevents an aversive stimulus from occurring at all, maintained by negative reinforcement: buckling up before the chime sounds, paying a bill before the late fee. Contrast escape, which ends an aversive stimulus that is already present. [Avoidance learning ›](https://operantconditioning.com/avoidance-learning/#discriminated-signaled-avoidance) - **Backward chaining**: Teaching a behavior chain by starting with the *last* step and adding earlier steps one at a time, so every training trial ends with the chain's natural reinforcer. A child learning to put on a shirt first does only the final tug down, then the last two steps, and so on. [Forward, backward and total-task chaining ›](https://operantconditioning.com/chaining/) - **Baseline**: Measurement of a behavior before any intervention, used as the comparison for judging whether the intervention worked. In single-case research designs the baseline is the "A" phase of an A-B or A-B-A-B design (unrelated to the "A" for antecedent). - **Behavior**: Anything an organism does that can be observed and measured — walking, talking, pressing, and, in radical behaviorism, private events such as thinking. A useful target behavior is active and specific, and passes the dead man's test. - **Behavior analysis**: The natural science of behavior founded on Skinner's work, with three branches: the experimental analysis of behavior (basic laboratory research), applied behavior analysis, and the conceptual branch, radical behaviorism. [B. F. Skinner ›](https://operantconditioning.com/bf-skinner/) - **Behavioral contrast**: A change in the rate of a behavior in one situation caused by a change in reinforcement in *another*. When reinforcement is reduced in the presence of one stimulus, responding often increases in the presence of a second stimulus where reinforcement has not changed (Reynolds, 1961). Practically: put a behavior on extinction at school and it may rise at home. - **Behavioral momentum**: John Nevin's metaphor for the persistence of behavior under disruption — extinction, distraction, satiation, or free reinforcers. The "mass" of a behavior depends on the rate of reinforcement obtained in the presence of a stimulus, so behavior from a richly reinforced context is harder to disrupt than behavior from a lean one. [Momentum and resistance to extinction ›](https://operantconditioning.com/extinction/#resistance-to-extinction-and-behavioral-momentum) - **Behaviorism**: The position that psychology should be a science of behavior. **Methodological behaviorism**, associated with John B. Watson, restricts the science to publicly observable events. **Radical behaviorism**, Skinner's version, treats private events — thoughts, feelings — as behavior subject to the same principles. [The history of behaviorism ›](https://operantconditioning.com/history/) - **Bridge (marker signal)**: A conditioned reinforcer, such as a clicker or the word "yes," delivered the instant the target behavior occurs to "bridge" the delay until the primary reinforcer arrives. The bridge is what makes precise timing possible when the treat is still in your pocket. [Clicker training ›](https://operantconditioning.com/dog-training/) - **Chaining**: Teaching a sequence of behaviors in which each response produces the discriminative stimulus for the next, with a reinforcer at the end of the chain. Built from a task analysis and taught forward, backward, or as a whole task. Contrast shaping, which builds a single new response. [Chaining in depth ›](https://operantconditioning.com/chaining/) - **Classical conditioning**: Pavlov's form of learning, in which a neutral stimulus paired with an unconditioned stimulus comes to elicit a reflexive response on its own: bell, then food, until the bell alone produces salivation. Also called respondent or Pavlovian conditioning. The key event comes *before* the response, whereas in operant conditioning it comes after. [Operant vs. classical conditioning ›](https://operantconditioning.com/operant-vs-classical-conditioning/) - **Compulsion loop**: A game-design pattern — anticipation, activity, reward, repeat — built on a variable schedule of reinforcement to keep players engaged. The term comes from the industry; the mechanism is the variable-ratio schedule. [Technology and product design ›](https://operantconditioning.com/applications/#technology) - **Conditioned punisher**: A previously neutral stimulus that has acquired punishing function by being paired with other punishers — the word "no," a frown, a warning light. Also called a secondary punisher. Like all punishers, it is defined by the fact that it decreases the behavior it follows. - **Conditioned reinforcer**: A previously neutral stimulus that has acquired reinforcing function through pairing with existing reinforcers: money, praise, grades, tokens, the click of a clicker. Also called a secondary reinforcer. Conditioned reinforcers make delayed real-world consequences workable. [Conditioned reinforcers in depth ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Consequence**: The stimulus change that follows a behavior — the "C" in A-B-C. A consequence that increases the behavior is reinforcement; one that decreases it is punishment; the absence of a consequence that used to arrive produces extinction. How much a consequence matters depends on contiguity and contingency. - **Contiguity**: Closeness in time between a behavior and its consequence. The shorter the delay, the stronger the effect; delays of even a few seconds sharply weaken operant learning, which is why a paycheck at month's end reinforces very little of any particular Tuesday's work. [Why timing matters ›](https://operantconditioning.com/positive-reinforcement/) - **Contingency**: The dependency between a behavior and a consequence: *if* this behavior, *then* this consequence. Learning requires that the consequence actually depend on the behavior; reinforcers that arrive regardless produce superstitious behavior instead. The word is also used for the whole antecedent–behavior–consequence relation, the three-term contingency. - **Contingency management**: A treatment, chiefly for substance use disorders, in which tangible reinforcers such as vouchers or prize draws are delivered contingent on objectively verified behavior — most often drug-negative urine samples or attendance. It is one of the best-supported psychosocial treatments for stimulant use disorders. [Operant conditioning in health care ›](https://operantconditioning.com/applications/#health) - **Continuous reinforcement (CRF)**: A schedule in which every occurrence of the behavior is reinforced. It produces the fastest acquisition and the fastest extinction, so it is used to build a behavior and then thinned to an intermittent schedule to make the behavior durable. [Continuous vs. intermittent reinforcement ›](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) - **Contrafreeloading**: Working for a reinforcer that is also freely available: a rat with a dish of food will still press a lever for the same food. Reported by Jensen in 1963 and found in most species tested, it shows that the opportunity to explore and manipulate is reinforcing in its own right. - **Cumulative record**: The graph produced by Skinner's cumulative recorder: paper feeds at a constant speed while a pen steps up with every response, so the slope of the line is the rate of responding. The signature patterns of each schedule — the fixed-interval scallop, the fixed-ratio break-and-run — were discovered on cumulative records. [The Skinner box and its recorder ›](https://operantconditioning.com/skinner-box/) - **Dead man's test**: A rule of thumb proposed by Ogden Lindsley: if a dead man can do it, it isn't behavior. "Not hitting" and "staying quiet" fail the test; "keeping hands on the desk" and "raising a hand" pass. It keeps target behaviors active and reinforceable. - **Delay discounting**: The decline in the present value of a reinforcer as the delay to receiving it increases. Steep discounting — preferring a small reward now to a larger one later — is associated with impulsivity and addiction, and it is the reason immediate consequences beat delayed ones in habit change. [Self-control and delay discounting ›](https://operantconditioning.com/matching-law/#self-control-when-the-alternatives-differ-in-time) - **Deprivation**: Going without a reinforcer for a period, which increases its effectiveness and increases behavior that has produced it. Hours without food make food a stronger reinforcer. Deprivation is the classic establishing operation; its opposite is satiation. - **Differential reinforcement**: Reinforcing one response or response class while withholding reinforcement from others. It is the procedure inside shaping and discrimination training, and the basis of the "DR" family of interventions: DRA (alternative behavior), DRI (incompatible behavior), DRO (other behavior), DRL (low rates, reinforcing only when responding is infrequent), and DRH (high rates). [Differential reinforcement in depth ›](https://operantconditioning.com/differential-reinforcement/) - **Differential reinforcement of alternative behavior (DRA)**: Reinforcing a desirable alternative to a problem behavior while placing the problem behavior on extinction — reinforcing "may I have a break?" instead of screaming. The most widely used function-based treatment; when the alternative is a communicative response, it is called functional communication training. [DRA, DRI, DRO and DRL ›](https://operantconditioning.com/differential-reinforcement/) - **Differential reinforcement of incompatible behavior (DRI)**: A form of DRA in which the reinforced alternative physically cannot occur at the same time as the problem behavior. Sitting is incompatible with running around the room; hands in pockets are incompatible with hitting. [DRA, DRI, DRO and DRL ›](https://operantconditioning.com/differential-reinforcement/) - **Differential reinforcement of other behavior (DRO)**: Delivering a reinforcer when the target behavior has *not* occurred for a set interval — a token for every five minutes without calling out. Despite the name, it reinforces the absence of a behavior rather than any specific alternative. Also called omission training. [DRA, DRI, DRO and DRL ›](https://operantconditioning.com/differential-reinforcement/) - **Discrete trial training (DTT)**: A teaching format in which learning is broken into short, clearly bounded trials: an instruction, a prompt if needed, the response, a consequence, and a brief pause. Associated with early intensive behavioral intervention for autism. Contrast the free operant, in which the learner can respond at any time. [ABA methods ›](https://operantconditioning.com/applications/#aba) - **Discrimination**: Responding differently in the presence of different stimuli, the result of reinforcement being available under one stimulus and not another. The rat presses when the light is on and not when it is off; you swear with friends and not with your grandmother. The opposite of generalization. [Stimulus control ›](https://operantconditioning.com/abc-model/) - **Discriminative stimulus (S^D)**: An antecedent stimulus in whose presence a behavior has been reinforced and in whose absence it has not, so that it now raises the probability of the behavior. A ringing phone is an S^D for answering; a green light for driving on. Its counterpart, the S-delta (S^Δ), signals that reinforcement is *not* available. [The discriminative stimulus explained ›](https://operantconditioning.com/abc-model/) - **Errorless learning**: Discrimination training arranged so the learner rarely or never responds to the wrong stimulus — for example, by introducing the SΔ so faintly and briefly at first that it is never responded to, then fading it in. Terrace (1963) showed it avoids the frustration and side effects of trial-and-error discrimination; it underlies prompting and fading in applied work. [Stimulus control ›](https://operantconditioning.com/abc-model/) - **Escape**: Behavior that terminates an aversive stimulus already present, maintained by negative reinforcement: taking an aspirin for a headache, leaving a loud room, handing a screaming child the candy. Contrast avoidance, which prevents the aversive stimulus from starting. [Escape, on the avoidance page ›](https://operantconditioning.com/avoidance-learning/#escape-the-easy-case) - **Establishing operation (EO)**: A motivating operation that temporarily *increases* the effectiveness of a reinforcer and increases the current frequency of behavior that has produced it. Food deprivation is the classic example; so are thirst, cold, and pain. The term was introduced by Jack Michael in 1982. [Motivating operations ›](https://operantconditioning.com/abc-model/) - **Extinction**: The procedure of no longer reinforcing a previously reinforced behavior, and the resulting decline in that behavior. Extinction is not one of the four quadrants: nothing is added or removed, the consequence that used to arrive simply stops. Extinguished behavior can return through spontaneous recovery, renewal, and resurgence. [Extinction in depth ›](https://operantconditioning.com/extinction/) - **Extinction burst**: A temporary increase in the frequency, intensity, or variability of a behavior when its reinforcer is first withheld, before the behavior declines. Pressing the elevator button harder and faster when it fails to respond. Giving in during the burst reinforces the more intense version of the behavior. [What is an extinction burst? ›](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) - **Fading**: Gradually removing prompts, or gradually changing a stimulus, so that the behavior comes under the control of the natural antecedent. Full physical guidance is faded to a light touch, then a gesture, then the spoken instruction alone. [Prompting and fading ›](https://operantconditioning.com/shaping/) - **Fixed interval (FI)**: A schedule in which the first response after a fixed period of time is reinforced (FI 60 s: the first press after a minute has elapsed). Produces the FI scallop: a pause after each reinforcer, then accelerating responding as the interval ends. [Fixed interval schedules in depth ›](https://operantconditioning.com/fixed-interval-schedule/) - **Fixed ratio (FR)**: A schedule in which reinforcement follows a set number of responses (FR 10: every tenth response). Produces a high, steady "break-and-run" rate with a post-reinforcement pause. Piecework pay and "buy ten, get one free" are fixed-ratio schedules. [Fixed ratio schedules in depth ›](https://operantconditioning.com/fixed-ratio-schedule/) - **Free operant**: A procedure in which the organism can respond at any time and at any rate — the lever in the operant chamber is always available — so that rate of responding becomes the primary measure. Skinner's central methodological innovation, in contrast to discrete-trial procedures such as mazes and puzzle boxes, where the experimenter starts each trial. - **Function of behavior**: What a behavior accomplishes for the organism: the reinforcer that maintains it. Functional assessment sorts problem behavior into four common functions — attention, escape or avoidance, access to tangibles, and automatic reinforcement — and treatment is matched to the function, not to what the behavior looks like. - **Functional analysis**: An experimental method for identifying the function of a behavior by systematically arranging conditions — attention, demand, alone, and a play control — and measuring which one produces the most behavior. Developed by Brian Iwata and colleagues (1982, reprinted 1994), it is the gold standard of assessment in applied behavior analysis. [ABA in practice ›](https://operantconditioning.com/applications/#aba) - **Functional behavior assessment (FBA)**: The broader process of identifying the antecedents and consequences maintaining a behavior, through interviews, rating scales, direct A-B-C observation and, when needed, a functional analysis. Widely required in schools before a behavior intervention plan is written. [A-B-C observation ›](https://operantconditioning.com/abc-model/) - **Generalization**: The spread of a learned behavior beyond the conditions in which it was trained. **Stimulus generalization**: responding to stimuli similar to the training stimulus. **Response generalization**: untrained but related responses also change. Generalization across settings, people, and time is a central goal of applied work (Stokes & Baer, 1977). The opposite of discrimination. - **Generalization gradient**: The curve relating response strength to how similar a test stimulus is to the training stimulus. A pigeon reinforced for pecking a 550 nm light pecks less and less as the color moves away from 550 nm (Guttman & Kalish, 1956). A steep gradient means sharp discrimination; a flat one means broad generalization. - **Generalized reinforcer**: A conditioned reinforcer that has been paired with many different reinforcers — money, tokens, social approval — so that it works regardless of any single deprivation state. Generalized reinforcers are why token economies work. [Generalized reinforcers in depth ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Habit**: In behavioral terms, a well-practiced operant under tight stimulus control, emitted with little deliberation when its cue appears. Habits are built by pairing a stable antecedent with a small behavior and an immediate consequence, and they persist because of intermittent reinforcement and behavioral momentum. In the neuroscience literature a habit is behavior that continues even after its outcome has lost value. [Building habits with operant conditioning ›](https://operantconditioning.com/habits/) - **Habituation**: The decline in responding to a stimulus that is simply repeated — you stop noticing the ticking clock. It is a form of non-associative learning and involves no consequences, so it is not operant conditioning. Distinct from extinction, which requires a history of reinforcement, and from satiation. - **Incentive salience**: The "wanting" that the brain's dopamine system attaches to reinforcers and to the cues that predict them, making them attention-grabbing and worth working for. Berridge and Robinson showed it is separate from "liking": animals without dopamine still enjoy sugar but will not seek it. [The neuroscience of operant conditioning ›](https://operantconditioning.com/neuroscience/) - **Instinctive drift**: The tendency of a trained behavior to drift toward the species-typical behavior the reinforcer naturally evokes, even at the cost of reinforcement. Keller and Marian Breland's raccoons, taught to drop coins in a box for food, began rubbing the coins together and refusing to let go — food-washing intruding on a trained response. Reported in "The misbehavior of organisms" (1961), it showed that reinforcement works within biological limits. [The Brelands in the history of the field ›](https://operantconditioning.com/history/) - **Instrumental conditioning**: The older term, from the Thorndike tradition, for what Skinner called operant conditioning: learning in which behavior is "instrumental" in producing an outcome. Some authors reserve "instrumental" for discrete-trial procedures such as mazes and puzzle boxes and "operant" for free-operant ones; in most modern usage the terms are interchangeable. [From Thorndike to Skinner ›](https://operantconditioning.com/history/) - **Intermittent reinforcement**: Any schedule in which only some occurrences of a behavior are reinforced — the fixed and variable ratio and interval schedules and their many variants. Also called partial reinforcement. Behavior on intermittent schedules is markedly more resistant to extinction than behavior on continuous reinforcement, the partial reinforcement extinction effect. [Continuous vs. intermittent reinforcement ›](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) - **Interval schedule**: A schedule in which reinforcement depends on the passage of time: the first response after an interval — fixed or variable — is reinforced. Responding faster does not bring reinforcement sooner, so interval schedules produce lower rates than ratio schedules. [Interval vs. ratio schedules ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Law of effect**: Edward Thorndike's principle, from his puzzle-box experiments with cats (1898; stated in full in 1911): responses followed by satisfaction are "stamped in" and become more likely, while responses followed by discomfort are "stamped out." The direct ancestor of reinforcement and punishment. [Thorndike and the law of effect ›](https://operantconditioning.com/history/) - **Learned helplessness**: The finding, first reported by Seligman and Maier in 1967, that animals exposed to inescapable aversive events later fail to escape even when escape becomes possible — as if they had learned that responding and outcomes are unrelated. Fifty years on, the same authors reframed it: passivity is the default response to prolonged aversive events, and what is learned is control. [Learned helplessness in depth ›](https://operantconditioning.com/learned-helplessness/) - **Learned industriousness**: Eisenberger's (1992) finding that reinforcing high effort in one task increases effort in unrelated tasks: the sensation of effort itself acquires conditioned reinforcing value. The counterpart of learned helplessness, and one reason to reinforce trying rather than only succeeding. - **Matching law**: Richard Herrnstein's 1961 finding that when two responses are available, the relative rate of each matches the relative rate of reinforcement it produces: pigeons on two keys distribute their pecks in proportion to the reinforcers each key delivers. It explains choice, and why a problem behavior persists when it is reinforced more richly than its alternative. [The matching law in depth ›](https://operantconditioning.com/matching-law/) - **Motivating operation (MO)**: An environmental variable that (1) alters the effectiveness of a stimulus as a reinforcer or punisher and (2) alters the current frequency of behavior that has produced that stimulus. Establishing operations increase both; abolishing operations decrease both. The concept was developed by Jack Michael. [Motivating operations ›](https://operantconditioning.com/abc-model/) - **Near miss**: A losing outcome that resembles a win — two cherries and a lemon on a slot machine. Near misses increase the urge to keep playing and recruit some of the same brain circuitry as wins, which is why machines are designed to produce them more often than chance would. [Gambling and games ›](https://operantconditioning.com/applications/#technology) - **Negative punishment**: The process in which a behavior is followed by the *removal* of a stimulus and decreases in future frequency as a result. Time-out, response cost, and losing privileges are the standard forms. "Negative" means removed, not bad. [Negative punishment in depth ›](https://operantconditioning.com/negative-punishment/) - **Negative reinforcement**: The process in which a behavior is followed by the *removal* (or avoidance) of a stimulus and increases in future frequency as a result. Buckling up to silence the chime; taking a painkiller to end a headache. It is not punishment: the behavior goes up. [Negative reinforcement in depth ›](https://operantconditioning.com/negative-reinforcement/) - **Noncontingent reinforcement (NCR)**: Delivering a reinforcer on a time-based schedule regardless of behavior, so the behavior no longer has to occur to obtain it. Used as a treatment for attention-maintained behavior — attention is given freely every few minutes — it breaks the contingency and acts as an abolishing operation for the reinforcer. Because no behavior is being strengthened, some analysts argue the procedure should not be called "reinforcement" at all (Poling & Normand, 1999). - **Operant**: A class of responses defined by its effect on the environment rather than by its form: "lever-pressing" includes pressing with the left paw, the right paw, or the nose. As an adjective, behavior that is emitted by the organism and controlled by its consequences. Skinner introduced the term in 1937. [Skinner's concept of the operant ›](https://operantconditioning.com/bf-skinner/) - **Operant conditioning**: Learning in which the future probability of a behavior is changed by the consequences that follow it: behaviors followed by reinforcement become more likely, behaviors followed by punishment less likely. Also called instrumental conditioning. [The complete guide to operant conditioning ›](https://operantconditioning.com/) - **Operant conditioning chamber (Skinner box)**: Skinner's apparatus for studying free-operant behavior: an enclosed space with a manipulandum (a lever or key), a device for delivering reinforcers (a food hopper or water dipper), optional stimuli such as lights and tones, and automatic recording. "Skinner box" was never Skinner's own term. [The Skinner box, part by part ›](https://operantconditioning.com/skinner-box/) - **Operant hoarding**: Letting earned reinforcers accumulate instead of consuming each one immediately. Cole (1990) found that rats on a schedule in which collecting pellets triggered a pause in reinforcement learned to let pellets pile up and collect them in batches — a form of self-control that simple impulsivity accounts do not predict. [Choice and self-control ›](https://operantconditioning.com/matching-law/) - **Operant variability**: Variability in behavior treated as a dimension that reinforcement controls. Page and Neuringer (1985) reinforced pigeons only for response sequences that differed from recent ones, and the birds became more variable; reinforcing repetition makes behavior stereotyped. New behavior in shaping comes from this variability. [Shaping ›](https://operantconditioning.com/shaping/) - **Overcorrection**: A positive punishment procedure developed by Richard Foxx and Nathan Azrin with two forms: **restitution** (repairing the environment beyond its original state — cleaning the whole table, not just the spill) and **positive practice** (repeatedly performing the correct form of the behavior). [Reprimands and overcorrection ›](https://operantconditioning.com/positive-punishment/#overcorrection) - **Partial reinforcement extinction effect (PREE)**: The finding that behavior reinforced intermittently is more resistant to extinction than behavior reinforced every time, first demonstrated systematically by Lloyd Humphreys in 1939. Paradoxically, less reinforcement produces more persistence. [The partial reinforcement extinction effect ›](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) - **Pavlovian-instrumental transfer (PIT)**: The increase in the rate of an operant behavior when a separately trained Pavlovian cue for the same or a similar reinforcer is presented — a rat presses faster for food while a tone that predicts food plays, although the tone was never part of the lever-press contingency. First reported by Estes in 1948; a laboratory model of how drug and food cues energize seeking. [How the two kinds of learning interact ›](https://operantconditioning.com/operant-vs-classical-conditioning/) - **Peak shift**: After discrimination training between an SD and a similar SΔ, the peak of the generalization gradient moves *away* from the SΔ: pigeons reinforced at 550 nm and extinguished at 555 nm respond most at about 540 nm (Hanson, 1959). Evidence that the SΔ carries an inhibitory gradient of its own. - **Positive punishment**: The process in which a behavior is followed by the *addition* of a stimulus and decreases in future frequency as a result. A burn after touching the stove; a reprimand after an interruption, if interruptions then decrease. "Positive" means added, not good. [Positive punishment in depth ›](https://operantconditioning.com/positive-punishment/) - **Positive reinforcement**: The process in which a behavior is followed by the *addition* of a stimulus and increases in future frequency as a result. A treat after a sit; praise after a chore; a like after a post. The most-used procedure in the field. [Positive reinforcement in depth ›](https://operantconditioning.com/positive-reinforcement/) - **Post-reinforcement pause**: The pause in responding that follows each reinforcer on fixed schedules (FR and FI). Its length grows with the size of the ratio or interval — the larger the requirement, the longer the break before the organism starts the next run. [The post-reinforcement pause ›](https://operantconditioning.com/fixed-ratio-schedule/) - **Premack principle**: David Premack's 1959 principle that a higher-probability behavior can reinforce a lower-probability behavior when access to it is made contingent: "first homework, then video games." Sometimes called Grandma's rule. [The Premack principle in depth ›](https://operantconditioning.com/premack-principle/) - **Primary reinforcer**: A stimulus that reinforces without any learning history because of its biological significance: food, water, warmth, sexual contact, relief from pain. Its effectiveness depends on deprivation. Also called an unconditioned reinforcer. [Primary and secondary reinforcers ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Prompt**: A supplementary antecedent stimulus added to evoke a correct response that the natural discriminative stimulus does not yet control: a verbal hint, a gesture, a model, physical guidance, or a highlighted stimulus. Prompts are meant to be faded. [Prompting ›](https://operantconditioning.com/shaping/) - **Prompt hierarchy**: An ordered set of prompts from least to most intrusive, used systematically in teaching. **Least-to-most** prompting waits for an independent response, then escalates — verbal, gestural, model, physical. **Most-to-least** begins with full guidance and fades it as the learner succeeds. - **Punisher**: A stimulus change that, when it follows a behavior, decreases the future frequency of that behavior. Defined entirely by effect: a consequence that fails to decrease the behavior is not a punisher, however unpleasant it looks. - **Punishment**: The process in which a behavior is followed by a stimulus change — the addition of a stimulus (positive punishment) or the removal of one (negative punishment) — and decreases in future frequency as a result. Punishment is defined by the decrease, not by intent or severity. [Punishment: types, evidence, alternatives ›](https://operantconditioning.com/punishment/) - **Radical behaviorism**: Skinner's philosophy of the science of behavior: private events such as thoughts and feelings are behavior, subject to the same principles as public behavior, and explanations should be sought in the organism's environmental history rather than in mental causes. "Radical" means thoroughgoing, not extreme. [Skinner's radical behaviorism ›](https://operantconditioning.com/bf-skinner/) - **Rate of response**: Responses per unit of time — Skinner's preferred measure of behavior, because it is continuous, sensitive, and reflects the probability of responding. It is read directly from the slope of a cumulative record. - **Ratio schedule**: A schedule in which reinforcement depends on the number of responses made, fixed or variable. Because faster responding produces reinforcement sooner, ratio schedules generate higher rates than interval schedules. [Ratio vs. interval schedules ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Ratio strain**: The breakdown of responding — long pauses, erratic rates, quitting — that occurs when a ratio schedule is thinned too quickly or set too high. The remedy is to lower the requirement and thin more gradually. [Ratio strain in depth ›](https://operantconditioning.com/fixed-ratio-schedule/) - **Reinforcement**: The process in which a behavior is followed by a stimulus change — the addition of a stimulus (positive reinforcement) or the removal of one (negative reinforcement) — and increases in future frequency as a result. Reinforcement is defined by the increase, not by whether the consequence seems rewarding. [Reinforcement: types, reinforcers, what makes it work ›](https://operantconditioning.com/reinforcement/) - **Reinforcer**: A stimulus change that, when it follows a behavior, increases the future frequency of that behavior. Not a synonym for "reward": a reward is something the giver thinks is nice, while a reinforcer is something that demonstrably strengthens the behavior it follows. [What counts as a reinforcer ›](https://operantconditioning.com/positive-reinforcement/) - **Renewal**: The return of an extinguished behavior when the context changes — a behavior extinguished at the clinic reappears at home. Extinction learning is unusually tied to the setting in which it happened, so extinction must be carried out in every context that matters. [Renewal ›](https://operantconditioning.com/extinction/#renewal) - **Resistance to extinction**: How long, and how much, a behavior persists once reinforcement stops. It is increased by intermittent schedules, by a rich history of reinforcement, and by a long history of the behavior paying off. [Resistance to extinction ›](https://operantconditioning.com/extinction/#resistance-to-extinction-and-behavioral-momentum) - **Respondent behavior**: Behavior elicited by a prior stimulus: reflexes and conditioned reflexes such as salivation, the startle response, the eye-blink, and conditioned fear. Skinner's term, chosen to contrast with operant behavior, which is emitted and controlled by its consequences. [Operant vs. respondent ›](https://operantconditioning.com/operant-vs-classical-conditioning/) - **Respondent conditioning**: Skinner's name for classical conditioning: a neutral stimulus paired with a stimulus that elicits a reflex comes to elicit the reflex itself. Both respondent and operant conditioning show extinction, spontaneous recovery, and renewal. [Full comparison ›](https://operantconditioning.com/operant-vs-classical-conditioning/) - **Response**: A single instance of behavior — one lever press, one spoken word, one glance at the phone. "Response" and "behavior" are used almost interchangeably; operant and response class refer to the group of responses that share a function. - **Response class**: A group of responses that share a function — the same effect on the environment — even when they look different. All the ways of getting a door open form one response class. This is why suppressing one form of a problem behavior often produces another form with the same function. - **Response cost**: A negative punishment procedure in which a specified amount of a reinforcer — tokens, points, money, minutes of screen time — is removed contingent on a behavior. Fines, penalties, and losing points are response cost. [Response cost ›](https://operantconditioning.com/negative-punishment/#response-cost) - **Response deprivation hypothesis**: Timberlake and Allison's (1974) rule for when one behavior will reinforce another: a contingency works if it restricts the contingent behavior below the level the organism performs when free. It generalizes the Premack principle and explains why a less-preferred activity can sometimes reinforce a more-preferred one. [Premack and response deprivation ›](https://operantconditioning.com/premack-principle/) - **Resurgence**: The reappearance of a previously reinforced behavior when a more recently reinforced behavior is placed on extinction. A child taught to ask politely instead of screaming goes back to screaming when polite requests stop being answered. The remedy is to keep the replacement behavior reliably reinforced. [Resurgence ›](https://operantconditioning.com/extinction/#resurgence) - **Reward prediction error**: The difference between the reinforcer received and the reinforcer predicted. Midbrain dopamine neurons fire a burst for a positive error, stay quiet for a fully predicted reward, and dip below baseline when an expected reward fails to arrive — the brain's teaching signal for operant learning and the quantity reinforcement-learning algorithms compute. [The neuroscience of operant conditioning ›](https://operantconditioning.com/neuroscience/) - **Satiation**: The reduced effectiveness of a reinforcer after it has been consumed or contacted in quantity, and the reduced behavior that follows: kibble is a weak reinforcer after dinner. Satiation is an abolishing operation; its opposite is deprivation. - **Scallop (fixed-interval scallop)**: The curved pattern a fixed-interval schedule leaves on a cumulative record: little responding just after a reinforcer, then a gradual acceleration to a high rate as the interval ends. Students who study little after an exam and cram before the next one trace the same curve. [The fixed-interval scallop ›](https://operantconditioning.com/fixed-interval-schedule/) - **Schedule of reinforcement**: The rule specifying which occurrences of a behavior will be reinforced: continuous, fixed ratio, variable ratio, fixed interval, variable interval, and many compound schedules built from them. Ferster and Skinner catalogued their effects in *Schedules of Reinforcement* (1957). [Schedules of reinforcement, with a simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) - **Secondary reinforcer**: Another name for a conditioned reinforcer: a stimulus that acquired its reinforcing power through pairing with other reinforcers, such as money, praise, or a clicker's sound. [Primary and secondary reinforcers ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Self-management**: Applying the principles of behavior to one's own behavior: arranging antecedents, defining small target behaviors, self-monitoring, and delivering one's own consequences. Skinner devoted a chapter of *Science and Human Behavior* (1953) to it, and it is the basis of evidence-based habit change. [Self-management and habits ›](https://operantconditioning.com/habits/) - **Setting event**: A broader condition, often distant in time, that alters how an antecedent–behavior–consequence relation plays out: poor sleep, illness, hunger, an argument earlier in the morning. Closely related to, and often treated as a kind of, motivating operation. [Antecedents and setting events ›](https://operantconditioning.com/abc-model/) - **Shaping**: Building a new behavior by reinforcing successive approximations toward it while withholding reinforcement from earlier approximations. Skinner used it to teach pigeons to play ping-pong; trainers use it to teach a dog to spin; speech therapists use it to build words from sounds. [How shaping works ›](https://operantconditioning.com/shaping/) - **Sidman avoidance (free-operant avoidance)**: Avoidance without a warning signal: shocks arrive on a timer unless the organism responds, and each response postpones the next shock for a set interval. Introduced by Murray Sidman in 1953; rats learn a steady response rate that keeps shocks rare. [Avoidance learning ›](https://operantconditioning.com/avoidance-learning/) - **Spontaneous recovery**: The reappearance of an extinguished behavior after a period away from the extinction setting, usually weaker than before and weaker with each recurrence if reinforcement is still withheld. First described by Pavlov for conditioned reflexes; it occurs just as reliably for operant behavior. [Spontaneous recovery ›](https://operantconditioning.com/extinction/#spontaneous-recovery) - **Stimulus**: Any event or energy change in the environment that can affect behavior: a sound, a light, a word, a touch, a food pellet. Stimuli that precede behavior are antecedents; stimuli that follow it are consequences. - **Stimulus control**: The condition in which a behavior occurs reliably in the presence of a particular stimulus and less often in its absence, established by discrimination training. Strong stimulus control is what makes a habit feel automatic and makes changing the cue an easier route to change than willpower. [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) - **Successive approximations**: The series of increasingly close versions of a target behavior that are reinforced, one after another, during shaping: any movement toward the lever, then touching it, then pressing it. [Shaping step by step ›](https://operantconditioning.com/shaping/) - **Superstitious behavior**: Behavior maintained by accidental reinforcement — a consequence that happened to follow the behavior without depending on it. In Skinner's 1948 experiment, pigeons fed at regular intervals regardless of what they did developed rituals such as turning and head-bobbing. Lucky socks work the same way. [Superstitious behavior in depth ›](https://operantconditioning.com/superstitious-behavior/) - **Target behavior**: The specific, observable, measurable behavior chosen for change in an intervention — "raises hand before speaking," not "is respectful." A good target behavior passes the dead man's test and can be counted. - **Task analysis**: Breaking a complex skill into its component steps, in order, as the basis for chaining. Hand-washing becomes eight discrete steps; each is taught and reinforced until the whole chain runs on its own. [Task analysis and chaining ›](https://operantconditioning.com/chaining/) - **Three-term contingency**: Skinner's basic unit of analysis: a discriminative stimulus sets the occasion, a response occurs, and a consequence follows. Written A → B → C, it is the same thing as the ABC model and the frame behind every entry in this glossary. [The three-term contingency ›](https://operantconditioning.com/abc-model/) - **Time-out**: Short for *time-out from positive reinforcement*: a negative punishment procedure in which access to reinforcement is removed for a brief period, contingent on a behavior. It works only when the "time-in" environment is actually reinforcing; a child sent from a hard lesson to a comfortable hallway has been negatively reinforced, not punished. [Time-out done correctly ›](https://operantconditioning.com/negative-punishment/#time-out-from-positive-reinforcement) - **Token economy**: A system in which tokens — points, stars, chips — are delivered contingent on target behaviors and later exchanged for backup reinforcers. Tokens are generalized conditioned reinforcers. Ayllon and Azrin's 1968 program on a psychiatric ward is the founding example. [Token economies in depth ›](https://operantconditioning.com/token-economy/) - **Topography**: The physical form of a behavior — what it looks like — as distinct from its function. A raised hand and a wave have similar topography and different functions; hitting and kicking have different topographies and may share one function. - **Two-factor theory of avoidance**: Mowrer's account of avoidance learning: a warning signal is first classically conditioned to elicit fear, then the avoidance response is operantly reinforced by escape from that fear. It explains why avoidance is so persistent — the animal never stays to learn the shock has stopped — but struggles with unsignaled avoidance and with the calm of well-trained avoiders. [Avoidance learning ›](https://operantconditioning.com/avoidance-learning/) - **Unconditioned reinforcer**: Another name for a primary reinforcer: a stimulus such as food, water, or warmth that reinforces without any prior learning. Contrast conditioned reinforcer. [Primary and secondary reinforcers ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) - **Variable interval (VI)**: A schedule in which the first response after an unpredictable interval, averaging some value, is reinforced (VI 60 s: intervals that average a minute). Produces a moderate, steady rate. Checking email is the everyday example: messages arrive on their own schedule, and the first check afterward is the one that pays off. [Variable interval schedules in depth ›](https://operantconditioning.com/variable-interval-schedule/) - **Variable ratio (VR)**: A schedule in which reinforcement follows an unpredictable number of responses, averaging some value (VR 10). Produces the highest, steadiest rate of any simple schedule and the greatest resistance to extinction — the schedule behind slot machines and infinite scroll. [Variable ratio schedules in depth ›](https://operantconditioning.com/variable-ratio-schedule/) - **Verbal behavior**: Skinner's 1957 analysis of language as operant behavior reinforced through the mediation of other people. It classifies verbal operants by function — the mand (request), tact (label), echoic (repetition), and intraverbal (reply to someone else's words), among others — and underlies most language programs in applied behavior analysis. Noam Chomsky's 1959 review of the book became a founding document of cognitive psychology. [Verbal Behavior and its critics ›](https://operantconditioning.com/history/) No terms match that search. ## Questions about operant conditioning terms **What's the difference between a reinforcer and reinforcement?** A reinforcer is the stimulus; reinforcement is the process. The treat is the reinforcer; the fact that giving the treat after a sit makes sitting more frequent is reinforcement. The same distinction separates a punisher (the stimulus) from punishment (the process). In both cases the stimulus earns its name only by its effect on behavior. **Why is it called positive punishment if it's bad?** Because "positive" in this vocabulary means *added*, like a plus sign, not pleasant. Positive punishment adds a stimulus after a behavior and the behavior decreases; negative punishment removes a stimulus and the behavior decreases. The same logic applies to reinforcement: positive reinforcement adds something, negative reinforcement removes something, and both make the behavior more likely. **What does S^D mean?** S^D, pronounced "ess-dee," stands for discriminative stimulus: an antecedent in whose presence a behavior has been reinforced, so that its presence now makes the behavior more likely. Its counterpart is S^Δ ("ess-delta"), a stimulus in whose presence the behavior has not been reinforced. The ringing phone is an S^D for answering; a phone that is switched off is an S^Δ. **Is a habit the same as an operant?** A habit is a kind of operant, but not every operant is a habit. An operant is any behavior controlled by its consequences. A habit is an operant that has been practiced so often in the presence of a stable cue that the cue alone triggers it with little deliberation — an operant under very strong stimulus control. That is why habits are built by fixing the antecedent and the consequence, not by relying on motivation. ## References Definitions follow the usage of the standard texts and primary sources below. 1. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 4. Catania, A. C. (2013). *Learning* (5th ed.). Sloan Publishing. 5. Michael, J. (1982). Distinguishing between discriminative and motivational functions of stimuli. *Journal of the Experimental Analysis of Behavior, 37*(1), 149–155. 6. Michael, J. (1993). Establishing operations. *The Behavior Analyst, 16*(2), 191–206. 7. Laraway, S., Snycerski, S., Michael, J., & Poling, A. (2003). Motivating operations and terms to describe them: Some further refinements. *Journal of Applied Behavior Analysis, 36*(3), 407–414. 8. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 9. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 10. Skinner, B. F. (1957). *Verbal Behavior*. Appleton-Century-Crofts. 11. Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. 12. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97. 13. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 14. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 15. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 16. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 17. Reynolds, G. S. (1961). Behavioral contrast. *Journal of the Experimental Analysis of Behavior, 4*(1), 57–71. 18. Seligman, M. E. P., & Maier, S. F. (1967). Failure to escape traumatic shock. *Journal of Experimental Psychology, 74*(1), 1–9. 19. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 20. Stokes, T. F., & Baer, D. M. (1977). An implicit technology of generalization. *Journal of Applied Behavior Analysis, 10*(2), 349–367. 21. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158. ## Related - [The ABC model](https://operantconditioning.com/abc-model/): Antecedent, behavior, consequence — the frame behind every term on this page. - [50+ examples](https://operantconditioning.com/examples/): See the terms in action, sorted by quadrant and setting. - [Can you spot it?](https://operantconditioning.com/quiz/): Twenty questions with instant explanations. Can you tell the quadrants apart? --- # Operant Conditioning Quiz: 20 Practice Questions With Instant Explanations > Free 20-question operant conditioning quiz with instant explanations: reinforcement and punishment scenarios, schedules, extinction, shaping, and more. - Source: https://operantconditioning.com/quiz/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-08 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Practice* Twenty questions, from a dog and a treat to motivating operations, with an explanation the moment you answer. Most people get the negative-reinforcement questions wrong. Will you? > **The two-question test** > > For any scenario, ask, in this order: **(1) Did the behavior become more or less likely?** More likely means reinforcement; less likely means punishment. **(2) Was something added or removed?** Added means positive; removed means negative. Ignore whether the consequence seems pleasant, ignore what anyone intended, and look only at what happened to the behavior. ## What this operant conditioning quiz covers This operant conditioning quiz has 20 multiple-choice practice questions that get harder as you go. The first eight ask you to identify the quadrant — positive or negative reinforcement, positive or negative punishment — from everyday scenarios that are written to catch the classic mistakes: the seat-belt chime, the candy that ends a tantrum, the scolding that turns out to be attention, the time-out that is really an escape. The rest cover [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/), the [extinction burst and spontaneous recovery](https://operantconditioning.com/extinction/), [shaping versus chaining](https://operantconditioning.com/shaping/), [operant versus classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/), a little [history](https://operantconditioning.com/history/), and the two ideas that separate people who understand this subject from people who have memorized it: the motivating operation and the functional definition of a reinforcer. ## How the quiz is scored One point per question, no penalty for wrong answers, and a short explanation after every answer whether you got it right or not. Your score appears at the end, and you can retake the quiz as often as you like. Nothing is recorded or sent anywhere. ## Answer key and explanations Every question from the quiz, with the correct answer and the reasoning. Open any question to check your thinking. **Q1. A dog sits when told to and immediately gets a treat. Over the next week, the dog sits on cue more reliably. Which process is this?** **Answer:** Positive reinforcement. Ask the two questions. Did the behavior increase? Yes, so it is reinforcement. Was something added or removed? A treat was added, so it is positive reinforcement. **Q2. You start the car and an irritating chime sounds until you buckle your seat belt. Over time you buckle up faster and more consistently. What is maintaining the buckling?** **Answer:** Negative reinforcement. Buckling increased, so this is reinforcement, and it increased because an aversive stimulus (the chime) was removed, so it is negative reinforcement. Nothing here is punishment: no behavior decreased. **Q3. A teenager comes home after curfew and loses the car keys for a week. Curfew violations become less frequent. Which quadrant is this?** **Answer:** Negative punishment. The behavior decreased, so it is punishment. It decreased because access to something (the car) was removed, so it is negative punishment. Grounding and losing privileges are the classic examples. **Q4. A teacher scolds a student every time he calls out without raising his hand. Over the month, calling out increases. What has happened?** **Answer:** Positive reinforcement, because the behavior increased. Consequences are defined by their effect, not by how they look. Scolding added a stimulus (attention) and the behavior went up, so the scolding functioned as a positive reinforcer. For a student who gets little attention, even a reprimand can be the payoff. **Q5. A child screams for candy in the checkout line. The parent hands over the candy and the screaming stops. On later shopping trips, screaming becomes more frequent. What is the candy, for the child?** **Answer:** A positive reinforcer for screaming. Candy was added after the screaming and the screaming increased: positive reinforcement of the tantrum. Call it a bribe if you like, but the contingency is still doing its work, and it is working on the child. **Q6. In the same checkout-line scenario, the parent finds herself handing over the candy faster and faster on each trip. What is happening to the parent’s behavior?** **Answer:** It is being negatively reinforced by the end of the screaming. The parent’s candy-giving increased because an aversive stimulus (the screaming) stopped when she did it. That is escape, which is negative reinforcement. Two behaviors were strengthened in one exchange, and neither person intended either of them. **Q7. A student hits the snooze button, the alarm stops, and she drifts back to sleep. Over the semester she hits snooze more and more often. Which process best explains the snooze-pressing?** **Answer:** Negative reinforcement: the alarm is removed. The immediate consequence of pressing the button is that the alarm stops: an aversive stimulus is removed and the behavior increases. That is negative reinforcement by escape. Extra sleep may add positive reinforcement on top, but the defining event is the removal of the noise. Being late is delayed and inconsistent, which is exactly why it fails to punish. **Q8. A boy who dislikes math worksheets is sent to the hallway for five minutes every time he swears during math. Swearing during math increases. What is the hallway time-out actually doing?** **Answer:** Negatively reinforcing swearing by removing the worksheet. Time-out is a punishment procedure only when the behavior decreases. Here swearing increased, and what changed when he swore was that the aversive task went away, so the time-out is functioning as negative reinforcement (escape). Time-out only works when the time-in environment is more reinforcing than the time-out. **Q9. A garment worker is paid a fixed amount for every 20 shirts she sews. Which schedule of reinforcement is this?** **Answer:** Fixed ratio (FR). Reinforcement depends on a set number of responses, so it is a ratio schedule, and the number is always the same, so it is fixed: FR 20. Fixed-ratio schedules produce a high rate with a brief pause after each reinforcer. **Q10. New email arrives at unpredictable times. Checking your inbox more often does not make messages arrive sooner, but the first check after a message lands is the one that pays off. Which schedule is this?** **Answer:** Variable interval (VI). Reinforcement depends on time passing, not on how many checks you make, so it is an interval schedule, and the time is unpredictable, so it is variable: VI. Variable-interval schedules produce moderate, steady responding. **Q11. Which schedule of reinforcement produces behavior that is most resistant to extinction?** **Answer:** Variable ratio. On a variable-ratio schedule the organism can never tell whether the next response will pay off, so responding persists long after reinforcement stops. It is the schedule behind slot machines and infinite scroll. Continuous reinforcement is at the other extreme: fast to learn, fast to extinguish. **Q12. On a cumulative record, one schedule shows a pause after each reinforcer followed by a gradually accelerating rate of responding as the next reinforcer approaches, a pattern called a scallop. Which schedule is it?** **Answer:** Fixed interval (FI). The FI scallop appears because responding early in a fixed interval is never reinforced, so the organism waits, then speeds up as the interval ends. Students who study little after an exam and cram before the next one show the same curve. Fixed ratio produces a pause too, but it is followed by an abrupt switch to a high, steady rate, not a gradual acceleration. **Q13. For months, a toddler’s bedtime crying has always brought a parent into the room. The parents decide to stop going in. On the first night the crying is louder, longer, and more varied than ever before. What is this?** **Answer:** An extinction burst. When a reinforcer is first withheld, behavior often becomes more frequent, more intense, and more variable before it declines. That is the extinction burst. Giving in at this point reinforces the louder version of the crying on an intermittent schedule, which makes it harder to extinguish later. **Q14. The parents hold firm and the crying stops within a week. Ten days later, with nothing else changed, the crying briefly returns one night at a lower intensity, then fades again. What is this called?** **Answer:** Spontaneous recovery. Spontaneous recovery is the reappearance of an extinguished behavior after time away from extinction, usually weaker than before and weaker each time if reinforcement is still withheld. It is not a sign that extinction failed. Resurgence is different: an old behavior returning when a newer replacement behavior stops being reinforced. **Q15. A trainer teaching a dog to spin in a circle first rewards a slight head turn, then a quarter turn, then a half turn, and finally a full spin, withholding rewards for earlier versions as each new step is learned. Which procedure is this?** **Answer:** Shaping. Reinforcing successive approximations of a target behavior while extinguishing earlier ones is shaping. Chaining is different: it links separate, already-learned behaviors into a sequence in which each step cues the next, usually built from a task analysis. **Q16. A dog starts to salivate when it hears the treat bag rustle. Which kind of learning best describes the salivation?** **Answer:** Classical (respondent) conditioning. Salivation is a reflex elicited by a stimulus, not a voluntary behavior strengthened by its consequences. The rustle was paired with food and now elicits the response on its own: classical conditioning. The dog’s running to the kitchen when it hears the bag, by contrast, is operant behavior, reinforced by getting the treat. **Q17. Which statement correctly distinguishes operant conditioning from classical conditioning?** **Answer:** Classical conditioning involves reflexive responses elicited by antecedent stimuli; operant conditioning involves emitted behavior controlled by its consequences. Classical conditioning is stimulus-stimulus learning: a neutral stimulus paired with an unconditioned stimulus comes to elicit a reflexive response, and the key event comes before the response. Operant conditioning is behavior-consequence learning: the organism emits a behavior and what follows changes its future probability. Both apply across species. **Q18. Whose puzzle-box experiments with cats produced the law of effect, the principle Skinner later developed into the concept of reinforcement?** **Answer:** Edward Thorndike. Thorndike timed cats escaping from latched boxes and found that responses followed by a satisfying outcome were “stamped in.” He published the law of effect in 1898. Skinner coined the term “operant” in 1937 and formalized the science in *The Behavior of Organisms* (1938); Pavlov studied reflexes; Watson launched behaviorism in 1913. **Q19. A rat has learned to press a lever for food pellets only when a light is on. One day the experimenter feeds the rat to fullness before the session. The light comes on, but the rat barely presses. What is the pre-session feeding?** **Answer:** An abolishing operation that reduces the value of food as a reinforcer. The light is the discriminative stimulus: it signals that pressing will pay off, and it still does. What changed is the rat’s motivation. Satiation is an abolishing operation, a motivating operation that lowers the effectiveness of a reinforcer and reduces the behavior that produces it. Discriminative stimuli signal availability; motivating operations change value. **Q20. A manager begins praising employees publicly whenever they submit reports early. Over the next month, early submissions decrease. Which statement is correct?** **Answer:** The praise was not a reinforcer for this behavior; whether a stimulus is a reinforcer is determined by its effect on behavior. A reinforcer is defined by its effect: a stimulus that follows a behavior and increases it. If early submissions fell after public praise was added, the praise was not reinforcing them; if the drop was caused by the praise, it functioned as a positive punisher, which is common when public attention is embarrassing. The only way to know what reinforces a behavior is to watch what happens to the behavior. ## Study guide: where each topic is explained Missed a few? Each question maps to one page on this site. Read the page, retake the quiz, and the same scenarios should feel obvious. | Questions | Topic | Where to study it | | --- | --- | --- | | 1, 4, 5, 20 | Positive reinforcement and the functional definition of a reinforcer | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | | 2, 6, 7, 8 | Negative reinforcement, escape, and avoidance | [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) | | 4, 20 | Positive punishment and why attention is not always punishment | [Positive punishment](https://operantconditioning.com/positive-punishment/) | | 3, 8 | Negative punishment, grounding, and time-out done correctly | [Negative punishment](https://operantconditioning.com/negative-punishment/) | | 9–12 | Fixed and variable, ratio and interval schedules; the scallop; resistance to extinction | [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) (with a simulator) | | 13, 14 | Extinction burst, spontaneous recovery, resurgence, renewal | [Extinction](https://operantconditioning.com/extinction/) | | 15 | Shaping, successive approximations, chaining | [Shaping](https://operantconditioning.com/shaping/) | | 16, 17 | Respondent vs. operant behavior | [Operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/) | | 18 | Thorndike, Watson, Skinner | [B. F. Skinner](https://operantconditioning.com/bf-skinner/) · [History](https://operantconditioning.com/history/) | | 19 | Discriminative stimuli and motivating operations | [The ABC model](https://operantconditioning.com/abc-model/) | For quick definitions of any term that appeared in the quiz, use the [operant conditioning glossary](https://operantconditioning.com/glossary/); for more scenarios to practice on, see the [50+ examples sorted by quadrant](https://operantconditioning.com/examples/). The [complete guide](https://operantconditioning.com/) covers everything on this page in one place. ## Frequently asked questions **Is this quiz free?** Yes. Nothing is recorded, you can take it as many times as you like, and the full answer key with explanations is printed on this page. **How many questions are in the operant conditioning quiz?** Twenty multiple-choice questions, each with four options and an instant explanation. Eight cover identifying the four quadrants from scenarios, four cover schedules of reinforcement, two cover extinction, and the rest cover shaping, operant versus classical conditioning, history, motivating operations, and the functional definition of a reinforcer. **Can I use this to study for AP Psychology or an intro psych exam?** Yes. The quiz covers the operant conditioning material that appears in AP Psychology and most introductory psychology courses: the four quadrants, the schedules and their response patterns, extinction and spontaneous recovery, shaping, and the difference between operant and classical conditioning. The scenario questions are written in the same style as exam items, and the study guide above links each topic to a full explanation. **What score is good?** Fourteen or more out of twenty (70%) indicates a solid grasp of the fundamentals; eighteen or more means you could teach it. If you score below ten, the quadrant pages and the schedules page will fix most of the gaps, because those topics account for more than half the questions. Most first-time mistakes are the negative-reinforcement items, which is exactly what the quiz is designed to expose. ## Related - [Glossary](https://operantconditioning.com/glossary/): 128 terms from "abolishing operation" to "variable ratio," defined. - [50+ examples](https://operantconditioning.com/examples/): More scenarios to practice on, sorted by quadrant and setting. - [The complete guide](https://operantconditioning.com/): Operant conditioning from definition to application, on one page. --- # Operant Conditioning Diagrams to Download and Reuse > Every diagram on this site — four quadrants, extinction curve, four schedules, shaping, escape vs. avoidance, peak shift, Skinner box — as SVG and PNG, CC BY. - Source: https://operantconditioning.com/diagrams/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-10 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reference · Diagram library* Every diagram on this site, as a crisp vector file for slides and handouts and as a PNG for anything else. All of them are free to reuse with attribution. Each image below is the diagram exactly as it appears on its page, redrawn as a standalone file with a white background and no external dependencies. Click a diagram to open the SVG, which scales to any size; use the PNG where a bitmap is required. The text under each one gives the page it comes from, and the description that screen readers hear. > **License: Creative Commons Attribution 4.0 (CC BY 4.0)** > > You may copy, adapt, and redistribute these diagrams, including in slides, handouts, worksheets, videos, and commercial textbooks, provided you credit **operantconditioning.com** and link to the diagram's page. A line such as "Diagram: operantconditioning.com (CC BY 4.0)" is enough. The license text is at [creativecommons.org/licenses/by/4.0](https://creativecommons.org/licenses/by/4.0/). ![A two-by-two grid. Columns: stimulus added (positive) and stimulus removed (negative). Rows: behavior increases (reinforcement) and behavior decreases (punishment). Top left, positive reinforcement: a treat after a sit. Top right, negative reinforcement: the seat-belt chime stops. Bottom left, positive punishment: a burn after touching a hot pan. Bottom right, negative punishment: losing the car keys after breaking curfew. Extinction sits outside the grid: nothing is added or removed, the reinforcer simply stops.](https://operantconditioning.com/assets/diagrams/the-four-quadrants-of-operant-conditioning.svg) **The four quadrants of operant conditioning** — The four quadrants. The column answers "added or removed?"; the row answers "more or less likely?" Positive and negative are arithmetic signs, not judgments. — From [Examples](https://operantconditioning.com/examples/) · [SVG](https://operantconditioning.com/assets/diagrams/the-four-quadrants-of-operant-conditioning.svg) · [PNG](https://operantconditioning.com/assets/diagrams/the-four-quadrants-of-operant-conditioning.png) ![A rectangular chamber. On the left wall: a stimulus light near the top, a lever at mid-height, and a food tray at the bottom. A speaker sits high on the right wall. The floor is a grid. Outside the chamber, a cumulative recorder draws a rising stepped line on a moving strip of paper.](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber.svg) **Schematic of an operant conditioning chamber** — An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a food pellet on any schedule; the light signals when pressing will pay off; the recorder draws responses over time. — From [The complete guide](https://operantconditioning.com/) · [SVG](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber.svg) · [PNG](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber.png) ![A chamber inside a sound-attenuating cubicle. On the left wall, from top to bottom: a cue light, a lever, and a food tray. A house light sits at the top of the chamber and a speaker high on the right wall. The floor is a grid of metal rods. Outside the chamber, a cumulative recorder draws a rising stepped line on a moving strip of paper, with small ticks marking reinforcers.](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber-for-a-rat.svg) **Schematic of an operant conditioning chamber for a rat** — An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a pellet on any schedule; the cue light signals when pressing will pay off; the recorder draws responses against time. — From [Skinner Box](https://operantconditioning.com/skinner-box/) · [SVG](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber-for-a-rat.svg) · [PNG](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber-for-a-rat.png) ![A line that rises from left to right. It begins with steep runs separated by flat pauses, each run ending in a diagonal tick. The pen then resets to the bottom and draws three curved scallops that start flat and accelerate into each tick. Finally, when reinforcement stops, the line rises briefly and then flattens out.](https://operantconditioning.com/assets/diagrams/an-annotated-cumulative-record.svg) **An annotated cumulative record** — Reading a cumulative record: slope is rate, ticks are reinforcers, flat stretches are pauses, and the curved "scallop" is the fixed-interval pattern. The line only ever goes up; a behavior that has stopped draws a horizontal line. — From [Skinner Box](https://operantconditioning.com/skinner-box/) · [SVG](https://operantconditioning.com/assets/diagrams/an-annotated-cumulative-record.svg) · [PNG](https://operantconditioning.com/assets/diagrams/an-annotated-cumulative-record.png) ![Four panels. Fixed ratio: steep runs separated by flat post-reinforcement pauses. Variable ratio: a steady steep line. Fixed interval: repeated scallops that start flat and curve upward before each reinforcer. Variable interval: a moderate, steady slope.](https://operantconditioning.com/assets/diagrams/idealized-cumulative-records-for-the-four-basic-schedules.svg) **Idealized cumulative records for the four basic schedules** — Idealized cumulative records. Each upward step is a response; green ticks mark reinforcers. Patterns after Ferster & Skinner (1957). — From [Schedules of Reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) · [SVG](https://operantconditioning.com/assets/diagrams/idealized-cumulative-records-for-the-four-basic-schedules.svg) · [PNG](https://operantconditioning.com/assets/diagrams/idealized-cumulative-records-for-the-four-basic-schedules.png) ![Response rate over time after reinforcement stops: a brief spike (the extinction burst), then a decline toward zero, a rest, and a smaller return of responding (spontaneous recovery) that fades again.](https://operantconditioning.com/assets/diagrams/the-extinction-curve.svg) **The extinction curve** — The shape of extinction. When reinforcement stops, responding briefly rises (the burst), then declines; after a rest it returns at a lower level (spontaneous recovery) and fades again. — From [Extinction](https://operantconditioning.com/extinction/) · [SVG](https://operantconditioning.com/assets/diagrams/the-extinction-curve.svg) · [PNG](https://operantconditioning.com/assets/diagrams/the-extinction-curve.png) ![Successive bell-shaped distributions of a behavior's form move rightward across the page; a criterion line steps upward so that only the upper tail of each distribution is reinforced, dragging the next distribution further toward the target.](https://operantconditioning.com/assets/diagrams/shaping-as-a-rising-criterion.svg) **Shaping as a rising criterion** — Shaping as a moving criterion. Reinforcing only the upper tail of what the animal currently does shifts the whole distribution; the criterion is then raised again. Variability (the width of each curve) is what makes the next step possible. — From [Shaping](https://operantconditioning.com/shaping/) · [SVG](https://operantconditioning.com/assets/diagrams/shaping-as-a-rising-criterion.svg) · [PNG](https://operantconditioning.com/assets/diagrams/shaping-as-a-rising-criterion.png) ![Two generalization gradients plotted against wavelength. A dashed curve for birds trained at 550 nanometers only peaks at 550. A taller solid curve for birds trained at 550 with extinction at 555 peaks at about 540, shifted away from the S-minus, and drops steeply on the 555 side.](https://operantconditioning.com/assets/diagrams/peak-shift-after-discrimination-training.svg) **Peak shift after discrimination training** — Schematic of Hanson's result. After discrimination training against a nearby S−, the peak of the gradient moves away from the S− (here from 550 to about 540 nm) and rises above the control gradient. — From [Stimulus Control](https://operantconditioning.com/stimulus-control/) · [SVG](https://operantconditioning.com/assets/diagrams/peak-shift-after-discrimination-training.svg) · [PNG](https://operantconditioning.com/assets/diagrams/peak-shift-after-discrimination-training.png) ![Two timelines. Escape: the aversive stimulus is on, the response occurs, the stimulus ends. Avoidance: a warning signal appears, the response occurs during it, and the aversive stimulus never arrives.](https://operantconditioning.com/assets/diagrams/escape-versus-avoidance.svg) **Escape versus avoidance** — Escape ends an aversive stimulus that is present; avoidance prevents one that would have come. Both increase the behavior, so both are negative reinforcement. — From [Negative Reinforcement](https://operantconditioning.com/negative-reinforcement/) · [SVG](https://operantconditioning.com/assets/diagrams/escape-versus-avoidance.svg) · [PNG](https://operantconditioning.com/assets/diagrams/escape-versus-avoidance.png) ![Left panel: a stimulus precedes a reflexive response; the organism does nothing to produce the stimulus. Right panel: a behavior is followed by a consequence, which changes the future rate of the behavior.](https://operantconditioning.com/assets/diagrams/classical-versus-operant-conditioning-as-sequences.svg) **Classical versus operant conditioning as sequences** — The two kinds of learning as sequences. In classical conditioning a stimulus comes to elicit a reflex; in operant conditioning a consequence changes how often a voluntary behavior recurs. — From [Operant vs. Classical](https://operantconditioning.com/operant-vs-classical-conditioning/) · [SVG](https://operantconditioning.com/assets/diagrams/classical-versus-operant-conditioning-as-sequences.svg) · [PNG](https://operantconditioning.com/assets/diagrams/classical-versus-operant-conditioning-as-sequences.png) ![A scatter plot with proportion of reinforcers on the horizontal axis and proportion of responses on the vertical axis. Data points lie close to the diagonal line where the two proportions are equal.](https://operantconditioning.com/assets/diagrams/the-matching-relation.svg) **The matching relation** — Schematic of the matching relation. Real data cluster around the diagonal; systematic departures from it are captured by the generalized matching law below. — From [Matching Law](https://operantconditioning.com/matching-law/) · [SVG](https://operantconditioning.com/assets/diagrams/the-matching-relation.svg) · [PNG](https://operantconditioning.com/assets/diagrams/the-matching-relation.png) ![Two panels of bar charts. Left: a water-deprived rat spends much more baseline time drinking than running, so drinking reinforces running. Right: a running-deprived rat spends more time running than drinking, so running reinforces drinking.](https://operantconditioning.com/assets/diagrams/the-premack-principle-as-a-reversal-of-baseline-probabilities.svg) **The Premack principle as a reversal of baseline probabilities** — Premack's reversal. Whichever behavior is more probable at baseline can reinforce the other; deprivation changes which one that is. — From [Premack Principle](https://operantconditioning.com/premack-principle/) · [SVG](https://operantconditioning.com/assets/diagrams/the-premack-principle-as-a-reversal-of-baseline-probabilities.svg) · [PNG](https://operantconditioning.com/assets/diagrams/the-premack-principle-as-a-reversal-of-baseline-probabilities.png) ## Using the diagrams in teaching - **Slides.** Drag the SVG into Keynote, PowerPoint, or Google Slides; it stays sharp at any projector size and inherits none of the site's fonts. - **Handouts and exams.** The extinction curve, the escape-versus-avoidance timeline, and the schedules diagram work as unlabeled prompts: crop the labels and ask students to supply them. - **Learning management systems.** Use the PNG where the platform will not render SVG. Every SVG carries an accessible description in its `` and `<desc>`; the same description appears under each image here, so paste it as the alt text wherever you re-upload. - **Attribution.** "operantconditioning.com, CC BY 4.0" with a link to the source page satisfies the license. Want a diagram that is not here? [Ask](https://operantconditioning.com/about/#contact); if it would help other teachers, it will be added to this page. ## Frequently asked questions **Can I use these diagrams in a textbook or a course I sell?** Yes. CC BY 4.0 permits commercial use. The only conditions are attribution to operantconditioning.com and a note if you changed the diagram. **Can I edit the diagrams?** Yes. The SVG files are plain text and open in Figma, Illustrator, Inkscape, or any text editor. Adaptations must be credited as adaptations. **Does the license cover the rest of the site?** No. The text, interactive tools, and other images remain © Operant Conditioning Inc. under the [terms of use](https://operantconditioning.com/terms/#copyright), which still allow quoting with attribution and non-commercial classroom use. ## Related - [For teachers](https://operantconditioning.com/#for-teachers): How to assign pages, the quiz, and the discussion questions. - [Glossary](https://operantconditioning.com/glossary/): Every term, defined and linked. - [Examples](https://operantconditioning.com/examples/): Fifty-plus scenarios sorted by quadrant. --- # For Teachers & Students: Cite, Assign, and Teach Operant Conditioning > Teach or study operant conditioning from this site: a suggested route, learning objectives, discussion questions, citation formats, and reuse permissions. - Source: https://operantconditioning.com/for-teachers/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-11 · Updated: 2026-09-11 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *For teachers & students* # Teach it, cite it, reuse it. Every page on this site can be assigned, quoted, printed, and linked without an account, a paywall, or a tracker. This page collects the parts built specifically for teaching and studying — the route through the material, the objectives it covers, discussion questions, citation formats, and exactly what you are allowed to reuse. > **The short version** > > Point people at **operantconditioning.com**. Nothing requires a login, every page prints cleanly, every diagram is CC BY 4.0, and the quiz gives instant explanations rather than a stored score. If you only assign one thing, assign the front page. ## A route through the guide The front page is written to be read straight through, but this order works well when it is split over two sessions. Session one runs about twenty-five minutes; the whole front page, read end to end with everything on it, is closer to forty. **Session one — about 25 minutes** 1. [Watch the four-minute video](https://operantconditioning.com/#watch) 2. [Read the one-paragraph version](https://operantconditioning.com/#operant-conditioning-in-one-paragraph) and [the A-B-C](https://operantconditioning.com/#how-operant-conditioning-works) 3. [Learn the four quadrants](https://operantconditioning.com/#the-four-quadrants-reinforcement-and-punishment), then work through the interactive checker 4. [Shape a lever press in the virtual Skinner box](https://operantconditioning.com/#lab) 5. [Read the schedules table](https://operantconditioning.com/#schedules-of-reinforcement) and [the three short sections after it](https://operantconditioning.com/#extinction-shaping-and-stimulus-control) 6. [Work through the eight-scenario check](https://operantconditioning.com/#can-you-spot-it) Session two: [the evidence on reinforcement versus punishment](https://operantconditioning.com/#reinforcement-versus-punishment-what-the-evidence-says), [the criticisms and limitations](https://operantconditioning.com/#criticisms-and-limitations), and the [full twenty-question quiz](https://operantconditioning.com/quiz/). **What the guide covers** - Define operant conditioning and distinguish it from classical conditioning. - Identify the antecedent, behavior, and consequence in any everyday example. - Classify a consequence into one of the four quadrants — and explain why negative reinforcement is not punishment. - Describe the five schedules of reinforcement and predict the response pattern each produces. - Explain shaping, extinction, the extinction burst, and spontaneous recovery. - State what the evidence shows about reinforcement versus punishment, and where the theory's limits lie. ## Assignable resources - [20-question quiz](https://operantconditioning.com/quiz/): Instant explanations after every answer. No account, no stored score, and a printable answer key. - [128-term glossary](https://operantconditioning.com/glossary/): Every term defined the way behavior analysts define it, with a live search box. - [50+ worked examples](https://operantconditioning.com/examples/): Sorted by quadrant and by setting, each with the reasoning spelled out. - [Diagrams to reuse](https://operantconditioning.com/diagrams/): Every diagram as SVG and PNG, released CC BY 4.0 for slides and handouts. - [The virtual Skinner box](https://operantconditioning.com/#lab): Magazine-train, shape a lever press, thin the schedule, then run extinction. Works on a phone. - [Schedule simulator](https://operantconditioning.com/schedules-of-reinforcement/): Watch a live cumulative record form under FR, VR, FI, and VI. The scallop is visible in about thirty seconds. ## Discussion questions Six that reliably produce disagreement rather than recitation. 1. A parent says, "I punished him by taking his phone, but he keeps doing it." Using the functional definition of punishment, what would you tell them — and what would you look at next? 2. Find three contingencies in your own day. For each, name the antecedent, the behavior, the consequence, and the quadrant. Which one surprised you? 3. Jim in *The Office* calls his Altoid experiment "Pavlov." Make the case that it is operant. Make the case that it is both. 4. Slot machines and social feeds run on variable-ratio schedules. Is there a line between designing for engagement and exploiting a learning mechanism? Where would you draw it? 5. Reinforce every time, or only sometimes? Explain to a new dog owner why the answer changes over the first month. 6. Chomsky said reinforcement cannot explain grammar. What can it explain about language — and what can it not? ## How to cite this site Citing the front page, updated 10 September 2026: APA 7 — Martinson, R. (2026, September 10). *Operant conditioning: The complete guide*. Operant Conditioning. https://operantconditioning.com/ MLA 9 — Martinson, Ryan. "Operant Conditioning: The Complete Guide." *Operant Conditioning*, Operant Conditioning Inc., 10 Sept. 2026, operantconditioning.com/. Chicago — Martinson, Ryan. "Operant Conditioning: The Complete Guide." Operant Conditioning. September 10, 2026. https://operantconditioning.com/. For any other page, substitute that page's own title and URL and use the "Updated" date printed under its heading. Where a page cites a primary source for a claim, cite the primary source rather than this site — every reference list here gives the original paper or book. ## What you may reuse - **Quoting.** Quote any passage with attribution and a link back to the page it came from. - **Classroom copies.** Any page may be reproduced for non-commercial teaching, with attribution. Printing is supported — the stylesheet strips navigation and interactive controls. - **Diagrams.** Every diagram on [the diagrams page](https://operantconditioning.com/diagrams/) is released under CC BY 4.0 as SVG and PNG. Use them in slides, handouts, and papers; credit *operantconditioning.com*. - **Linking.** Link anywhere, including directly to a section anchor. Nothing here is gated and no link will rot behind a login. - **Full terms.** [The copyright terms](https://operantconditioning.com/terms/#copyright) spell out the rest. ## Corrections and requests If something is wrong, imprecise, or out of date, please say so — corrections are made promptly and the page's "updated" date changes with them. Requests for a diagram, a worked example, or a page that does not exist yet are welcome too. [Contact details are here](https://operantconditioning.com/about/#contact). The site is also seeking a credentialed reviewer — a BCBA-D or a PhD in behavior analysis or learning — to read it critically. [More on that](https://operantconditioning.com/about/#expert-review). --- # Virtual Skinner Box: Free Operant Conditioning Simulator > Run a virtual Skinner box in your browser: magazine-train a rat, shape a lever press, put it on a schedule, then watch extinction. Free, no account. - Source: https://operantconditioning.com/lab/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Interactive · Simulation* A hungry, untrained rat, a lever, a food tray, and you holding the pellet button. Magazine-train it, shape a lever press one approximation at a time, put the press on a schedule of reinforcement, then take the food away and watch extinction. About three minutes, no account, works on a phone. **In brief** - The lab is a stylized model of an [operant conditioning chamber](https://operantconditioning.com/skinner-box/): the rat's tendencies shift with what you reinforce, drift back when you don't, and follow the schedule patterns Ferster and Skinner documented. It is not a replay of real data. - You will do the four things every operant experiment does: establish the reinforcer (magazine training), [shape](https://operantconditioning.com/shaping/) a new response, thin reinforcement onto a [schedule](https://operantconditioning.com/schedules-of-reinforcement/), and run [extinction](https://operantconditioning.com/extinction/). - The counters under the chamber are a cumulative record in miniature: every pellet you deliver and every press the rat makes, from the first step on. *(The interactive simulator runs at https://operantconditioning.com/lab/ and needs a browser with JavaScript.)* ## How to run the lab The lab walks you through the same sequence a first-year student runs with a live rat, compressed from a week into a few minutes. Each step is a real procedure with a real name. 1. **Magazine training.** Before the rat can learn that pressing produces food, it has to learn that the click of the food magazine *means* food. Deliver a few free pellets. The click becomes a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer), which is what lets you reinforce a movement the instant it happens instead of waiting for the rat to find the tray.[1] 2. **Shaping.** Now reinforce successive approximations: facing the lever, approaching it, rising toward it, touching it, pressing it. Each time the rat reliably produces the current step, stop paying for that step and pay only for the next one. Reinforce too early and you get a rat that stands still; too late and the behavior you were building drifts away. This is [shaping by successive approximations](https://operantconditioning.com/shaping/), and it is the single most useful skill on this site.[2] 3. **Put the press on a schedule.** Once the rat presses on its own, move from continuous reinforcement to an intermittent schedule — a fixed ratio, a variable ratio, a fixed interval, a variable interval. Watch the rate and the pattern change: the pause after each reinforcer on a fixed ratio, the steady grind of a variable ratio, the scallop of a fixed interval.[3] 4. **Run extinction.** Stop delivering pellets and watch. Rate usually rises first — the [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) — then falls, and if you come back later it partly returns (spontaneous recovery). A rat trained on a variable ratio takes far longer to give up than one trained on continuous reinforcement, which is the partial reinforcement extinction effect.[[3]](#ref-3)[[4]](#ref-4) ## What to watch for The point of a Skinner box was never the box; it was the **rate of response** as a dependent variable, and the way that rate changes moment to moment under different contingencies. A few things worth noticing while you run it: - **Timing.** A pellet delivered one second late reinforces whatever the rat was doing one second later. If your rat starts doing something odd, look at what you reinforced, not at the rat.[2] - **Criterion drift.** In shaping, the criterion has to move. If you keep paying for "approach the lever," you get a rat that approaches the lever forever. - **Rate versus pattern.** Two schedules can produce similar average rates with very different patterns. Ratio schedules produce pausing after reinforcement and fast runs; interval schedules produce steadier, lower rates. The counters show rate; the animation shows the pattern.[3] - **Extinction is not forgetting.** When the presses stop, the learning has not been erased. Come back and the response recovers on its own for a while, and it returns almost instantly if a single pellet is delivered.[4] Where the real thing differs Real rats take days, not minutes: magazine training is typically a session on its own, and shaping a clean lever press can take an hour of careful work by an experienced trainer. Real rats also bring their own repertoire — sniffing, grooming, rearing — that a model rat does not, and that repertoire is exactly what makes shaping an art. Skinner's own account of discovering shaping, with a pigeon and a wooden ball, is a good reminder that the method was found by doing it.[[2]](#ref-2)[[5]](#ref-5) ## For teachers and students The lab is free to assign and link. It runs in any modern browser, on phones, without an account, and nothing a student does is tracked. Suggested uses: - **Before the lecture on shaping.** Ask students to shape a lever press and write down each criterion they used. Compare lists: the number of steps, where people got stuck, and what happened when a criterion was raised too fast. - **Schedules lab.** Have each student run the same rat on two schedules and describe the difference in pattern in their own words before they learn the terms. Then match their descriptions to the four [idealized cumulative records](https://operantconditioning.com/schedules-of-reinforcement/). - **Extinction and persistence.** Train one rat on continuous reinforcement and one on a variable ratio, extinguish both, and count presses to extinction. This is the partial reinforcement extinction effect, done by hand. Citation formats and learning objectives for the whole site are on the [teachers page](https://operantconditioning.com/for-teachers/). If you want the apparatus itself explained part by part — the lever, the magazine, the cumulative recorder — read [The Skinner box](https://operantconditioning.com/skinner-box/). ## Key takeaways - Magazine training makes the feeder click a conditioned reinforcer, which is what makes precise shaping possible. - Shaping is differential reinforcement of successive approximations with a moving criterion: pay for the current step until it is reliable, then only for the next one. - Schedules change the pattern of behavior, not just its amount: ratio schedules produce pauses and runs, interval schedules produce steadier rates and scallops. - Extinction produces a burst, then decline, then partial spontaneous recovery, and behavior trained on intermittent reinforcement is far more persistent. - The model is stylized. Real rats are slower, more varied, and more surprising, which is why shaping them is a skill. ### Check yourself **Your rat rears up near the lever and you deliver a pellet a full second later, just as it drops back down and turns away. What did you reinforce?** Whatever was happening when the pellet arrived: dropping down and turning away. Reinforcement acts on the behavior it follows immediately, not on the behavior you intended. Shorten your reaction time, or use the magazine click as a marker. **After shaping, you switch the rat to a fixed-ratio 10 and notice it stops for a moment after every pellet before pressing quickly again. Is something wrong?** No. That is the post-reinforcement pause characteristic of fixed-ratio schedules, followed by a high-rate run. It is the signature of the schedule, not a fault in training. ## Frequently asked questions **Is the virtual Skinner box based on real data?** No. It is a stylized model. The rat's tendencies shift with what you reinforce, drift when you stop, and follow the schedule patterns described by Ferster and Skinner, but it is not a replay of recorded sessions. Real shaping takes longer and real rats vary more. **Is it free, and can I assign it to a class?** Yes. There is no account, no paywall, and no tracking. Link to https://operantconditioning.com/lab/ directly; citation formats are on the teachers page. **What is a Skinner box?** An operant conditioning chamber: a small enclosure with a lever or key, a device that delivers reinforcers such as food, and a recorder that counts responses. Skinner built it to measure how consequences change the rate of a freely emitted behavior. The full explanation is on the Skinner box page. **What is magazine training?** The first step of any operant experiment: delivering free reinforcers so the animal learns that the sound of the food magazine predicts food. The click becomes a conditioned reinforcer that can then mark the exact moment a response occurs. **Why does the rat press faster when the food stops?** That is the extinction burst: a temporary increase in the rate, force, or variability of a behavior when its reinforcer is first withheld. It is followed by a decline, and later by some spontaneous recovery. **Does the lab work on a phone?** Yes. It is a small JavaScript simulation that runs in any modern browser, including iOS Safari and Android Chrome. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Peterson, G. B. (2004). A day of great illumination: B. F. Skinner's discovery of shaping. *Journal of the Experimental Analysis of Behavior, 82*(3), 317–328. 3. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 4. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 5. Skinner, B. F. (1956). A case history in scientific method. *American Psychologist, 11*(5), 221–233. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [The Skinner box](https://operantconditioning.com/skinner-box/): The apparatus, part by part, and how to read a cumulative record. - [Shaping](https://operantconditioning.com/shaping/): Successive approximations, step by step, with a protocol. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Ratio and interval, fixed and variable, with a simulator. # Variable Ratio Schedule: Definition, Examples, and Evidence > A variable ratio schedule reinforces after an unpredictable number of responses. VR notation, random ratio vs. VR, slot machines, thinning, and ratio strain. - Source: https://operantconditioning.com/variable-ratio-schedule/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Schedules · Ratio* A slot machine, a fishing rod, and a well-trained dog have one thing in common: the payoff comes after an unpredictable number of tries. That arrangement is the variable ratio schedule, and it produces the fastest, steadiest, and most stubborn behavior in operant conditioning. This page explains how it is built, what it does to behavior, and where the evidence stops. > **Definition** > > A variable ratio schedule (VR) is a schedule of reinforcement in which a reinforcer follows an unpredictable number of responses that varies around an average. On a VR 10 schedule a reinforcer might come after 4 responses, then 17, then 9, then 10; over many reinforcers the count averages 10. The requirement depends only on responses, never on time, and nothing in the situation tells the organism which response will be the one that pays.[1] > > Variable ratio schedules produce the highest and steadiest rates of any basic schedule and the greatest persistence once reinforcement stops. Ferster and Skinner mapped them in *Schedules of Reinforcement* (1957), and Skinner had already identified them as the schedule behind gambling.[1][2] **In brief** - A variable ratio schedule reinforces after a number of responses that varies unpredictably around an average; VR 10 means ten responses per reinforcer on average, and VR 1 is continuous reinforcement. - Because the very next response could be the one that pays, VR schedules produce a high, steady rate with little or no [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause), and behavior built on them is the hardest to [extinguish](https://operantconditioning.com/extinction/). - The way to reach a VR schedule is to reinforce every response first and then thin gradually and irregularly; thin too fast and you get [ratio strain](https://operantconditioning.com/glossary/#ratio-strain) instead of persistence. ## How a variable ratio schedule works Every [schedule of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) answers one question: which responses earn the reinforcer? A ratio schedule answers it with a count. On a [fixed ratio](https://operantconditioning.com/fixed-ratio-schedule/) the count is always the same. On a variable ratio the count changes from one reinforcer to the next, and only its average is fixed. The notation gives that average: VR 5, VR 20, VR 100. A VR 1 reinforces every response and is [continuous reinforcement](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) under another name.[1] Two features follow. First, the reinforcer depends on responding and on nothing else: waiting accomplishes nothing, and responding faster brings the next reinforcer sooner. Second, no signal marks the response that will pay. After a reinforcer on a fixed ratio of 20, the organism has nineteen unpaid responses ahead of it and pauses before starting the next run. After a reinforcer on a VR 20, the next reinforcer might be one response away.[1][3] Procedure and process "VR 10" names a procedure: the rule by which reinforcers are arranged. The high, steady rate is a process: what the rule does to behavior once the consequence is actually reinforcing. You can arrange a flawless VR 10 and get nothing if the consequence is not a reinforcer for that organism at that moment. ## Two ways to build one: arranged variable ratio and random ratio A variable ratio is built in one of two ways, and the difference matters. The first is the method Ferster and Skinner used. The experimenter writes a list of ratios whose mean is the value wanted — for a VR 10, something like 1, 3, 5, 8, 10, 12, 14, 17, and 20 — and the apparatus works through it in an irregular order. This is an **arranged variable ratio**. It has a smallest ratio and a largest one, so the longest unpaid run the subject will ever meet is known in advance.[1] The second is a **random ratio** (RR). Each response is reinforced with a fixed probability, independently of every other response. On an RR 10, every response has a one-in-ten chance of paying, whatever happened on the previous nine or ninety. The mean is still 10, but the ratios are unbounded: runs of fifty or a hundred unpaid responses are rare but certain to occur eventually. A random ratio has no memory: a long losing run does not make the next response any more likely to pay, which is what the gambler who feels a win is "due" gets wrong. Modern slot machines are random-ratio devices: each spin is an independent draw at a fixed probability set by the program. Haw has argued that calling them variable-ratio schedules blurs a real distinction, because the unbounded losing runs and occasional early wins of a random ratio are part of what the player experiences.[4] The two produce similar behavior, and this page uses "variable ratio" for both except where the difference matters. ## The behavior it produces ### A high, steady rate The signature of a VR schedule is a high rate of responding that runs on without breaks. The [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause) that punctuates fixed-ratio performance is short or absent. Schlinger, Derenne, and Baron's review of fifty years of research on pausing explains why: the pause on ratio schedules is better understood as a pre-ratio pause, controlled by the size of the ratio ahead rather than by the reinforcer just received. On a fixed ratio the ratio ahead is always the full count, and the pause grows with it. On a variable ratio the ratio ahead might be one, so the pause shrinks, lengthening again only when the average ratio becomes large.[5] Ratio schedules also produce higher rates than interval schedules at the same rate of reinforcement. Zeiler's review of the controlling variables and Baum's comparison of pigeons on ratio and interval schedules both trace this to the **feedback function**, the relation between response rate and reinforcement rate. On a ratio schedule it is a straight line: double the rate and you double the reinforcers. On an interval schedule it flattens; past a modest rate, responding faster earns almost nothing extra.[6][7] ### The greatest resistance to extinction When reinforcement stops, behavior that was reinforced intermittently persists far longer than behavior that was reinforced every time. This is the [partial reinforcement extinction effect](https://operantconditioning.com/glossary/#partial-reinforcement-extinction-effect), and among the basic schedules the variable ratio produces the most of it.[1][3] The effect was first shown in classical conditioning: Humphreys found in 1939 that eyeblink responses conditioned with the air puff on only half the trials extinguished more slowly than responses conditioned with the puff on every trial.[8] Mowrer and Jones then trained rats to press a lever for food on ratios of one to four and on an irregular pattern; the more presses had been required per reinforcer in training, the more presses the rats made in extinction.[9] The usual explanation is discrimination. On continuous reinforcement, the first unreinforced response is unlike anything in the organism's history. On a VR 50, a run of eighty unreinforced responses is within normal experience, so nothing has visibly changed, and behavior continues until the run grows long enough to be discriminated from training.[3] ### Ratio strain The high rate has a limit. If the ratio is stretched too far, or too fast, responding breaks down: long pauses, bursts separated by inactivity, and eventually extinction even though reinforcement is still available. Ferster and Skinner saw this in pigeons when ratios were pushed too high, and trainers see it whenever reinforcement is thinned faster than the behavior can bear.[1][3] [Ratio strain](https://operantconditioning.com/glossary/#ratio-strain) is the schedule failing to maintain the behavior, not the organism being stubborn, and the remedy is a richer schedule and a slower stretch. ## What Ferster and Skinner's cumulative records show The evidence for all of this is a stack of cumulative records. Skinner's cumulative recorder, described in *The Behavior of Organisms*, made rate visible: a pen steps up once for every response as the paper rolls, so a steep line means a high rate, a flat line means no responding, and a small diagonal tick marks each reinforcer.[10] Ferster and Skinner's 1957 book is hundreds of such records, mostly from pigeons pecking a lighted key.[1] A fixed-ratio record is a staircase: a flat pause after each reinforcer, then a steep run to the next. A fixed-interval record shows the scallop, a curve that accelerates toward each reinforcer. A variable-interval record is a straight line at a moderate slope. A variable-ratio record is a straight line at a steep slope, with the reinforcer ticks scattered irregularly along it and almost no visible reaction to any of them. When the mean ratio was pushed high enough, the record broke into runs separated by flat stretches, which is what ratio strain looks like on paper.[1] The site's [virtual Skinner box](https://operantconditioning.com/lab/) produces a stylized VR record, and the [Skinner box page](https://operantconditioning.com/skinner-box/) explains how to read one. ## Variable ratio vs. fixed ratio, variable interval, and fixed interval The [schedules hub](https://operantconditioning.com/schedules-of-reinforcement/) gives the two-question rule for telling the basic schedules apart: count or time, fixed or variable? Here the point is narrower — what each of the other three does that a variable ratio does not. | Schedule | Reinforcer follows | Pattern on the record | Resistance to extinction | How it differs from VR | | --- | --- | --- | --- | --- | | **Variable ratio (VR)** | An unpredictable number of responses, averaging *n* | Steep, straight line; little or no pause | Highest | — | | **Fixed ratio (FR)** | Every *n*th response | Staircase: pause, then run; pause grows with *n* | High | The count ahead is known, so the organism pauses before each ratio | | **Variable interval (VI)** | First response after an unpredictable time, averaging *t* | Straight line, moderate slope | High | Responding faster does not bring the reinforcer sooner, so the rate is moderate | | **Fixed interval (FI)** | First response after a fixed time | Scallop: pause, then acceleration | Moderate | Time, not effort, sets up the reinforcer; early responses are wasted | Two contrasts matter most. Against the fixed ratio, the variable ratio trades predictability for steadiness: the same reinforcers per hundred responses, delivered on an irregular count, removes the pauses and produces a smoother and usually higher rate.[1][5] Against the [variable interval](https://operantconditioning.com/variable-interval-schedule/), the variable ratio rewards speed: on VI, a pigeon that pecks twice as fast earns barely any more food and settles into a moderate rate, while on VR twice the pecks means twice the food.[7] The [fixed interval](https://operantconditioning.com/fixed-interval-schedule/) is the schedule least like VR: the clock sets up the reinforcer, and the organism does little until the time is nearly up. ## Where variable ratio schedules show up In each row below, a response is followed by a reinforcer after an unpredictable count, and more responses mean more reinforcers. The schedule column is honest about how close the fit is. | Setting | Response | Reinforcer | Schedule | Note | | --- | --- | --- | --- | --- | | Casino | A spin of a slot machine | A payout | Random ratio, set by the program | Each spin independent; a near miss is a loss | | Fishing | A cast | A fish on the line | Approximately VR | More casts, more fish, on an unpredictable count | | Sales | A cold call | An appointment or a sale | Approximately VR | "A numbers game" describes a ratio schedule | | Dog training | Sit on cue | A treat | VR 3 to VR 5 after thinning | The standard way to maintain a trained behavior | | Phone | Pulling to refresh, scrolling | A message, a like, a good post | Resembles a variable schedule | Ratio and interval features mixed; the contingency is not public | ### The slot machine Skinner made the connection himself. In *Science and Human Behavior* he observed that the effectiveness of variable-ratio schedules in generating high rates had long been known to the proprietors of gambling establishments, and he described the compulsive gambler as the result of such a schedule.[2] Schüll's ethnography of machine gambling in Las Vegas shows the industry's side: designers working explicitly to extend "time on device," with fast play, small frequent payouts, and features tuned to keep players seated.[11] ### The near miss One feature of slot machines is not a schedule effect at all. A **near miss** is a loss that looks almost like a win — two matching symbols on the payline and the third just above or below it. Reid argued in 1986 that near misses encourage continued play even though they pay nothing, and that a game can be built to produce them.[12] In one laboratory test, people playing a simulated slot machine whose losing spins were near misses about 30 percent of the time kept playing longer after the machine stopped paying than people who saw near misses more rarely or more often.[13] On a pure random ratio every loss is the same event. The near miss is a stimulus added on top of the schedule, a reminder that the schedule is only one of the variables at work. ### Feeds and notifications Refreshing a social feed is now the textbook example after the slot machine, and it deserves more care. The resemblance is real: each pull or scroll is a response, most produce nothing, and the count between payoffs is unpredictable. But new posts arrive with time whether or not you refresh, which is an interval feature, and no one outside the companies knows the actual contingency. The careful statement is that feed-checking *resembles* behavior on a variable schedule; the studies behind this page used pigeons, rats, and slot machines, not feeds. ## How to use a variable ratio schedule A variable ratio is a maintenance schedule, not a teaching schedule: a learner who has never been reinforced for a behavior will not persist through nine unpaid attempts to reach the tenth. The sequence, in [dog training](https://operantconditioning.com/dog-training/), classrooms, and self-management alike, is continuous reinforcement first and then a gradual stretch.[3] 1. **Establish the behavior on continuous reinforcement.** Reinforce every correct response until it is fluent — a dog that sits on cue nine times in ten, a child who starts homework when asked. If you are still shaping the behavior, you are not ready to thin. 2. **Stretch to a small, variable ratio.** Reinforce about two responses in three, then one in two, irregularly rather than in a pattern. Write the sequence down in advance or draw it from a hat; a "variable" schedule that always pays on the third response is a fixed ratio with extra steps. 3. **Keep some short ratios in the mix.** Include ratios of one and two, so that a reinforcer is sometimes followed almost immediately by another. If a reinforcer reliably means the next one is far off, you have rebuilt the fixed-ratio pause. 4. **Stretch slowly, and watch the behavior rather than the calendar.** Raise the mean only when responding at the current ratio is steady. Pauses, refusals, and a drop in quality are the early signs of ratio strain; when they appear, drop back to the last ratio that worked and hold it longer.[3] 5. **Keep the reinforcer worth working for.** Thinning changes how often the reinforcer arrives, not how much it has to matter. A full dog or a token that buys nothing will defeat any schedule. 6. **Hand the behavior off to natural consequences.** A contrived VR keeps the behavior alive until the world's own reinforcers take over — the recall that ends in a game, the homework habit that pays off at school. The [habits page](https://operantconditioning.com/habits/) covers the same sequence for your own behavior. ## Common mistakes - **Thinning too fast.** The commonest error. Jumping from continuous reinforcement to "every fifth time or so" produces ratio strain, and the trainer concludes that intermittent reinforcement does not work. It does; the stretch was too steep. - **Starting on a variable ratio.** A behavior that has never been reinforced cannot be maintained by a schedule that pays one attempt in ten. Continuous reinforcement comes first. - **Being predictable.** Every third response, or a strict alternation, is a fixed schedule, and the learner will find the pattern. Randomize. - **Calling an interval schedule a ratio schedule.** If the reinforcer becomes available with time and responding faster does not bring it sooner, it is an interval schedule. Waiting for a reply is variable interval; the slot machine and the cold call are variable ratio. - **Arranging one by accident.** Giving in "just this once," now and then, is the most effective way known to make an unwanted behavior permanent. Before trying to extinguish anything, ask what schedule it is on and who is delivering the reinforcer, and expect an [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) when the schedule ends. ## What the evidence does not show **It does not show that variable ratio schedules are addictive.** A schedule is a description of when reinforcers arrive. It explains why a behavior persists through unpaid runs; it does not, by itself, make a consequence reinforcing or a behavior harmful. Most people who fish, sell, or play an occasional slot machine develop no problem, and Schüll's account makes clear that the schedule is embedded in design choices — speed of play, credits in place of coins, near misses, the ergonomics of the seat — and in the circumstances players bring with them, none of which is a schedule.[11] "Variable ratio" is one variable in that analysis, not a diagnosis. **It does not show that the partial reinforcement extinction effect is simple.** Nevin argued in 1988 that the effect, though reliably found when groups trained on different schedules are compared, sits awkwardly beside another well-established finding: within a single organism, behavior maintained by more frequent reinforcement is more resistant to disruption. How the effect comes out depends on how extinction is measured and on how easily the change from training to extinction can be discriminated.[14] Mowrer and Jones suggested something similar: the effective unit after ratio training may be the whole run ending in food, so counting single responses in extinction overstates the persistence.[9] **It does not show that adults behave like pigeons.** Lowe found that adults on simple schedules often produce patterns unlike the animal records, and argued that the reason is verbal: people describe the schedule to themselves and then follow the description, which may be wrong. His evidence came mostly from fixed-interval schedules, but the implication for ratio schedules is direct.[15] A person who has decided that a machine is "due" is responding to a rule, not to the random ratio in front of them, and a person who has correctly concluded that the odds are fixed can stop in a way no pigeon can. ## Key takeaways - A variable ratio schedule reinforces after a number of responses that varies unpredictably around a mean; VR 10 averages ten responses per reinforcer, and VR 1 is continuous reinforcement. - Because the ratio ahead might be one, a variable ratio produces a high, steady rate with little or no post-reinforcement pause, and higher rates than interval schedules at the same rate of reinforcement. - Behavior on a variable ratio is the most resistant to extinction of any basic schedule: a long unpaid run looks like ordinary experience, so nothing signals that the contingency has changed. - An arranged variable ratio draws from a bounded list; a random ratio pays each response with a fixed probability and has no memory. Slot machines are random-ratio devices, and a win is never "due." - Reinforce every response first, then thin gradually and irregularly, keeping some short ratios in the mix; thinning too fast produces ratio strain rather than persistence. - The schedule is a description, not a diagnosis: it does not make a consequence reinforcing or gambling harmful by itself, and adult humans layer rules over it that pigeons do not. ### Check yourself **A trainer moves a dog straight from a treat for every sit to a treat for roughly every fifth sit. Within one session the dog starts wandering off between cues. What happened, and what should the trainer do?** Ratio strain. The schedule was thinned faster than the behavior could bear, so responding broke down: pausing, wandering, and, if it continues, extinction. The trainer should drop back to a richer schedule — every sit, or every other sit — hold it until sitting is steady again, then stretch in smaller steps, keeping some ratios of one and two in the mix. **A friend says that checking for text replies is a variable ratio schedule, "like a slot machine." Is it?** No. A reply becomes available with the passage of time, not with the number of checks, and checking faster does not make it arrive sooner; only the first check after it arrives is reinforced. That is a variable interval schedule, which produces a moderate, steady rate. The slot machine is a ratio schedule because every spin is a chance to win and more spins mean more wins. **After forty losing spins someone says the machine "must be about to pay." On a random ratio schedule, is the next spin more likely to win?** No. On a random ratio each response is reinforced with the same fixed probability, independently of everything before it, so forty losses change nothing about the forty-first spin. The belief that a win is due is a rule the person has formed, not a property of the schedule, and it is a good example of how human persistence on ratio schedules can be governed by self-talk that misdescribes the contingency. ## Frequently asked questions **What is a variable ratio schedule in simple terms?** A variable ratio schedule pays off after a number of responses that changes unpredictably, averaging some value. On a VR 10, a reinforcer might come after 3 responses, then 15, then 12, averaging 10 over time. Because the next response could always be the one that pays, the behavior stays fast and steady, and it keeps going for a long time after the payoffs stop. **What is an example of a variable ratio schedule?** A slot machine: each spin is a response, most spins pay nothing, and a payout arrives after an unpredictable number of them. Other cases are casting a fishing line, making sales calls, submitting job applications, and a trained dog that gets a treat for an unpredictable one sit in four. In every case, more responses mean more reinforcers, but the count between reinforcers varies. **What does VR 10 mean?** VR stands for variable ratio, and the number is the average ratio of responses to reinforcers. VR 10 means that, averaged over many reinforcers, ten responses are required for each one; the actual counts vary, perhaps from one to twenty. VR 1 is continuous reinforcement. A random ratio 10 (RR 10) is the related schedule in which every response has a one-in-ten chance of being reinforced. **What is the difference between a variable ratio and a variable interval schedule?** Both are unpredictable, but a variable ratio depends on how many responses you make, while a variable interval depends on how much time has passed. On VR, responding faster earns reinforcers faster, so the rate is high. On VI, only the first response after the interval counts, so responding faster gains nothing and the rate is moderate. A slot machine is VR; checking for a text reply is VI. **What is the difference between a variable ratio and a fixed ratio schedule?** A fixed ratio reinforces after the same number of responses every time; a variable ratio reinforces after a number that varies around an average. Fixed ratios produce a pause after each reinforcer followed by a fast run, because the organism knows the full count lies ahead. Variable ratios produce a steady rate with little pausing, because the next reinforcer might be one response away. **Why is the variable ratio schedule the most resistant to extinction?** Because a long run without reinforcement looks like ordinary experience. On continuous reinforcement, the first unreinforced response is a clear signal that something has changed. On a lean variable ratio, dozens of unreinforced responses are normal, so the organism has no way to tell that reinforcement has stopped and keeps responding. This is the partial reinforcement extinction effect, and variable ratio schedules produce the most of it. **Is a slot machine a variable ratio or a random ratio schedule?** Strictly a random ratio. Each spin is an independent draw with a fixed probability of paying, so the number of spins between wins is unbounded and a win is never due. A laboratory variable ratio draws from a fixed list of ratios. The behavior the two produce is similar, a high and steady rate, but the random ratio's long losing runs and occasional back-to-back wins are part of what machine gambling is. **Is social media a variable ratio schedule?** It resembles one. Pulling to refresh or scrolling is a response, most produce nothing, and something interesting turns up after an unpredictable number. But new posts and messages arrive with time regardless of your checking, which is an interval feature, and the actual contingency inside any app is not public. The careful statement is that feed-checking resembles behavior on a variable schedule, not that it has been shown to be one. ## References 1. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 2. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 3. Mazur, J. E. (2017). *Learning and Behavior* (8th ed.). Routledge. 4. Haw, J. (2008). Random-ratio schedules of reinforcement: The role of early wins and unreinforced trials. *Journal of Gambling Issues, 21*, 56–67. 5. Schlinger, H. D., Derenne, A., & Baron, A. (2008). What 50 years of research tell us about pausing under ratio schedules of reinforcement. *The Behavior Analyst, 31*(1), 39–60. 6. Zeiler, M. D. (1977). Schedules of reinforcement: The controlling variables. In W. K. Honig & J. E. R. Staddon (Eds.), *Handbook of Operant Behavior* (pp. 201–232). Prentice-Hall. 7. Baum, W. M. (1993). Performances on ratio and interval schedules of reinforcement: Data and theory. *Journal of the Experimental Analysis of Behavior, 59*(2), 245–264. 8. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158. 9. Mowrer, O. H., & Jones, H. (1945). Habit strength as a function of the pattern of reinforcement. *Journal of Experimental Psychology, 35*(4), 293–311. 10. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 11. Schüll, N. D. (2012). *Addiction by Design: Machine Gambling in Las Vegas*. Princeton University Press. 12. Reid, R. L. (1986). The psychology of the near miss. *Journal of Gambling Behavior, 2*(1), 32–39. 13. Kassinove, J. I., & Schare, M. L. (2001). Effects of the "near miss" and the "big win" on persistence at slot machine gambling. *Psychology of Addictive Behaviors, 15*(2), 155–158. 14. Nevin, J. A. (1988). Behavioral momentum and the partial reinforcement effect. *Psychological Bulletin, 103*(1), 44–56. 15. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): All five basic schedules, the two-question rule, and the interactive simulator. - [Fixed ratio schedule](https://operantconditioning.com/fixed-ratio-schedule/): The pause-and-run pattern, piece-rate pay, and why the pause grows with the ratio. - [Extinction](https://operantconditioning.com/extinction/): What happens when reinforcement stops: bursts, recovery, and resurgence. # Fixed Ratio Schedule: Definition, Examples & Break-and-Run > A fixed ratio schedule reinforces every nth response. Definition, FR notation, the break-and-run pattern, ratio strain, piece-rate pay, and how to use it. - Source: https://operantconditioning.com/fixed-ratio-schedule/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Schedules · Ratio* Every tenth coffee free, pay per windshield installed, a sticker after five books: whenever a reinforcer arrives after a set number of responses, behavior settles into the pattern that Ferster and Skinner's pigeons drew on paper seventy years ago — a pause, then a run. Knowing that pattern is the difference between a ratio that builds a behavior and one that breaks it. > **Definition** > > A fixed ratio schedule (FR) is a [schedule of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) in which a reinforcer is delivered after a fixed number of responses. The number is the ratio: on FR 20, the twentieth response is reinforced and the nineteen before it are not, and the count then starts again from zero.[1] > > FR 1, on which every response is reinforced, is the same thing as continuous reinforcement. Every ratio above 1 is an intermittent schedule, and the pattern it produces — a pause after each reinforcer, then a fast, steady run to the next — is called break-and-run.[1] **In brief** - On a fixed ratio schedule a reinforcer follows every *n*th response. FR 1 is continuous reinforcement, FR 20 pays every twentieth response, and because only the count matters, the organism's own rate of responding sets the rate of reinforcement. - The signature pattern is break-and-run: a pause after each reinforcer, then an abrupt switch to a high, steady run that lasts until the next one. The pause lengthens as the ratio grows; the running rate hardly changes. - Raise the ratio too fast and you get ratio strain — long pauses, broken runs, and eventually extinction. Start at FR 1, stretch in small steps, and back off at the first sign of strain. ## How a fixed ratio schedule works Every schedule answers one question: which responses will be followed by a reinforcer? On a fixed ratio schedule the answer is a count. Each response advances a counter; the response that completes the count is reinforced; the counter resets and the cycle begins again. The requirement is a number of responses, not an amount of time, and the number is the same on every cycle.[1] Skinner studied fixed-ratio reinforcement in rats in the 1930s and compared it, in *The Behavior of Organisms*, with periodic reconditioning, his early name for what is now called a [fixed interval](https://operantconditioning.com/fixed-interval-schedule/). The ratio schedule produced a markedly higher rate of lever pressing.[2] The reason is built into the rule. On an interval schedule, pressing faster does not make the reinforcer come sooner; only the first press after the interval counts. On a ratio schedule, pressing faster brings the reinforcer sooner in exact proportion, so the time between reinforcers is set not by the experimenter but by how fast the organism works.[3] Baum's review found that at the same rate of reinforcement, ratio schedules maintain the higher rate of responding.[4] Reading the notation FR 1 reinforces every response and is identical to [continuous reinforcement](https://operantconditioning.com/glossary/#continuous-reinforcement) (CRF). FR 10 reinforces every tenth response. In the laboratory the count is of lever presses or key pecks; outside it, of worksheets completed, units installed, or coffees bought. The responses between reinforcers count only toward the next one. ## Break and run: the post-reinforcement pause Put a pigeon on FR 50 and let the [cumulative recorder](https://operantconditioning.com/skinner-box/) run. The record looks like a staircase. After each reinforcer the bird stops pecking and the pen draws a flat line. Then, with no external signal, it starts again and pecks at a high, nearly constant rate — a steep, straight segment — until the fiftieth peck delivers grain and the pattern repeats. Ferster and Skinner reported hundreds of such records in 1957; the two parts, the [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause) and the run, make up the break-and-run pattern.[1] Set the [virtual Skinner box](https://operantconditioning.com/lab/) to FR 20 to watch it appear. The pause is lawful. Felton and Lyon trained pigeons on a series of fixed ratios and found that the pause grew longer as the ratio grew larger, while the rate of pecking during the run stayed roughly constant.[5] Powell showed that the same relation holds when the ratio is raised in small sequential steps rather than in jumps.[6] Ferster and Skinner had already found that at ratios in the hundreds the pauses could stretch to many minutes, with the runs between them still fast.[1] The bigger the job, the longer the organism waits before starting it; once started, it works at about the same speed regardless. ### Why the pause belongs to the next ratio, not the last reinforcer The name suggests that the reinforcer causes the pause: the bird has eaten and is resting. Schlinger, Derenne, and Baron reviewed fifty years of research on pausing and concluded that this reading is mistaken. When pigeons work on schedules that alternate between a small ratio and a large one, with a signal showing which is coming, the pause depends on the ratio ahead rather than the one just completed: after the same reinforcer, the bird pauses briefly if a small ratio comes next and at length if a large one does. The pause is better described as a pre-ratio pause, and it behaves less like rest than like avoidance — the start of a long stretch of unreinforced work is aversive, and the organism postpones it.[7] Anyone who has stared at a blank page for twenty minutes and then written steadily for an hour knows the shape. ## Ratio strain: what happens when the ratio is raised too fast Ratios cannot be raised at will. Move a pigeon from FR 10 to FR 200 in one step and responding breaks down: the pause stretches, the run is interrupted by further pauses, and responding may stop altogether. Ferster and Skinner described this straining at high ratios,[1] and applied behavior analysts call it [ratio strain](https://operantconditioning.com/glossary/#ratio-strain).[8] From the organism's side, the schedule has become hard to tell apart from [extinction](https://operantconditioning.com/extinction/): a long string of unreinforced responses is exactly what extinction looks like, and behavior weakens accordingly.[9] Whether a ratio is too high depends on more than the number. Ferster and Skinner maintained food-deprived pigeons on very large ratios, but only by arriving at them gradually.[1] The ratio that holds up for a hungry pigeon fails for a sated one, and the ratio a bonus will carry is not the ratio a sticker will carry. Strain is a relation among the requirement, the reinforcer, and the organism's motivation. The signs appear outside the laboratory in the same order: longer delays before starting, work abandoned partway through the count, and finally quitting. A student who stops mid-worksheet after the requirement was doubled and a dog that wanders off from a training session are both ratios stretched faster than the behavior could follow. The remedy is the same in every case: drop back to the last ratio that worked and raise it more slowly.[8] ### What happens when reinforcement stops Behavior maintained on a fixed ratio outlasts behavior maintained on continuous reinforcement once the reinforcer is withdrawn — the [partial reinforcement extinction effect](https://operantconditioning.com/glossary/#partial-reinforcement-extinction-effect), covered on the [continuous versus intermittent](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) page. Ferster and Skinner's extinction records after fixed-ratio training show the schedule's pattern persisting for a while: bursts at the old running rate, separated by pauses that lengthen until responding stops.[1] ## Examples of fixed ratio schedules Any arrangement in which a reinforcer arrives after a set number of countable responses is a fixed ratio. Three with the best evidence are taken up after the table. | Setting | Response counted | Schedule | Reinforcer | What to expect | | --- | --- | --- | --- | --- | | Workplace | Windshield installed | FR 1 (pay per unit) | Wages per unit | Output rises; care is not counted | | Coffee shop | Purchase stamped on a card | FR 10 | Free drink | Faster buying as the tenth stamp nears; a lull after it | | Elementary classroom | Book logged on a reading chart | FR 5 | Sticker or small prize | A flurry before the fifth book; a quiet spell after | | Sales | Sale closed | FR 10 (bonus after ten sales) | Bonus | Sales are FR; the calls behind them are variable ratio | | Dog training | Sit on cue | FR 3 during thinning | Treat | A sound step up from FR 1 on the way to a variable ratio | | Video games | Enemy defeated or item collected | FR 10 ("collect ten") | Quest reward | Steady grinding, then a break after turning it in | ### Piece-rate pay Skinner pointed to piecework as the industrial form of the fixed ratio and warned that it can drive rates to exhausting levels.[10] The best-documented modern case is Safelite Glass, which in the mid-1990s moved its windshield installers from hourly wages to pay per unit installed, with a guaranteed hourly minimum. Lazear analyzed the company's records and reported that output per worker rose by about 44 percent, roughly half from existing installers installing more and the rest from sorting, as the new pay attracted and kept more productive workers.[11] It is one of the cleanest field demonstrations that tying reinforcement to a count of responses raises the rate of those responses; what it does not measure is taken up below. ### Loyalty punch cards "Buy ten, get one free" is FR 10 on purchases. Kivetz, Urminsky, and Zheng tracked café customers using stamp cards and found that the time between purchases shrank as the card filled, and that after redeeming a card and starting a fresh one customers slowed down again, a lull the authors called post-reward resetting.[12] The acceleration toward the goal is not a pigeon's flat run, but the lull after the reward and the rush before it are the same shape. ### Reading logs and sticker charts A classroom reading chart that pays a sticker after every five books is FR 5, and it behaves like one: a flurry of reading as the fifth book approaches, a quiet spell after the sticker. Both are the schedule working, not the child slacking. If the quiet spell matters, lower the ratio, make the reinforcer larger, or move to a [variable ratio](https://operantconditioning.com/variable-ratio-schedule/) so that there is no predictable moment at which the next reinforcer is far away.[7] One caution. The schedule is defined by what is counted, and the counted unit is often not the behavior you care about. A commission per sale is FR 1 on sales, but the salesperson's behavior is making calls, and a sale follows an unpredictable number of calls, so the calling is on a variable ratio. ## Fixed ratio vs. variable ratio, fixed interval, and variable interval All four basic schedules are defined by two questions — count or time, fixed or variable — and the fixed ratio is the "count, fixed" cell. | Feature | Fixed ratio (FR) | Variable ratio (VR) | Fixed interval (FI) | Variable interval (VI) | | --- | --- | --- | --- | --- | | Reinforcer follows | A fixed number of responses | A varying number of responses, averaging *n* | The first response after a fixed time | The first response after a varying time | | Does responding faster bring the reinforcer sooner? | Yes, in proportion | Yes, in proportion | No | No | | Is the next reinforcer predictable? | Yes, by count | No | Yes, by time | No | | Pattern on the cumulative record | Break-and-run: flat, then steep | High, steady, little pausing | Scallop: flat, then curving upward | Moderate, steady | | Pause after the reinforcer | Yes; grows with the ratio | Brief or none | Yes; then gradual acceleration | Little | | Everyday example | Punch card; piece rate | Slot machine; sales calls | Watching the oven timer | Checking email | The contrast with the variable ratio matters most, because both are ratio schedules and both maintain high rates. What the fixed count adds is predictability: the organism, in effect, knows how far away the next reinforcer is, and it is the certainty that it is far away that produces the pause. On a variable ratio the very next response might be the one that pays, and Ferster and Skinner's records show the pauses largely disappearing.[1][7] The contrast with the fixed interval is a matter of shape. Both begin with a pause, but the fixed-interval record curves upward — the scallop — as the rate builds gradually, while the fixed-ratio record breaks abruptly from flat to steep, and on the interval schedule responding faster does not help.[1][9] The [variable interval](https://operantconditioning.com/variable-interval-schedule/) shares neither the count nor the predictability and produces a moderate, steady rate with little pausing. ## How to use a fixed ratio schedule The fixed ratio is the natural first step away from continuous reinforcement, because it is the easiest schedule to count and to explain. The protocol is the one applied behavior analysts use to thin a schedule, whoever the learner is.[8] 1. **Define the unit.** Decide what counts as one response — one completed problem, one sit, one paragraph — and make it unambiguous. The schedule reinforces exactly what it counts. 2. **Start at FR 1.** Reinforce every response until the behavior is fluent. A new behavior put straight onto a lean ratio extinguishes before it is learned. 3. **Raise the ratio in small steps.** FR 1 to FR 2, FR 2 to FR 3, FR 3 to FR 5, moving up only after the behavior is stable at the current step. Even small increases lengthen the pause; large jumps invite strain.[6] 4. **Watch for strain and back off.** Longer delays before starting, runs abandoned partway, quitting. At the first sign, return to the last ratio that worked and hold there longer.[8] 5. **Match the reinforcer to the ratio.** A large ratio needs a reinforcer worth the work, or a conditioned reinforcer — a token, a point, a check mark — that bridges to one. That is what a [token economy](https://operantconditioning.com/token-economy/) is for. 6. **Decide whether you want the pause.** If the goal is steady behavior with no dead time, switch to a variable ratio once the behavior is established. If the pause is harmless, the fixed ratio is simpler and more transparent. 7. **Hand off to natural reinforcers.** The contrived count is scaffolding. Fade it once the behavior is maintained by what it produces on its own, as in [building habits](https://operantconditioning.com/habits/). ## Common mistakes - **Jumping the ratio.** FR 1 to FR 20 in one move is how many home and classroom programs fail. The behavior looks as if it stopped working; the schedule broke it. - **Counting the wrong thing.** A count of pages read reinforces turning pages; a count of problems attempted reinforces attempting. Make the unit the behavior itself, not a proxy. - **Treating the pause as laziness.** The pause is the schedule's signature, governed by the size of the ratio ahead.[7] Scolding a child for it adds punishment to a situation the schedule created. Shrink it by lowering the ratio or switching to a variable ratio. - **Counting quantity when quality matters.** If the count is units, care is not reinforced unless it is also measured. Build quality into the definition of a countable response or arrange a separate consequence for it. - **Calling anything with a number a fixed ratio.** "Read for ten minutes, then play" is time-based, not count-based, and pays the same whether one page or twenty is read. A salary is neither; it is paid on time regardless of the number of responses. - **Expecting the fixed ratio to be the most persistent schedule.** It is more durable than continuous reinforcement and less durable than a variable ratio. If persistence is the goal, it is a step on the way, not the destination. ## What the evidence does not show Break-and-run, the growth of the pause with ratio size, and ratio strain are among the most reliable findings in the experimental analysis of behavior. They were established with pigeons and rats, and three limits apply when the schedule is used with people. **People follow rules as well as schedules.** Lowe showed that adult human performance on simple schedules often departs from the animal patterns: on fixed-interval schedules, adults tend to respond either at a very low rate or at a high, steady one rather than producing the scallop, and the difference tracks the rules they form about the contingency.[13] A customer who reads "every tenth coffee free" is responding to the rule as much as to the count, and the schedule alone does not predict what the rule will do. Field results like the stamp-card study are consistent with the laboratory pattern, but they are not a test of it.[12] **Piece rates raise output, and the evidence is about output.** Lazear's Safelite study measured units installed per worker and reported that they rose. He reports no sign that quality fell, since a broken windshield had to be replaced on the installer's own time, but quality was not the outcome the study measured.[11] The general concern is a matter of contingency, not of any particular finding: a schedule reinforces what it counts, and if the count is units, care is reinforced only insofar as it is also counted or separately consequated. Skinner's related worry was that the schedule can push rates to exhausting levels.[10] Neither point is an argument against piece rates; both are arguments for counting what you want. **The pause is not rest, and the fixed ratio is not the most durable schedule.** The evidence reviewed by Schlinger and colleagues counts against the intuitive story that the organism pauses because it has just eaten; the pause is governed by the work ahead.[7] A fixed ratio is more persistent than continuous reinforcement, but as the [hub page](https://operantconditioning.com/schedules-of-reinforcement/) explains, the variable ratio is the schedule to choose when persistence is the goal.[9] ## Key takeaways - A fixed ratio schedule delivers a reinforcer after a set number of responses. FR 1 is continuous reinforcement, FR 20 pays every twentieth response, and the organism's own rate sets how often reinforcers arrive. - The pattern is break-and-run: a pause after each reinforcer, then a fast, steady run. The pause lengthens as the ratio grows while the running rate stays about the same, and it is controlled by the ratio ahead rather than by the reinforcer just received. - Ratios raised too fast produce ratio strain — long pauses, broken runs, and quitting — because a long string of unreinforced responses is what extinction looks like. Drop back and raise the ratio more gradually. - Piece-rate pay, loyalty punch cards, reading logs, and "collect ten" quests are fixed ratios, and field evidence such as the productivity gain at Safelite shows count-based reinforcement doing what the laboratory predicts. - To use one, define the unit, start at FR 1, raise the ratio in small steps, watch for strain, match the reinforcer to the ratio, and switch to a variable ratio if the pause is a problem. The schedule reinforces only what it counts, and people also follow the rules they are given. ### Check yourself **A teacher gives a sticker after every five completed math problems. A student finishes five, collects the sticker, and then sits doing nothing for several minutes before starting the next five. Is the program failing?** No. The idle stretch is the post-reinforcement pause, the signature of a fixed ratio schedule, and the evidence says it is governed by the size of the ratio ahead rather than by the sticker just received. If the dead time matters, lower the ratio, make the reinforcer larger, or switch to a variable ratio so there is no predictable moment when the next sticker is far away. **A café changes its loyalty card overnight from "buy ten, get one free" to "buy twenty-five, get one free." What does the schedule predict?** Ratio strain. The requirement was raised in one large step, so expect longer gaps between purchases, cards abandoned partway, and some customers dropping the program altogether, because from the customer's side a long run of unrewarded purchases looks like extinction. The schedule-informed alternative is to raise the ratio in smaller steps or make the reward worth the larger count. **A company pays a commission on every sale. Is the salesperson's work on an FR 1 schedule?** Only if the counted unit is the sale. The behavior the salesperson emits all day is making calls, and a sale follows an unpredictable number of calls, so the calling is on a variable ratio. The commission is FR 1 on sales and VR on calls, and it is the variable count that shapes the pattern of the working day. ## Frequently asked questions **What is a fixed ratio schedule in simple terms?** A rule that delivers a reward after a set number of actions. Every tenth coffee is free, every windshield installed earns a set amount, every five books earn a sticker. The number is always the same, so the behavior comes in bursts: a pause after each reward, then a fast run until the next one. **What is an example of a fixed ratio schedule?** A loyalty punch card that gives a free drink after ten purchases is FR 10. Piece-rate pay, in which a worker is paid per unit produced, is FR 1 with money as the reinforcer. In the laboratory, a pigeon on FR 50 receives grain after every fiftieth peck and shows a pause after each delivery followed by a rapid run of pecking. **What is the difference between a fixed ratio and a variable ratio schedule?** Both deliver reinforcement after a number of responses, but on a fixed ratio the number is always the same, while on a variable ratio it varies around an average. Because the fixed count is predictable, it produces a pause after each reinforcer; the variable ratio removes the pause, maintains a steadier rate, and is more resistant to extinction. **Is FR 1 the same as continuous reinforcement?** Yes. FR 1 reinforces every response, which is the definition of continuous reinforcement. It is the schedule to use while a behavior is being learned. Any ratio larger than 1 is an intermittent schedule, and the usual practice is to start at FR 1 and raise the ratio gradually once the behavior is reliable. **What is the post-reinforcement pause?** The pause in responding that follows each reinforcer on a fixed ratio schedule, before the next run begins. It grows longer as the ratio grows larger while the run itself stays about as fast. Research reviewed by Schlinger, Derenne, and Baron shows the pause is controlled by the size of the ratio ahead, not by the reinforcer just received, so it is more accurately called a pre-ratio pause. **What is ratio strain?** The breakdown of responding that occurs when a ratio schedule is raised too quickly or set too high for the reinforcer: long pauses, runs that stop partway, and eventually quitting. It happens because a long string of unreinforced responses is what extinction looks like. The remedy is to drop back to a ratio that worked and raise it more gradually. **Is piece-rate pay a fixed ratio schedule?** Yes. Pay per unit produced is a fixed ratio with money as the reinforcer, and Skinner described piecework as the industrial form of the schedule. Lazear's study of Safelite Glass found that output per worker rose by about 44 percent after a switch from hourly wages to piece rates. Because the schedule reinforces only what it counts, quality has to be measured separately. ## References 1. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Zeiler, M. D. (1977). Schedules of reinforcement: The controlling variables. In W. K. Honig & J. E. R. Staddon (Eds.), *Handbook of Operant Behavior* (pp. 201–232). Prentice-Hall. 4. Baum, W. M. (1993). Performances on ratio and interval schedules of reinforcement: Data and theory. *Journal of the Experimental Analysis of Behavior, 59*(2), 245–264. 5. Felton, M., & Lyon, D. O. (1966). The post-reinforcement pause. *Journal of the Experimental Analysis of Behavior, 9*(2), 131–134. 6. Powell, R. W. (1968). The effect of small sequential changes in fixed-ratio size upon the post-reinforcement pause. *Journal of the Experimental Analysis of Behavior, 11*(5), 589–593. 7. Schlinger, H. D., Derenne, A., & Baron, A. (2008). What 50 years of research tell us about pausing under ratio schedules of reinforcement. *The Behavior Analyst, 31*(1), 39–60. 8. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 9. Mazur, J. E. (2017). *Learning and Behavior* (8th ed.). Routledge. 10. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 11. Lazear, E. P. (2000). Performance pay and productivity. *American Economic Review, 90*(5), 1346–1361. 12. Kivetz, R., Urminsky, O., & Zheng, Y. (2006). The goal-gradient hypothesis resurrected: Purchase acceleration, illusionary goal progress, and customer retention. *Journal of Marketing Research, 43*(1), 39–58. 13. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): All four basic schedules, side by side, with a simulator. - [Variable ratio schedule](https://operantconditioning.com/variable-ratio-schedule/): The same count, made unpredictable: why the pause disappears and the behavior persists. - [Continuous vs. intermittent reinforcement](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/): When to reinforce every response and how to thin. # Fixed Interval Schedule: Definition, the Scallop & Examples > A fixed interval schedule reinforces the first response after a set time. The scallop, FI vs. fixed time, Congress, cramming, why adults differ, and mistakes. - Source: https://operantconditioning.com/fixed-interval-schedule/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Schedules · Interval* Nothing for most of the interval, then a rush at the end. That is what a fixed interval schedule does to a pigeon, to Congress, and to anyone with a deadline, and understanding why is the first step to designing around it. > **Definition** > > A fixed interval (FI) schedule is a schedule of reinforcement in which the first response after a fixed amount of time has elapsed since the previous reinforcer is reinforced. Responses made before the interval has run out have no effect: they neither bring the reinforcer sooner nor delay it. Written as FI 60 s, the schedule reinforces a response once 60 seconds have passed since the last reinforcer, and not before.[1] > > Skinner studied it in rats in the 1930s as *periodic reconditioning*; Ferster and Skinner's 1957 catalogue of schedules gave it its modern name and documented its signature, the **fixed-interval scallop**: a pause after each reinforcer, then an accelerating run of responses as the interval nears its end.[2][1] **In brief** - On a fixed interval schedule only the first response after a set time since the last reinforcer is reinforced; responses during the interval do nothing, and responding faster does not bring the reinforcer sooner. - Animals on FI schedules produce the scallop, little responding after each reinforcer and an accelerating rate as the interval ends, because time since reinforcement comes to act as a [discriminative stimulus](https://operantconditioning.com/discriminative-stimulus/). - Adults often do not scallop, because they form rules about the schedule; the clearest scallops in human life come from institutions with fixed deadlines, such as Congress and courses with widely spaced tests. ## How a fixed interval schedule works Two conditions have to be met before a fixed interval schedule delivers a reinforcer: a fixed amount of time has to have passed since the last reinforcer, and after that time a response has to occur. In a [Skinner box](https://operantconditioning.com/skinner-box/) running FI 60 s, a timer starts when a pellet is delivered. For the next 60 seconds the lever is effectively disconnected; the rat can press once or a hundred times and nothing happens. When the timer runs out, the schedule sets up a reinforcer and holds it. The next press, whenever it comes, delivers the pellet and restarts the timer.[1] Three consequences follow. - **Responding faster does not pay.** On a [ratio schedule](https://operantconditioning.com/fixed-ratio-schedule/) every response brings the reinforcer closer. On FI, one well-timed response per interval collects everything the schedule offers, which is why interval schedules maintain lower rates than ratio schedules delivering the same reinforcers per hour.[3] - **Waiting too long does cost.** Because the timer restarts at delivery, a late response delays the reinforcer and the next interval. The organism cannot speed things up, but it can slow them down. - **The reinforcer itself is a signal.** Each reinforcer marks the start of a period in which no response can be reinforced, so time since the last reinforcer becomes the cue the organism uses. The FI pattern is a matter of [stimulus control](https://operantconditioning.com/stimulus-control/) as much as of reinforcement. Laboratory variants add a **limited hold**, a window after which an uncollected reinforcer is withdrawn and the interval restarts, which pushes responding still closer to the end of the interval.[1] ### Not "reinforcement every N seconds" The commonest misreading of FI drops the response requirement. "The rat gets a pellet every minute" describes a **fixed time (FT)** schedule, on which the reinforcer arrives on the clock whether or not any behavior occurs. On FI the reinforcer is *available* every minute; it is *delivered* only when the organism responds. That is the difference between a contingent and a noncontingent reinforcer, and in everyday terms between a salary and a deadline.[3][4] The two schedules produce different behavior. Skinner's 1948 "superstition" experiment fed pigeons on FT 15 s, and instead of one strong response the birds developed idiosyncratic, accidentally reinforced rituals.[5] Zeiler's review of schedule performance treats the response requirement as a controlling variable: remove it, as FT does, and responding generally falls, though behavior that happens to precede deliveries can persist for a while on adventitious reinforcement.[4] [More on superstitious behavior ›](https://operantconditioning.com/superstitious-behavior/) ## The scallop: what the cumulative record shows Skinner saw the pattern first in rats. Under periodic reconditioning his rats at first pressed at a fairly even rate, but with continued exposure what he called a temporal discrimination developed: pressing dropped off after each pellet and picked up as the next came due.[2] Ferster and Skinner's pigeons produced the same thing across a wide range of interval lengths, and their cumulative records, flat after each reinforcer and then curving upward with increasing steepness, gave the pattern its name: a series of them looks like the scalloped edge of a shell.[1] You can watch one form in the site's [virtual Skinner box](https://operantconditioning.com/lab/). The pause is not idleness the schedule tolerates; it is behavior the schedule shapes. Responses just after a reinforcer are never reinforced, so they undergo [extinction](https://operantconditioning.com/extinction/); responses near the end of the interval sometimes are, so they persist. Elapsed time, unlike a light or a tone, is a cue the organism reads imprecisely, and the scallop is what imprecise timing looks like: the nearer the end of the interval, the likelier the next response is to pay, and the faster the organism responds.[3] The [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause) grows with the interval, so FI 5 min produces a much longer pause than FI 30 s.[1] ### Curve or break-and-run Whether individual intervals show a smooth acceleration is debated. Schneider trained pigeons on FI schedules from 16 s to 512 s and found that responding within each interval was well described by two states, almost none and then an abrupt switch to a steady high rate, with the switch coming about two-thirds of the way through the interval whatever its length.[6] On that reading the averaged scallop comes from stacking many break-and-run intervals whose breaks fall at slightly different times. Dews, examining how responses were distributed within intervals, argued that the rate rises progressively and that the curvature is genuine rather than an artifact of averaging.[7] Both agree on what matters outside the laboratory: after a reinforcer there is a pause, and it ends well before the interval does. See the [glossary entry for scallop](https://operantconditioning.com/glossary/#scallop). ## Examples of fixed interval schedules An everyday situation is a fixed interval schedule when three things are true: the reinforcer becomes available after a fixed time, a response is still required to collect it, and responding before the time does nothing. Skinner argued in *Science and Human Behavior* that laboratory schedules run throughout human affairs, and the FI pattern is easy to spot once you know the test.[8] | Setting | Interval | Response | Reinforcer | Pattern | | --- | --- | --- | --- | --- | | Laboratory | FI 2 min | Lever press | Food pellet | Pause after each pellet, accelerating presses toward the end of the interval | | Kitchen | A 45-minute bake | Opening the oven to check | A finished cake; opening early yields nothing | No checking early, more and more as the timer runs down | | Home | Mail delivered around 11 a.m. | Walking to the mailbox | The mail | Few trips in the morning; trips cluster around eleven | | Dog training | Dinner at 6 p.m. | Coming to the kitchen and sitting by the bowl | Dinner | Hovering begins around 5:30 and builds | | Classroom | A test every three weeks | Studying | The grade | Little studying after a test, cramming before the next | | Workplace | A report due on the first of each month | Working on the report | The report accepted, the pressure gone | Slow start, sprint at month's end | | Not FI: payday | Every two weeks | None required on the day | The paycheck | Fixed time, not fixed interval: the check arrives whether or not you did anything on payday | The classroom and workplace rows are approximations. On a true FI schedule early responses are wasted, whereas studying in week one does count toward the grade in week three. What those rows share with the laboratory schedule is a fixed, known time at which the consequence arrives and a response required to collect it. The practical tell is the pause. If you do almost nothing right after the reinforcer and a great deal just before the next one, you are on something close to a fixed interval, whatever it is called. ## Fixed interval vs. variable interval, fixed ratio, and variable ratio The four basic [intermittent schedules](https://operantconditioning.com/schedules-of-reinforcement/) differ on two questions: is the requirement a count of responses or a passage of time, and is it fixed or does it vary around an average? Fixed interval is the time-based, fixed cell.[1][3] | Feature | Fixed interval (FI) | Variable interval (VI) | Fixed ratio (FR) | Variable ratio (VR) | | --- | --- | --- | --- | --- | | What is required | First response after a set time | First response after a time that varies around an average | A set number of responses | A number of responses that varies around an average | | Does responding faster pay? | No; one well-timed response collects everything | No | Yes; every response counts | Yes | | Pause after reinforcement | Yes, growing with the interval | Little or none | Yes, growing with the ratio | Little or none | | Pattern | Scallop: pause, then acceleration | Moderate, steady rate | Break and run: pause, then a fast steady burst | High, steady rate | | Resistance to extinction | Moderate | High | Moderate to high | Highest | | Everyday example | Oven timer, a test every three weeks | Checking email | Piece-rate pay | Slot machine | FI against VI is the comparison that matters most in application. On a [variable interval schedule](https://operantconditioning.com/variable-interval-schedule/) the interval is unpredictable, so no moment after a reinforcer is safe to ignore; the pause disappears and the organism checks at a steady, moderate rate. Replace a scheduled Friday inspection with unannounced inspections averaging one a week and the "clean up on Thursday" scallop flattens into steady tidiness. FI against FR is the comparison students most often get wrong, because both produce a pause after reinforcement. The difference is what ends the pause. On a fixed ratio the organism ends it by working, and running to the next reinforcer at top speed is then the best strategy, so FR gives a break and then a burst; on a fixed interval nothing the organism does shortens the wait, so the run-up is gradual. When you want the most behavior, the [variable ratio schedule](https://operantconditioning.com/variable-ratio-schedule/) removes both the pause and the predictability. ## Fixed interval in the wild: Congress and the classroom **Congress.** In 1972 Weisberg and Waldrop plotted the cumulative number of bills passed by the United States Congress across each session and found scallops: few bills in the early months and a steep acceleration as adjournment approached.[9] Critchfield and colleagues extended the analysis to half a century of sessions, from the late 1940s to 2000, and found the same shape session after session.[10] The adjournment date works like the end of an interval: the pressure to act is felt only as it approaches. **Students.** Mawhinney and colleagues measured college students' studying directly, by timing how long they worked with course materials available only in a monitored study room, under daily, weekly, and three-week testing. With tests three weeks apart, studying was low just after each test and rose steeply in the days before the next, a scallop drawn in study hours. With daily tests, studying was spread steadily across days.[11] [Operant principles in the classroom ›](https://operantconditioning.com/classroom/) Neither is a laboratory FI: early work in the session or the term is not wasted the way early lever presses are, and nothing is set up and held at a fixed moment. What the studies show is that a known, fixed deadline produces in legislators and undergraduates the shape a fixed interval produces in a pigeon: little early, much late. ## Why people often do not scallop Put an adult in front of a button on FI 30 s with points as the reinforcer and you usually do not get a scallop. You get one of two things: a high, steady rate, as if the person were on a ratio schedule, or a very low rate, often a single response just after the interval ends.[12] Harold Weiner showed how much of this depends on history: adults whose earlier sessions had reinforced fast responding on fixed ratio schedules responded at high rates on FI, adults with a history of reinforcement for slow, spaced responding responded at low rates, and a small cost per response pushed people toward the efficient low-rate pattern.[13] Fergus Lowe argued that the deeper reason is verbal. Adults describe the situation to themselves ("it pays about every half minute, so wait" or "press as fast as you can") and then follow the description. The schedule shapes the rule, and the rule rather than the schedule shapes the responding, which is why adult FI performance can look nothing like a pigeon's and can persist even when the rule is wrong.[12] It follows that people who do not yet have language should perform like animals, and they do. Infants under a year old, responding on FI schedules, produced the pausing and accelerating patterns of the animal laboratory.[14] Across childhood, children under about five again performed like animals, children of about seven and older like adults, and the ages between showed a mixture, with the children's own descriptions of the schedule tracking their patterns.[15] For anyone applying schedules to people, the lesson is that the description of the contingency is itself a variable: a deadline you understand produces a plan, and a deadline you cannot see coming produces a pigeon's scallop. ## How to use, and recognize, a fixed interval schedule FI is rarely the schedule you want to design. It produces less behavior than a ratio schedule and lumpier behavior than a variable interval. Its practical importance is that the world imposes it on you through deadlines, scheduled events, and periodic reviews. 1. **Check the contingency.** Does the reinforcer arrive on the clock regardless of behavior (fixed time), only after a response once the time has passed (fixed interval), or in proportion to work (ratio)? A salary is closer to fixed time; a report due after month's end is closer to fixed interval; commission is ratio. 2. **Decide what you want, and change the schedule to match.** For steady behavior, make the interval variable: unannounced checks instead of Friday inspections, spot quizzes instead of the scheduled quiz. For more behavior, tie the reinforcer to output (pages written, calls made), which is a ratio schedule and maintains a higher rate than any interval schedule. 3. **If the interval is fixed and cannot be changed, shorten it.** Daily tests in place of a three-week test produced steady studying in Mawhinney's students.[11] Weekly check-ins instead of a quarterly review do the same thing: many short scallops add up to more even work than one long one. 4. **Reinforce early-interval responding directly.** Pay something for early responses (comments on an early draft, credit for a plan submitted in week one) and the contingency is no longer a fixed interval. 5. **Recognize your own scallops.** If nothing happens the day after a deadline and everything happens the day before the next, you are on a fixed interval. Set shorter intervals with real consequences attached, or move a reinforcer closer to the behavior. [Building habits with operant principles ›](https://operantconditioning.com/habits/) ## Common mistakes - **Confusing fixed interval with fixed time.** "A pellet every minute," "an allowance every Saturday," and "a paycheck every two weeks" describe reinforcers delivered on the clock. Without a required response there is no fixed interval schedule. - **Expecting steady work from a fixed deadline.** A single known deadline predicts a scallop. When a team does nothing for three weeks and everything in the fourth, the schedule, not the team, is the first thing to change. - **Calling every time-based situation fixed interval.** Checking texts, refreshing a feed, and waiting for a reply are variable interval schedules: the time varies, so there is no pause. - **Assuming a person will scallop like a pigeon.** Adults usually do not; their performance depends on the rules they form and on their history, so check what the people actually do.[12] - **Treating the scallop as a character flaw.** Cramming, end-of-session legislation, and last-week report writing are what a fixed interval produces in any organism. Changing the schedule changes the behavior; lecturing about discipline generally does not. ## What the evidence does not show The scallop is one of the best-replicated findings in the animal laboratory, and that reputation gets stretched to cover claims the data do not support. - **It does not show that adults reliably scallop.** Human laboratory studies find high-rate and low-rate patterns far more often than the animal curve, depending on history and on the rule the person forms.[13][12] - **It does not show that Congress or students are on a fixed interval schedule.** The records are scallop-shaped, but the contingency differs: early work is not wasted and no reinforcer is set up at a fixed moment.[10][11] - **It does not settle whether the scallop is a curve or a step.** Schneider's two-state analysis and Dews's within-interval analyses describe the same records differently, and how much of the curvature belongs to individual intervals rather than to averaging remains open.[6][7] ## Key takeaways - On a fixed interval schedule the first response after a set time since the last reinforcer is reinforced; responses during the interval do nothing, so responding faster does not bring the reinforcer sooner. - The schedule produces the scallop, a pause after each reinforcer and an accelerating rate as the interval ends, because early responses are never reinforced and elapsed time is a cue the organism reads imprecisely. - A fixed interval is not "reinforcement every N seconds." That is a fixed time schedule, on which the reinforcer arrives regardless of behavior; a salary is closer to fixed time, a deadline closer to fixed interval. - Fixed deadlines produce scallop-shaped records in institutions: bills passed by Congress accelerate toward the end of each session, and students tested every three weeks study little until the days before each test. - Adults usually do not scallop in the laboratory, because they form rules about the schedule and follow the rule; infants and preschool children, who cannot, perform like animals. - Fixed interval is rarely the schedule you want. For steady behavior make the interval variable; for more behavior tie the reinforcer to responses; if the deadline cannot be moved, shorten it. ### Check yourself **A parent gives a child five dollars every Saturday morning, whether or not any chores were done that week. The parent calls it a fixed interval schedule. Is it?** No. A fixed interval schedule requires a response after the interval has elapsed; here the money arrives on the clock with no response required, which makes it a fixed time schedule. Whatever the child happens to be doing on Saturday morning may be adventitiously reinforced, but no particular behavior is being maintained. Making the money contingent on presenting a completed chore chart after Saturday would turn it into something closer to a fixed interval. **A rat on FI 2 min presses forty times in one interval and six times in the next. Both intervals end with one pellet. Why did the forty presses earn no more than the six?** Because on an interval schedule only the first response after the interval has elapsed is reinforced. The other thirty-nine presses in the first interval were made before the timer ran out and had no effect. Over sessions, exactly this fact shapes the scallop: the wasted early presses extinguish, and the rat learns to pause after each pellet and press mostly toward the end of the interval. **A teacher replaces a midterm with a short quiz every class period and finds that students study more evenly across the term. Which principle is at work?** Shortening the interval. A midterm is one long fixed interval, and studying under it scallops: little early, a lot just before. Daily quizzes are many short intervals, and Mawhinney and colleagues found that daily testing produced steady studying where three-week testing produced cramming. The kind of contingency is the same; the interval is shorter, so the pauses are shorter too. ## Frequently asked questions **What is a fixed interval schedule in simple terms?** A rule for reinforcement in which a behavior pays off only after a fixed amount of time has passed since it last paid off, and only if the behavior occurs then. Checking the oven works once the cake has had its 45 minutes; checking earlier does nothing. The result is a pattern of pausing after each payoff and doing more and more as the next one comes due. **What is an example of a fixed interval schedule?** A rat on FI 2 min gets a pellet for the first lever press after two minutes have passed since the last pellet; presses before that do nothing. Outside the lab: walking to the mailbox when mail arrives at the same time each day, checking a slow cooker, or a course with a test every three weeks, where studying falls after each test and rises before the next. **What is the difference between a fixed interval and a variable interval schedule?** Both reinforce the first response after a period of time, but on a fixed interval the period is always the same and on a variable interval it changes unpredictably around an average. A fixed interval produces a pause after each reinforcer and an accelerating run before the next. A variable interval produces a steady, moderate rate, because there is never a safe period to ignore. **Is a paycheck a fixed interval schedule?** Not really, although textbooks often say so. A fixed interval schedule requires a response after the interval; a salary arrives on the clock as long as you remain employed, which makes it closer to a fixed time, or noncontingent, schedule. A deadline is a better everyday example: the report is accepted only after the month ends and only if you submit it. **What is the fixed interval scallop?** The curved pattern in a cumulative record produced by a fixed interval schedule: almost no responding just after a reinforcer, then a gradually accelerating rate until the next reinforcer, after which the pause repeats. Stacked interval after interval, the curves look like the edge of a scallop shell. The pause reflects the fact that early responses are never reinforced. **What is the difference between fixed interval and fixed ratio?** On a fixed ratio the reinforcer depends on a count of responses, so working faster pays and the pattern is a pause followed by a fast, steady run. On a fixed interval the reinforcer depends on time, so working faster does not pay and the pattern is a pause followed by a gradual acceleration. Both have a post-reinforcement pause that grows with the size of the requirement. **Do humans show the fixed interval scallop?** Often not. Adults tend to respond either at a high steady rate or at a very low rate with a response just after the interval ends, depending on the rule they form about the situation and on their history. Infants and preschool children, who cannot yet describe the schedule, perform much more like laboratory animals. Fixed deadlines in institutions, such as Congress, do produce scallop-shaped records. **What does FI 60 s mean?** A fixed interval schedule of 60 seconds. Once a reinforcer has been delivered, sixty seconds must pass before another can be earned; the first response after that point is reinforced and restarts the timer. The notation gives the schedule type and the interval length; FI 2 min and FI 5 min work the same way with longer waits. ## References 1. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Mazur, J. E. (2017). *Learning and Behavior* (8th ed.). Routledge. 4. Zeiler, M. D. (1977). Schedules of reinforcement: The controlling variables. In W. K. Honig & J. E. R. Staddon (Eds.), *Handbook of Operant Behavior* (pp. 201–232). Prentice-Hall. 5. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 6. Schneider, B. A. (1969). A two-state analysis of fixed-interval responding in the pigeon. *Journal of the Experimental Analysis of Behavior, 12*(5), 677–687. 7. Dews, P. B. (1978). Studies on responding under fixed-interval schedules of reinforcement: II. The scalloped pattern of the cumulative record. *Journal of the Experimental Analysis of Behavior, 29*(1), 67–75. 8. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 9. Weisberg, P., & Waldrop, P. B. (1972). Fixed-interval work habits of Congress. *Journal of Applied Behavior Analysis, 5*(1), 93–97. 10. Critchfield, T. S., Haley, R., Sabo, B., Colbert, J., & Macropoulis, G. (2003). A half century of scalloping in the work habits of the United States Congress. *Journal of Applied Behavior Analysis, 36*(4), 465–469. 11. Mawhinney, V. T., Bostow, D. E., Laws, D. R., Blumenfeld, G. J., & Hopkins, B. L. (1971). A comparison of students' studying-behavior produced by daily, weekly, and three-week testing schedules. *Journal of Applied Behavior Analysis, 4*(4), 257–264. 12. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley. 13. Weiner, H. (1969). Controlling human fixed-interval performance. *Journal of the Experimental Analysis of Behavior, 12*(3), 349–373. 14. Lowe, C. F., Beasty, A., & Bentall, R. P. (1983). The role of verbal behavior in human learning: Infant performance on fixed-interval schedules. *Journal of the Experimental Analysis of Behavior, 39*(1), 157–164. 15. Bentall, R. P., Lowe, C. F., & Beasty, A. (1985). The role of verbal behavior in human learning: II. Developmental differences. *Journal of the Experimental Analysis of Behavior, 43*(2), 165–181. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): All four basic schedules side by side, with an interactive cumulative-record simulator. - [Variable interval schedule](https://operantconditioning.com/variable-interval-schedule/): The unpredictable interval: why it flattens the scallop into a steady rate. - [Fixed ratio schedule](https://operantconditioning.com/fixed-ratio-schedule/): The other schedule with a post-reinforcement pause, and why its pause ends differently. # Variable Interval Schedule: Definition, Examples & Uses > A variable interval schedule reinforces the first response after an unpredictable time: VI 30 s notation, the steady rate it produces, matching, and uses. - Source: https://operantconditioning.com/variable-interval-schedule/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Schedules · Interval* Messages arrive when they arrive, and only the check that comes after one pays off. That is a variable interval schedule. It produces the steadiest behavior in operant psychology, steady enough that laboratories use it as the baseline for almost every experiment on choice and persistence. > **Definition** > > A variable interval (VI) schedule is a schedule of reinforcement in which the first response after a variable, unpredictable period of time has elapsed since the previous reinforcer is reinforced. The intervals vary around an average, and the average names the schedule: on VI 30 s the wait is sometimes 5 seconds and sometimes 70, but it averages 30. Responses made before the interval has run out earn nothing, and responding faster does not make the next reinforcer come sooner.[1] > > VI is one of the four basic [intermittent schedules](https://operantconditioning.com/schedules-of-reinforcement/), alongside fixed interval, fixed ratio, and variable ratio. It produces a moderate, remarkably steady rate of responding with no pauses and no scallop. **In brief** - A variable interval schedule reinforces the first response after an unpredictable interval that averages a set value (VI 30 s). Responses before the interval ends earn nothing, and responding faster does not shorten the wait. - It produces a moderate, very steady rate with no post-reinforcement pause and no scallop: behavior that keeps checking but never races, and that persists when reinforcement is thinned or withdrawn. - Because a VI schedule pays at nearly the same rate however fast the organism responds, it is the standard baseline for laboratory work on choice (the matching law) and on persistence (behavioral momentum). ## How a variable interval schedule works The clock on an interval schedule decides when a reinforcer becomes *available*, not when it is delivered. On a VI schedule the clock starts when a reinforcer is delivered and runs for an interval drawn from a list: 12 seconds this time, 48 the next, 3 after that. While it runs, responses do nothing. When it runs out, the next response is reinforced, the clock resets, and a new interval begins.[1][2] A pigeon on VI 1 min collects close to 60 reinforcers an hour whether it pecks 20 times a minute or 80, because each reinforcer, once set up, waits for the next peck to collect it.[1] Some procedures add a **limited hold**, a window (say, two seconds) after which an uncollected reinforcer is cancelled.[3] Responding faster buys almost nothing On a ratio schedule, doubling the response rate doubles the reinforcement rate. On a VI schedule it does almost nothing once the organism responds often enough to collect each reinforcer soon after it becomes available; plotted, reinforcers per hour climb steeply at very low response rates and then go flat at the programmed rate. That flat feedback function removes the incentive to race, and the unpredictability of the next reinforcer removes pausing.[4] ### How the intervals are generated Ferster and Skinner built their VI schedules from an arithmetic series of intervals, evenly spaced lengths from short to long, arranged in an irregular order.[1] That construction has a flaw: with evenly spaced lengths, the longer an interval has already lasted, the more likely it is to end in the next few seconds, so the probability that a response will be reinforced rises with time since the last reinforcer, and a well-trained animal can learn to speed up as the interval ages.[5] Fleshler and Hoffman solved this in 1962 with a formula that generates intervals distributed so that the probability of a reinforcer becoming available in the next moment is the same however long it has been since the last one: a **constant-probability VI**, the standard laboratory construction since.[5] Catania and Reynolds later showed that pigeons' moment-to-moment rate tracks the moment-to-moment probability of reinforcement: an arithmetic VI produces a mild acceleration within intervals, a constant-probability VI a roughly flat rate.[6] ## The behavior a variable interval schedule produces On a [cumulative record](https://operantconditioning.com/skinner-box/), VI responding is a nearly straight line of moderate slope, ticked at irregular intervals by reinforcers that leave no visible mark on the rate. There is no [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause) worth the name, and no [scallop](https://operantconditioning.com/glossary/#scallop), because nothing about the passage of time predicts the next reinforcer.[1] Skinner emphasized the stability of VI behavior and its resistance to extinction.[7] How fast is "moderate"? Catania and Reynolds answered this in 1968. They ran pigeons on VI schedules spanning a very wide range of reinforcement rates and found that response rate was a **negatively accelerated** function of reinforcement rate: it rose steeply as the schedule went from very lean to moderately rich and then flattened, so that across the middle and upper range, large changes in reinforcement rate produced comparatively small changes in pecking.[6] Herrnstein fitted a hyperbola to these data two years later and made it the quantitative form of the law of effect, which is why their study is the empirical foundation of the [matching law](https://operantconditioning.com/matching-law/).[8] The steadiness has a second source. On any interval schedule, the longer it has been since the last response, the more likely a reinforcer has become available in the meantime, so a response after a long pause is more likely to be reinforced than one after a short pause. The schedule quietly reinforces moderate spacing, whereas a ratio schedule pays most to whoever responds fastest.[9] The result is a rate reliably lower than a ratio schedule with the same reinforcement rate would maintain, as Baum showed with pigeons on VR and VI schedules matched for reinforcers per hour.[4] ## Variable interval vs. fixed interval, variable ratio, and fixed ratio Two questions place any basic schedule: does reinforcement depend on a count of responses or on the passage of time, and is the requirement fixed or variable? VI is time-based and variable. Its nearest neighbors are the [fixed interval schedule](https://operantconditioning.com/fixed-interval-schedule/), which shares its indifference to response rate but adds predictability, and the [variable ratio schedule](https://operantconditioning.com/variable-ratio-schedule/), which shares its unpredictability but pays for every response. | Schedule | Reinforcer depends on | Requirement | Pattern | Faster responding pays? | Everyday example | | --- | --- | --- | --- | --- | --- | | **[Fixed ratio (FR)](https://operantconditioning.com/fixed-ratio-schedule/)** | A count of responses | Fixed (FR 10) | High rate; a pause, then a run | Yes | Piece-rate pay | | **Variable ratio (VR)** | A count of responses | Varies around a mean (VR 10) | Very high, steady rate | Yes | Slot machine | | **Fixed interval (FI)** | Time since the last reinforcer | Fixed (FI 60 s) | Scallop: a pause, then acceleration | No | Checking the oven as the timer runs down | | **Variable interval (VI)** | Time since the last reinforcer | Varies around a mean (VI 60 s) | Moderate, steady rate; no scallop | No | Checking for messages | The confusion that matters most is VI with VR, because both are called variable and both are described as unpredictable. The test is what the unpredictability is attached to. On VR an unpredictable *number of responses* is required, and every response moves the count forward: on a slot machine every pull is another draw at the payout, so more pulls mean more payouts, even though no pull is ever 'due'. On VI an unpredictable *amount of time* must pass, and responses do not move the clock: checking the mailbox for the twentieth time this morning does not bring the mail. VR produces high rates because effort is rewarded; VI produces moderate rates because only timing is.[2] The confusion with FI is the reverse: same clock, different predictability. On FI the interval is always the same, so the organism learns to wait and then accelerate. On VI the interval is different every time, so waiting is never safe and the scallop disappears.[1] ## Why VI is the workhorse of choice experiments VI schedules appear in a large share of the operant literature because they make such a good baseline. The rate is steady, so any change stands out. The obtained reinforcement rate is nearly independent of the response rate, so an experimenter who programs 60 reinforcers an hour gets about 60 whatever the animal does. And the schedule tolerates interruption: a reinforcer that becomes available while the animal is doing something else is still there when it comes back.[4][6] That last property made the modern study of choice possible. In 1961 Richard Herrnstein gave pigeons two keys, each paying on its own VI schedule, a **concurrent VI VI** schedule, and varied how the total reinforcement was divided between them. Because a reinforcer set up on the unattended key waits to be collected, a bird could work mostly on the richer key and still profit by visiting the other now and then. The result was the matching law: the proportion of pecks on a key equalled the proportion of reinforcers it delivered.[10] Concurrent VI VI remains the standard procedure for studying choice.[2] Ratio schedules cannot do this job. On a concurrent VR VR schedule every response on the poorer key is one that could have advanced the count on the richer one, and pigeons settle on the richer alternative almost exclusively, which leaves little to measure.[11] Interval schedules keep both alternatives alive and make the division of behavior the thing under study. ## Why behavior trained on a VI schedule is so persistent Behavior maintained on a VI schedule is hard to disrupt. In extinction it keeps going long after continuously reinforced behavior has stopped, the [partial reinforcement extinction effect](https://operantconditioning.com/glossary/#partial-reinforcement-extinction-effect), and the reason is plain from the animal's side: on VI, long unreinforced stretches are normal, so the moment reinforcement is switched off is not a signal that anything has changed.[2] Skinner made the same point about everyday variable schedules: they build persistence because they never announce that reinforcement has ended.[7] John Nevin turned this into a general theory, with VI schedules as the tool. In 1974 he trained pigeons on multiple schedules, two VI components alternating, each with its own signal and its own rate of reinforcement, and then disrupted responding in both at once by presenting free food between components or by extinction. Responding in the richer component fell less, in proportion to its baseline, than responding in the leaner one.[12] Resistance to change, he argued, is the proper measure of a behavior's strength. Nevin and Grace's 2000 synthesis, **behavioral momentum theory**, made the analogy explicit: response rate is like velocity and is governed by the response–reinforcer contingency, while resistance to change is like mass and is governed by the stimulus–reinforcer relation, that is, by how much reinforcement the situation has delivered, whatever response earned it. VI schedules were essential because they let the experimenter set the reinforcement rate without it being hostage to the animal's response rate.[13] The lesson cuts both ways. Praise on a VI schedule builds classroom behavior that survives a bad day; a problem behavior that has paid off unpredictably in a rich context will survive a good deal of treatment. [Momentum and resistance to extinction ›](https://operantconditioning.com/extinction/#resistance-to-extinction-and-behavioral-momentum) ## Variable interval schedules in everyday life Pure VI schedules are rare outside the laboratory, but the structure is everywhere: something becomes available at unpredictable times, and only the first check after it does is rewarded. | Setting | The hidden clock | The response | Why it is an interval schedule | | --- | --- | --- | --- | | Messaging and email | A reply arrives at an unpredictable time | Glancing at the phone or opening the inbox | Checking faster does not hurry the reply; only the first look after it arrives pays | | Fishing with a set line | A fish bites when it bites | Checking the line or watching the bobber | The bite comes on the fish's time; the angler only has to check often enough not to miss it | | Workplace | A manager drops by at unpredictable times | Being at work on the task when she appears | Praise is available only at the moment of the visit: a VI with a limited hold | | Classroom | Quiz days are unannounced | Studying | Only preparation done before the quiz pays, and there is no date to cram for | Two rows deserve a note. Fishing is the textbook example of a variable *ratio*: each cast is a response, and the fish arrive after an unpredictable number of casts. A line left in the water and checked from time to time is a variable *interval*: the bite comes on the fish's clock, and the angler's responses only collect what time has set up. The manager's walk-through is a VI with a limited hold, which is why it maintains steady work rather than the burst-then-slump of a scheduled inspection.[3] Skinner observed that much everyday behavior is maintained on intermittent schedules that nobody designed, and that its persistence follows from the schedule rather than from anything in the person.[7] ## How to use a variable interval schedule Interval schedules are the easiest intermittent schedules to run in a [classroom](https://operantconditioning.com/classroom/), a home, or a workplace, because they need a timer rather than a count. The teacher does not have to track how many problems each child has finished, only whether the child is working when the timer goes off. Applied behavior analysts use VI to maintain established behavior at a steady rate and to thin reinforcement without the pausing of fixed schedules.[3] 1. **Establish the behavior first.** A VI schedule maintains behavior; it does not build it. Start with [continuous reinforcement](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) until the behavior is reliable, then thin. 2. **Pick an average and write out the intervals.** For VI 5 min of praise, list intervals that average five minutes (1, 3, 4, 6, 7, and 9, say) and use them in shuffled order. The learner must not be able to predict the next check. 3. **Reinforce the first target response after each interval.** When the interval ends, wait for the next instance of the behavior, the child returning to the worksheet or the dog lying quietly, and reinforce that, not whatever happens to be going on when the timer sounds. For a check-in style, add a short limited hold. 4. **Thin gradually.** Move from VI 1 min to VI 2 min to VI 5 min as the behavior holds up. If the rate drops, go back a step; because VI behavior is steady, a drop is easy to spot. 5. **Hand the behavior to natural reinforcers.** A VI of praise is a bridge to behavior that continues because the work itself, the peers, or the finished product reinforce it. ### Pop quizzes and scheduled tests The classic classroom application is testing. A midterm on a known date is, loosely, a fixed interval, and it produces the scallop of cramming. Unannounced quizzes make the interval variable: because a quiz could come at any class, the only study pattern that pays is a steady one. The analogy is loose, since a quiz reinforces studying done days earlier and grades are only one consequence of studying, but the direction of the prediction follows from the schedules. ## Common mistakes and misconceptions - **Confusing VI with VR.** Both are unpredictable, but VR pays per response and VI pays per unit of time. If doing more of the behavior brings the reinforcer sooner, it is a ratio schedule. Slot machines and casting a lure are VR; checking messages and checking a set line are VI. - **Confusing VI with VT.** On a variable time (VT) schedule the reinforcer arrives after a variable interval *whether or not a response occurs*. VT is noncontingent reinforcement, the procedure behind Skinner's ["superstition" experiment](https://operantconditioning.com/superstitious-behavior/); a VI requires a response.[3] - **Expecting VI to produce a high rate.** It produces a moderate one. For more of a behavior per hour, use a ratio schedule; for steady and durable, use VI.[4] - **Calling "giving in every so often" a VI.** A parent who gives in to a tantrum now and then is not running a VI. Each tantrum is a response, and giving in after an unpredictable number of them is a variable ratio. - **Delivering the reinforcer when the timer sounds instead of after the next response.** That is VT, not VI, and it reinforces whatever the learner happens to be doing at that instant. - **Using a predictable "variable" series.** Intervals of 2, 4, 6, and 8 minutes in that order are a pattern, and learners find patterns. Shuffle them. - **Starting on VI.** A new behavior on a lean VI is not emitted often enough to contact the reinforcer. Build first, then thin. ## What the evidence does not show The cumulative records are from pigeons and rats, and the everyday examples are classifications by structure, not measurements. **Adult humans often do not produce the textbook patterns.** People given interval schedules in the laboratory frequently respond at a very high steady rate or a very low one, depending on what they have been told or have concluded about the rule, and instructions can override the programmed contingency for a long time.[14] In adults, the description of the schedule is itself a variable, and no one should assume that praise on a VI will produce a pigeon's straight line. **The insensitivity of rate to reinforcement rate is bounded.** Catania and Reynolds's curve is flat in the middle and upper range, not everywhere. At very lean schedules, response rate falls steeply and behavior can be lost altogether.[6] Thinning is safe only within limits found by watching the behavior. **Pop quizzes and steady studying is a prediction, not a finding reported here.** Studying is maintained by many consequences at once, and the mapping of response and reinforcer is loose. The same goes for messaging: checking is also reinforced by the content itself and by escape from boredom, and nothing on this page shows that making notifications predictable would reduce it. **VI does not make behavior persistent by itself.** Nevin's finding is that persistence tracks the rate of reinforcement in a context. VI was the instrument, not the cause; a lean VI in a lean context produces behavior that is easy to disrupt.[13] ## Key takeaways - A variable interval schedule reinforces the first response after an unpredictable interval that averages a set value. Responses before the interval ends earn nothing, and responding faster does not shorten the wait. - VI produces a moderate, very steady rate with no post-reinforcement pause and no scallop, because time since the last reinforcer predicts nothing and speed buys nothing. Laboratory VI schedules use the Fleshler–Hoffman progression to keep the probability of reinforcement constant from moment to moment. - Response rate on VI rises with reinforcement rate and then flattens, which is why VI behavior is stable across a range of reinforcement rates and why Herrnstein could fit the hyperbola of the matching law to VI data. - VI is the workhorse of the laboratory because its reinforcement rate is nearly independent of response rate and its reinforcers wait to be collected. Concurrent VI VI produced the matching law, and multiple VI schedules produced behavioral momentum theory. - VR pays per response and VI pays per unit of time: slot machines are VR, while checking messages and a set fishing line are VI. To use VI, build the behavior first, shuffle the intervals, reinforce the first response after each one, and thin gradually. ### Check yourself **A teacher sets a timer to go off at unpredictable times averaging five minutes and, when it sounds, praises whichever students are working at that instant. Is this a variable interval schedule of praise for working?** Nearly. On a true VI the first working response after the interval ends would be reinforced whenever it occurred; here praise is available only at the moment the timer sounds, so it is a VI with a very short limited hold, a momentary check. It still maintains steady work, but a student who looks up at the wrong instant misses the reinforcer, and an off-task student who glances at the worksheet at the right instant may be praised. To make it a VI, wait after the timer for the next instance of working and praise that. **A slot-machine player and a person refreshing an inbox are both being rewarded unpredictably. Which is on a variable ratio and which on a variable interval, and what pattern would you predict for each?** The slot machine pays after an unpredictable number of pulls, so each pull advances the count: variable ratio, and a high, steady rate. The inbox pays only when a message has arrived, and refreshing does not make messages arrive sooner: variable interval, and a moderate, steady rate of checking. Both persist when reinforcement stops; the difference between them is in rate, not persistence. **You switch a dog from a treat every time she lies on her mat to a treat for the first lie-down after intervals averaging three minutes. Within a day she has stopped going to the mat. What went wrong?** You thinned too far too fast. A jump from continuous reinforcement to VI 3 min means long stretches of unreinforced behavior before the dog has learned that the mat still pays, so the behavior never contacted the new schedule. Go back to continuous reinforcement, then thin in steps, VI 20 s, then VI 45 s, then VI 90 s, watching the rate at each step and backing up if it falls. ## Frequently asked questions **What is a variable interval schedule in simple terms?** A rule for when a behavior pays off: the first response after an unpredictable amount of time has passed is reinforced, and responses before that earn nothing. Checking for messages is the everyday case. A message arrives on its own schedule, the first look after it arrives finds it, and looking more often does not make the next one come sooner. **What does VI 30 s mean?** VI 30 s is a variable interval schedule whose intervals average 30 seconds. After each reinforcer, a timer runs for an unpredictable interval, sometimes a few seconds and sometimes a minute or more, and the first response after it ends is reinforced. The number states the mean of the intervals, not a maximum or a minimum. **What is an example of a variable interval schedule?** Checking email. Messages arrive at unpredictable times, only a check made after one has arrived is rewarded, and checking faster does not bring mail sooner. Others: glancing at a fishing bobber, looking down the street for an overdue bus, a dog watching the window for the owner's car, and a manager's unannounced walk-throughs. **What is the difference between a variable interval and a variable ratio schedule?** Both are unpredictable, but a variable ratio pays after an unpredictable number of responses, so every response brings the reinforcer closer, while a variable interval pays for the first response after an unpredictable time, so responding faster does nothing. Variable ratio produces a high rate, as in slot machines; variable interval produces a moderate, steady rate, as in checking messages. **What is the difference between variable interval and fixed interval?** Both reinforce the first response after a period of time. On fixed interval the period is always the same, so the learner pauses after each reinforcer and speeds up as the deadline approaches, producing the scallop. On variable interval the period changes unpredictably, so there is no safe time to pause and no deadline to rush for, and the rate stays steady. **Why do researchers use variable interval schedules so often?** Because VI behavior is a steady baseline against which changes are easy to see, and because the reinforcement rate on VI is nearly independent of how fast the animal responds, so experimenters can set it precisely. Herrnstein's matching law came from pigeons on two concurrent VI schedules, and Nevin's behavioral momentum research used multiple VI schedules for the same reason. **Is a pop quiz a variable interval schedule?** Loosely, yes. A test on a known date resembles a fixed interval and produces cramming; unannounced quizzes make the interval unpredictable, so only steady studying pays. The mapping is imperfect, since a quiz reinforces studying done earlier and grades are only one consequence of studying, so treat it as an analogy that predicts the direction of the effect rather than a measured result. **Does a variable interval schedule produce a scallop?** No. The fixed-interval scallop appears because the animal learns that reinforcement is never available soon after the last one and always available at a fixed time. On a variable interval schedule the next reinforcer could become available at any moment, so nothing about elapsed time is worth waiting for, and the cumulative record is a nearly straight line. ## References 1. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 2. Mazur, J. E. (2017). *Learning and Behavior* (8th ed.). Routledge. 3. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 4. Baum, W. M. (1993). Performances on ratio and interval schedules of reinforcement: Data and theory. *Journal of the Experimental Analysis of Behavior, 59*(2), 245–264. 5. Fleshler, M., & Hoffman, H. S. (1962). A progression for generating variable-interval schedules. *Journal of the Experimental Analysis of Behavior, 5*(4), 529–530. 6. Catania, A. C., & Reynolds, G. S. (1968). A quantitative analysis of the responding maintained by interval schedules of reinforcement. *Journal of the Experimental Analysis of Behavior, 11*(3, Suppl.), 327–383. 7. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 8. Herrnstein, R. J. (1970). On the law of effect. *Journal of the Experimental Analysis of Behavior, 13*(2), 243–266. 9. Zeiler, M. D. (1977). Schedules of reinforcement: The controlling variables. In W. K. Honig & J. E. R. Staddon (Eds.), *Handbook of Operant Behavior* (pp. 201–232). Prentice-Hall. 10. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. *Journal of the Experimental Analysis of Behavior, 4*(3), 267–272. 11. Herrnstein, R. J., & Loveland, D. H. (1975). Maximizing and matching on concurrent ratio schedules. *Journal of the Experimental Analysis of Behavior, 24*(1), 107–116. 12. Nevin, J. A. (1974). Response strength in multiple schedules. *Journal of the Experimental Analysis of Behavior, 21*(3), 389–408. 13. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 14. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): All four basic schedules, with a cumulative-record simulator. - [The matching law](https://operantconditioning.com/matching-law/): What concurrent VI VI schedules revealed about choice. - [Variable ratio schedule](https://operantconditioning.com/variable-ratio-schedule/): The count-based cousin: slot machines and the highest rates. # Continuous vs. Intermittent Reinforcement: When to Use Each > Continuous reinforcement pays every response; intermittent pays only some. When to use each, the partial reinforcement extinction effect, and how to thin. - Source: https://operantconditioning.com/continuous-vs-intermittent-reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Schedules · Fundamentals* Reinforce every response and a behavior is learned quickly and abandoned quickly. Reinforce only some and it is learned slowly and kept for a very long time. The choice between the two is the most practical decision in operant conditioning, and it hides one of the field's oldest puzzles. > **Definition** > > Continuous reinforcement (CRF) is a schedule on which every occurrence of a behavior is followed by the reinforcer. **Intermittent reinforcement**, also called partial reinforcement, is any schedule on which only some occurrences are. Ferster and Skinner treated continuous reinforcement as the baseline from which every other schedule departs.[1] > > The two schedules do different jobs. Continuous reinforcement establishes a new behavior fastest, because every response contacts its consequence. Intermittent reinforcement makes an established behavior more durable: it delays satiation, it costs fewer reinforcers, and above all it persists far longer when reinforcement stops — the [partial reinforcement extinction effect](https://operantconditioning.com/glossary/#partial-reinforcement-extinction-effect).[2] **In brief** - On a **continuous** schedule every response is reinforced; on an **intermittent** schedule only some are. Continuous reinforcement teaches a new behavior fastest, and intermittent reinforcement keeps an established one going. - Behavior built on intermittent reinforcement persists far longer once reinforcement stops, even though it received fewer reinforcers: the partial reinforcement extinction effect, first demonstrated by Humphreys in 1939. - The working rule is continuous first, then thin in small steps toward a variable schedule. Thin too fast and responding breaks down (ratio strain); never thin and the behavior collapses the first time the reinforcer fails to arrive. ## How continuous and intermittent reinforcement differ Every schedule answers one question: which responses will be reinforced? On continuous reinforcement the answer is all of them (FR 1 is the same thing under another name). A rat whose every lever press produces a pellet, a toddler whose every "please" produces the cookie, a light switch that works every time — each is on CRF, and the learner cannot fail to notice the relation.[1] On an intermittent schedule some responses go unreinforced by design, under a rule that counts responses (fixed and [variable ratio](https://operantconditioning.com/variable-ratio-schedule/)) or clock time (fixed and variable interval); the [schedules page](https://operantconditioning.com/schedules-of-reinforcement/) covers those four patterns. Skinner argued that almost all reinforcement outside the laboratory is intermittent — the angler does not catch a fish on every cast — whereas continuous reinforcement is mostly something people arrange on purpose: a trainer with a pouch of treats, a vending machine.[3] ### Two schedules, two jobs During **acquisition** — while a behavior is new, weak, or being [shaped](https://operantconditioning.com/shaping/) — reinforce every occurrence; an unreinforced response early on is a lost opportunity. During **maintenance** — once the behavior is reliable — move to an intermittent schedule, which holds the behavior with fewer reinforcers and makes it robust when reinforcement is occasionally missed.[4] Ferster and Skinner followed the same rule with pigeons: establish the response on continuous reinforcement, then introduce the schedule.[1] Three differences drive it. - **Information.** On CRF, one unreinforced response is news; on an intermittent schedule it is ordinary. This is why continuous reinforcement teaches fast and [extinguishes](https://operantconditioning.com/extinction/) fast. - **Satiation.** Every reinforcer brings satiation closer. A rat on CRF fills up within a short session; on a lean schedule a few pellets sustain a long one.[1] - **Rate and pattern.** Continuous reinforcement produces a steady, moderate rate limited by consumption time. Ratio schedules produce higher rates, and each intermittent schedule has its own pattern of pausing and running.[1] ### The comparison at a glance | Feature | Continuous reinforcement (CRF) | Intermittent (partial) reinforcement | | --- | --- | --- | | Which responses are reinforced | Every one | Only some, by a ratio or interval rule | | Speed of acquisition | Fastest | Slower; may never be acquired on a lean schedule | | Response rate once established | Moderate and steady | Higher on ratio schedules | | Satiation | Reached quickly | Delayed | | Resistance to extinction | Low; extinction is rapid | High; highest on variable ratio | | Best used for | Teaching a new behavior; shaping | Maintaining an established behavior | ## Examples of continuous and intermittent reinforcement In each row the behavior is the same; only the rule for reinforcing it differs. | Setting | Behavior | Continuous | Intermittent | | --- | --- | --- | --- | | Machines | Pressing a button | Vending machine: every purchase delivers the snack | Slot machine: a payout after an unpredictable number of plays | | Dog training | Sitting on cue | Week one: a treat for every sit | Later: a treat for some sits, praise for all of them | | Classroom | Raising a hand before speaking | New routine: the teacher calls on the student every time | Established routine: called on some of the time, never for calling out | | Parenting | Asking politely | Every polite request granted while the phrase is being learned | Polite requests granted often but not always | | Fishing | Casting | A stocked pond, briefly | Some casts catch a fish; the angler keeps casting for hours | The dog-training row shows the usual sequence: continuous reinforcement while the sit is being taught, then intermittent treats with praise on every trial as a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer). The slot-machine row shows a designer who wants persistence rather than learning: no one needs to be taught to press a button, so the machine can start lean and stay lean — Skinner's gambler, held through long droughts by a variable ratio.[3] ## The partial reinforcement extinction effect When reinforcement is withdrawn, behavior that was reinforced only some of the time keeps going longer than behavior that was reinforced every time. This is the partial reinforcement extinction effect (PRE), and it is paradoxical on its face: the intermittently reinforced behavior received fewer reinforcers and yet is harder to get rid of. Skinner came to intermittent reinforcement by accident, when a shortage of hand-made pellets led him to reinforce his rats only once a minute rather than for every press.[5] In *The Behavior of Organisms* he reported that rats reinforced on this "periodic" schedule produced far larger extinction curves than rats reinforced for every press.[6] The first experiment designed to isolate the effect was Lloyd Humphreys's in 1939, and it was not operant at all. Humphreys paired a light with a puff of air to the eye until his human subjects blinked at the light; one group received the puff on every trial, another on a random half. When the puff was discontinued, the half-reinforced group went on blinking for far longer.[2] Mowrer and Jones extended the finding to operant behavior in 1945. Rats pressed a bar for food on schedules that paid every press, every second, third, or fourth press, or an irregular mixture; the leaner the schedule, the more presses the rats made in extinction. They also noticed that if each rat's "response" was counted as the run of presses its schedule required, the groups extinguished after roughly the same number of runs — the response-unit hypothesis.[7] The effect has since been reproduced in runways and Skinner boxes, in rats, pigeons, and people, and Mackintosh's 1974 review treated it as one of the most reliable phenomena in the study of extinction.[8] ## Why less reinforcement produces more persistence The effect took decades to explain, and each explanation captures something the others miss. ### The discrimination hypothesis The oldest account, from Mowrer and Jones, is that extinction has to be noticed before it can take hold. After continuous reinforcement the first unreinforced response is unmistakable evidence that conditions have changed; after intermittent reinforcement a run of unreinforced responses looks like more of the same.[7] It is correct as far as it goes. But in two 1962 experiments, animals given partial reinforcement, then a block of continuous reinforcement, and only then extinction still showed the effect, although the shift to extinction should have been just as detectable for them as for animals trained on continuous reinforcement alone.[9][10] Something learned during partial reinforcement survived the intervening block. ### Amsel's frustration theory Abram Amsel argued that nonreward is not a neutral event. When an organism expects a reinforcer and does not get it, the omission produces **frustrative nonreward**, an aversive state that disrupts behavior and drives the organism away. On a continuous schedule, frustration is first met in extinction, where it does exactly that. On an intermittent schedule the organism meets frustration during training, keeps responding, and is reinforced, so the cues of anticipated frustration become signals for continuing rather than quitting.[11] This explains why the effect survives a block of continuous reinforcement. ### Capaldi's sequential theory E. J. Capaldi's account is about memory rather than emotion. On each trial the organism carries a trace of what happened on the last one, and that trace is part of the situation in which the next response occurs. When a nonrewarded trial is followed by a rewarded one, responding is reinforced in the presence of the memory of nonreward. In extinction that memory is the only kind on offer; a subject whose responding has been reinforced in its presence keeps going, and a subject on continuous reinforcement, which never has, stops.[12] The theory predicts effects of the *sequence* of trials, not just their proportion, and it handles effects that appear after only a handful of trials, where frustration has little time to develop. ### Nevin's behavioral momentum John Nevin came at the problem from the other direction. In his research on [behavioral momentum](https://operantconditioning.com/glossary/#behavioral-momentum), resistance to disruption is measured as the proportional decline from baseline, in the same subject, across situations that differ in rate of reinforcement. Measured this way, richer reinforcement produces *more* resistance to extinction, not less. Nevin argued that the classic group comparison counts absolute responses without controlling for the different baseline rates the schedules produce, and that what remains of the PRE reflects the size of the change from training to extinction: from continuous reinforcement to none is a large change, from lean reinforcement to none a small one.[13] In the fuller theory, persistence depends on the rate of reinforcement in a context, and the speed of extinction also on how detectable the change is.[14] What the theories agree on Intermittent reinforcement teaches something continuous reinforcement never can: that unreinforced responses are ordinary and that responding through them pays. Whether that lesson is called a discrimination, a counterconditioned frustration, a memory, or a small stimulus change, the prescription is the same. An organism that has never met an unreinforced response has no defense against the first one. ## How to thin reinforcement from continuous to intermittent The transition is called **schedule thinning**, and it is where most applied failures happen: the goal is to reinforce fewer responses without the behavior faltering.[4][15] 1. **Establish the behavior on continuous reinforcement.** Stay on CRF until the behavior occurs promptly and reliably whenever the occasion arises: the dog sits on the first cue, the student raises a hand without a reminder. 2. **Take the smallest step.** From every response, go to two of every three, then every other (FR 2), then two of every five — steps small enough that the learner hardly notices. Big jumps are the usual cause of collapse. 3. **Make it variable as soon as it is intermittent.** Past every-other, vary the requirement around the average rather than fixing it. Variable schedules produce steadier responding with fewer pauses, and they are what the world will impose anyway.[1] 4. **Move only when performance is stable.** Raise the requirement after the behavior has held steady at the current step, not on a timetable. 5. **Watch for ratio strain and back off.** Pausing, slowing, errors, and emotional behavior after a step up are [ratio strain](https://operantconditioning.com/glossary/#ratio-strain): the schedule has been stretched faster than the behavior can bear. Return to the last requirement that worked and stretch again more gradually.[4] 6. **Keep a conditioned reinforcer on every response.** The tangible reinforcer becomes intermittent; praise, a click, or a check mark can stay continuous. In a [token economy](https://operantconditioning.com/token-economy/), the tokens thin and the exchange for backup reinforcers is spaced out separately.[15] 7. **Hand the behavior to natural reinforcers.** The endpoint is no contrived schedule at all: the sit that produces the walk, the hand-raise that produces the turn to speak; the [habits page](https://operantconditioning.com/habits/) covers this handoff for your own behavior. ## Why intermittently reinforced problem behavior is so hard to extinguish The partial reinforcement extinction effect has a dark side. Problem behavior in homes, classrooms, and institutions is almost never reinforced continuously: a tantrum is given in to on some occasions and not others; a self-injurious act draws attention from one staff member and is ignored by the next. By the time anyone tries extinction, the behavior has a long history on a lean, variable schedule — the condition that produces the greatest persistence. Lerman and Iwata's 1996 review of basic and applied extinction research made the point directly: laboratory findings predict that intermittently reinforced behavior will resist extinction, the natural environment supplies intermittent reinforcement almost by default, and applied studies that varied the schedule before extinction were scarce. Their practical conclusions still hold: expect extinction to be slow, expect an [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst), and combine extinction with [differential reinforcement](https://operantconditioning.com/differential-reinforcement/) of an alternative behavior rather than relying on withholding alone.[16] Behavioral momentum adds a warning: reinforcing an alternative behavior in the same context enriches it, and Nevin and Grace's theory predicts that a richer context makes all behavior in it more resistant to change, the problem behavior included.[14] Practitioners therefore deliver the alternative reinforcement in a clearly different situation where they can, and above all they stay consistent: a single give-in during extinction is a reinforcer delivered on the leanest schedule yet. ## "Intermittent reinforcement" in relationships The term has escaped the laboratory. In articles about dating and abuse, "intermittent reinforcement" describes a partner who is warm one day and cold or cruel the next, and the argument is that the unpredictability itself keeps the other person attached, often under the heading of "trauma bonding." It is worth being precise about what the schedule concept supports. What it supports is a claim about behavior. If someone's texting, apologizing, or effort to please has been reinforced by affection on an unpredictable schedule, that behavior will persist longer when the affection stops than it would have under reliable affection. People are not exempt from the effect; Humphreys's original subjects were people.[2] What it does not support is any claim about the *bond*. The laboratory effect is measured in lever presses and eyeblinks after reinforcement stops; it says nothing about whether unpredictable affection produces stronger attachment than reliable affection, and no schedule experiment has tested that. The "trauma bonding" claims come from clinical and popular writing, not schedule research, and the evidence for them as operant findings is thin. Real relationships also contain much that a schedule analysis leaves out: relief when tension breaks ([negative reinforcement](https://operantconditioning.com/negative-reinforcement/)), fear, money, children, and the practical difficulty of leaving. What the term fairly says is that behavior reinforced unpredictably will be hard to stop, and will get worse before it fades. ## Common mistakes - **Starting intermittent.** Reinforcing "every few" sits before the puppy knows what a sit is. A behavior that has not been acquired cannot be maintained. - **Thinning too fast.** Jumping from every response to one in ten and concluding that the learner "lost motivation." The behavior did not fail; the schedule did. - **Never thinning.** The learner satiates, the behavior comes to depend on the trainer's presence, and it collapses the first time the reinforcer is unavailable. - **Confusing intermittent with inconsistent.** Intermittent reinforcement is a rule applied on purpose to an established behavior. Inconsistency is an accidental schedule applied to whatever happens to be occurring, which is often the behavior nobody wanted. - **Confusing intermittent with noncontingent.** A treat that arrives whether or not the dog sat is not a schedule of reinforcement for sitting; intermittent means some responses pay, not that reinforcers arrive at random. - **Reading the effect as "intermittent is stronger."** It is more [resistant to extinction](https://operantconditioning.com/glossary/#resistance-to-extinction), which is not the same thing. On Nevin's measures, richly reinforced behavior is more resistant to other disruptions, such as satiation and free reinforcers, than leanly reinforced behavior.[13] ## What the evidence does not show - **It does not show that intermittent reinforcement produces better learning.** Acquisition is slower, and a behavior started on a lean schedule may never be acquired. The advantage is durability once reinforcement stops. - **It does not show a single, unconditional law.** The size of the effect, and in some analyses its direction, depends on how resistance is measured — absolute responses across groups or proportional decline within a subject — and on the training history.[13][8] - **It does not show that intermittently reinforced behavior cannot be extinguished.** Extinction is slower, sometimes by a great deal; it is not prevented. Skinner's periodically reinforced rats stopped pressing in the end.[6] - **It does not identify an optimal ratio or a recipe for thinning.** The applied literature recommends small steps and stability at each step; the numbers depend on the learner, the behavior, and the reinforcer.[4][15] - **It does not say much about clinical populations directly.** Lerman and Iwata found few applied studies that manipulated the schedule before extinction; the clinical rule of thumb is an extrapolation from the laboratory.[16] - **It does not explain attachment.** Claims that unpredictable affection creates a distinctive bond are not schedule findings. ## Key takeaways - Continuous reinforcement follows every response with the reinforcer; intermittent reinforcement follows only some. The first teaches a new behavior fastest, and the second maintains an established one with fewer reinforcers and less satiation. - Behavior built on intermittent reinforcement persists far longer when reinforcement stops, even though it received fewer reinforcers: the partial reinforcement extinction effect, shown by Humphreys in 1939 and extended to rats by Mowrer and Jones in 1945. - The explanations — a harder-to-detect change, counterconditioned frustration (Amsel), responding conditioned to the memory of nonreward (Capaldi), a smaller stimulus change (Nevin) — agree that an organism which has never met an unreinforced response has no defense against the first one. - To thin a schedule, establish the behavior on continuous reinforcement, take the smallest step, go variable as soon as the schedule is intermittent, move only when performance is stable, and back off at the first sign of ratio strain. - Problem behavior is almost always reinforced intermittently by accident, which is why it resists extinction in clinics and homes; expect a slow extinction and a burst, and pair extinction with reinforcement of an alternative behavior. - The schedule concept describes the persistence of a behavior, not the strength of a bond; the popular use of "intermittent reinforcement" to explain relationships holds only as a claim about behavior. ### Check yourself **A new owner starts teaching a puppy to sit and gives a treat for roughly every third sit "so the behavior will be strong." Two weeks later the puppy still does not sit on cue. What went wrong?** The schedule was thinned before the behavior existed. On an intermittent schedule most sits go unreinforced, so the puppy rarely contacts the relation between sitting and the treat and the behavior never gets established. Intermittent reinforcement maintains a behavior that has already been acquired; it does not build one. Start on continuous reinforcement, and thin only once the sit is prompt and reliable. **A teacher has kept a class working quietly for a month by praising them after every good five-minute block. She decides that is too much praise and switches to praising once per lesson. Within a week the class is noisy again. What happened, and what should she do?** She stretched the schedule far too fast, from every block to roughly one in ten, and the behavior showed ratio strain and then broke down. The class did not lose motivation; the schedule did not hold. She should return to a requirement close to the one that worked, hold there until quiet work is stable, then thin in small, variable steps, keeping a brief acknowledgment on most blocks as a conditioned reinforcer while the tangible reinforcer thins. **A friend says her partner's unpredictable affection has "intermittently reinforced" her into staying. Is the schedule concept being used correctly?** Partly. The concept applies to behavior: if her efforts to reach out have been reinforced by affection only unpredictably, those efforts will persist longer than they would have under reliable affection, and there is no reason to think people are exempt. But the laboratory effect says nothing about attachment or the strength of a bond, and staying in a relationship involves negative reinforcement, fear, and practical constraints that a schedule analysis leaves out. The fair conclusion is only that the behavior will be hard to stop and will get worse before it fades. ## Frequently asked questions **What is the difference between continuous and intermittent reinforcement in simple terms?** With continuous reinforcement, every time the behavior happens it is reinforced: every sit earns a treat, every coin in the vending machine delivers a snack. With intermittent reinforcement, only some occurrences are reinforced: some sits earn a treat, some pulls on a slot machine pay out. Continuous reinforcement teaches a behavior fastest; intermittent reinforcement keeps a learned behavior going and makes it much harder to stop. **Continuous vs. intermittent reinforcement: which is better?** It depends on the stage. Continuous reinforcement is better for teaching a new behavior, because the learner cannot miss the connection between the behavior and its consequence. Intermittent reinforcement is better for maintaining a behavior that is already reliable: it needs fewer reinforcers, delays satiation, and produces behavior that survives when reinforcement is occasionally missed. The standard sequence is continuous first, then a gradual shift to intermittent. **What is an example of continuous reinforcement, and what is an example of intermittent reinforcement?** A light switch is continuous reinforcement: flip it and the light comes on every time. A treat for every sit during a puppy's first week of training is another. A slot machine is intermittent reinforcement: it pays after an unpredictable number of plays. So is fishing, where only some casts catch a fish, and a parent who gives in to whining some of the time. **What is the partial reinforcement extinction effect?** The finding that behavior reinforced only some of the time persists longer after reinforcement stops than behavior reinforced every time. Lloyd Humphreys first demonstrated it in 1939 with human eyeblink conditioning, and Mowrer and Jones extended it to rats pressing a bar in 1945. It is paradoxical because the intermittently reinforced behavior received fewer reinforcers yet is harder to extinguish. **Why does intermittent reinforcement make behavior harder to extinguish?** Several explanations are in use. The change to extinction is harder to detect, because unreinforced responses were always common. Amsel argued that the organism learns to keep responding through the frustration of nonreward; Capaldi that responding becomes conditioned to the memory of nonreward; Nevin that the shift from a lean schedule to none is a smaller change than from continuous reinforcement to none. All agree that an organism that has never met an unreinforced response has no defense against it. **How do you switch from continuous to intermittent reinforcement?** Gradually, and only after the behavior is reliable. Go from every response to two of every three, then every other, then vary the requirement around a slowly growing average. Move to the next step only when performance is stable at the current one. If responding slows, pauses, or falls apart, that is ratio strain: drop back to the last step that worked. Keep praise or another conditioned reinforcer on every response while the tangible reinforcer thins. **Is intermittent reinforcement the same as being inconsistent?** No. Intermittent reinforcement is a rule applied deliberately to a behavior that has already been learned, to maintain it economically. Inconsistency is an accidental schedule applied to whatever behavior happens to be occurring, which is often the behavior nobody wanted. A parent who gives in to a tantrum one time in five has put the tantrum on a lean variable schedule, which is the most persistent kind of behavior there is. **What does "intermittent reinforcement" mean in relationships?** In popular writing it describes a partner whose affection is unpredictable and the claim that the unpredictability itself keeps the other person attached. The operant concept supports only part of this: behavior that has been reinforced unpredictably, such as repeatedly reaching out, will persist longer when the affection stops. It says nothing about the strength of a bond, and the "trauma bonding" claims are not established findings from schedule research. ## References 1. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 2. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158. 3. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 4. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 5. Skinner, B. F. (1956). A case history in scientific method. *American Psychologist, 11*(5), 221–233. 6. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 7. Mowrer, O. H., & Jones, H. (1945). Habit strength as a function of the pattern of reinforcement. *Journal of Experimental Psychology, 35*(4), 293–311. 8. Mackintosh, N. J. (1974). *The Psychology of Animal Learning*. Academic Press. 9. Jenkins, H. M. (1962). Resistance to extinction when partial reinforcement is followed by regular reinforcement. *Journal of Experimental Psychology, 64*(5), 441–450. 10. Theios, J. (1962). The partial reinforcement effect sustained through blocks of continuous reinforcement. *Journal of Experimental Psychology, 64*(1), 1–6. 11. Amsel, A. (1958). The role of frustrative nonreward in noncontinuous reward situations. *Psychological Bulletin, 55*(2), 102–119. 12. Capaldi, E. J. (1966). Partial reinforcement: A hypothesis of sequential effects. *Psychological Review, 73*(5), 459–477. 13. Nevin, J. A. (1988). Behavioral momentum and the partial reinforcement effect. *Psychological Bulletin, 103*(1), 44–56. 14. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 15. Kazdin, A. E. (2013). *Behavior Modification in Applied Settings* (7th ed.). Waveland Press. 16. Lerman, D. C., & Iwata, B. A. (1996). Developing a technology for the use of operant extinction in clinical settings: An examination of basic and applied research. *Journal of Applied Behavior Analysis, 29*(3), 345–382. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): The four basic intermittent schedules, with a live simulator. - [Extinction](https://operantconditioning.com/extinction/): What happens when reinforcement stops: bursts, recovery, resurgence. - [Variable ratio schedule](https://operantconditioning.com/variable-ratio-schedule/): The schedule that produces the most persistent behavior of all. # Token Economy: Definition, Examples, and How to Set One Up > A token economy pays tokens for target behaviors, exchanged later for backup reinforcers. Definition, history, components, evidence, setup steps, mistakes. - Source: https://operantconditioning.com/token-economy/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applied · Reinforcement systems* Poker chips for chimpanzees, points on a hospital ward, stickers on a refrigerator, miles on an airline card: the token economy is the most widely used reinforcement system there is, and one of the best studied. Here is how it works, what six decades of evidence show, and why the hard part is the day the tokens stop. > **Definition** > > A token economy is a reinforcement system in which tokens — points, chips, stars, check marks — are delivered immediately after specified **target behaviors** and later exchanged for **backup reinforcers** the earner actually wants: privileges, activities, goods, time. The token has no value of its own. It works because it has been paired with many different reinforcers, which makes it a generalized conditioned reinforcer, the same class of stimulus as money.[1][2] > > The name became standard with Ayllon and Azrin's 1968 book, but any system that pays in a currency redeemable for something else — wages, loyalty points, a sticker chart — has the same structure.[1] **In brief** - A token economy delivers tokens contingent on target behaviors and lets the earner exchange them later for backup reinforcers. The token is a generalized conditioned reinforcer, which is what money is. - Every token economy has the same parts — target behaviors, tokens, backup reinforcers, an exchange rate, a schedule of exchange, and optionally fines — and most failures trace to one of them. - The evidence that token economies change behavior while they run is strong across wards, group homes, classrooms, and addiction treatment. The evidence that gains survive the program is weak unless the fade is planned from the start. ## How a token economy works Reinforcement works best when the consequence arrives within seconds, and most consequences people care about — a day pass off a ward, a paycheck — cannot be delivered that fast. A token solves the timing problem. It can be handed over the instant the behavior occurs, and it carries the value of whatever it will later buy, the way a clicker bridges the gap between a dog's sit and the treat.[1] What gives a token its value is its history. A poker chip is nothing to a chimpanzee until it has gone into a vending machine and produced a grape. Paired with one backup reinforcer, the token becomes a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer); paired with *many*, it becomes a [generalized conditioned reinforcer](https://operantconditioning.com/positive-reinforcement/#generalized-reinforcers), and that is what makes a token economy work. A single reinforcer satiates; a token that buys candy, or free time, or a phone call home, is worth something in almost any state the person is in. Skinner made the same point about money.[2] [Primary and secondary reinforcers ›](https://operantconditioning.com/primary-and-secondary-reinforcers/) Behavior analysts describe a token economy as three interlocking [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/).[3] The **token-production schedule** says how much behavior earns a token: one point per completed problem. The **exchange-production schedule** says when exchange becomes available: at the end of the lesson, after twenty tokens. The **token-exchange schedule** is the price list. Procedure and process A token economy is a *procedure*: rules about what earns tokens and what tokens buy. Whether reinforcement, the *process*, has occurred is a question about behavior. If the target behaviors do not increase, the tokens are not functioning as reinforcers, however carefully the chart was designed — usually because the backup reinforcers were chosen by the staff rather than by the earners.[1] ## The components of a token economy Every token economy has the same parts. Ivy and colleagues found that many published programs leave one or more unspecified — most often the schedules of token production and exchange, and how backup reinforcers were identified — which makes them hard to replicate.[4] | Component | What it specifies | Classroom | Home sticker chart | | --- | --- | --- | --- | | **Target behaviors** | Countable behaviors that earn tokens | "Starts work within one minute of the bell"; "hand raised before speaking" | "Teeth brushed before 7:30"; "shoes by the door" | | **Tokens** | The currency: instant to deliver, hard to counterfeit, visible to the earner | Points on a card; plastic chips; a stamp | Stickers; marbles in a jar | | **Backup reinforcers** | What tokens buy; a menu drawn from what the earner chooses when free | Free-choice time; computer time; a homework pass; lunch with the teacher | Choosing dinner; ten extra minutes before bed; a game with a parent | | **Token-production schedule** | How much behavior earns one token | One point per completed problem; one chip per ten minutes on task | One sticker per completed routine | | **Exchange rate** | The prices: how many tokens each backup reinforcer costs | 5 points for a homework pass; 30 for choosing the class music | 3 stickers for an extra story; 15 for a trip to the park | | **Schedule of exchange** | When and how often tokens can be spent | End of each lesson at first, then end of day, then end of week | Same day at first, then twice a week | | **Response cost (fine)** | Optional loss of tokens for specified behaviors | 2 points for leaving the seat without permission | Usually best omitted for young children | The fine runs in the opposite direction. Removing tokens contingent on a behavior is [negative punishment](https://operantconditioning.com/negative-punishment/), which behavior analysts call [response cost](https://operantconditioning.com/glossary/#response-cost). Token loss reliably suppresses behavior in the laboratory,[3] and Achievement Place used fines throughout.[5] But a fine only works while the person has tokens to lose. ## History: from poker chips to hospital wards ### Chimpanzees and poker chips The first token economies had no name and no human subjects. In the 1930s John Wolfe taught chimpanzees to insert poker chips into a vending machine, the "Chimp-o-mat," that delivered a grape per chip, and then to earn chips by lifting a weighted handle. The animals worked for chips much as for grapes, chose chips that bought two grapes over chips that bought one, ignored brass slugs that bought nothing, and kept working when the machine was unavailable, though less the longer the exchange was delayed.[6] John Cowles then showed that chips could teach as well as maintain: chimpanzees learned new discriminations when the only immediate consequence of a correct choice was a token to be spent later, and performed almost as well as for food itself.[7] ### The ward at Anna State Hospital Teodoro Ayllon and Nathan Azrin built the first systematic token economy for people in the early 1960s, on a ward of Anna State Hospital in Illinois whose patients were women, most hospitalized for years with diagnoses of schizophrenia. Patients earned tokens for self-care and ward jobs — serving meals, cleaning, laundry — and spent them on privacy, walks on the grounds, trips into town, time with staff, and commissary items.[8] Six experiments tested whether the tokens were doing the work. When the tokens were moved from the job each patient preferred to the one she had avoided, the patients moved with them; when tokens were handed out daily regardless of work, the ward's total labor fell from about 45 hours a day to about one, and recovered when the contingency was restored.[8] Their 1968 book laid out the method as rules, two of which every later program has had to relearn: let people sample a backup reinforcer before asking them to work for it, and reinforce only behaviors that will go on being reinforced once the program ends.[1] ### Achievement Place and the social-learning program In 1967 a house opened in Lawrence, Kansas, for boys the courts had labeled "pre-delinquent," run by a married couple as teaching-parents. Points were earned for chores, homework, punctuality, and appropriate speech; lost for aggressive statements, poor grammar, and lateness; and exchanged for the following week's privileges — allowance, snacks, television, permission to go downtown. Elery Phillips's 1968 report used reversal designs to show that each behavior rose or fell with the points attached to it.[5] The Teaching-Family Model that grew from it was replicated widely, and its follow-up evaluation, discussed below, is the clearest demonstration of the token economy's limits. The most rigorous test remains Gordon Paul and Robert Lentz's multi-year comparison of three treatments for long-term psychiatric inpatients: a social-learning program built around a token economy and skills training, milieu therapy, and standard hospital care. The social-learning program produced the largest improvements in functioning and the most successful community discharges, at the lowest cost.[9] Classrooms adopted tokens almost as soon as the ward did; by 1972 the published programs spanned schools, wards, prisons, and homes.[10] [Tokens and the Good Behavior Game in the classroom ›](https://operantconditioning.com/classroom/) ## What the evidence shows ### The applied reviews Kazdin and Bootzin's 1972 review, and Kazdin's follow-up a decade later, judged the effect of token economies *while they were running* to be well established across psychiatric wards, institutions for people with intellectual disabilities, delinquency programs, and classrooms. They also named the problems that have not gone away: behavior fell when tokens were withdrawn, did not transfer to settings without tokens, some individuals never responded, and staff often stopped delivering tokens consistently.[10][11] By 1982 the psychiatric programs were in decline, as deinstitutionalization emptied the wards and court rulings made meals, beds, and ground privileges rights rather than things to be earned.[11] Two recent reviews of the classroom literature reach a mixed verdict. Maggin and colleagues applied the What Works Clearinghouse standards to studies with students with challenging behavior and found the effects generally positive but the studies too weak in design and reporting to qualify as evidence-based under those standards.[12] Soares and colleagues' meta-analysis of single-case classroom studies found positive effects in most cases, varying with the students and how the systems were built.[13] The gap in the record is not evidence that the programs failed; it is evidence that too many were never described well enough to be copied.[4] ### The laboratory evidence Timothy Hackenberg's reviews of token reinforcement in the laboratory — pigeons earning lights, rats earning marbles, chimpanzees earning chips — establish what a practitioner needs to know.[3][14] Tokens function as reinforcers: behavior that produces them increases, and behavior that costs them decreases. Their power depends on the exchange schedule: responding is weakest when exchange is far away and rises as it approaches, and tokens that cannot be exchanged lose their effect. Tokens also serve as [discriminative stimuli](https://operantconditioning.com/discriminative-stimulus/) marking distance from exchange, which is one reason a visible token board works where an invisible tally does not. And animals given the choice will often accumulate tokens before exchanging them, the laboratory version of saving.[3] ## Where token economies are used | Setting | Tokens | Target behaviors | Backup reinforcers | | --- | --- | --- | --- | | Classroom | Points, chips, marbles in a class jar | On task, hand raised, work completed, quiet transitions | Free time, privileges, small prizes, class rewards | | Applied behavior analysis programs | A token board with five hook-and-loop tokens | Task completion, requesting, tolerating a demand | A preferred toy, activity, or break | | Contingency management for substance use | Vouchers or prize draws | Drug-negative urine samples, attendance, medication taken | Goods and services bought with vouchers; prizes | | Home | Stickers, marbles, points | Morning routine, chores, homework started | Same-day privileges, time with a parent, small purchases | | Apps and loyalty programs | Points, miles, streaks, badges | Purchases, check-ins, logged workouts, daily use | Discounts, free flights, status, unlocked features | The clinical case with the strongest trial evidence is [contingency management](https://operantconditioning.com/glossary/#contingency-management): people in treatment for substance use earn vouchers or prize draws for drug-negative urine samples, with the value escalating across consecutive negatives and resetting after a positive — a token economy with a laboratory test as the target behavior. A meta-analysis of the controlled trials found moderate, reliable effects, strongest for stimulants and opioids, that shrink after the incentives end.[15] [All applications, with the evidence rated ›](https://operantconditioning.com/applications/) The last row is the one most people live in. A stamp card, an airline's miles, and a language app's streak all have the token economy's structure, and whether they change behavior depends on the same variables: whether the token arrives immediately, buys something the person wants, and can be spent soon enough to matter.[14] [Sticker charts and token systems at home ›](https://operantconditioning.com/parenting/) ## How to set up a token economy The steps follow Ayllon and Azrin's rules and the maintenance strategies Kazdin's reviews recommended.[1][10] 1. **Choose three to five target behaviors, stated as things to do.** "Starts the worksheet within a minute," not "stays focused." Prefer behaviors the world will eventually reinforce on its own — Ayllon and Azrin's relevance-of-behavior rule — because those will survive the program. 2. **Pick a token you can deliver in a second.** A chip, a tally, a sticker. For young children, keep the token tangible and the board in view. 3. **Build the menu from observation, not guesses.** Watch what the person does when free to choose — the [Premack principle](https://operantconditioning.com/premack-principle/) in use — and let people sample items they have never had. Price it so a single good hour buys something, and refresh it, because backup reinforcers satiate. 4. **Deliver immediately, with specific praise every time.** "You started on your own — one point." The praise inherits the token's power and keeps working when the tokens are gone. 5. **Exchange early and often at first.** Same day, or same lesson for young children; laboratory tokens far from exchange are weak ones.[3] Lengthen the interval only after the behavior is steady. 6. **Use fines rarely, if at all.** Start without them. If you add response cost, define the fined behaviors in advance, keep each fine small relative to a day's earnings, and never let a balance go below zero. 7. **Count before you start, review weekly, and plan the fade from day one.** A baseline is the only way to know whether the tokens did anything. Then thin the token-production schedule, space the exchanges, raise prices, shift from tokens to praise to natural consequences, and hand the counting to the earner. ## Common mistakes - **Tokens with no backup value.** A chart with nothing to buy, or a menu of things the earner never chose, is a list of tallies; tokens that cannot be exchanged stop working.[3] - **Exchange delayed too long.** A weekly prize for a six-year-old asks the child to work for something five days away, and behavior early in a long exchange cycle is the weakest in the system.[3] - **Fines that exceed earnings.** A person at zero has nothing to work for and nothing to lose. When fines dominate, the token economy has become a punishment procedure with a bookkeeping layer, and the rational response is to stop participating. - **Never fading.** If the system is identical in June and September, the behavior belongs to the system and disappears with the tokens.[10] - **Inconsistent delivery.** Staff who stop handing out tokens are a failure mode Kazdin's reviews found repeatedly.[10][11] A system that depends on adults remembering needs prompts for the adults. - **Paying for behavior that was already happening for its own sake.** Tokens are for building behavior that is not occurring; see the next section. ## The intrinsic-motivation debate The standing objection to any token system is that paying people for a behavior makes them stop doing it for its own sake. The evidence is a long-running dispute between two meta-analyses of largely the same experiments. Deci, Koestner, and Ryan's analysis of 128 experiments found that *expected, tangible* rewards for an activity people already found interesting reduced their later free-choice engagement with it and their self-reported interest, most strongly when the reward was given merely for engaging in or completing the task. Verbal rewards — praise — increased intrinsic motivation, and unexpected rewards did not undermine it.[16] Cameron and Pierce, five years earlier, concluded that rewards do not, on the whole, reduce intrinsic motivation. The one reliable negative effect they found was the same one: expected tangible rewards for simply doing an interesting task, regardless of how well, reduced free-choice time afterward. Rewards tied to performance did not.[17] Read side by side, the two agree on the boundary of the effect and disagree on how much it matters: on how large and general it is, which studies belong in the analysis, and which outcome measures count. For a token economy the practical conclusions follow from either reading. Use tokens for behavior that is not happening, not for activities the person already does for pleasure. Pay for quality or completion rather than mere engagement, pair every token with specific praise, and fade the tokens as the behavior meets the consequences the world already provides. ## What the evidence does not show **It does not show that gains survive the program.** This has been the weak point since the first review.[10][11] The clearest case is Achievement Place. When Teaching-Family group homes were compared with other community homes for juvenile offenders, the youths in Teaching-Family homes had fewer recorded offenses *during* treatment; in the year after they left, the difference was gone.[18] Contingency management shows the same shape: effects that shrink after the vouchers stop.[15] Behavior follows the contingencies in force; when the tokens end and nothing replaces them, the behavior is on [extinction](https://operantconditioning.com/extinction/). Maintenance has to be engineered: by fading, by pairing tokens with praise, by training where the behavior must eventually occur, and by choosing target behaviors the natural environment will pay for.[1][10] **It does not show that tokens teach skills.** Ayllon and Azrin subtitled their book "a motivational system," and that is what a token economy is. It increases behavior already in the repertoire; a student who cannot do the problems will not do them for points. Skills need [shaping](https://operantconditioning.com/shaping/) and instruction; the token economy pays for the practice. **It does not show large effects in ordinary classrooms.** Most of the classroom evidence comes from single-case studies of students with challenging behavior, and the best of it is methodologically modest.[12][13] The strongest controlled trial concerns long-term psychiatric inpatients in a hospital system that no longer exists in that form.[9] And some people do not respond to any menu on offer; the reviews have said so from the beginning.[10] What the evidence does show is narrower and still useful: while a well-built token economy is running, target behaviors rise, and the components that make it work are the ones the chimpanzees demonstrated in 1936. ## Key takeaways - A token economy delivers tokens immediately after target behaviors and lets the earner exchange them later for backup reinforcers. The token works because it is a generalized conditioned reinforcer, like money. - The components are target behaviors, tokens, backup reinforcers, an exchange rate, a schedule of exchange, and optionally response cost. Most failures trace to a menu nobody chose, an exchange too far away, fines that exceed earnings, or a fade that never happened. - The lineage runs from Wolfe and Cowles's chimpanzees working for poker chips, through Ayllon and Azrin's ward and Achievement Place, to Paul and Lentz's social-learning program, still the most rigorous test. - Expected tangible rewards for an already interesting activity can reduce later free-choice engagement; praise does not. Build token economies for behavior that is not happening, pair tokens with praise, and fade them. - Gains do not survive the program by default. Maintenance has to be engineered: fade the tokens, pair them with praise, train where the behavior must occur, and choose behaviors the world will go on reinforcing. ### Check yourself **A teacher runs a point system all year. Points are tallied on a chart the students cannot see and exchanged for prizes on the last Friday of each month. On-task behavior has not changed. Which components are most likely at fault?** The schedule of exchange and the visibility of the token. A month between exchanges puts most of the behavior far from the reinforcer, where laboratory tokens are weakest, and a tally the students cannot see gives them no signal of how close they are. Move exchange to the same day, put the tally where the student can watch it, and check that the prizes are things the students actually choose when free. **A group home fines residents heavily for rule violations. Several residents are in debt and have stopped doing chores at all. What has gone wrong?** Fines have exceeded earnings, so the token economy has become a punishment procedure. A resident below zero has nothing to lose and nothing to work toward, and the rational response is to stop participating. Fines should be rare, small relative to a day's earnings, and never allowed to take a balance below zero; the system should run on earning, not on loss. **A residential program for adolescents shows excellent behavior while youths are enrolled, but a year after discharge they are no different from youths who went elsewhere. Does this mean the program did not work?** It means the program changed behavior while its contingencies were in force and did not arrange for anything to replace them, which is what the Achievement Place evaluation found. A token economy is a set of contingencies; when they end without a fade, without pairing to praise, and without training in the settings where the behavior must occur, the behavior is on extinction. The procedure worked; maintenance was never engineered. ## Frequently asked questions **What is a token economy in simple terms?** A system in which someone earns tokens, such as points, chips, or stickers, immediately after doing specific behaviors, and later trades the tokens for things they actually want, such as privileges, activities, or small prizes. The token works because it can be exchanged for many different things, which makes it a kind of money. **What is an example of a token economy?** A classroom in which students earn a point each time they start work within a minute of the bell and can spend five points on a homework pass or thirty on choosing the class music. Other examples: a sticker chart that buys an extra bedtime story, a psychiatric ward where tokens for self-care buy walks on the grounds, and vouchers earned for drug-negative urine tests in addiction treatment. **What is the difference between a token economy and contingency management?** Contingency management is a token economy used in medical and addiction treatment. The tokens are vouchers or prize draws, the target behavior is usually an objectively verified one such as a drug-negative urine sample or attendance, and the value often escalates with consecutive successes. The structure, the strengths, and the weaknesses are the same as in any other token economy. **Is money a token economy?** Money is the everyday example of a generalized conditioned reinforcer, which is what a token is. Wages are delivered contingent on work and exchanged later for almost anything, so a paycheck has the structure of a token economy with a long exchange delay. Skinner used money as his standard illustration of a generalized reinforcer. **Who invented the token economy?** Teodoro Ayllon and Nathan Azrin built the first systematic one at Anna State Hospital in Illinois in the early 1960s and published the method in 1965 and, as a book, in 1968. The idea is older: in the 1930s John Wolfe and John Cowles showed that chimpanzees would work for poker chips exchangeable for grapes, and would learn new tasks for them. **Do token economies work in the classroom?** While they are running, yes: reviews and meta-analyses of classroom studies find positive effects in most cases, particularly for students with challenging behavior. The caveats are that the studies are mostly single-case designs of uneven quality, that many programs are described too poorly to copy, and that gains fade when the tokens are withdrawn unless the system has been faded toward praise and natural consequences. **What is response cost in a token economy?** A fine: the removal of tokens contingent on a specified behavior. Because a reinforcer is taken away and the behavior decreases, it is negative punishment. It works only while the earner has tokens to lose, so fines should be rare, small relative to earnings, defined in advance, and never allowed to push a balance below zero. **How do you fade a token economy?** Gradually and on a plan. Deliver tokens for fewer occurrences of the behavior, lengthen the time between exchanges, raise the prices, pair every token with specific praise from the start so the praise takes over, hand the counting to the earner, and choose target behaviors that the classroom, workplace, or family will keep reinforcing on their own. Measure the behavior throughout so you can slow the fade if it drops. ## References 1. Ayllon, T., & Azrin, N. H. (1968). *The Token Economy: A Motivational System for Therapy and Rehabilitation*. Appleton-Century-Crofts. 2. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 3. Hackenberg, T. D. (2009). Token reinforcement: A review and analysis. *Journal of the Experimental Analysis of Behavior, 91*(2), 257–286. 4. Ivy, J. W., Meindl, J. N., Overley, E., & Robson, K. M. (2017). Token economy: A systematic review of procedural descriptions. *Behavior Modification, 41*(5), 708–737. 5. Phillips, E. L. (1968). Achievement Place: Token reinforcement procedures in a home-style rehabilitation setting for "pre-delinquent" boys. *Journal of Applied Behavior Analysis, 1*(3), 213–223. 6. Wolfe, J. B. (1936). Effectiveness of token-rewards for chimpanzees. *Comparative Psychology Monographs, 12*(5), 1–72. 7. Cowles, J. T. (1937). Food-tokens as incentives for learning by chimpanzees. *Comparative Psychology Monographs, 14*(5), 1–96. 8. Ayllon, T., & Azrin, N. H. (1965). The measurement and reinforcement of behavior of psychotics. *Journal of the Experimental Analysis of Behavior, 8*(6), 357–383. 9. Paul, G. L., & Lentz, R. J. (1977). *Psychosocial Treatment of Chronic Mental Patients: Milieu versus Social-Learning Programs*. Harvard University Press. 10. Kazdin, A. E., & Bootzin, R. R. (1972). The token economy: An evaluative review. *Journal of Applied Behavior Analysis, 5*(3), 343–372. 11. Kazdin, A. E. (1982). The token economy: A decade later. *Journal of Applied Behavior Analysis, 15*(3), 431–445. 12. Maggin, D. M., Chafouleas, S. M., Goddard, K. M., & Johnson, A. H. (2011). A systematic evaluation of token economies as a classroom management tool for students with challenging behavior. *Journal of School Psychology, 49*(5), 529–554. 13. Soares, D. A., Harrison, J. R., Vannest, K. J., & McClelland, S. S. (2016). Effect size for token economy use in contemporary classroom settings: A meta-analytic review of single-case research. *School Psychology Review, 45*(4), 379–399. 14. Hackenberg, T. D. (2018). Token reinforcement: Translational research and application. *Journal of Applied Behavior Analysis, 51*(2), 393–435. 15. Prendergast, M., Podus, D., Finney, J., Greenwell, L., & Roll, J. (2006). Contingency management for treatment of substance use disorders: A meta-analysis. *Addiction, 101*(11), 1546–1560. 16. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 17. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423. 18. Kirigin, K. A., Braukmann, C. J., Atwater, J. D., & Wolf, M. M. (1982). An evaluation of Teaching-Family (Achievement Place) group homes for juvenile offenders. *Journal of Applied Behavior Analysis, 15*(1), 1–16. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Operant conditioning in the classroom](https://operantconditioning.com/classroom/): Praise, token economies, the Good Behavior Game, and what the evidence says. - [Primary and secondary reinforcers](https://operantconditioning.com/primary-and-secondary-reinforcers/): How a neutral stimulus becomes a reinforcer, and why generalized ones work everywhere. - [Operant conditioning in parenting](https://operantconditioning.com/parenting/): Sticker charts, time-out, tantrums, and what the evidence says. # Differential Reinforcement: Definition, Types & Examples > Differential reinforcement reinforces one behavior and withholds reinforcement from another. DRA, DRI, DRO, DRL, DRH, function, FCT, evidence, how to run it. - Source: https://operantconditioning.com/differential-reinforcement/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Procedures · Reinforcement-based reduction* Most problem behavior is not a failure of discipline; it is a behavior that works. Differential reinforcement is the family of procedures that makes something else work better — an alternative behavior, or simply the absence of the problem — while the problem behavior stops paying. It is the reason modern behavior analysis reaches for reinforcement before punishment. > **Definition** > > Differential reinforcement is the procedure of reinforcing one response, or one class of responses, while withholding reinforcement from another. The reinforced class becomes more frequent; the unreinforced class is placed on [extinction](https://operantconditioning.com/extinction/) and declines. Every use of the term has both halves — reinforcement of something and extinction of something else — running at the same time.[1] > > It is the mechanism inside [shaping](https://operantconditioning.com/shaping/) and discrimination training, and in applied work it names a family of procedures — DRA, DRI, DRO, DRL, and DRH — that reduce a problem behavior by paying for something other than it, rather than by punishing it.[2] **In brief** - Differential reinforcement always has two halves: one response class is reinforced and another is placed on extinction. Drop either half and the procedure is not differential. - The five applied procedures differ in what earns the reinforcer: an alternative behavior (DRA), an incompatible behavior (DRI), the absence of the behavior for an interval (DRO), a low rate (DRL), or a high rate (DRH). - It works when a functional assessment has identified what reinforces the problem behavior, so that the same reinforcer can be withheld for the problem and delivered for the alternative. ## How differential reinforcement works [Reinforcement](https://operantconditioning.com/positive-reinforcement/) strengthens whatever it follows, and left to itself it is not selective: if the pellet comes after hard presses and soft presses alike, both persist. Differential reinforcement adds a criterion. Responses that meet it are reinforced, responses that do not are not, and because the two classes now have different consequences their frequencies diverge. Skinner described the procedure in *The Behavior of Organisms* under two headings. In **differentiation** the criterion is a property of the response itself: reinforce only lever presses above a certain force and the whole distribution of forces shifts upward. In **discrimination** the criterion is the stimulus present when the response occurs: reinforce presses when a light is on and never when it is off, and the rat comes to press only in the light.[1] Both are the same operation — pay for one class, not the other — applied to different dimensions of behavior. Three central procedures are differential reinforcement under other names. [Shaping](https://operantconditioning.com/shaping/) is differential reinforcement of successive approximations, with the criterion moved step by step. [Stimulus control](https://operantconditioning.com/stimulus-control/) is built by differential reinforcement with respect to a [discriminative stimulus](https://operantconditioning.com/discriminative-stimulus/). And the reduction procedures below are differential reinforcement with respect to a problem behavior: something else is reinforced, and the problem behavior is placed on extinction.[2] Procedure and process "Differential reinforcement" names what you do: deliver a consequence after one class of responses and withhold it after another. What happens to behavior — one class rising, the other declining — is the process, and the only test that the procedure worked. If the "reinforced" behavior does not increase, the consequence was not a reinforcer for that person; if what you withheld was not the reinforcer maintaining the problem behavior, nothing was put on extinction. ## The five procedures: DRA, DRI, DRO, DRL, and DRH Applied behavior analysis uses differential reinforcement mainly to reduce behavior without [punishment](https://operantconditioning.com/punishment/). The five procedures differ in what has to happen for the reinforcer to arrive.[2] | Procedure | What is reinforced | What is withheld | Example | Best for | | --- | --- | --- | --- | --- | | **DRA** — alternative behavior | A specific appropriate behavior that produces what the problem behavior produced | The reinforcer for the problem behavior | A child who screams for help is taught to tap the adult's arm; tapping is answered, screaming is not | Behavior with a clear function that an acceptable behavior can serve | | **DRI** — incompatible behavior | An alternative that physically cannot occur at the same time as the problem behavior | The reinforcer for the problem behavior | A dog that jumps on guests is greeted only while sitting | Behavior with an obvious physical opposite | | **DRO** — other behavior | The absence of the problem behavior for a set interval | The reinforcer for the problem behavior; a response usually resets the interval | A point for every five minutes without calling out | Behavior with no obvious alternative | | **DRL** — low rates | Responding at or below a limit, or spaced far enough apart | Reinforcement when responding is too frequent | A class earns free time if there are five or fewer talk-outs in a period | Behavior that is fine in moderation | | **DRH** — high rates | Responding at or above a set rate | Reinforcement when responding is too slow | A break for finishing twenty math facts in a minute | Fluency in a skill that is accurate but slow | ### DRA and DRI DRA is the workhorse. The alternative can be anything the person can do, or can be taught, that gets them what the problem behavior got them: asking instead of grabbing, raising a hand instead of shouting, requesting a break instead of shoving the worksheet away. DRI adds one constraint — the alternative cannot coexist with the problem behavior — which makes the extinction half easier to keep, since a sitting dog is not jumping. Not every behavior has a useful opposite.[2] ### DRO DRO is the odd one out, because nothing in particular is reinforced. The reinforcer is delivered when an interval passes without the target behavior, whatever else the person was doing. Reynolds coined the term in a 1961 pigeon experiment on behavioral contrast, in which one schedule delivered food only when the bird had refrained from pecking for a set time.[3] In interval DRO the behavior must be absent for the whole interval; in momentary DRO only at the instant the interval ends. The whole-interval version is the stronger treatment; momentary DRO is easier to run and useful for maintenance. A response usually resets the clock.[2] ### DRL and DRH DRL and DRH act on rate rather than on which behavior occurs. DRL reinforces responding only when it is infrequent: either the total for a session is at or below a limit (full-session DRL), or each response is reinforced only if enough time has passed since the last (spaced-responding DRL). It is the tool for behavior that should be reduced, not eliminated. Deitz and Repp's 1973 classroom study is the standard example: a boy in a special-education class earned candy when talk-outs in a period stayed at five or fewer, a whole class earned it as a group on the same terms, and high-school students earned a free period for keeping off-topic remarks under a limit lowered in stages toward zero. In each case the behavior fell to the criterion.[4] DRH is the mirror image, reinforcing only when responding is fast enough, and belongs to fluency training rather than behavior reduction.[2] Both are [schedules](https://operantconditioning.com/schedules-of-reinforcement/) before they are treatments. ## Why the function of the behavior matters The extinction half only works if the reinforcer you withhold is the one maintaining the behavior, and that is not something you can see by looking. The same tantrum can be maintained by attention, by escape from a demand, by access to an item, or by the sensation it produces, and each calls for a different thing to be withheld and paid.[5] Iwata and colleagues' [functional analysis](https://operantconditioning.com/glossary/#functional-analysis), first published in 1982, made the function testable. Children who injured themselves were observed under a series of conditions — an adult who responded to self-injury with attention, an adult who withdrew task demands when it occurred, a room with nothing to do, and a play condition as a control — and for most of them the behavior was reliably higher in one condition than the others.[5] Vollmer and Iwata's 1992 review drew the consequence for treatment: differential reinforcement should use the *functional* reinforcer, the one identified by the analysis, rather than an arbitrary one that merely seems appealing. Withholding it for the problem behavior is extinction; delivering it for the alternative gives the alternative the job the problem behavior used to do.[6] | Function | What is withheld | What the alternative earns | Example alternative | | --- | --- | --- | --- | | Attention | Reactions to the problem behavior | Attention, promptly | Tapping an arm; a raised hand | | Escape from demands | Removal of the task | A break, help, or an easier step | "Break, please"; "help" | | Access to items or activities | The item | The item | Asking; pointing; a picture card | | Automatic (sensory) | Difficult; the behavior produces its own reinforcer | A matched sensory alternative | Chewing a safe object instead of a sleeve | A 1993 study by Vollmer, Iwata, and colleagues shows the logic at full strength. Three women whose self-injury had been shown by functional analysis to be attention-maintained were treated with DRO using attention as the reinforcer, and with noncontingent attention delivered on a time schedule regardless of behavior. Both reduced self-injury. The authors noted an advantage of the noncontingent schedule worth remembering when designing a DRO: attention arrived densely from the start, whereas DRO begins with stretches in which the person earns nothing, which is the condition that produces bursts.[7] The escape case is the one that catches parents and teachers: ignoring a tantrum maintained by getting out of a task is not extinction, because the task still went away. [Escape-maintained behavior ›](https://operantconditioning.com/negative-reinforcement/) ## Functional communication training: DRA with a request The most studied form of DRA teaches the person to *ask* for what the problem behavior produced. Carr and Durand's 1985 study established the method. Four children with developmental disabilities were observed while task difficulty and adult attention were varied; for some, problem behavior rose when attention was scarce, for others when tasks were hard. Each child was then taught a phrase — "Am I doing good work?" for attention, "I don't understand" for help — and the phrase was answered whenever it was used. Problem behavior fell when the phrase matched the child's function and did not fall when the child was taught the other one.[8] The reinforcer, not the words, was the active ingredient. Tiger, Hanley, and Bruzek's practical guide lays out the modern package: a functional analysis; a communicative response chosen for the learner — vocal, signed, a card, a switch — and easy enough to beat the problem behavior; teaching by prompting and reinforcing it every time; and then, once the problem behavior is low, thinning the schedule with signals for when requests will and will not be honored, or with gradually longer delays.[9] Thinning is where FCT most often fails: a request refused too often stops being worth making, and the problem behavior returns. Whether extinction is required has been tested directly. In a summary of 21 inpatient cases, functional communication training without extinction did not reduce problem behavior; adding extinction produced large reductions in many cases but not all, and the remainder required punishment components before the behavior came down.[10] Teaching the request is necessary; it is not sufficient while the scream still works. ## What the research shows DRA has the strongest evidence base of the family. Petscher, Rey, and Bailey's 2009 review concluded that DRA has substantial empirical support as a treatment for problem behavior in people with developmental disabilities, and observed that most studies combined it with extinction, so the effect of reinforcing an alternative *without* withholding reinforcement for the problem behavior is much less well established.[11] A methodological review of the adult literature reached a similar verdict with a caution: the procedures worked in most studies, but the studies were often small and short, so durability in adults is less well documented than the initial reductions.[12] DRO's evidence is broad but its mechanism is unsettled. Jessel and Ingvarsson's 2016 summary of recent DRO research notes that the procedure may reduce behavior less by reinforcing "other behavior" — which is not measured, and need not increase in any specific form — than through the extinction and the response-contingent postponement of reinforcement built into it.[13] For practice this matters little: DRO reduces behavior. For understanding, it means the name is partly a misnomer. Two laboratory findings are cautions. Reynolds' pigeons showed **behavioral contrast**: cutting reinforcement for pecking in one component of a multiple schedule raised the rate of pecking in the other, where reinforcement was unchanged.[3] Put a behavior on extinction in one setting and it may rise in another. And research on [behavioral momentum](https://operantconditioning.com/glossary/#behavioral-momentum) shows that adding reinforcement to a situation — which DRA does — makes all behavior there more resistant to change, including the problem behavior when its reinforcement is later withheld.[14] Neither is a reason not to use the procedure; both are reasons to run it everywhere and expect persistence. ## Examples across settings In every row the two halves are named: what the alternative earns, and what the problem behavior no longer earns. The function column is a guess; in real life, check it first. | Setting | Problem behavior | Likely function | Procedure | | --- | --- | --- | --- | | Classroom | Calling out | Teacher attention | DRA: raised hands are answered promptly; called-out answers get no response | | Parenting | Whining for snacks | Access to the item | DRA: a plain request in a normal voice gets the snack, or a clear answer; whining never does | | Parenting | Screaming during homework | Escape from the task | FCT: "break, please" earns a two-minute break; screaming does not end the homework | | [Dog training](https://operantconditioning.com/dog-training/) | Jumping on guests | Attention from the guest | DRI: sitting is greeted and petted; jumping is turned away from | | [Self-management](https://operantconditioning.com/habits/) | Checking your phone mid-task | Escape from boredom or difficulty | DRA: a short walk or a glass of water for the same relief, with notifications off so checking pays less | ## How to run differential reinforcement 1. **Define the problem behavior and count it.** Observable, measurable, with a baseline rate. A DRO interval or a DRL limit depends on how often the behavior occurs now. 2. **Find the function.** A functional assessment, formal or informal: what arrives right after the behavior, and what goes away when it starts? The answer tells you what to withhold and what to pay with.[5][6] 3. **Choose the alternative.** Something the person can do or can be taught quickly, that produces the same reinforcer, and that takes less effort than the problem behavior. If nothing fits, use DRO; if the behavior is acceptable in moderation, use DRL. 4. **Start the schedule dense.** Reinforce the alternative every time at first. For DRO, set the first interval a little shorter than the average gap between problem behaviors at baseline, so the person contacts the reinforcer early and often.[2] 5. **Withhold the maintaining reinforcer for the problem behavior.** This is extinction, and it has to be as consistent as the reinforcement. If the function is escape, the demand stays; if attention, the reaction stops; if an item, the item is not delivered. 6. **Plan for the burst.** Expect the problem behavior to increase briefly when it stops working. The [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) is less likely when extinction is combined with reinforcement of an alternative, but everyone involved should know it may come and agree not to give in.[15] 7. **Thin the schedule, then generalize.** Once the problem behavior is low, lengthen DRO intervals, tighten DRL limits, and move the alternative onto intermittent reinforcement with clear signals for when it will pay.[9] Then run the procedure everywhere the behavior occurs, with everyone it occurs with. ## Common mistakes - **Paying the alternative less than the problem behavior was paid.** If screaming brought a parent to the room in three seconds and asking gets "just a minute," screaming wins. The [matching law](https://operantconditioning.com/matching-law/) predicts it: behavior goes where reinforcement is richer, faster, and more reliable.[9] - **DRO intervals that are too long.** An interval the person never completes delivers no reinforcement, and a procedure that delivers no reinforcement is extinction alone. Start below the baseline gap and lengthen from there.[2] - **Forgetting the extinction half.** Reinforcing the alternative while the problem behavior still works gives the person two ways to earn the same thing, and the old one is better practiced. The evidence for DRA is largely evidence for DRA with extinction.[10][11] - **Withholding the wrong reinforcer.** Ignoring escape-maintained behavior is not extinction; it delivers the escape. Only the reinforcer that matches the function can be withheld.[6] - **An alternative that is harder than the problem, or thinned too fast.** A five-word sentence competes badly with a scream, and moving from every time to occasionally in one step puts the alternative on extinction. Start with the easiest response that does the job, and thin in small steps.[9] - **Running it in one setting only.** Behavioral contrast and plain failure to generalize both mean the behavior can rise wherever the procedure is not running.[3] ## What the evidence does not show Most of the evidence comes from single-case experimental designs with people with developmental disabilities, in clinics and schools with trained staff. The reductions are large and replicated. What the literature does not establish is how well the same procedures work when run by parents and teachers without support, how often the effects hold months later, and how much of the effect is due to reinforcing the alternative rather than to the extinction that almost always accompanies it.[11][12] Nor does the evidence show that differential reinforcement is free of side effects. Extinction bursts, contrast in other settings, and the momentum problem — a richer context making the problem behavior more persistent once its reinforcement is withheld — are all documented.[15][3][14] They are milder than the side effects of punishment, and manageable when expected. Differential reinforcement is the default not because it is perfect but because it builds something while it removes something, and because its failures are usually failures of function, schedule, or consistency that better design can fix. ## Key takeaways - Differential reinforcement is two procedures at once: one response class is reinforced and another is placed on extinction. Shaping, discrimination training, and the DR family of treatments are the same operation on different dimensions of behavior. - DRA reinforces a specific alternative, DRI an incompatible one, DRO the absence of the behavior for an interval, DRL a low rate, and DRH a high rate. DRA, especially as functional communication training, has the strongest evidence. - The reinforcer you withhold must be the one maintaining the behavior, and the reinforcer you deliver should be the same one. A functional assessment tells you which it is; guessing is how "ignoring" ends up delivering the escape. - Start dense and thin slowly: reinforce the alternative every time, set DRO intervals shorter than the baseline gap, and expect a burst when the problem behavior stops working. - The side effects — bursts, contrast in other settings, and momentum — are documented, but they are milder and more manageable than those of punishment, which is why reinforcement-based reduction comes first. ### Check yourself **A teacher decides to reduce a student's calling out by praising him whenever he raises his hand, but keeps answering his called-out questions so he does not fall behind. Two weeks later calling out is unchanged. What is missing?** The extinction half. Both responses still produce the same reinforcer, and the older, easier one is better practiced, so it persists. DRA requires that the reinforcer be withheld for the problem behavior: called-out questions go unanswered, and raised hands are answered promptly and every time. **A child's screaming during homework has been shown to be escape-maintained. A parent sets up a DRO in which five minutes without screaming earns a sticker, but screaming continues. Why might the DRO be failing?** Two likely reasons. The sticker is an arbitrary reinforcer competing with escape, which is the functional one, and the parent may still be ending homework when screaming occurs, so the problem behavior is not on extinction. The better design uses the functional reinforcer: a request for a break earns one, and screaming no longer does. And check the interval; if screaming occurs every two minutes at baseline, five minutes is never reached and no reinforcement is delivered. **A trainer uses DRI for a dog that jumps on people: at home the dog is greeted only while sitting, and jumping at home stops. At the park, where the trainer does not run the procedure, jumping increases. What is happening?** Jumping is still reinforced at the park, and behavioral contrast means that reducing reinforcement in one setting can raise the behavior in another where reinforcement is unchanged. Generalization has to be programmed: run the procedure across settings and with the other people the dog meets. ## Frequently asked questions **What is differential reinforcement in simple terms?** Reinforcing one behavior while no longer reinforcing another. The reinforced behavior becomes more common and the other fades. It is how shaping works, how animals learn to respond only to certain cues, and, in applied settings, the main way to reduce a problem behavior without punishment: pay for something else, and stop paying for the problem. **What is the difference between DRA and DRO?** DRA reinforces a specific alternative behavior that serves the same purpose as the problem behavior, such as asking instead of grabbing. DRO reinforces the absence of the problem behavior for a set interval, whatever else the person does. DRA teaches a replacement; DRO does not, which is why DRA is preferred when a suitable alternative exists. **What is an example of differential reinforcement?** A child who whines for snacks is given the snack, or a clear answer, whenever she asks in a normal voice, and never when she whines. Asking increases and whining declines. In dog training, a dog that jumps on guests is greeted only while sitting, so sitting replaces jumping. **Is differential reinforcement the same as extinction?** No, but extinction is half of it. Extinction alone withholds the reinforcer for a behavior and leaves a gap where the behavior was. Differential reinforcement withholds that reinforcer and at the same time delivers it for something else, so the person still gets what they were after, through the behavior you chose. The combination produces smaller bursts and more durable change. **What is DRL, and when is it used?** Differential reinforcement of low rates reinforces a behavior only when it occurs infrequently, either below a limit for the session or with enough time between responses. It is used for behavior that is acceptable in moderation, such as asking questions in class or eating quickly, where the goal is fewer, not none. Deitz and Repp reduced classroom talk-outs this way in 1973. **What is functional communication training?** A form of DRA in which the person is taught a communicative response, such as a word, sign, or card, that produces the same reinforcer the problem behavior produced, while the problem behavior no longer produces it. Introduced by Carr and Durand in 1985, it is the most studied function-based treatment for problem behavior. **Does differential reinforcement work without extinction?** Rarely, in the published research. Reviews of DRA find that nearly all successful studies also placed the problem behavior on extinction, and a series of clinical cases found that functional communication training without extinction did not reduce problem behavior. If the problem behavior still works, teaching an alternative gives the person two routes to the same reinforcer. **Is differential reinforcement a type of punishment?** No. Punishment adds or removes a stimulus after a behavior to make it less likely. Differential reinforcement reduces a behavior by reinforcing something else and withholding the reinforcer for the problem behavior, which is extinction, not punishment. It is the standard alternative to punishment in applied behavior analysis. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 3. Reynolds, G. S. (1961). Behavioral contrast. *Journal of the Experimental Analysis of Behavior, 4*(1), 57–71. 4. Deitz, S. M., & Repp, A. C. (1973). Decreasing classroom misbehavior through the use of DRL schedules of reinforcement. *Journal of Applied Behavior Analysis, 6*(3), 457–463. 5. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. *Journal of Applied Behavior Analysis, 27*(2), 197–209. (Reprinted from *Analysis and Intervention in Developmental Disabilities, 2*, 3–20, 1982.) 6. Vollmer, T. R., & Iwata, B. A. (1992). Differential reinforcement as treatment for behavior disorders: Procedural and functional variations. *Research in Developmental Disabilities, 13*(4), 393–417. 7. Vollmer, T. R., Iwata, B. A., Zarcone, J. R., Smith, R. G., & Mazaleski, J. L. (1993). The role of attention in the treatment of attention-maintained self-injurious behavior: Noncontingent reinforcement and differential reinforcement of other behavior. *Journal of Applied Behavior Analysis, 26*(1), 9–21. 8. Carr, E. G., & Durand, V. M. (1985). Reducing behavior problems through functional communication training. *Journal of Applied Behavior Analysis, 18*(2), 111–126. 9. Tiger, J. H., Hanley, G. P., & Bruzek, J. (2008). Functional communication training: A review and practical guide. *Behavior Analysis in Practice, 1*(1), 16–23. 10. Hagopian, L. P., Fisher, W. W., Sullivan, M. T., Acquisto, J., & LeBlanc, L. A. (1998). Effectiveness of functional communication training with and without extinction and punishment: A summary of 21 inpatient cases. *Journal of Applied Behavior Analysis, 31*(2), 211–235. 11. Petscher, E. S., Rey, C., & Bailey, J. S. (2009). A review of empirical support for differential reinforcement of alternative behavior. *Research in Developmental Disabilities, 30*(3), 409–425. 12. Chowdhury, M., & Benson, B. A. (2011). Use of differential reinforcement to reduce behavior problems in adults with intellectual disabilities: A methodological review. *Research in Developmental Disabilities, 32*(2), 383–394. 13. Jessel, J., & Ingvarsson, E. T. (2016). Recent advances in applied research on DRO procedures. *Journal of Applied Behavior Analysis, 49*(4), 991–995. 14. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. *Behavioral and Brain Sciences, 23*(1), 73–90. 15. Lerman, D. C., & Iwata, B. A. (1996). Developing a technology for the use of operant extinction in clinical settings: An examination of basic and applied research. *Journal of Applied Behavior Analysis, 29*(3), 345–382. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Extinction](https://operantconditioning.com/extinction/): The other half of every differential reinforcement procedure: bursts, recovery, and how to use it. - [Shaping](https://operantconditioning.com/shaping/): Differential reinforcement of successive approximations, step by step. - [Punishment](https://operantconditioning.com/punishment/): What differential reinforcement replaces, and why it comes first. # Behavior Chaining: Forward, Backward & Total-Task Methods > Chaining links responses into a sequence where each step cues the next. Task analysis, forward vs. backward vs. total-task chaining, evidence, and mistakes. - Source: https://operantconditioning.com/chaining/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Procedures · Building sequences* Brushing your teeth is not one behavior but a dozen, performed in an order you never think about. Chaining is how operant conditioning explains sequences like that — each step cueing and paying for its neighbors — and how trainers, teachers, and therapists build them, from a rat's lever press to a teenager's first phone call. > **Definition** > > Chaining is a procedure for teaching a **behavior chain**: a sequence of responses in which each response produces a stimulus change that serves as the [discriminative stimulus](https://operantconditioning.com/discriminative-stimulus/) for the next response and as a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer) for the response that produced it, until the final response produces the terminal reinforcer.[1][2] The chain is the product; chaining is the set of methods — forward, backward, or total-task — for linking the steps identified by a **task analysis**.[3] > > Unlike [shaping](https://operantconditioning.com/shaping/), which builds a response that does not yet exist, chaining assembles responses the learner can already perform into an order that runs on its own. Washing hands, making a phone call, a dog's retrieve, and a rat's lever press are all chains. **In brief** - A behavior chain holds together because each link stimulus does two jobs: it is the discriminative stimulus for the next response and a conditioned reinforcer for the last one. Only the final response contacts the terminal reinforcer. - Chaining begins with a task analysis and a baseline probe, then links the steps forward from the first, backward from the last, or as a total task with prompts on whichever steps fail. Comparative studies find no single method best for every learner. - Shaping changes the form of one response; chaining changes the order of many. A step the learner cannot perform is shaped or prompted on its own, then put back in the chain. ## How a behavior chain works Skinner's analysis of the rat in the [operant chamber](https://operantconditioning.com/skinner-box/) already treated the "simple" lever press as a chain: the rat faces the lever, approaches, raises a paw, presses, hears the food magazine operate, lowers its head to the tray, seizes the pellet, and eats. Each response changes what the rat sees, feels, or hears, and each change sets the occasion for the next.[1] Keller and Schoenfeld drew the sequence out link by link in 1950.[2] The stimulus at each junction plays two roles. The sound of the magazine is the discriminative stimulus for going to the tray: in its presence, tray-approach is reinforced; in its absence, it is not. The same sound is a conditioned reinforcer for pressing: a stimulus that reliably precedes reinforcement acquires the power to reinforce the response that produces it, and Skinner showed in 1938 that the magazine sound alone, once paired with food, could condition a new response.[1] Nothing but the last link touches food, yet every link is reinforced, because every link produces the stimulus the next link needs. Procedure and process Chaining is the procedure — task analysis, prompting, and reinforcement in a set order. The process that makes it work is [stimulus control](https://operantconditioning.com/stimulus-control/) plus conditioned reinforcement: each response comes under the control of the stimulus the previous response produced, and that stimulus reinforces the response before it. When a chain fails, one of those relations has broken. The laboratory version is the **chained schedule**: a pigeon completes several [schedule](https://operantconditioning.com/schedules-of-reinforcement/) components in sequence, each signaled by its own key color, with food only after the last. Kelleher and Gollub's 1962 review drew a lesson from these schedules that every trainer eventually rediscovers: responding is strongest in the link closest to food and weakest in the links farthest from it, and in long chains the early links can fail outright.[4] The idea is older than the vocabulary. In 1890 William James wrote that in habitual action "what instigates each new muscular contraction to take place in its appointed order is not a thought or a perception, but the sensation occasioned by the muscular contraction just finished."[5] That is a chain described from the inside ([the chapter](https://operantconditioning.com/library/james-principles-of-psychology-vol-1/#chapter-iv-habit) is in the library); what Skinner added is that the sensation does not merely trigger the next movement, it pays for the last one. ## Task analysis: the first step A chain cannot be taught until it has been written down. A [task analysis](https://operantconditioning.com/glossary/#task-analysis) breaks a skill into its component responses in the order they occur, each ending in a detectable stimulus change. Cooper, Heron, and Heward describe three ways to build one: watch a competent person perform the task, ask people who do it well, or perform it yourself and note each thing that changes.[3] The third catches the steps experts no longer notice. | Step | Response | Stimulus it produces (the cue for the next step) | | --- | --- | --- | | 1 | Turn on the tap | Running water | | 2 | Put hands under the water | Wet hands | | 3 | Press the soap pump | Soap on the palm | | 4 | Rub hands together | Lather covering both hands | | 5 | Rinse under the water | No lather; wet hands | | 6 | Turn off the tap | Silence; wet hands | | 7 | Dry hands on the towel | Dry hands — end of chain; the terminal reinforcer follows (lunch, "all done," a token) | The right grain depends on the learner: seven steps suit a preschooler, while for a learner with a severe disability "rub hands together" may need to become four steps. Cuvo, Leaf, and Borakove's 1978 restroom-cleaning study broke the job into subtasks — mirror, sink, toilet, floor — and each subtask into component responses with a definable start and finish.[6] The test of a task analysis is whether a trainer who has never met the learner could score each step as done or not done. A task analysis also makes a **baseline probe** possible. In a *single-opportunity* probe the trial stops at the first error and everything after it is scored as not done; in a *multiple-opportunity* probe the trainer quietly completes any failed step so the learner can attempt the next, which shows exactly which steps are already in the repertoire.[3] The probe decides where teaching starts and is the baseline against which progress — usually the percentage of steps done independently — is measured. ## Forward chaining, backward chaining, and total-task presentation Once the steps are written there are three ways to link them; the difference is where the learner starts and where the reinforcer falls. In **forward chaining** the learner masters the first step, then the first two, and so on, with the trainer completing the rest. In **backward chaining** the trainer performs every step except the last, the learner performs the last step and contacts the natural reinforcer, and earlier steps are added one at a time, so every trial ends with the finished task. In **total-task presentation** the learner attempts every step on every trial and the trainer prompts the steps that fail. A variant, **backward chaining with leaps ahead**, skips steps the probe showed the learner can already do.[3] Reviews of the applied literature have found all three effective and none consistently superior.[7] | Method | How it works | Who it suits | Evidence | | --- | --- | --- | --- | | **Forward chaining** | Step 1 to mastery, then 1–2, then 1–3; reinforce after the last mastered step; trainer finishes the rest | Learners who already have the early steps; tasks where the order carries meaning | About as efficient as backward chaining in a comparison with four children;[8] slower than backward chaining for a keyboard sequence in adults[9] | | **Backward chaining** | Trainer does all but the last step; learner finishes and gets the natural reinforcer; add earlier steps one at a time | Long chains; learners who need every trial to end in success; animal training; tasks whose payoff sits at the end | Better than forward chaining for a keyboard sequence;[9] taught internet skills to adults with developmental disabilities;[10] the standard method in animal training[11] | | **Total-task presentation** | Every step every trial; prompts on failed steps, faded over time; reinforcer at the end | Learners who can already do most steps; short chains; routines practiced whole | The method of Bellamy, Horner, and Inman's vocational program for adults with severe intellectual disabilities;[12] effective, with no consistent advantage over the alternatives[7] | Backward chaining's advantage is structural: the new step always leads directly into steps already mastered and then into the reinforcer. Forward chaining's is that the learner initiates and practices the chain in the order it will be used; in the [classroom](https://operantconditioning.com/classroom/), where most academic procedures are chains whose order carries meaning, it is the default.[3] ## What the research shows Spooner and Spooner reviewed the chaining studies available in 1984, most with learners with severe disabilities, and found no method that consistently beat the others; the studies differed in learners, tasks, prompting, and measures.[7] Slocum and Tiger taught four children arbitrary sequences of simple actions, some by forward chaining and some by backward chaining, then assessed which procedure each child preferred. Both methods worked, neither was reliably faster across the four children, and the children had preferences that differed from child to child.[8] Ash and Holding, with adults learning a sequence of key presses, found that the group trained by backward chaining outperformed the group trained by forward chaining.[9] The honest summary: all three methods teach chains, backward chaining has an edge in some tasks, and the learner's preference is a legitimate tie-breaker. The applied literature is where chaining earns its keep. Cuvo and colleagues taught adolescents with intellectual disabilities to clean a restroom; the skills generalized to an untrained restroom and were maintained at follow-up.[6] Bellamy, Horner, and Inman built a vocational program for adults with severe intellectual disabilities around task-analyzed assembly work presented as a total task.[12] Wilson, Reid, Phillips, and Burgio taught residents of an institution family-style dining — passing serving dishes, serving themselves — and reported, as the title says, both effects and noneffects: the skills were acquired; some broader changes the authors looked for did not follow.[13] Test, Spooner, Keul, and Grossi task-analyzed the public pay-phone call and taught it to adolescents with severe disabilities.[14] Jerome, Frantino, and Sturmey used backward chaining with errorless prompting to teach adults with developmental disabilities to reach a chosen website.[10] The successful studies did not stop at acquisition: they probed generalization, checked maintenance, and trained where the chain would be used.[6][13] [More on applied behavior analysis ›](https://operantconditioning.com/applications/) ## Chaining in dog and animal training Animal trainers work with chains constantly, and they almost always build them backward. A formal retrieve is a chain: sit and wait, run out on cue, pick up the object, return, hold, release to the hand. Trainers teach the hold and release first, then the return, then the pick-up, then the run-out, so that every new link leads into links the dog already performs fluently.[11] Karen Pryor recommends the same for people: learn a piece of music or a speech from the end, so that practice always moves from the unfamiliar into the familiar.[11] Trainers also see what happens when an unwanted response gets into a chain: if a dog barks and then sits, and the sit is reinforced, the bark is now the first link and will be strengthened with the sit. The cure is [differential reinforcement](https://operantconditioning.com/differential-reinforcement/) of the clean sequence: the sit that follows no bark is reinforced, the sit that follows a bark is not. A marker signal — a click or a word — is useful in chains for the same reason: it is a conditioned reinforcer that can be delivered at the exact link that needs strengthening without stopping the chain to deliver food.[11] [Reinforcement-based dog training ›](https://operantconditioning.com/dog-training/) ## How to build a behavior chain 1. **Write the task analysis.** List the steps in order, each ending in a stimulus change you can see. Check it by doing the task yourself and watching someone competent do it. If a step contains two stimulus changes, split it.[3] 2. **Run a baseline probe.** A multiple-opportunity probe tells you which steps the learner already has; a single-opportunity probe tells you how far they get unaided. Record the percentage of steps performed independently; that is your progress measure.[3] 3. **Choose a method.** Most steps already present: total-task presentation. A long chain, a learner who tires easily, or a reinforcer that is the natural end of the task: backward chaining. Early steps present and an order that carries meaning: forward chaining. If two seem equal, let the learner's preference decide.[8] 4. **Prompt and fade.** Decide the prompt hierarchy before the first trial — most-to-least for a learner who makes many errors, least-to-most or time delay for one who can sometimes beat the prompt — and fade every prompt until the product of the previous step, not your voice or hand, is the cue.[3] 5. **Reinforce the link being taught, then let the learner finish.** During acquisition the new step is too far from the end of the chain for the terminal reinforcer to reach it. Mark it immediately with praise, a token, or a click — [positive reinforcement](https://operantconditioning.com/positive-reinforcement/) delivered where the link happens — then let the learner run the rest of the chain to the natural reinforcer, and drop the added reinforcers one at a time as the link stimuli take over.[4] 6. **Thin, vary, and check back.** Once every step is independent, thin any added reinforcement, practice with different materials and settings so the cue is "wet hands" and not "this sink," and probe again after a few weeks.[6] ## Shaping vs. chaining The [shaping](https://operantconditioning.com/shaping/) page gives the one-line contrast: shaping builds a new form of a single response, chaining links existing responses into a sequence. The deeper difference is what the trainer manipulates. In shaping, the criterion for reinforcement moves along a dimension of one response — louder, longer, closer. In chaining, the criterion is fixed by the task analysis; what moves is the number of steps the learner performs without help, and the tools are prompting, fading, and conditioned reinforcement. | Question | Shaping | Chaining | | --- | --- | --- | | What is missing at the start? | The response itself, in any usable form | The order; the individual responses exist | | What changes across training? | The criterion for reinforcement | The number of steps done unaided | | Main tool | Differential reinforcement of successive approximations | Prompting and fading; each link's stimulus as a conditioned reinforcer | | What an error looks like | The form drifts back, or variability collapses | A step is skipped, done out of order, or waits for a prompt | | Example | A first word; a lever press; a fuller range of arm motion | Brushing teeth; a phone call; a retrieve; cleaning a restroom | The two meet constantly. A learner working through a chain will hit a step they cannot perform in any form — the child who can do every part of shoe-tying except pulling the loop through — and that step is shaped or physically prompted on its own, then put back into the chain. Ask, of any skill, whether the problem is *form* or *order*. Form problems are shaping problems; order problems are chaining problems; and attaching a new habit to the end of an old one — coffee, then vitamins — is chaining under another name, with the end of the old routine as the cue for the new one. [Building habits ›](https://operantconditioning.com/habits/) ## Common chaining mistakes - **Steps that are too big.** A step that hides two or three stimulus changes stalls the learner in the middle, where nothing cues the next move. If a learner keeps failing the same step, it is usually two steps. - **Prompts that are never faded.** The learner performs the whole chain — as long as the trainer says "now rinse." The trainer's voice, not the lather, has become the discriminative stimulus, and the chain will not run when the trainer leaves.[3] - **Reinforcing only the end.** A single reinforcer at the end of a long chain reaches the last few links well and the first few barely, as chained schedules show in the laboratory.[4] Reinforce the link being taught where it happens; thin later. - **Reinforcing an unclean chain.** Whatever happened just before the reinforcer is in the chain, wanted or not — the bark before the sit, the flourish before the signature. Reinforce only the clean sequence. - **Skipping the probe.** Teaching steps the learner already has wastes sessions; assuming steps they do not have puts them on extinction at that step every trial. - **Changing the cues without noticing.** A new soap dispenser, a different sink, a phone with the buttons elsewhere: each removes a stimulus the chain depended on. If a chain breaks at one step in a new setting, look for the stimulus that changed.[6] - **Stopping at acquisition.** A chain that runs on trainer-delivered praise and tokens is not finished until it runs on its natural end, where it is needed, weeks later. ## What the evidence does not show It does not show that one method is best. The comparative studies are few, small, and mixed: Spooner and Spooner found no consistent winner, Slocum and Tiger found the answer varied by child, and Ash and Holding's backward-chaining advantage was for one keyboard task in adults.[7][8][9] Nor does the evidence say how to predict which method a given learner will do best with; pick one, measure, and switch if the data say so. It does not show that the dual-function account is the whole story. That each link stimulus is both discriminative stimulus and conditioned reinforcer is inferred from performance, and Kelleher and Gollub's review made clear that conditioned reinforcing strength falls off with distance from the terminal reinforcer, so "every link reinforces the one before" is true in kind but not in degree.[4] It does not show that chains generalize or persist on their own. The studies that found generalization and maintenance had built them in, and Wilson and colleagues' "noneffects" are a reminder that teaching a dining chain teaches a dining chain; anything else hoped for has to be measured.[6][13] And nearly all of the applied evidence comes from learners with developmental disabilities. Chaining is used every day in classrooms, sports coaching, and workplace training, but the controlled evidence for those uses is thin. ## Key takeaways - A behavior chain is a sequence in which each response produces the discriminative stimulus for the next and a conditioned reinforcer for the one before. Only the last link contacts the terminal reinforcer. - Chaining begins with a task analysis — steps in order, each ending in a detectable stimulus change — and a baseline probe of which steps the learner already has. - Forward chaining teaches from the first step, backward chaining from the last, and total-task presentation runs the whole chain with prompts on failed steps. Studies find all three effective and none consistently best. - Fade prompts until the product of each step, not the trainer, is the cue, and reinforce the new link where it happens; a reinforcer at the end of a long chain barely reaches its beginning. - Shaping changes the form of one response; chaining changes the order of many. Ask whether the problem is form or order, and combine the two when a chain contains a step the learner cannot yet perform. - Generalization and maintenance are programmed, not automatic: vary materials and settings, thin added reinforcers, and probe again weeks later. ### Check yourself **A child performs every step of hand-washing correctly, but only after a parent names each step aloud. Has the chain been learned?** Not yet. The parent's instructions, not the products of each step, are the discriminative stimuli, so the chain will not run when the parent is absent. The prompts have to be faded until running water cues wetting the hands and lather cues rinsing. Score a step as mastered only when it follows the previous step's result rather than a prompt. **A trainer teaching a dog a ten-obstacle agility sequence reinforces only at the finish. The last three obstacles are fast and clean; the first three are slow and often skipped. What is going on?** A reinforcer at the end of a long chain reaches the links closest to it and barely reaches the earliest ones, the same pattern seen in laboratory chained schedules, where responding is weakest in the links farthest from food. During training the early links need their own conditioned reinforcement, such as a marker, and back-chaining the sequence would let each newly added obstacle lead directly into obstacles that already pay off. **A teenager can lift the receiver, insert coins, and speak on a pay phone but cannot press the buttons firmly enough to register. Shaping or chaining?** Both. The order problem is a chaining problem, but the button press does not yet exist in a usable form, which is a form problem. That step is shaped, or physically prompted and the prompt faded, on its own, and then put back into the chain. ## Frequently asked questions **What is chaining in simple terms?** Chaining is a way of teaching a skill made of several steps done in order, such as washing hands or making a phone call. The steps are listed, the learner is taught to link them one at a time, and the result of each step becomes the signal for the next. The learner ends up performing the whole sequence without help. **What is a behavior chain?** A behavior chain is a sequence of responses in which each response produces a stimulus that cues the next response and reinforces the one just completed, with a reinforcer at the end. Skinner analyzed a rat's lever press as a chain: approach, press, hear the magazine, go to the tray, eat. Everyday chains include tying shoes, brushing teeth, and driving to work. **What is an example of chaining?** Teaching a child to put on a coat by backward chaining: the parent does everything except the final zip, the child zips and is done; then the child starts from pulling the coat closed and zipping; then from the second sleeve; and so on until the child does the whole task. In animal training, a dog's retrieve is taught the same way, starting with the release. **What is the difference between shaping and chaining?** Shaping builds a single response that does not yet exist by reinforcing successively closer approximations. Chaining links responses the learner can already perform into a sequence, using a task analysis, prompts that are faded, and reinforcement. If the problem is the form of a response, shape it; if the problem is the order of several responses, chain them. The two are often combined. **What is the difference between forward and backward chaining?** In forward chaining the learner is taught the first step first and the trainer completes the rest; steps are added in order until the learner does all of them. In backward chaining the trainer does all but the last step, the learner completes the chain and gets the reinforcer, and earlier steps are added one at a time. Both work; studies find neither reliably faster for every learner. **What is task analysis in ABA?** A task analysis is the written list of the component responses in a skill, in the order they occur, with each step ending in an observable result. It is built by watching a competent performer, asking people who do the task well, or doing it yourself. In applied behavior analysis it is the first step of chaining and the basis for measuring progress as the percentage of steps done independently. **What is total-task chaining?** In total-task presentation the learner attempts every step of the chain on every trial. The trainer prompts whichever steps fail, fades the prompts as the learner improves, and delivers reinforcement at the end. It suits learners who already perform most of the steps and short chains, and it was the method of Bellamy, Horner, and Inman's vocational training program for adults with severe disabilities. **Why does a behavior chain keep going when only the last step is rewarded?** Because the stimulus each step produces is a conditioned reinforcer for that step as well as the cue for the next one. Running water reinforces turning on the tap; lather reinforces rubbing. This strength is not uniform: laboratory chained schedules show that steps far from the final reinforcer are weaker, which is why long chains need extra reinforcement during training. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Keller, F. S., & Schoenfeld, W. N. (1950). *Principles of Psychology: A Systematic Text in the Science of Behavior*. Appleton-Century-Crofts. 3. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 4. Kelleher, R. T., & Gollub, L. R. (1962). A review of positive conditioned reinforcement. *Journal of the Experimental Analysis of Behavior, 5*(S4), 543–597. 5. James, W. (1890). *The Principles of Psychology* (Vol. 1). Henry Holt. 6. Cuvo, A. J., Leaf, R. B., & Borakove, L. S. (1978). Teaching janitorial skills to the mentally retarded: Acquisition, generalization, and maintenance. *Journal of Applied Behavior Analysis, 11*(3), 345–355. 7. Spooner, F., & Spooner, D. (1984). A review of chaining techniques: Implications for future research and practice. *Education and Training of the Mentally Retarded, 19*(2), 114–124. 8. Slocum, S. K., & Tiger, J. H. (2011). An assessment of the efficiency of and child preference for forward and backward chaining. *Journal of Applied Behavior Analysis, 44*(4), 793–805. 9. Ash, D. W., & Holding, D. H. (1990). Backward versus forward chaining in the acquisition of a keyboard skill. *Human Factors, 32*(2), 139–146. 10. Jerome, J., Frantino, E. P., & Sturmey, P. (2007). The effects of errorless learning and backward chaining on the acquisition of internet skills in adults with developmental disabilities. *Journal of Applied Behavior Analysis, 40*(1), 185–189. 11. Pryor, K. (1999). *Don't Shoot the Dog! The New Art of Teaching and Training* (rev. ed.). Bantam. 12. Bellamy, G. T., Horner, R. H., & Inman, D. P. (1979). *Vocational Habilitation of Severely Retarded Adults: A Direct Service Technology*. University Park Press. 13. Wilson, P. G., Reid, D. H., Phillips, J. F., & Burgio, L. D. (1984). Normalization of institutional mealtimes for profoundly retarded persons: Effects and noneffects of teaching family-style dining. *Journal of Applied Behavior Analysis, 17*(2), 189–201. 14. Test, D. W., Spooner, F., Keul, P. K., & Grossi, T. (1990). Teaching adolescents with severe disabilities to use the public telephone. *Behavior Modification, 14*(2), 157–171. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Shaping](https://operantconditioning.com/shaping/): Building a response that does not yet exist, one approximation at a time. - [Stimulus control](https://operantconditioning.com/stimulus-control/): How discriminative stimuli come to govern behavior — the process inside every link. - [Dog training](https://operantconditioning.com/dog-training/): Reinforcement-based training, including back-chained sequences and marker signals. # Law of Effect: Thorndike's Cats, the 1911 Law, and Skinner > Thorndike's law of effect explained: the 1898 puzzle-box cats, the exact 1911 statement, the 1932 revision that dropped punishment, and Skinner's rebuild. - Source: https://operantconditioning.com/law-of-effect/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *History · Foundations* A hungry kitten in a wooden box, food outside, and a stopwatch. From that arrangement Edward Thorndike drew the first experimental law of learning, then cut it in half on his own evidence. Everything on this site about reinforcement descends from it, and so does the algorithm that trains modern AI. > **Definition** > > The law of effect is Edward L. Thorndike's principle that, of the several responses an animal makes to a situation, those "accompanied or closely followed by satisfaction" become more firmly connected with that situation and are more likely to recur, while those "accompanied or closely followed by discomfort" have their connections weakened and are less likely to recur.[1] Thorndike found the effect in his puzzle-box experiments of 1898, gave it this formal statement in 1911, and dropped the second half in 1932 when his own data showed that discomfort does not weaken what satisfaction strengthens.[2][3] > > It is the ancestor of [reinforcement](https://operantconditioning.com/reinforcement/). Skinner kept the finding, discarded the theory of connections and satisfactions that came with it, and rebuilt it as operant conditioning. **In brief** - Thorndike's hungry cats escaped from latched boxes by acts that began as accidents; escape times fell gradually across trials, and he concluded that the resulting pleasure stamped the successful act in and the useless ones out. - The 1911 law has two halves, satisfaction strengthens and discomfort weakens, and Thorndike himself dropped the second in 1932 after finding that "wrong" did not weaken human responses the way "right" strengthened them. - Skinner replaced "satisfaction" with a reinforcer defined by its effect on rate, Herrnstein turned the law into an equation, and reinforcement learning in AI is its direct descendant. ## The puzzle-box experiments of 1898 Thorndike began the work as William James's graduate student at Harvard, testing chicks in improvised pens, and finished it at Columbia, where the cats and dogs were added and the whole was published in 1898 as his doctoral thesis.[4][2] The method was plain. A hungry animal was shut in a box with food outside in sight and could get out only by some simple act: pulling a loop of cord, pressing a lever, turning a wooden button, stepping on a platform. Thorndike timed the escape, put the animal back, and repeated the trial until the time was short and constant; an animal that failed was taken out but, as he underlined, *not fed*. The subjects were about a dozen kittens, three small dogs, and chicks a few days old, whose motive was not hunger but, in his phrase, dislike of loneliness.[2] A kitten dropped into a box did not study the latch. It squeezed at every opening, clawed and bit the bars, and thrust its paws through the gaps at anything within reach. Somewhere in that scramble a paw caught the loop, the door fell open, and the cat ate. On later trials the useless movements dropped away and the successful one came sooner, until the cat clawed the loop the moment it was put in. Thorndike's description gave the law its first vocabulary: the non-successful impulses are "stamped out," and the impulse leading to the successful act is "stamped in by the resulting pleasure."[2] What made this an experiment rather than an anecdote was the time-curve, each animal's escape time plotted against trial number. Cat 12 in box A took 160 seconds on the first trial, 30 on the second, 90 on the third, and 6 or 7 seconds by the twenty-fourth.[2] The times, he said, were "facts which may be obtained by any observer who can tell time." In the easiest boxes every cat succeeded; in the thumb latch and the boxes requiring three separate acts, some never did.[2] ### The argument against reasoning Thorndike's target was the anecdotal literature, George Romanes in particular, which took a cat's opening a latched door as proof of reasoning. His answer was that all his cats opened latched doors, by accident, and that the curve showed what kind of process was at work: a cat that grasped how the box worked should show "a sudden vertical descent," and none did. Thorndike read the gradual slope as "the wearing smooth of a path in the brain, not the decisions of a rational consciousness."[2] Cats that watched a trained cat escape learned nothing from it, and cats whose paws he pressed onto the mechanism did not learn by being put through the act.[2] It was neither imitation nor insight but selection among the cat's own movements by what followed them. The monograph is republished as Chapter II of the 1911 book: [read the 1898 monograph](https://operantconditioning.com/library/thorndike-animal-intelligence/#chapter-ii-animal-intelligence-an-experimental-study-of-the-associative-processes-in-animals). ## The law of effect as Thorndike stated it in 1911 The 1898 monograph described the process; the 1911 book, *Animal Intelligence*, stated it as a law. Chapter VI, "Laws and Hypotheses for Behavior," gives two provisional laws of learning. The first is the law of effect. The law of effect (Thorndike, 1911, p. 244) "Of several responses made to the same situation, those which are accompanied or closely followed by satisfaction to the animal will, other things being equal, be more firmly connected with the situation, so that, when it recurs, they will be more likely to recur; those which are accompanied or closely followed by discomfort to the animal will, other things being equal, have their connections with that situation weakened, so that, when it recurs, they will be less likely to occur. The greater the satisfaction or discomfort, the greater the strengthening or weakening of the bond."[1] The second is the **law of exercise**: a response becomes more strongly connected with a situation in proportion to how often, how vigorously, and how long it has been connected with it.[1] Repetition strengthens; consequences select. Thorndike held that effect was the more fundamental, since an animal that mostly makes one response and only occasionally another can end up making the rare one every time if it alone is followed by satisfaction: "the law of effect is primary, irreducible to the law of exercise."[1] ### What "satisfaction" meant The word invites a reading in terms of feelings, and Thorndike headed it off. A satisfying state of affairs is "one which the animal does nothing to avoid, often doing such things as attain and preserve it"; an annoying one is "one which the animal commonly avoids and abandons."[1] Satisfiers "cannot be determined with precision and surety save by observation," and what satisfies is not what is good for the animal; he listed overeating and intoxication among the most potent satisfiers of man.[1] This is a functional definition, nearly three decades before Skinner's, and it is why the law survived the change of vocabulary. Three "other things" had to be equal: exercise; closeness in time, which Thorndike illustrated rather than measured (a button that opened the door after one, five, fifty, or five hundred seconds would be learned fastest in the first case and in the fourth "almost certainly never"); and attention, whether the successful movement was "an eminent, emphatic part" of what the animal was doing.[1] Beneath the law sat a neural hypothesis, learning as a change in what he called the intimacy of the synapse, and the image of a path worn smooth in the brain echoes his teacher William James's [chapter on habit](https://operantconditioning.com/library/james-principles-of-psychology-vol-1/#chapter-iv-habit).[1] Read the statement in context: [Chapter VI](https://operantconditioning.com/library/thorndike-animal-intelligence/#chapter-vi-laws-and-hypotheses-for-behavior). ## Before Thorndike: Bain, Morgan, and trial and error The idea was not new; the experiment was. In *The Senses and the Intellect* (1855) Alexander Bain described how an organism's spontaneous movements are sorted by their consequences: a movement that happens to coincide with pleasure is kept up and repeated, and one that coincides with pain is dropped.[5] That is the law of effect without the box, the curve, or the name, and the phrase "trial and error" is already in Bain.[5] C. Lloyd Morgan supplied the rule of evidence: his canon of 1894 holds that no animal action should be explained by a higher mental faculty if a lower one will account for it.[6] Thorndike quoted Morgan at length in 1898 as "the least offender" among the theorists he was attacking.[2] What he added was everything that turns a plausible principle into a law: a repeatable situation, many animals, a controlled motive, a measure independent of the observer, and a test that could have come out the other way. Where this fits in the longer story is on the [history page](https://operantconditioning.com/history/). ## The 1932 revision: the truncated law of effect In a 1927 paper titled simply "The law of effect" Thorndike restated the law for adult human subjects and argued that an after-effect strengthens a connection directly and automatically, whether or not the learner thinks about it.[7] The experiments that followed changed the law. In a typical one a subject was shown a word, chose a number to go with it, and was told "right" or "wrong." Right strengthened. Wrong did little or nothing.[3] In *The Fundamentals of Learning* (1932) he concluded that annoyers do not act on connections the way satisfiers do: a punished response is not weakened in proportion to the punishment; at most, the annoyer leads the learner to vary, and the variation may then be rewarded.[3] The second half of the 1911 law was gone. This is the **truncated law of effect**. The law of exercise went with it: subjects who drew lines of a set length hundreds of times without being told how they were doing did not improve.[3] The asymmetry became one of the field's most durable findings. Skinner reported in 1938 that when a rat's lever presses were punished at the start of extinction by having the lever slap back against its paws, responding was suppressed only while the slap was in force, and the punished rats made about as many responses as unpunished ones before they quit.[8] Estes extended the result with shock in 1944: suppression, but temporary.[9] The modern statement, that [punishment](https://operantconditioning.com/punishment/) is not reinforcement with the sign reversed, descends from Thorndike's own retraction; whether stronger or better-timed punishers do more than a spoken "wrong" is taken up on that page. ## The criticisms ### Is the law circular? The most persistent objection is that the law explains nothing, because a satisfier is identified by the very strengthening it is supposed to explain. Leo Postman's 1947 review of the law's first half-century laid out this argument along with the others then in play, among them whether an after-effect is needed for learning or only for performance.[10] Paul Meehl's answer, in 1950, is the one still taught. Separate the **weak law**, a definition (a reinforcer is whatever strengthens the response it follows), from the **strong law**, an empirical claim (all learning requires such an event). Then notice that the weak law is not empty either, because reinforcers are **trans-situational**: an event shown to strengthen one response in one situation will strengthen other responses in other situations. That is a prediction, and it could fail; if food had strengthened loop-pulling in box A but not lever-pressing in box I, the law would have been in trouble.[11] Thorndike's definition of a satisfier by approach and avoidance had already pointed the same way.[1] ### Was it trial and error at all? In 1946 Edwin Guthrie and George Horton published *Cats in a Puzzle Box*, in which cats escaped by touching a pole in the center of the box and were photographed at the moment of escape. Each cat settled on its own way of hitting the pole and repeated it almost exactly, and Guthrie took the stereotypy as evidence for his rival theory: whatever movement the animal happens to be making when the situation changes stays attached to it by contiguity alone.[12] In 1979 Bruce Moore and Susan Stuttard pointed out that rubbing a flank or cheek against an upright object is the domestic cat's species-typical greeting. In a replica of Guthrie's box, cats rubbed the pole when a person was visible whether or not rubbing opened the door, and rarely did so when no one was in view. The response was not learned in the box; it was elicited by the experimenters standing in front of it.[13] The paper's subtitle, "Tripping over the cat," is fair. The complication cuts both ways. Thorndike's boxes required acts that are no part of a greeting, and his cats' times fell across dozens of trials, so his curves are not explained away. But what an animal brings to the box is never a blank repertoire, and Thorndike had glimpsed this himself: cats released from box Z whenever they licked themselves learned to lick as soon as they were put in, but the lick shrank with practice to "a mere quick turn of the head," and he confessed himself ignorant of why.[2] Consequences work on a repertoire shaped by evolution, a point the [shaping](https://operantconditioning.com/shaping/) page returns to. ## How Skinner rebuilt it: reinforcement and selection by consequences [B. F. Skinner](https://operantconditioning.com/bf-skinner/) read Thorndike's law as a fact in search of a better description, and *The Behavior of Organisms* (1938) supplied one. The "response" became the **operant**, a class of acts defined by their common effect rather than their form. "Satisfaction" was dropped: a **reinforcer** was any event that, following a response, raised its future rate, and nothing was said about how it felt. And the measure changed: Skinner's rats lived with the lever and pressed whenever they liked, so rate of responding, drawn by the cumulative recorder, replaced time to escape.[8] He called this Type R conditioning, to separate it from Pavlov's Type S (see [operant versus classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/)), and the apparatus is on the [Skinner box](https://operantconditioning.com/skinner-box/) page.[8] In "Selection by Consequences" (1981) Skinner argued that the law of effect is one instance of a causal mode peculiar to living things, in which variation is followed by selection: natural selection shapes the species, operant conditioning shapes the individual's behavior within its lifetime, and cultural practices are selected by their consequences for the group.[14] Read this way, the law of effect is Darwinian, and Thorndike had used the language in 1898: from among the cat's movements, "one is selected by success."[2] Richard Herrnstein made it quantitative. In "On the Law of Effect" (1970) he extended the [matching law](https://operantconditioning.com/matching-law/) to the single response in a Skinner box by treating even that as a choice between the measured response and everything else the animal could be doing. The result is a hyperbola: response rate rises with reinforcement rate at diminishing returns, to a ceiling set by the reinforcement available for everything else.[15] ### Thorndike's law of effect vs. Skinner's reinforcement | Aspect | Thorndike's law of effect (1898–1932) | Skinner's reinforcement (1938 onward) | | --- | --- | --- | | **Unit of analysis** | A connection between a situation and a specific movement | The operant: a class of responses defined by its common consequence | | **Mechanism** | Satisfaction "stamps in" a connection, physically a change at the synapse; later, a direct and automatic after-effect | None claimed; reinforcement is a functional relation between consequence and rate, explained as selection by consequences | | **Punishment** | 1911: the mirror image of satisfaction. 1932: dropped, since annoyers do not weaken connections the way satisfiers strengthen them | Not the mirror image of reinforcement: suppresses responding, often temporarily and with side effects; a separate process | | **Measurement** | Time to escape on each discrete trial; a learning curve | Rate of responding in the free operant; a cumulative record | | **How the consequence is defined** | A state of affairs the animal does nothing to avoid and often acts to attain | Whatever increases the rate of the response it follows | ## Why the law of effect still matters Strip away the neurons and the word "satisfaction," and the law of effect is the working rule behind every reinforcement procedure on this site. | Setting | Situation | Response | After-effect | Next time | | --- | --- | --- | --- | --- | | Dog training | Dog at the closed door | Sits | Door opens | Sitting at doors comes sooner | | Classroom | Teacher asks a question | Student raises a hand | Is called on and answers | Hand-raising increases | | Parenting | Toddler in the grocery cart | Whines | Parent hands over a snack | Whining in carts increases | | Self-management | Sitting down to write | Opens the file and types one sentence | The dread lifts | Starting gets easier | Three of Thorndike's lessons transfer unchanged. Satisfiers are found by observation, not assumption, the insight behind the [Premack principle](https://operantconditioning.com/premack-principle/). Closeness in time matters: the button that opens the door after fifty seconds teaches almost nothing. And learning is a curve, not a step; the [trainer](https://operantconditioning.com/dog-training/) who expects a perfect sit after one treat is expecting the sudden vertical descent Thorndike never saw. A fourth lesson is negative: since "wrong" does not undo what "right" does, the tools for weakening a response are [extinction](https://operantconditioning.com/extinction/) and reinforcement of something else.[3] ### Reinforcement learning in AI Reinforcement learning, the branch of machine learning in which an agent learns by acting in an environment and receiving a numerical reward, descends from Thorndike's law directly, and its founders say so. Sutton and Barto open the field's history with the law of effect and note that it contains the two things trial-and-error learning requires: it is selectional, since among the actions tried those with better outcomes are kept, and associative, since the kept actions are tied to the situations in which they were tried.[16] An agent that tries actions, keeps the ones that pay, and learns which states they pay in is doing what the cat did in box A, with the reward written down as a number. The differences are real: an engineer chooses the reward, the agent can run millions of trials, and the algorithms keep explicit estimates of value. But the core claim is Thorndike's: a system with no knowledge of a task can acquire competent behavior from nothing but the consequences of its own actions. ## What the evidence does not show - **That animals cannot reason.** The curves rule out sudden mastery in those cats, in those boxes. Thorndike limited the claim to "just these particular animals" and suspected the primates would differ.[2] - **That a feeling of satisfaction does the work.** The satisfier is defined by approach and avoidance; the law says nothing about experience, and Meehl's analysis shows what content it has without it.[1][11] - **That punishment never works.** The truncated law rests on a spoken "wrong," and Skinner's demonstration on a brief, mild slap.[3][8] They show that punishment is not the mirror image of reinforcement, not that suppression is impossible. - **That every response was free to be selected.** Guthrie and Horton's cats were performing a greeting, and Thorndike's licking cats produced a shrunken lick he could not explain.[13][2] Consequences select from a repertoire the species supplies. - **That all learning requires an after-effect.** This is Meehl's strong law, and Postman's review had already listed the findings against it.[10][11] The weak law, that consequences change the probability of the responses they follow, is the part that has never failed. - **That the curves were smooth.** Thorndike warned that the slope of any particular stretch of a time-curve "may be due to accident," and his second trials were often slower than his first because the early successes were still accidents.[2] ## Key takeaways - Thorndike's law of effect says that responses accompanied or closely followed by satisfaction become more firmly connected with the situation and more likely to recur; it was found in the 1898 puzzle boxes and stated formally in 1911, with a functional definition of satisfaction that anticipated Skinner's reinforcer. - The evidence was the time-curve: escape times fell gradually across trials with no sudden drop, and cats learned nothing from imitation or from being put through the act, so Thorndike read the slope as selection among accidental movements rather than reasoning. - Thorndike cut the law in half in 1932 after finding that "wrong" did not weaken human responses the way "right" strengthened them, the first evidence that punishment is not reinforcement with the sign reversed. - The charge of circularity is answered by trans-situationality, a prediction that could fail and does not; Guthrie and Horton's cats turned out to be greeting the experimenters, a reminder that consequences select from a repertoire the species supplies. - Skinner kept the finding and replaced the theory with the operant, rate of response, a reinforcer defined by its effect, and selection by consequences; Herrnstein made the law quantitative, and reinforcement learning in AI is its direct descendant. ### Check yourself **A trainer says her dog "figured out" that sitting opens the back door, because it now sits the instant it reaches the door. What would Thorndike want to see before agreeing?** The curve. If the dog grasped the rule, the latency to sit should have dropped from long to short in a single trial and stayed there, the sudden vertical descent Thorndike looked for and never found in his cats. A gradual fall over many door-approaches, with the sit emerging from a scramble of other behaviors, is the law of effect at work: a response selected by its consequence. Thorndike would also note that a single, simple, definite act can be stamped in by one experience without any inference, so even a fast curve is not proof of understanding. **A student writes that the law of effect is circular because reinforcers are defined by what they reinforce. How does Meehl's analysis answer this?** By separating two laws. The weak law is indeed a definition: a reinforcer is whatever strengthens the response it follows. The strong law, that all learning requires reinforcement, is an empirical claim that could be false. And the weak law still has content, because reinforcers are trans-situational: an event that strengthens one response in one situation is predicted to strengthen other responses elsewhere. That prediction could fail and does not, which is what makes the law a law rather than a tautology. **A teacher marks every wrong answer with a red X and expects the errors to disappear. What did Thorndike's own data suggest, and what should the teacher do instead?** In the experiments behind the 1932 revision, telling a learner "wrong" did little or nothing to weaken the response, while "right" reliably strengthened the correct one. The X is the second half of the 1911 law, the half Thorndike withdrew. The teacher should make sure correct answers are followed promptly by something the student works for, and treat the errors by withholding that consequence and reinforcing the alternative, rather than expecting the mark itself to subtract the error. ## Frequently asked questions **What is the law of effect in simple terms?** Behavior that is followed by a satisfying result becomes more likely in that situation; behavior followed by an unpleasant result becomes less likely. Thorndike found it with hungry cats learning to escape from latched boxes: the movement that happened to open the door was repeated sooner on every trial. The first half of the law is the basis of reinforcement; Thorndike himself later dropped the second half. **Who discovered the law of effect?** Edward L. Thorndike, in puzzle-box experiments with cats, dogs, and chicks published in 1898 as his Columbia doctoral thesis. He stated the law formally in Animal Intelligence (1911). Alexander Bain had described the same principle in 1855, and Lloyd Morgan had argued for trial-and-error explanations in 1894, but Thorndike was the first to test it with repeated trials, many animals, and learning curves. **What is an example of the law of effect?** A cat in Thorndike's box A clawed at everything, happened to pull a loop, and the door opened; after two dozen trials it pulled the loop within seconds of being put in. Everyday examples: a dog that sits at the door and is let out sits sooner next time; a child who whines in the store and gets a snack whines more; a kick that unjams a vending machine is repeated the next time it jams. **What is the difference between the law of effect and reinforcement?** Reinforcement is Skinner's rebuilding of the law of effect. Thorndike connected a specific response to a situation through "satisfaction" and measured time to escape on separate trials. Skinner defined the operant as a class of responses, defined a reinforcer purely by its effect on the rate of responding, and measured rate in a free-operant chamber. Skinner also treated punishment as a separate process rather than the mirror image of reinforcement, which Thorndike had concluded in 1932. **What is the truncated law of effect?** The revised law Thorndike proposed in The Fundamentals of Learning (1932). In experiments with human learners told "right" or "wrong" after each response, "right" strengthened responses but "wrong" did little to weaken them. He concluded that annoyers do not weaken connections the way satisfiers strengthen them, and dropped the punishment half of his 1911 statement. The reward half is the truncated law. **What is the difference between the law of effect and the law of exercise?** Thorndike stated both in 1911. The law of exercise says a response becomes more strongly connected to a situation the more often, vigorously, and lastingly it has been made in it; repetition strengthens. The law of effect says consequences decide which responses are strengthened. Thorndike argued that effect was primary, and by 1932 he had concluded that repetition without an after-effect strengthens nothing, abandoning the law of exercise as an independent law. **Is the law of effect circular?** Only in its weak form, which is a definition: a reinforcer is whatever strengthens the response it follows. Paul Meehl showed in 1950 that the law still makes a testable claim, because reinforcers are trans-situational: an event that strengthens one response in one situation will strengthen other responses in other situations. That prediction could fail and has not. Thorndike had also defined satisfiers independently, by what the animal approaches and does not avoid. **How is the law of effect related to reinforcement learning in AI?** Reinforcement learning is the branch of machine learning in which an agent learns by acting and receiving a numerical reward. Sutton and Barto's textbook opens the field's history with Thorndike's law and notes that it contains both things trial-and-error learning needs: selection of actions by their outcomes, and association of those actions with the situations in which they occurred. An RL agent is a cat in a puzzle box with the reward written down as a number. ## References 1. Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. Full text in the library: /library/thorndike-animal-intelligence/ 2. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. Reprinted as Chapter II of Thorndike (1911). 3. Thorndike, E. L. (1932). *The Fundamentals of Learning*. Teachers College, Columbia University. 4. Chance, P. (1999). Thorndike's puzzle boxes and the origins of the experimental analysis of behavior. *Journal of the Experimental Analysis of Behavior, 72*(3), 433–440. 5. Bain, A. (1855). *The Senses and the Intellect*. John W. Parker. 6. Morgan, C. L. (1894). *An Introduction to Comparative Psychology*. Walter Scott. 7. Thorndike, E. L. (1927). The law of effect. *American Journal of Psychology, 39*(1/4), 212–222. 8. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 9. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40. 10. Postman, L. (1947). The history and present status of the law of effect. *Psychological Bulletin, 44*(6), 489–563. 11. Meehl, P. E. (1950). On the circularity of the law of effect. *Psychological Bulletin, 47*(1), 52–75. 12. Guthrie, E. R., & Horton, G. P. (1946). *Cats in a Puzzle Box*. Rinehart. 13. Moore, B. R., & Stuttard, S. (1979). Dr. Guthrie and *Felis domesticus* or: Tripping over the cat. *Science, 205*(4410), 1031–1033. 14. Skinner, B. F. (1981). Selection by consequences. *Science, 213*(4507), 501–504. 15. Herrnstein, R. J. (1970). On the law of effect. *Journal of the Experimental Analysis of Behavior, 13*(2), 243–266. 16. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [The history of operant conditioning](https://operantconditioning.com/history/): The timeline from Aristotle to reinforcement learning, and what Skinner changed. - [Reinforcement](https://operantconditioning.com/reinforcement/): What the law of effect became: both types, the kinds of reinforcers, and what makes them work. - [Thorndike's Animal Intelligence (1911)](https://operantconditioning.com/library/thorndike-animal-intelligence/): The full text, including the 1898 monograph and the chapter that states the law. # Learned Helplessness: Definition, Experiments & the Reversal > Learned helplessness explained: the triadic dog experiments, the human studies, the attributional reformulation, the depression model, and the 2016 reversal. - Source: https://operantconditioning.com/learned-helplessness/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Aversive control · Controllability* Dogs that had once received shocks they could not stop later lay down and took shocks they could have ended with a single jump. Seligman and Maier concluded that the dogs had learned that nothing they did mattered. Fifty years later the same two authors announced, on the strength of the neuroscience, that they had had it backward. > **Definition** > > Learned helplessness is the finding that an organism exposed to aversive events it cannot control later fails to escape or avoid aversive events it could control, and is slow to learn even when a response does succeed. Martin Seligman, Steven Maier, and Bruce Overmier reported it in dogs in 1967 and named it for their explanation: the animal had learned that its responses and the shocks were independent.[1][2][3] > > The name has outlived the explanation. In 2016 Maier and Seligman concluded from the neuroscience that passivity under prolonged aversive stimulation is the unlearned default, and that what animals given control learn is control.[4] This page uses the term for the phenomenon, as the literature still does, and says at each step which explanation is on the table. **In brief** - Dogs given shocks they could not stop later lay down under shocks they could have ended with a jump; dogs given the same shocks, but able to turn them off, escaped normally. Uncontrollability, not shock, produced the deficit. - The original theory said the animals learned that responding was useless; the attributional reformulation said that in people, what they conclude about *why* they are helpless governs how broad and how lasting the effect is. - Fifty years on, the neuroscience reversed the picture: passivity is the default response to prolonged aversive events, and what is learned, by the medial prefrontal cortex, is control. ## The dog experiments and the triadic design The phenomenon was noticed by accident in Richard Solomon's laboratory at the University of Pennsylvania: dogs that had received Pavlovian conditioning with shock while restrained in a hammock were later moved to a shuttle box to learn [escape and avoidance](https://operantconditioning.com/avoidance-learning/), and they would not learn. Bruce Overmier and Martin Seligman made it an experiment. Dogs in the hammock received 64 shocks that nothing they did could end. A day later each was placed in a two-compartment shuttle box where a signal was followed by shock, jumping the barrier ended the shock, and an unescaped shock ended on its own after a minute. Naive dogs jumped within a few trials. Most of the pre-shocked dogs ran about and howled during the first shock or two, then lay down, whined, and took the full minute, trial after trial; a dog that happened to jump and end a shock often failed to jump again. Two further details mattered later: the deficit from a single session was gone when the test was delayed two days, and dogs paralyzed with curare during the hammock shocks, who could not have learned any competing movement, showed it anyway.[1] The decisive experiment is Seligman and Maier's, published the same year, and it used what became known as the **triadic design**. One group of dogs could end each hammock shock by pressing a panel with its head. A second group was *yoked* to the first: each yoked dog received exactly the shocks its partner received, starting and stopping at the same moments, but its own presses did nothing. A third group received no hammock shocks. In the shuttle box the next day, the escape dogs and the naive dogs learned normally; most of the yoked dogs did not.[2] The escape and yoked groups had received identical shocks; the only difference was whether responding had mattered. Shock had not produced the deficit. Uncontrollable shock had. The same paper reported the finding that would matter most fifty years later. Dogs given escape training in the shuttle box first, and only then inescapable shock in the hammock, did not become helpless; Seligman and Maier called this **immunization**.[2] Later work in the same laboratory found that helpless dogs recovered if they were dragged across the barrier on a leash, again and again, until moving reliably ended the shock and they began to move on their own.[3] What was done to the dogs These experiments hurt the animals, and the papers do not hide it: Seligman and Maier's title is "Failure to escape traumatic shock." The shocks were intense, repeated, and, for the yoked dogs, impossible to end; the design depended on exactly that. The later work, including all of the neuroscience below, was done in rats and also involved inescapable pain. A study like this would face far stricter review today. The findings are here because they changed how psychology understands control, not because the methods would be acceptable now. ## The original theory: learning that responses do not matter Maier and Seligman's 1976 review laid the theory out. When shock ends whether or not the animal responds, the animal learns that the two are independent and carries that expectation into the next situation, where it does three things. It lowers the tendency to initiate responses at all (a **motivational deficit**); it makes the animal slow to notice that a response has worked when one does (an **associative deficit**); and it produces agitation followed by a flat passivity (an **emotional deficit**).[3] This was a cognitive claim made in an operant laboratory, and much of the review defends it against plainer alternatives. Exhaustion or physical damage? The yoked dogs had received the same shocks as dogs that escaped normally. A learned habit of holding still, accidentally reinforced by shock ending? Curarized dogs, who could not hold still or do anything else, became helpless too. Mere inactivity? An inactive animal that stumbles into a successful response should learn from it, and helpless animals did not. Immunization was the strongest card: one expectation appeared to protect against another. The theory also implied that animals can learn the *absence* of a relation between behavior and outcome, which a strict reading of the [law of effect](https://operantconditioning.com/law-of-effect/), in which consequences stamp responses in and nothing else is learned, has no room for.[3] ## Learned helplessness in people Donald Hiroto brought the design to college students in 1974, with loud noise through headphones in place of shock. One group could switch the noise off with a button; a yoked group heard the same noise, and its presses did nothing; a third group heard no noise. The test was a "finger shuttle box," in which sliding the hand from one side to the other stopped the noise. Students who had heard escapable noise, or none, learned to slide their hands; those who had heard inescapable noise were slower to escape and more often never did. Students who scored as "externals" on a locus-of-control scale, and students told the test was a matter of chance rather than skill, were more likely to fail.[5] Hiroto and Seligman then crossed two kinds of uncontrollable pretreatment, inescapable noise and unsolvable concept-formation problems, with two kinds of test, noise escape and anagram solving. Helplessness crossed over: inescapable noise impaired anagrams, and unsolvable problems impaired noise escape. Whatever people had acquired behaved like a general expectation.[6] Rats, meanwhile, became the standard animal: a rat in a restraining tube receives tailshock it can end by turning a small wheel, its yoked partner receives the identical shock and can do nothing, and both are later tested in a shuttle box.[7] The human effects were real but modest, and they moved with what subjects were told and believed about the task, which nothing in the dog data had led anyone to expect. Explaining that took a different kind of theory. ## The attributional reformulation and hopelessness theory Even before the reformulation, Carol Dweck had shown that what a child concludes from failure can be changed. In 1975 she worked with twelve children whom school staff had identified as reacting to failure with a collapse in performance. Over twenty-five sessions, half received "success only": problems they could solve, and nothing else. The other half received **attribution retraining**: a few problems were arranged to be failed, and after each failure the trainer told the child that it came from not trying hard enough. When failure was then reintroduced, the success-only children still fell apart, some worse than before; the retrained children held or improved their performance.[8] A diet of success had taught nothing about failure. It was a small study, and the reformulation gave it a theory. Abramson, Seligman, and Teasdale published the **attributional reformulation** in 1978. The original theory, they argued, could not say why helplessness in people sometimes cost self-esteem and sometimes did not, why it was sometimes confined to one task and sometimes spread to everything, or why it sometimes passed in an hour and sometimes lasted. Their answer: when people find an outcome uncontrollable they ask why, and the attribution determines what follows. The reformulation also distinguished *universal* helplessness, in which no one could have controlled the outcome, from *personal* helplessness, in which others could have and I could not.[9] | Dimension | What it governs | After a failed exam | | --- | --- | --- | | Internal vs. external | Whether self-esteem suffers | "I am not smart" vs. "the exam was unfair" | | Stable vs. unstable | How long the helplessness lasts | "I never could" vs. "I was exhausted" | | Global vs. specific | How far the helplessness spreads | "I am bad at everything" vs. "I am bad at chemistry" | A habit of explaining bad events by internal, stable, global causes was proposed as a vulnerability: not a cause of depression by itself, but a style that turns bad events into broad and lasting helplessness.[9] In 1989 Abramson, Metalsky, and Alloy revised the theory again into **hopelessness theory**. The proximal cause of a proposed subtype, hopelessness depression, is hopelessness itself: the expectation that desired outcomes will not occur, or aversive ones will, and that nothing one does will change it. Farther back sit negative life events and a tendency to read them as stable and global, as having severe consequences, and as saying something bad about the self. Helplessness became one ingredient of hopelessness, and internal attributions were demoted to a role in self-esteem.[10] It is a diathesis-stress theory about people, tested with questionnaires and prospective studies, a long way from a hammock in Philadelphia. ## Learned helplessness and depression: a model, not a diagnosis Seligman proposed learned helplessness as a laboratory model of depression in his 1975 book *Helplessness*, on the strength of parallels: helpless animals and depressed people both initiate less behavior, are slow to recognize that their actions have had effects, and show a flattening of mood; both conditions, he argued, can follow loss of control over important outcomes; and both, in the model, are relieved by experiences in which action reliably produces results and prevented by prior mastery.[11] Peterson, Maier, and Seligman revisited the model two decades later and concluded, with care, that it captures some features of some depressions, in particular those that follow uncontrollable bad events and are accompanied by hopelessness.[12] Several things should be kept straight. A model is an analogy that generates hypotheses; an animal cannot be diagnosed with a mood disorder. The human laboratory effects come from a few minutes of noise or unsolvable puzzles and are small and short-lived. Hopelessness theory is about a subtype, and its distal cause is a cognitive style *plus* bad events. And Maier and Seligman noted in 2016 that the animal pattern, which includes exaggerated fear and reduced social exploration as well as passivity, looks at least as much like anxiety as like depression.[4] A model, not a diagnosis Nothing in this literature says that depression is a failure of effort, that a depressed person has "learned to be helpless," or that a change of attitude is a treatment. The theory describes one route by which uncontrollable events can produce passivity and hopelessness, and offers it as one contributor among many. Questions about a person's own mood belong with a clinician, not with a page about dogs and shuttle boxes. ## The 2016 reversal: passivity is the default, control is learned From the 1990s Maier's laboratory took the triadic design into the rat brain. Uncontrollable tailshock strongly activates the serotonin neurons of the **dorsal raphe nucleus** in the brainstem, far more than identical controllable shock does. The activation leaves the nucleus sensitized for a period afterward, and the serotonin it then releases in its targets produces the behavioral effects: failure to escape in a shuttle box, exaggerated conditioned fear, reduced social exploration. Blocking the activation during inescapable shock prevented these effects, and activating the nucleus pharmacologically, with no shock at all, produced them.[7] The passivity had a circuit, and the circuit required no learning. But the dorsal raphe cannot tell whether a shock is controllable; it receives the shock either way. Amat and colleagues found the detector in 2005 in the ventral medial prefrontal cortex. When they inactivated it during *escapable* shock, escapable shock behaved like inescapable shock: the dorsal raphe was activated and the rats later showed the full set of deficits, even though every shock had ended by their own wheel turn.[13] Later work found the converse: activating the region during inescapable shock prevented the deficits.[4] The prefrontal cortex detects that a response controls the stressor and inhibits the dorsal raphe. Without that signal, the raphe's response is simply what happens. Immunization got a mechanism too. An experience of control changes the prefrontal circuit so that, in a later stressor no response can end, the dorsal raphe is inhibited anyway; and because control over one stressor protects against a different one later, what the circuit learns is not tied to the wheel. That is why the immunized dogs of 1967 could take inescapable shock in a hammock and still jump in a shuttle box.[4] In the fiftieth-anniversary paper Maier and Seligman drew the conclusion. The dogs in the inescapable condition had not learned that responses were useless; they had not needed to learn anything. Prolonged aversive stimulation activates the dorsal raphe, and the passivity and fear that follow are the unlearned default. What was learned, in the escapable condition, was that a response worked, and that learning inhibited the default. The original theory, they wrote, had it backward.[4] The shuttle box could never have shown this, because a test that measures only whether the animal jumps cannot tell a learned expectation of no control from an unlearned failure to engage. The wider [neuroscience of reinforcement](https://operantconditioning.com/neuroscience/) has its own page. | Question | Original theory (1967–1976) | Neuroscience account (2016) | | --- | --- | --- | | During inescapable shock | The animal learns that responses and outcomes are independent | Nothing is learned; the shock activates the dorsal raphe by default | | During escapable shock | An ordinary escape response is learned | The medial prefrontal cortex detects control and inhibits the dorsal raphe; this is the learning | | Source of later passivity | A learned expectation of uncontrollability | An unlearned serotonergic default never switched off | | Why prior control protects | An earlier expectation competes with the new one | Prefrontal plasticity keeps the inhibitory circuit available under later uncontrollable stress | | Why it transfers | The expectation generalizes | The learning is about control itself, not a particular response | | What is learned | Helplessness | Control (the old name stayed) | ## Where helplessness sits in operant conditioning The shuttle box is an escape and avoidance task, so what a helpless animal fails to acquire is behavior maintained by [negative reinforcement](https://operantconditioning.com/negative-reinforcement/): a jump that removes the shock. Procedure and process should be kept apart. Inescapable shock is a procedure, aversive stimulation delivered independently of behavior; learned helplessness names the process, the later failure of escape and avoidance learning, and, until 2016, the theory of it. Inescapable shock is not [punishment](https://operantconditioning.com/punishment/), because punishment is a consequence of a response and inescapable shock is a consequence of nothing. Helplessness and avoidance are mirror-image failures of contact with a contingency. Solomon, Kamin, and Wynne's well-trained avoiders kept jumping after the shock was switched off, because a dog that jumps early never finds out that the contingency has changed; what stopped them was a barrier that prevented the jump.[14] A helpless dog never finds out that a contingency exists, because it does not respond, and when a response does end the shock the relief fails to strengthen it; what worked was a leash that made the dog move.[3] In both cases behavior is insulated from its consequence, and in both the remedy was forced contact. See [avoidance learning](https://operantconditioning.com/avoidance-learning/) and [extinction](https://operantconditioning.com/extinction/). The other mirror image is [learned industriousness](https://operantconditioning.com/glossary/#learned-industriousness). Robert Eisenberger's 1992 review concluded that reinforcing high effort on one set of tasks increases effort and persistence on unrelated ones, because the sensation of effort, repeatedly paired with reinforcement, loses aversiveness and becomes a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer); reinforcing low effort does the opposite.[15] Helplessness follows when effort has not mattered; industriousness follows when it has. ## What this means in practice The literature supports a small number of practical statements. First, **controllability matters separately from the amount of stress**: two animals with identical shock histories differ afterward according to whether their responses worked. Second, **experience of control is protective**: it prevented helplessness in 1967, it undid it, and on the 2016 account the prefrontal learning that a response works stays available in later stressors the animal cannot control. Maier and Seligman treat such experiences, especially early ones, as the practical lesson of the whole literature.[4] Turned into arrangements, the suggestion looks like this. In a [classroom](https://operantconditioning.com/classroom/), a student who is behind meets tasks at a level where effort produces success, which is what [shaping](https://operantconditioning.com/shaping/) is for, rather than a run of failures no effort can change. In [dog training](https://operantconditioning.com/dog-training/), the dog's behavior is what produces the outcomes. In a [workplace](https://operantconditioning.com/workplace/), feedback follows what a person did rather than arriving on its own timetable. For [parents](https://operantconditioning.com/parenting/), the point is not to remove hardship but to keep a child's actions connected to results. None of these is an experiment; they are the shape the evidence suggests, and the evidence comes from rats, dogs, and a few minutes of noise. Ordinary setbacks are not inescapable shock. ## Common misreadings, and what the evidence does not show ### Common misreadings - **"Helplessness is learned."** The name teaches this, and on the evidence it is backward: passivity under prolonged uncontrollable stress is the default, and control is what has to be learned.[4] - **"Helpless means lazy."** Dogs paralyzed with curare became helpless without moving at all, dogs dragged across the barrier recovered, and failing to learn from a successful response is not a matter of effort.[1][3] - **"It is a kind of punishment."** Punishment is a consequence of a response. Inescapable shock is a consequence of nothing; that is the point of the design. - **"Any failure produces it."** The procedure is uncontrollability, not failure. Dweck's retraining used deliberate failure, paired with the message that effort could change it, and it helped.[8] - **"It is permanent."** After one session the dogs' deficit was gone within two days; immunization prevented it; forced exposure reversed it.[1][2][3] - **"Depressed people have learned helplessness."** The model captures some features of some depressions. It is not a diagnosis.[12] ### What the evidence does not show - That the human laboratory effect is the same process as the animal one. The human effects are small, brief, and sensitive to instructions and beliefs; the reformulation exists because the dog theory did not fit them.[5][9] - That an attributional style causes depression on its own. Hopelessness theory is a diathesis-stress proposal about a subtype.[10] - That the circuit found in rats maps directly onto human mood disorders. Its authors note that its behavioral signature looks as much like anxiety as like depression.[4] - That telling someone they have control is the same as having it. Instructions do move the human laboratory effect, but the animal evidence and immunization concern actual contingencies between responses and outcomes.[5][2] - That passivity in general is learned helplessness. Most passivity has a history of ordinary reinforcement and extinction behind it; the term adds nothing unless uncontrollable aversive events are in that history. ## Key takeaways - Dogs given identical shocks differed afterward only by whether their responses had ended them, and only the dogs without control later failed to escape. Uncontrollability, not shock, produced the deficit. - The original theory held that the animals learned that responses and outcomes were independent; immunization by prior control and recovery through forced exposure were its strongest supports. - In people the effect was modest and moved by beliefs, which led to the attributional reformulation (internal, stable, global) and then to hopelessness theory, a diathesis-stress account of a proposed subtype of depression. - Learned helplessness is a laboratory model of some features of some depressions, not a diagnosis. - The neuroscience reversed the theory: uncontrollable stress activates the dorsal raphe by default, and the medial prefrontal cortex learns control and inhibits it. Helplessness is not learned; control is. - Controllability matters separately from the amount of stress, experiences of control are protective, and learned industriousness is the mirror image. ### Check yourself **Two dogs are strapped side by side. Every shock the first dog receives is copied exactly to the second. The first can end each shock by pressing a panel; the second cannot. The next day both are tested in a shuttle box. What is the design called, what does each dog do, and what does the design rule out?** The triadic design, once a third, unshocked group is added. The first dog escapes normally; the second mostly does not. Because the two received identical shocks, the design rules out the shock itself, exhaustion, and physical damage as explanations, and leaves only the difference in control. **A colleague says, "The dogs learned to be helpless, so we should teach people that they aren't." What does the 2016 account change about the first half of that sentence, and what does it leave intact?** It reverses the first half: the dogs in the inescapable condition did not learn helplessness; passivity was the default, and it was the escapable dogs that learned something, namely control. It leaves the practical direction intact and strengthens it: experiences in which responses actually work, before or after uncontrollable stress, are what protect, and the prefrontal circuit that carries that learning is what inhibits the default. **A student fails a quiz and says, "I'm just bad at this." Using the 1978 reformulation, which features of that attribution predict a broad and lasting effect, and what did Dweck's retraining change?** "Bad at this" is internal (about me) and stable (a trait); whether it is global depends on how much "this" covers. Internal attributions bear on self-esteem, stable ones on how long the helplessness lasts, global ones on how far it spreads. Dweck taught children to attribute failure to insufficient effort, which is unstable and changeable, after which their performance survived failure, while a diet of success alone did not help. ## Frequently asked questions **What is learned helplessness in simple terms?** After enough experience with bad events that nothing you do can stop, you stop trying even when trying would work. In the original experiments, dogs given shocks they could not end later lay down and took shocks they could have escaped with a jump. The 2016 revision adds that this passivity is the default, and that what has to be learned is that your actions matter. **What was the learned helplessness experiment?** Seligman and Maier (1967) gave one group of dogs shocks they could end by pressing a panel, gave a yoked group exactly the same shocks with no way to end them, and gave a third group none. The next day, in a shuttle box where a jump ended the shock, the first and third groups learned to jump and most of the yoked dogs did not. Only the lack of control produced the failure. **Is learned helplessness actually learned?** According to its own authors, no. Maier and Seligman concluded in 2016 that passivity under prolonged uncontrollable stress is an unlearned default driven by serotonin neurons in the dorsal raphe nucleus, and that what animals with control learn, through the medial prefrontal cortex, is control, which inhibits the default. The phenomenon is real; the name describes the wrong direction. **What is the difference between learned helplessness and hopelessness theory?** Learned helplessness is the laboratory finding and the original theory that organisms learn their responses do not matter. Hopelessness theory (Abramson, Metalsky, and Alloy, 1989) is a theory about people: a tendency to explain bad events by stable, global causes, combined with negative life events, produces hopelessness, which is proposed as the proximal cause of a subtype of depression. Helplessness is one ingredient of hopelessness, not the whole of it. **Does learned helplessness cause depression?** It is a laboratory model of some features of some depressions, proposed by Seligman in 1975, not a demonstrated cause and not a diagnosis. The human versions, the attributional reformulation and hopelessness theory, are diathesis-stress accounts of a proposed subtype. Maier and Seligman also note that the animal pattern, passivity plus heightened fear, looks at least as much like anxiety as like depression. **What is an example of learned helplessness?** The clearest examples are from the laboratory: dogs that lie still through escapable shock after inescapable shock, and students who sit through loud noise they could switch off after hearing noise they could not. Everyday cases are harder to verify, because the term applies only when uncontrollable aversive events are actually in the history; a person who has stopped trying may simply have a history of ordinary extinction. **What is the attributional reformulation of learned helplessness?** Abramson, Seligman, and Teasdale (1978) proposed that when people find an outcome uncontrollable they ask why, and the answer shapes what follows. Internal attributions affect self-esteem, stable ones make the helplessness last, and global ones make it spread. A habit of explaining bad events by internal, stable, global causes was proposed as a vulnerability, not a cause on its own. **Can learned helplessness be reversed?** In the animal work, yes. Dogs given prior experience of control did not become helpless (immunization), and helpless dogs recovered when they were repeatedly moved across the barrier until moving reliably ended the shock. In children, Dweck (1975) found that teaching a small group to attribute failure to insufficient effort helped where success alone did not. These are laboratory results, not treatment advice. ## References 1. Overmier, J. B., & Seligman, M. E. P. (1967). Effects of inescapable shock upon subsequent escape and avoidance responding. *Journal of Comparative and Physiological Psychology, 63*(1), 28–33. 2. Seligman, M. E. P., & Maier, S. F. (1967). Failure to escape traumatic shock. *Journal of Experimental Psychology, 74*(1), 1–9. 3. Maier, S. F., & Seligman, M. E. P. (1976). Learned helplessness: Theory and evidence. *Journal of Experimental Psychology: General, 105*(1), 3–46. 4. Maier, S. F., & Seligman, M. E. P. (2016). Learned helplessness at fifty: Insights from neuroscience. *Psychological Review, 123*(4), 349–367. 5. Hiroto, D. S. (1974). Locus of control and learned helplessness. *Journal of Experimental Psychology, 102*(2), 187–193. 6. Hiroto, D. S., & Seligman, M. E. P. (1975). Generality of learned helplessness in man. *Journal of Personality and Social Psychology, 31*(2), 311–327. 7. Maier, S. F., & Watkins, L. R. (2005). Stressor controllability and learned helplessness: The roles of the dorsal raphe nucleus, serotonin, and corticotropin-releasing factor. *Neuroscience & Biobehavioral Reviews, 29*(4–5), 829–841. 8. Dweck, C. S. (1975). The role of expectations and attributions in the alleviation of learned helplessness. *Journal of Personality and Social Psychology, 31*(4), 674–685. 9. Abramson, L. Y., Seligman, M. E. P., & Teasdale, J. D. (1978). Learned helplessness in humans: Critique and reformulation. *Journal of Abnormal Psychology, 87*(1), 49–74. 10. Abramson, L. Y., Metalsky, G. I., & Alloy, L. B. (1989). Hopelessness depression: A theory-based subtype of depression. *Psychological Review, 96*(2), 358–372. 11. Seligman, M. E. P. (1975). *Helplessness: On Depression, Development, and Death*. W. H. Freeman. 12. Peterson, C., Maier, S. F., & Seligman, M. E. P. (1993). *Learned Helplessness: A Theory for the Age of Personal Control*. Oxford University Press. 13. Amat, J., Baratta, M. V., Paul, E., Bland, S. T., Watkins, L. R., & Maier, S. F. (2005). Medial prefrontal cortex determines how stressor controllability affects behavior and dorsal raphe nucleus. *Nature Neuroscience, 8*(3), 365–371. 14. Solomon, R. L., Kamin, L. J., & Wynne, L. C. (1953). Traumatic avoidance learning: The outcomes of several extinction procedures with dogs. *Journal of Abnormal and Social Psychology, 48*(2), 291–302. 15. Eisenberger, R. (1992). Learned industriousness. *Psychological Review, 99*(2), 248–267. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Avoidance learning](https://operantconditioning.com/avoidance-learning/): Escape, avoidance, and the paradox; where the shuttle box comes from. - [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/): The quadrant the shuttle-box jump belongs to, and why it is not punishment. - [The neuroscience of reinforcement](https://operantconditioning.com/neuroscience/): Dopamine, prediction error, and the circuits that learn from consequences. # Discriminative Stimulus (SD): Definition, Examples & S-Delta > A discriminative stimulus (SD) is a cue in whose presence a behavior has been reinforced. Definition, S-delta, examples, SD vs. CS and MO, how to train one. - Source: https://operantconditioning.com/discriminative-stimulus/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Antecedents · Stimulus control* A phone rings and you pick it up; the light goes off and the rat stops pressing. Nothing forced either response. The discriminative stimulus is the antecedent that makes a behavior likely without compelling it — here is how it is defined, how it is built, and how to tell it from the two things it is most often confused with. > **Definition** > > A discriminative stimulus, written SD and pronounced "ess-dee," is an antecedent stimulus in whose presence a response has been reinforced and in whose absence it has not, and which, because of that history, raises the probability of the response when it is present. Its counterpart, the **S-delta** (SΔ), is a stimulus in whose presence the same response has gone unreinforced.[1] > > B. F. Skinner introduced the term and the notation in *The Behavior of Organisms* (1938). An SD does not elicit behavior the way a puff of air elicits a blink; it sets the occasion for it, and only because of what has followed the behavior in its presence before.[1][2] **In brief** - A discriminative stimulus is defined by history, not appearance: a cue is an SD for a behavior only if that behavior has been reinforced in its presence and not in its absence. - An SD evokes an operant rather than eliciting a reflex, and it signals that reinforcement is *available*; a motivating operation is the different thing that changes how much the reinforcer is *worth*. - To put a behavior under an SD, reinforce it in the presence of the cue, withhold reinforcement in its absence, and keep the difference consistent until responding diverges. ## How a discriminative stimulus works Skinner's original demonstration is still the cleanest. A rat presses a lever for food, and the food is then made available only while a light is on: presses in the light produce a pellet, presses in the dark produce nothing. At first the rat presses at about the same rate either way. Over sessions the rates separate, until the rat presses briskly when the light comes on and barely touches the lever when it goes off. The light has become a discriminative stimulus, and the response is a **discriminated operant**.[1] Three things in that description carry the definition. The SD is an antecedent, present before the response rather than delivered after it. It works by *changing probability*: the rat can press in the dark and sometimes does, and the gap between the two rates measures how much control the light has. And its power is borrowed entirely from the consequence. Take the food away in the light and the light stops mattering within a few sessions.[1][2] ### Notation: SD, SΔ, S+, S−, and SDp Skinner wrote the discriminative stimulus as SD and the stimulus correlated with extinction as SΔ, "S-delta."[1] The generalization literature usually writes the pair as S+ and S−. Applied texts add SDp for a stimulus in whose presence a response has been *punished* and in whose absence it has not.[3] In each case the superscript names the consequence correlated with the stimulus. ### The three-term contingency The SD is the first term of the [three-term contingency](https://operantconditioning.com/abc-model/) — in the presence of SD, response R produces reinforcer SR — which Skinner made the basic unit of analysis in *Science and Human Behavior*.[2] The consequence explains why the behavior exists; the discriminative stimulus explains when and where it appears. ## How a discriminative stimulus is established A stimulus becomes an SD through **discrimination training**: the response is reinforced in the presence of one stimulus and undergoes [extinction](https://operantconditioning.com/extinction/) in the presence of another until responding in the two conditions diverges.[1] The procedure is [differential reinforcement](https://operantconditioning.com/differential-reinforcement/) applied to the antecedent; the process it produces, responding differently in the two conditions, is discrimination. Both halves are necessary. Reinforcement in the light builds the response; extinction in the dark is what attaches it to the light. The second half is easy to forget, and one experiment shows what happens without it. Jenkins and Harrison trained pigeons to peck a key with a tone sounding continuously through every session. The tone was present for every reinforced peck, yet when the birds were later tested with other pitches they pecked about equally at all of them: the tone controlled nothing. Birds for whom the tone had been on during reinforcement and off during extinction responded sharply to the training pitch and little to the others.[4] A stimulus that predicts no difference in consequences gains no control, however often it accompanies reinforcement. Ordinary discrimination training involves errors: the learner responds to the SΔ, is not reinforced, and gradually stops. Herbert Terrace showed in 1963 that the errors can be engineered out. He trained pigeons to peck a red key and not a green one, but introduced the green key from the first session so dim and so brief that the birds never pecked it, then raised its brightness and duration in small steps. Birds trained this way made few or no errors, where conventionally trained birds made thousands.[5] The SD was established just as surely; only the way the SΔ was introduced changed. [Errorless learning and its side effects ›](https://operantconditioning.com/stimulus-control/) Dinsmoor's review adds a requirement the contingency alone does not guarantee: the organism has to actually look at or listen to the stimulus. A cue the learner never attends to cannot become an SD.[6] The stimulus is not special; the history is The same light is an SD for a rat that has been fed for pressing in its presence and nothing at all for a rat that has not. Before calling a cue an SD, ask what the behavior has produced in its presence and in its absence. If the answer is "the same thing," it is not an SD yet. ## Examples of discriminative stimuli In each row the behavior has paid in the presence of the SD and not in its absence. | Setting | SD | Behavior | What has followed it | SΔ | | --- | --- | --- | --- | --- | | Everyday | Phone ringing | Picking it up and saying hello | A voice on the line | A silent phone | | Driving | Green light | Pressing the accelerator | Getting through the intersection | Red light | | Classroom | Teacher's raised hand (the quiet signal) | Stopping talking and raising a hand | Praise; the activity the class was waiting for begins | Teacher's hand down, working with a small group | | Dog training | The word "sit" | Sitting | A treat, or the leash going on | Any other word, or silence | | Retail | A lit "Open" sign | Pulling the door | The door opens | A dark sign | | Social | A particular friend | Telling a particular kind of joke | Laughter | Your manager, who has never laughed at it | The same event is usually an SD for one behavior and an SΔ for another: the red light is an SΔ for accelerating and an SD for braking; the raised hand is an SD for hand-raising and an SΔ for calling out. A stimulus has a discriminative function *for a particular response*, and the function has to be stated with the response attached.[2] The "particular friend" row deserves a pause. Skinner devoted a chapter of *Verbal Behavior* to the audience as a discriminative stimulus: the person you are speaking to controls not only whether you speak but what you say, because different listeners have reinforced different things.[7] The child who "behaves differently for the substitute" is discriminating between two people with two histories of consequences. ## Discriminative stimulus vs. conditioned stimulus The nearest neighbor is the conditioned stimulus (CS) of classical conditioning. Both are learned antecedents, but they are established by different histories and do different things. A CS acquires its function by being paired with an unconditioned stimulus — a tone followed by food, whatever the animal does in between. Rescorla's summary of the modern view is that the animal learns the *predictive relation* between the two events: a CS works to the extent that it carries information about the unconditioned stimulus, and contiguity without prediction produces little conditioning.[8] Nothing the animal does is part of that relation. The tone predicts food whether the dog salivates or not, and the salivation that comes to follow the tone is *elicited*, on a reflex model, without depending on its own consequences. An SD acquires its function through a contingency that runs through the behavior. The light predicts food only *if the rat presses*; without the press there is no pellet, light or no light. The response is emitted in the stimulus's presence at a rate the stimulus modulates, and Skinner's formulation was that the SD sets the occasion for the response rather than eliciting it.[1][2] An elicited response follows its stimulus with roughly fixed form and latency; an evoked operant varies in speed and force and can fail to occur at all. | Aspect | Discriminative stimulus (SD) | Conditioned stimulus (CS) | | --- | --- | --- | | **Established by** | Reinforcement of a response in its presence and extinction in its absence | Predicting an unconditioned stimulus, regardless of behavior | | **What it does** | Evokes an operant: raises its probability | Elicits a respondent: triggers a conditioned reflex | | **Does the behavior matter?** | Yes; the reinforcer depends on the response | No; the unconditioned stimulus comes anyway | | **Example** | Lit key: pecking now produces grain | Tone before food: salivation follows the tone | One event often carries both functions: the click of the food magazine is a CS that elicits approach, a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer) for the press before it, and an SD for going to the tray. Ask which function an event serves for which response. [Operant vs. classical conditioning in depth ›](https://operantconditioning.com/operant-vs-classical-conditioning/) ## Discriminative stimulus vs. motivating operation The confusion that matters more in practice is with the motivating operation. In 1982 Jack Michael pointed out that "SD" was being used for two different antecedent functions. One is discriminative: the stimulus is correlated with the *availability* of a reinforcer, obtainable in its presence and withheld in its absence. The other is motivational: the event changes the *effectiveness* of a reinforcer, how much it is worth right now, and in doing so evokes the behavior that has produced it. He called the second an **establishing operation** and argued that the two should not share a name.[9] The test is what happens to the reinforcer in the stimulus's absence. A drinking fountain is an SD for walking over and pressing the bar: water has been available there and not elsewhere. A bag of salted peanuts is not an SD for the same behavior, though it makes the behavior far more likely; water was no less available before the peanuts, what changed is how much it is worth. Likewise, the vending machine says a snack is available; six hours without food says it is worth having. Both are usually needed before the behavior occurs.[9][10] Michael drew a consequence that still surprises people. The onset of a painful stimulus is routinely called an SD for escape, but it fails the test: the reinforcer for escape, the pain ending, is not merely unavailable before the pain starts, it does not exist. Pain onset is an establishing operation for its own removal. The SD for escape, if there is one, is whatever signals that a particular response will work, such as the lever that has ended the shock before.[9] Michael's 1993 paper set out the two effects every establishing operation has, on the value of a consequence and on the current frequency of behavior that has produced it.[10] Laraway and colleagues completed the vocabulary in 2003: [motivating operation](https://operantconditioning.com/glossary/#motivating-operation) is the umbrella term, an *establishing* operation raises a reinforcer's effectiveness and an *abolishing* operation lowers it, and the behavioral effect is *evocative* or *abative*.[11] | Aspect | Discriminative stimulus (SD) | Motivating operation (MO) | | --- | --- | --- | | **What it changes** | The probability of the response, given a history of differential reinforcement | The value of the reinforcer, and with it the probability of the response | | **What it signals** | That the reinforcer is available now | Nothing about availability; it makes the reinforcer worth more or less | | **Test** | Was the reinforcer obtainable in its presence and withheld in its absence? | Is the reinforcer worth more or less after this event? | | **Example** | The drinking fountain; the vending machine; the light above the lever | Salted peanuts; six hours without food; food deprivation before a session | A cue with an impeccable history does nothing if the reinforcer it signals is currently worth nothing, which is why "sit" fails after a bowl of treats. When a cue stops working, ask whether the problem is the cue's history or the reinforcer's value. ## Discrimination, generalization, and natural concepts An SD rarely controls in isolation; control spreads to stimuli that resemble it. Guttman and Kalish reinforced pigeons for pecking a key lit with a single wavelength and then, in extinction, presented nearby wavelengths. Responding was highest at the training wavelength and fell away smoothly on either side, a generalization gradient.[12] Discrimination training steepens the gradient, and when the SΔ lies close to the SD on the same dimension it moves the peak away from the SΔ, the [peak shift](https://operantconditioning.com/glossary/#peak-shift) Hanson demonstrated in 1959.[13] Dinsmoor's review covers the mechanics, and both effects are treated in full on the [stimulus control page](https://operantconditioning.com/stimulus-control/).[6] For the definition, the point is that "the SD" is the center of a region of stimuli that share its control, and training sets how sharp that region is. At the other extreme, the "stimulus" can be a category no single feature defines. Herrnstein, Loveland, and Cable reinforced pigeons for pecking at color slides that contained a tree, or a body of water, or one particular young woman, and not at slides that did not. The birds learned each discrimination and transferred it to slides they had never seen.[14] The SD here is "tree," a class with no common wavelength, shape, or size: the discriminative stimulus is as abstract as the history of reinforcement makes it. ## How to bring a behavior under a discriminative stimulus The procedure is the same whether the learner is a rat, a dog, a child, or you. If the behavior does not yet occur, [shape](https://operantconditioning.com/shaping/) it first; a cue cannot be attached to a response that is never emitted. 1. **Choose the cue.** Pick a stimulus that is distinct, easy to perceive, under your control to present and withhold, and present where you want the behavior. "Sit" said once in a normal voice; a raised hand; one desk at one hour. Avoid a cue that is already an SD for something else. 2. **Reinforce the behavior in the presence of the cue.** Present the cue, wait for the response, reinforce immediately. Early on, reinforce every correct response and, if needed, use a [prompt](https://operantconditioning.com/glossary/#prompt) to get the response going, then fade the prompt so the cue is what remains.[3] 3. **Ensure extinction in the cue's absence.** This is the half most people skip. If the dog is fed for sitting whenever it sits, the word "sit" predicts nothing and will control nothing. 4. **Add the SΔ explicitly.** Present other stimuli — other words, a different gesture, the same word from a different posture — and withhold reinforcement when the behavior follows them. This sharpens control from "anything the trainer says" to "this word." Introduce the SΔ gradually to keep errors low, as Terrace did.[5] 5. **Thin the schedule and vary the setting.** Once responding is reliable in the cue's presence and rare in its absence, move to an [intermittent schedule](https://operantconditioning.com/schedules-of-reinforcement/) so the behavior persists, and train in other rooms, with other people, and around distractions so that control belongs to the cue rather than the kitchen.[3] The same steps describe how a [dog trainer](https://operantconditioning.com/dog-training/) adds a cue and how a [classroom](https://operantconditioning.com/classroom/) quiet signal is taught in the first week of school. Reinforcement in the cue's presence is what people remember to arrange; extinction in its absence is what makes the cue work. ## Common mistakes - **Calling any cue an SD.** A stimulus is a discriminative stimulus for a response only if that response has been differentially reinforced in its presence. A sign nobody has ever followed, an instruction never backed by a consequence, or a light that is always on is a stimulus, not an SD.[1][9] - **Confusing the SD with a prompt.** A prompt is a supplementary antecedent added because the natural SD does not yet evoke the behavior: the pointed finger, the "What do you say?", the hand guiding the dog into a sit. It is meant to be faded; the SD is meant to stay.[3] - **Confusing the SD with a motivating operation.** Hunger, thirst, pain, and boredom make behavior more likely without signaling that anything is available. If removing the event would change only how much the reinforcer is wanted, not whether it can be obtained, it is an MO.[9] - **Treating the SΔ as a punisher.** An SΔ signals that the response will go unreinforced, not that it will be punished. Extinction and punishment are different procedures with different side effects. - **Assuming the cue you intended is the cue that controls.** Learners discriminate whatever actually predicts reinforcement, which is often the trainer's posture, the treat pouch, or the room rather than the word. If the behavior vanishes when those change, they were the SD. ## What the evidence does not show The literature on discriminative stimuli is among the oldest and most replicated in psychology, and its limits are worth stating. - **It does not show that stimuli control behavior without a consequence history.** Every demonstration of an SD depends on differential reinforcement, and the control disappears when the differential is removed. A claim that a cue "triggers" a habit on its own is a claim about a history.[1][4] - **It does not show that an SD works regardless of motivation.** A discriminative stimulus evokes a response only when the reinforcer it signals currently has value; under an abolishing operation the same stimulus evokes nothing.[9][11] - **It does not show that the organism learned the stimulus you had in mind.** Generalization tests routinely reveal control by a feature the experimenter did not intend, and a pigeon can respond to a category with no single defining feature. What was learned is inferred from the gradient, not observed.[12][14] - **It does not settle how much of an everyday cue's power is discriminative.** Many everyday stimuli carry eliciting and evoking functions at once, and the experiments do not say which is doing more of the work when your phone buzzes.[8] ## Key takeaways - A discriminative stimulus is an antecedent in whose presence a response has been reinforced and in whose absence it has not. The definition is a history; no stimulus is an SD on its own. - An SD evokes an operant by raising its probability; it does not elicit a reflex. A conditioned stimulus predicts an unconditioned stimulus regardless of behavior; an SD works through a contingency that runs through the behavior. - An SD signals that a reinforcer is available; a motivating operation changes how much it is worth. The drinking fountain and the salted peanuts both make you drink, and only the first is an SD. - Discrimination training — reinforcement in the cue's presence, extinction in its absence — establishes an SD, and the extinction half is what makes the cue predictive. - Control spreads to similar stimuli and can be trained onto a category as abstract as "tree." Learners discriminate whatever actually predicts reinforcement, not necessarily the cue the teacher intended. - To put behavior under an SD: choose the cue, reinforce only in its presence, ensure extinction in its absence, add the SΔ explicitly, then thin the schedule and vary the setting. ### Check yourself **A dog sits every time its owner says "sit" in the kitchen, and also sits every time the owner reaches toward the treat pouch, whether or not the word is said. What is the SD?** The reach toward the pouch, or the pouch itself. The word has never predicted a difference in consequence that the pouch did not already predict, so control has settled on the pouch. To make "sit" the discriminative stimulus, the owner needs to reinforce sits that follow the word alone and stop reinforcing sits that follow the reach without the word. **A rat is shocked through the floor, and pressing a lever turns the shock off. A textbook calls the shock onset the SD for pressing. Is that right?** By Michael's analysis, no. An SD is a stimulus in whose presence the reinforcer has been available and in whose absence it has been withheld. Before the shock, its removal is not withheld; it does not exist. Shock onset is an establishing operation: it makes shock termination a reinforcer and evokes the behavior that has produced it. If there is an SD for pressing, it is the lever, which signals which response will work. **A sign in the office kitchen says "Please wash your own mug." Nobody does. Is the sign a discriminative stimulus for mug-washing?** No. An SD is defined by a history of differential reinforcement, and mug-washing has never been reinforced in the sign's presence more than in its absence; the sign predicts no difference in consequence. It is a stimulus, perhaps a prompt, but not yet a discriminative stimulus. It could become one if washing in its presence began to produce something that not washing did not. ## Frequently asked questions **What is a discriminative stimulus in simple terms?** A discriminative stimulus is a cue that tells you a behavior will pay off right now, because in the past that behavior has been rewarded when the cue was present and not when it was absent. A ringing phone is a discriminative stimulus for answering; a lit "Open" sign is one for pulling the door. The cue does not force the behavior; it makes it more likely. **What is the difference between an SD and an S-delta?** An SD (discriminative stimulus) is a stimulus in whose presence a response has been reinforced, so the response becomes more likely when it appears. An S-delta is a stimulus in whose presence the same response has gone unreinforced, so the response becomes less likely. The lit "Open" sign is the SD for pulling the door; the dark sign is the S-delta. Discrimination training alternates the two. **What is an example of a discriminative stimulus?** The word "sit" for a trained dog: sitting after the word has produced treats, and sitting at other times has not, so the word raises the probability of sitting. Others: a green light for accelerating, a teacher's raised hand for going quiet, a lit key for a pigeon's peck, and a particular friend for a particular joke. In each case the behavior has been reinforced in the cue's presence and not in its absence. **What is the difference between a discriminative stimulus and a conditioned stimulus?** A conditioned stimulus, from classical conditioning, gets its power by predicting an unconditioned stimulus regardless of what the organism does, and it elicits a reflex such as salivation. A discriminative stimulus gets its power from a contingency that runs through the behavior, since the reinforcer arrives only if the response occurs, and it evokes an operant by raising its probability rather than triggering it. **What is the difference between a discriminative stimulus and a motivating operation?** A discriminative stimulus signals that a reinforcer is available: the drinking fountain, the vending machine, the lit sign. A motivating operation changes how much the reinforcer is worth: salty food, hours without eating, a hot day. Jack Michael drew the distinction in 1982. The test is whether removing the event would change the reinforcer's availability (discriminative) or only its value (motivational). **Is a discriminative stimulus the same as a prompt?** No. A prompt is an extra antecedent added because the natural discriminative stimulus does not yet evoke the behavior, such as pointing at the answer or guiding a dog into a sit. Prompts are meant to be faded so that control transfers to the discriminative stimulus, which stays. If the learner still needs the prompt, the discriminative stimulus has not been established. **How is a discriminative stimulus established?** Through discrimination training: the behavior is reinforced when the stimulus is present and not reinforced when it is absent, over many repetitions, until the behavior occurs far more often in its presence. Both halves are necessary. A stimulus that is present during reinforcement and during extinction alike predicts nothing and gains no control, which Jenkins and Harrison demonstrated with pigeons and a continuous tone. **Who introduced the term discriminative stimulus?** B. F. Skinner, in The Behavior of Organisms (1938), where he described the discriminated operant and introduced the symbols SD and S-delta. He held that the discriminative stimulus sets the occasion for a response rather than eliciting it, which separated operant discrimination from Pavlov's conditioned reflex. Jack Michael later (1982) separated the discriminative function of antecedent events from their motivational function. ## References 1. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 2. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 3. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson. 4. Jenkins, H. M., & Harrison, R. H. (1960). Effect of discrimination training on auditory generalization. *Journal of Experimental Psychology, 59*(4), 246–253. 5. Terrace, H. S. (1963). Discrimination learning with and without "errors." *Journal of the Experimental Analysis of Behavior, 6*(1), 1–27. 6. Dinsmoor, J. A. (1995). Stimulus control: Part I. *The Behavior Analyst, 18*(1), 51–68. 7. Skinner, B. F. (1957). *Verbal Behavior*. Appleton-Century-Crofts. 8. Rescorla, R. A. (1988). Pavlovian conditioning: It's not what you think it is. *American Psychologist, 43*(3), 151–160. 9. Michael, J. (1982). Distinguishing between discriminative and motivational functions of stimuli. *Journal of the Experimental Analysis of Behavior, 37*(1), 149–155. 10. Michael, J. (1993). Establishing operations. *The Behavior Analyst, 16*(2), 191–206. 11. Laraway, S., Snycerski, S., Michael, J., & Poling, A. (2003). Motivating operations and terms to describe them: Some further refinements. *Journal of Applied Behavior Analysis, 36*(3), 407–414. 12. Guttman, N., & Kalish, H. I. (1956). Discriminability and stimulus generalization. *Journal of Experimental Psychology, 51*(1), 79–88. 13. Hanson, H. M. (1959). Effects of discrimination training on stimulus generalization. *Journal of Experimental Psychology, 58*(5), 321–334. 14. Herrnstein, R. J., Loveland, D. H., & Cable, C. (1976). Natural concepts in pigeons. *Journal of Experimental Psychology: Animal Behavior Processes, 2*(4), 285–302. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Stimulus control](https://operantconditioning.com/stimulus-control/): Discrimination, generalization gradients, peak shift, and errorless learning in depth. - [The ABC model](https://operantconditioning.com/abc-model/): The three-term contingency the discriminative stimulus is the first term of, plus motivating operations. - [Operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/): Why a stimulus that evokes is not a stimulus that elicits. # Primary and Secondary Reinforcers: Definition & Examples > Primary reinforcers work without learning; secondary reinforcers earn their power by predicting them. Definitions, evidence, tokens, clickers, and mistakes. - Source: https://operantconditioning.com/primary-and-secondary-reinforcers/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Reinforcement · Types of reinforcers* Food reinforces a hungry rat because of what a rat is. A click reinforces the same rat because of what the click has come to predict. That difference, built in versus learned, sorts every reinforcer into one of two kinds, and it explains why praise fails with some children, why poker chips can teach a chimpanzee, and why a clicker in a trainer's hand can mark a behavior faster than any treat. > **Definition** > > A primary reinforcer (also called an unconditioned reinforcer) is a stimulus that strengthens the behavior it follows without any prior learning: food, water, warmth, sleep, sexual contact, and escape from pain are the standard examples. A **secondary reinforcer** (also called a conditioned reinforcer) is a stimulus that has acquired the power to strengthen behavior through its relation to other reinforcers: a magazine click, praise, a grade, money, a token. Both are reinforcers only if they do the job: the test is the effect on behavior, not the biology or the history.[1] > > "Secondary" does not mean weaker. Most of what reinforces adult human behavior is conditioned, and money keeps working even when nothing behind it is currently wanted. **In brief** - A primary reinforcer works without a learning history, but not without conditions: its power rises with deprivation and falls with satiation, so "unlearned" never means "always effective." - A secondary reinforcer earns its function by predicting another reinforcer, keeps it only while the prediction holds, and loses it by extinction when the pairing stops. - Generalized reinforcers such as money, tokens, and attention are backed by many reinforcers at once, which is why they work almost regardless of the person's state, and why a [token economy](https://operantconditioning.com/token-economy/) collapses when the tokens stop buying anything. ## Primary reinforcers: unlearned, but not unconditional A primary reinforcer works the first time. A rat that has never seen a lever will eat the pellet that arrives after it presses, and pressing will increase; no pairing is needed, because the reinforcing effect of food is part of what a rat is. The standard list is short: food, water, warmth when cold and cooling when hot, sleep, sexual contact, and the removal of pain or other aversive stimulation, which is the [negative reinforcement](https://operantconditioning.com/negative-reinforcement/) side of the same list.[1] For some species physical contact belongs on it too. Unlearned does not mean always effective. Food does not reinforce a rat that has just eaten. Skinner treated the rat's hunger as an experimental variable in its own right, set by hours of food deprivation, and reported that the rate of lever pressing rose and fell with it.[2] A primary reinforcer's power is a property of the stimulus and the organism's current state together, and a trainer who works after dinner with kibble has changed the second without noticing. The modern name for that state is the [motivating operation](https://operantconditioning.com/glossary/#motivating-operation). Jack Michael defined an establishing operation as an event that does two things at once: it raises the effectiveness of some stimulus as a reinforcer, and it raises the frequency of whatever behavior has produced that stimulus in the past.[3] Deprivation is the establishing operation for food; satiation is its opposite, an abolishing operation; the two were later grouped under the term motivating operation.[4] A list of effects, not of pleasures The list of primary reinforcers is a list of things shown to strengthen behavior without a learning history, not of things that feel good. Escape from pain reinforces without anyone enjoying the pain, and Olds and Milner's rats, below, pressed for a pulse of electricity that answers no known need. ## Secondary reinforcers: how a click comes to count A secondary reinforcer starts as a stimulus that does nothing. The sound of the food magazine in a [Skinner box](https://operantconditioning.com/skinner-box/) means nothing to a naive rat. In *The Behavior of Organisms* Skinner described what happens once the sound has been followed by food a number of times: the rat comes to the tray at the click, and the click itself, delivered after a lever press with no food behind it, is enough to strengthen pressing, for a while. As clicks keep arriving without food, the effect fades.[2] That fading is the signature of a conditioned reinforcer: its function is borrowed, and it is withdrawn when the relation that produced it is broken. The *procedure* here is pairing: a neutral stimulus is followed by a reinforcer. The *process* is the change in what the stimulus does, and it is the only evidence that a conditioned reinforcer exists. Trainers who "charge" a clicker are running the procedure; the dog's head snapping toward them at the sound is the first sign of the process. Magazine training is step one in the [virtual Skinner box](https://operantconditioning.com/lab/). Paul Bersh measured how much pairing it takes. Rats received a light followed by food, with the number of pairings and the light–food interval varied across groups; the light was then tested by making it the only consequence of lever pressing. More pairings produced a stronger conditioned reinforcer, up to a limit, and a light that came on shortly before the food worked far better than one that came on long before it.[5] The order is the one classical conditioning requires. [How the two kinds of conditioning combine ›](https://operantconditioning.com/operant-vs-classical-conditioning/) Kelleher and Gollub's review sorted the evidence by method: maintaining an old response in [extinction](https://operantconditioning.com/extinction/) with the stimulus still following it, or teaching a new response with the stimulus as its only consequence. Both methods demonstrated conditioned reinforcement, and both showed it to be short-lived when the stimulus was never paired again, because the test itself extinguishes the pairing.[6] Zimmerman's answer was to establish the stimulus intermittently: a buzzer followed by water only some of the time, for thirsty rats, went on to reinforce a new bar-pressing response that persisted far longer than earlier methods had achieved.[7] The reason is familiar from [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): a relation that was never reliable is slow to be detected as broken. ## Pairing is not enough: prediction, information, and dopamine Pairing is the procedure, but not the variable that matters. Egger and Miller gave rats a food pellet after a sequence of two stimuli. For one group the first stimulus reliably predicted the food, so the second added nothing; for another the first stimulus was unreliable, so the second carried the news. The second stimulus had been paired with food equally often in both groups, and it was a far stronger reinforcer when it was the informative one.[8] A stimulus becomes a conditioned reinforcer to the extent that it predicts a reinforcer, not merely accompanies one. Edmund Fantino's delay-reduction hypothesis makes the point quantitative: a stimulus reinforces to the extent that its onset signals a reduction in the time to the next primary reinforcer, relative to the average wait in that situation. A stimulus that means "food soon" when food is usually far off is a strong conditioned reinforcer; one that means "no food" is not a reinforcer at all, however informative. In the observing-response experiments he reviewed, pigeons work to produce a stimulus that tells them which schedule is in effect only when it can announce the better outcome.[9] Good news reinforces; news does not. Ben Williams's review reached the same place: conditioned reinforcement is a sound and necessary concept, but it is a statement about predictive relations.[10] ### The brain agrees, up to a point Schultz, Dayan, and Montague recorded midbrain dopamine neurons in monkeys while a cue came to predict a squirt of juice. Early in training the neurons fired at the juice; once the cue reliably predicted it, the burst moved to the cue, and the fully predicted juice evoked nothing.[11] A predictive cue inherits the signal, which is close to a physiological description of a conditioned reinforcer. Olds and Milner's rats, which pressed a lever for nothing but a brief pulse of electrical stimulation to the septal area, mark the other boundary: a reinforcer that needs no learning history and serves no known biological need.[12] The full account is on the [neuroscience page](https://operantconditioning.com/neuroscience/). ## Generalized reinforcers: money, tokens, and attention A conditioned reinforcer backed by one primary reinforcer inherits that reinforcer's motivating operations: a click backed only by food is worth little to a full dog. Skinner's explanation of why money, attention, and approval work almost all the time was the [generalized reinforcer](https://operantconditioning.com/glossary/#generalized-reinforcer), a conditioned reinforcer paired with many different reinforcers, so that at any moment at least one of the deprivations behind it is likely to be in force. Money buys food, warmth, shelter, and company; attention is the precondition for everything another person might provide. Skinner added that a generalized reinforcer can eventually work even when none of the reinforcers behind it is currently relevant, which is as close as behavior analysis comes to explaining the miser.[1] The laboratory version is the token. John Wolfe taught chimpanzees to operate a weighted lever to earn poker chips, which they could insert into a vending machine, the "Chimp-O-Mat," for grapes. The chimps worked for chips, preferred a chip that bought two grapes to one that bought one, learned to ignore a chip that bought nothing, and kept working when the chips could not be cashed in until later, though longer delays weakened the effect.[13] John Cowles then showed that tokens alone, exchanged for food only after a run of trials, could sustain the learning of new discriminations.[14] A token is a conditioned reinforcer that can be carried, counted, and saved, which is most of what money is. Timothy Hackenberg's review draws the modern picture: a token system is three schedules at once, one for earning tokens, one for reaching the exchange period, and one for exchanging, and the tokens' power depends on all three, above all on how reliably and how soon exchange happens.[15] A [token economy](https://operantconditioning.com/token-economy/) in a classroom or a ward is the same machinery, and it fails the same way: when the exchange stops, the tokens become paper, and the behavior they were maintaining goes with them. ## Three cases that strain the categories The primary/secondary distinction is clean in the laboratory and leaky everywhere else. Three cases show the seams. ### Social reinforcers and Harlow's caution Praise, attention, a smile, being listened to: social reinforcers are usually filed under "conditioned," on the theory that a caregiver was paired with food and warmth from birth. Harry Harlow's surrogate-mother experiments showed that this cannot be the whole story. Infant rhesus monkeys raised with two artificial mothers, one of bare wire that held the milk bottle and one covered in soft cloth that gave no milk, spent most of their time clinging to the cloth mother and ran to her when frightened, whichever mother had fed them.[16] Contact comfort reinforced on its own, not as a by-product of feeding. The category of a reinforcer is an empirical question: some social reinforcers are primary for some species, some are conditioned, and some are not reinforcers at all for a particular person. ### The clicker A clicker is a conditioned reinforcer built on purpose. Karen Pryor, who learned the method training dolphins with a whistle, describes charging the clicker by pairing it with food and then using the click to mark the exact instant of the behavior, with the food following a moment later.[17] Feng, Howell, and Bennett compared the three explanations on offer, that the click reinforces, that it marks, and that it bridges, and concluded that the published studies cannot yet tell them apart; direct comparisons of clicker-plus-food against food alone are few and mixed.[18] What is not in doubt is how the click acquires its function. [Marker training in practice ›](https://operantconditioning.com/dog-training/) ### Activity reinforcers David Premack's work sits across the distinction. In his account a reinforcer is not a stimulus but a behavior, eating rather than food, and any behavior can reinforce a less probable one.[19] Running needs no pairing to reinforce drinking in a rat that has been kept from running, so it behaves like a primary reinforcer without being a biological necessity. The [Premack principle](https://operantconditioning.com/premack-principle/) page covers the experiments; here it is a reminder that the primary/secondary sorting was built for stimuli, and a good deal of what reinforces is not a stimulus. ## Types of reinforcers and how each fails When a reinforcer stops working, its type tells you where to look; the last column is the practical one. | Type | Where the function comes from | Examples | How it fails | | --- | --- | --- | --- | | Primary (unconditioned) | Biology; no learning history needed | Food, water, warmth, sleep, sexual contact, escape from pain, contact comfort in infant primates | Satiation, or the wrong motivating operation: kibble after dinner, a blanket in July | | Secondary (conditioned) | Pairing with, and prediction of, another reinforcer | The magazine click, a clicker, a marker word, praise, a grade, a check mark | Extinction when the pairing stops; satiation on the reinforcer behind it; redundancy when it predicts nothing new | | Generalized | Pairing with many reinforcers | Money, tokens, points, attention, approval | Loss of backing: a token that never buys anything, points with an empty menu, attention that is free anyway | | Social | Delivered by another person; primary or conditioned depending on the reinforcer and the species | A smile, thanks, being listened to, physical contact | Not a reinforcer for this person; delivered whether or not the behavior occurs; paired with criticism until it becomes a warning | | Activity | The opportunity to perform a more probable behavior | Play after homework, the walk after the sit | The requirement deprives nobody; the activity has become less probable through satiation | "Social" and "activity" describe what form a reinforcer takes; "primary," "secondary," and "generalized" describe where its function comes from. A smile can be any of the latter three. ## How to build a conditioned reinforcer The procedure is the same for a rat, a dog, a first-grader, or you, and the same variables govern it: order, interval, number of pairings, and prediction. 1. **Choose a stimulus that is brief, distinct, and otherwise meaningless.** A click, a short word you do not use in conversation, a specific check mark, a chip; not a word the learner already hears forty times a day for free. 2. **Confirm that the backing reinforcer works right now.** A conditioned reinforcer can only be as strong as what stands behind it at the moment of pairing. Pair before the meal, not after it, and use something the learner would actually choose. 3. **Stimulus first, then reinforcer, within about a second.** A light that came on shortly before food became a stronger conditioned reinforcer than one that came on long before, and more pairings made a stronger one up to a limit.[5] A few dozen pairings across two or three short sessions is a reasonable start; it has worked when the learner orients to the stimulus before the reinforcer appears. 4. **Make it predictive, not merely frequent.** No backing reinforcer without the stimulus, no stimulus without the backing reinforcer. A redundant stimulus, one whose reinforcer was already predicted by something else, acquires little function no matter how often it is paired.[8] 5. **Use it to mark, then pay.** The stimulus goes at the instant the behavior occurs; the backing reinforcer follows. A conditioned reinforcer can be delivered with a precision the primary reinforcer cannot match, which is what makes [shaping](https://operantconditioning.com/shaping/) possible. 6. **Keep it backed.** Every click earns a treat while a behavior is being taught; tokens are exchanged on a schedule the learner can count on. Intermittent pairing makes a conditioned reinforcer more durable, but nothing makes it permanent.[7] If the learner stops orienting to the stimulus, recharge it. ## Common mistakes - **Assuming praise is reinforcing for everyone.** Praise is a conditioned reinforcer, which means it was conditioned somewhere, for some learners. For a teenager in front of friends it may be a punisher; for a child whose praise has always come bundled with a demand it may be a warning. The only test is what the behavior does afterward. - **Letting a token lose its backing.** A sticker chart with nothing at the end, classroom points that never buy anything, a bonus scheme that stops paying out: each is a conditioned reinforcer on extinction, and it takes the behavior down with it. - **Confusing a reward with a reinforcer.** A reward is something the giver thinks is nice; a reinforcer is something shown to increase the behavior it follows. A "secondary reward" that changes no behavior is not a secondary reinforcer, or any other kind. [The functional definition, in full ›](https://operantconditioning.com/reinforcement/) - **Paying without marking, or marking without paying.** Treats without the click make the click redundant; clicks without treats extinguish it. Either way the stimulus stops predicting, and prediction is the whole mechanism. - **Hearing "secondary" as "weaker."** Money outcompetes a sandwich for most adults most of the time, precisely because it is backed by everything. ## What the evidence does not show The literature on conditioned reinforcement is large, and several claims made in its name go beyond it. - **It does not show that every conditioned reinforcer traces back to food.** Harlow's infants preferred the cloth mother that never fed them.[16] How much human social reinforcement is primary is not settled. - **It does not show that a conditioned reinforcer can be made permanent.** Zimmerman's buzzer was durable, not permanent, and reviewers from Kelleher and Gollub onward have noted that conditioned reinforcers maintained for long periods in the laboratory are almost always still being paired, however thinly, within chained schedules.[7][6][10] - **It does not show that the number of pairings is what matters.** More pairings helped in Bersh's study, but an equally paired stimulus was much weaker in Egger and Miller's when it was redundant.[5][8] - **It does not show that a clicker beats a word, or that a marker beats food alone.** The comparisons are few and inconsistent.[18] The clicker's advantages are practical: it is fast, consistent, and free of the tone of voice a word carries. - **It does not show that dopamine is the reinforcer, or that it is pleasure.** The transfer of the dopamine burst to a predictive cue fits the predictive account, but the signal is a prediction error, not a reward in itself.[11] ## Key takeaways - A primary reinforcer strengthens behavior without any learning history; a secondary reinforcer strengthens behavior because it has come to predict another reinforcer. Both are reinforcers only if the behavior actually increases. - "Unlearned" does not mean "always effective." A primary reinforcer's power depends on the motivating operation in force, and satiation switches it off as surely as deprivation switches it on. - A conditioned reinforcer is established by pairing but governed by prediction: stimulus first, reinforcer within about a second, and no reinforcer arriving without it. It fades by extinction when the pairing stops. - Generalized reinforcers such as money, tokens, and attention are backed by many reinforcers, so they work almost regardless of the person's state; they fail when the backing is withdrawn. - The categories leak: contact comfort is a primary social reinforcer for infant monkeys, an activity can reinforce without being paired with anything, and the clicker's mechanism is settled while its superiority over alternatives is not. ### Check yourself **A teacher hands out points every day for on-task behavior. For the first month, points buy ten minutes of free time on Friday; then the free-time period is dropped but the points continue. On-task behavior falls back to where it started. What happened to the points?** The points were a conditioned reinforcer backed by free time. When the exchange stopped, the pairing stopped, and the points' function extinguished; they went on being delivered but predicted nothing. This is the standard failure mode of a token system, and the remedy is to restore the backing, not to hand out more points. **A trainer charges a clicker with treats for two sessions, then starts using it, but also keeps tossing the dog treats "to keep him motivated" whether or not a click has sounded. Within a week the dog barely reacts to the click. Why?** The click stopped predicting anything. Treats that arrive without a click make the click redundant, and Egger and Miller showed that a redundant stimulus acquires little reinforcing function even when it has been paired with food just as often as an informative one. Restore the rule that treats follow clicks and only clicks, and recharge. **Is a warm blanket a primary or a secondary reinforcer?** Warmth is a primary reinforcer: it strengthens behavior in a cold organism with no learning history. But whether the blanket reinforces anything right now depends on the motivating operation, and in a warm room it will not. "Primary" describes where the function comes from, not whether it is currently in force. ## Frequently asked questions **What are primary and secondary reinforcers in simple terms?** A primary reinforcer works without being learned: food, water, warmth, sleep, escape from pain. A secondary (conditioned) reinforcer has to be learned; it works because it has come to predict something else that reinforces, the way a clicker predicts a treat or money predicts everything money buys. Both are defined by their effect: if the behavior does not increase, neither one is a reinforcer. **What is the difference between a primary and a secondary reinforcer?** History. A primary reinforcer strengthens behavior the first time it is delivered, because of the organism's biology. A secondary reinforcer starts out neutral and acquires its power by being paired with, and coming to predict, another reinforcer. Both depend on conditions: a primary reinforcer needs the right deprivation to be in force, and a secondary reinforcer needs the reinforcer behind it to keep arriving. **What is an example of a secondary reinforcer?** The click of a clicker in dog training. It means nothing to an untrained dog; after being followed by treats a few dozen times it comes to strengthen whatever behavior it follows, and the trainer can use it to mark a sit at the exact instant it happens. Praise, grades, money, tokens, and a check mark in a habit app are secondary reinforcers in the same way. **Is money a primary or secondary reinforcer?** Secondary, and specifically a generalized conditioned reinforcer. Money has no biological value of its own, but it has been paired with nearly everything a person needs and wants, so it works almost regardless of which need is currently active. Skinner used it as the standard example of a generalized reinforcer, and the chimpanzee token studies of the 1930s showed the same thing in animals working for poker chips. **Is praise a primary or secondary reinforcer?** Usually secondary: praise acquires its power from what has followed it in a person's history, which is why it reinforces some people strongly, others weakly, and some not at all. Whether social contact in general is primary is less clear. Harlow's monkeys preferred a cloth mother that never fed them, so some social reinforcement, at least contact comfort in infant primates, appears to be unlearned. **Why do secondary reinforcers stop working?** Two reasons. First, extinction: a conditioned reinforcer keeps its function only while it goes on predicting the reinforcer behind it, so a token that no longer buys anything, or a click no longer followed by food, gradually loses its power. Second, satiation of the backing reinforcer: a food-backed clicker is weak after a large meal. Generalized reinforcers resist the second problem, not the first. **What is a generalized reinforcer?** A conditioned reinforcer that has been paired with many different reinforcers rather than one. Money, tokens, attention, and approval are the standard examples. Because several deprivations are likely to be in force at any moment, a generalized reinforcer works almost regardless of the person's state, which is why a token economy can run all day in a classroom without anyone being hungry. **Is a conditioned reinforcer the same as a discriminative stimulus?** No, but the same stimulus often does both jobs. A discriminative stimulus comes before a behavior and signals that the behavior will be reinforced; a conditioned reinforcer comes after a behavior and strengthens it. In a behavior chain, each link's stimulus reinforces the response that produced it and sets the occasion for the next one, which is why chaining works. ## References 1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 3. Michael, J. (1993). Establishing operations. *The Behavior Analyst, 16*(2), 191–206. 4. Laraway, S., Snycerski, S., Michael, J., & Poling, A. (2003). Motivating operations and terms to describe them: Some further refinements. *Journal of Applied Behavior Analysis, 36*(3), 407–414. 5. Bersh, P. J. (1951). The influence of two variables upon the establishment of a secondary reinforcer for operant responses. *Journal of Experimental Psychology, 41*(1), 62–73. 6. Kelleher, R. T., & Gollub, L. R. (1962). A review of positive conditioned reinforcement. *Journal of the Experimental Analysis of Behavior, 5*(S4), 543–597. 7. Zimmerman, D. W. (1957). Durable secondary reinforcement: Method and theory. *Psychological Review, 64*(6, Pt. 1), 373–383. 8. Egger, M. D., & Miller, N. E. (1962). Secondary reinforcement in rats as a function of information value and reliability of the stimulus. *Journal of Experimental Psychology, 64*(2), 97–104. 9. Fantino, E. (1977). Conditioned reinforcement: Choice and information. In W. K. Honig & J. E. R. Staddon (Eds.), *Handbook of Operant Behavior* (pp. 313–339). Prentice-Hall. 10. Williams, B. A. (1994). Conditioned reinforcement: Experimental and theoretical issues. *The Behavior Analyst, 17*(2), 261–285. 11. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599. 12. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427. 13. Wolfe, J. B. (1936). Effectiveness of token-rewards for chimpanzees. *Comparative Psychology Monographs, 12*(5), 1–72. 14. Cowles, J. T. (1937). Food-tokens as incentives for learning by chimpanzees. *Comparative Psychology Monographs, 14*(5), 1–96. 15. Hackenberg, T. D. (2009). Token reinforcement: A review and analysis. *Journal of the Experimental Analysis of Behavior, 91*(2), 257–286. 16. Harlow, H. F. (1958). The nature of love. *American Psychologist, 13*(12), 673–685. 17. Pryor, K. (1999). *Don't Shoot the Dog! The New Art of Teaching and Training* (rev. ed.). Bantam. 18. Feng, L. C., Howell, T. J., & Bennett, P. C. (2016). How clicker training works: Comparing reinforcing, marking, and bridging hypotheses. *Applied Animal Behaviour Science, 181*, 34–40. 19. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Reinforcement: the hub](https://operantconditioning.com/reinforcement/): Both types, the kinds of reinforcers, and what makes reinforcement work. - [Token economies](https://operantconditioning.com/token-economy/): Generalized conditioned reinforcers put to work in classrooms and wards. - [Dog training](https://operantconditioning.com/dog-training/): Marker and clicker training, timing, and the evidence on methods. # Superstitious Behavior: Skinner's Pigeons, Rituals & Luck > Superstitious behavior is behavior maintained by accidental reinforcement. Skinner's pigeons, the reanalysis, human rituals, lucky charms, and a self-test. - Source: https://operantconditioning.com/superstitious-behavior/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Experiments · Accidental reinforcement* Feed a hungry pigeon every fifteen seconds no matter what it does and it will, within minutes, start doing something odd and keep doing it. Skinner called the result superstition and argued that a batter's ritual and a gambler's lucky charm are built the same way. Later work showed his pigeons were doing something more interesting than he thought, and that the lesson for people is subtler than "reinforcement makes us dumb." > **Definition** > > Superstitious behavior is behavior maintained by **accidental reinforcement**: a reinforcer that happened to follow the behavior without depending on it. The organism acts as though its behavior produces the outcome when the outcome would have arrived anyway. B. F. Skinner introduced the term in 1948 to describe the rituals pigeons developed when food was delivered on a timer, and he argued that human superstitions have the same form.[1] > > Behavior analysts also call the process **adventitious reinforcement**. The procedure that produces it, a reinforcer delivered on a timer regardless of behavior, is a fixed-time or variable-time schedule, now usually called noncontingent reinforcement. **In brief** - Skinner fed hungry pigeons every fifteen seconds no matter what they did. Six of eight developed a distinctive ritual, which he explained as reinforcement of whatever the bird happened to be doing when food arrived. - When Staddon and Simmelhag recorded behavior through the whole interval, the rituals turned out to be largely species-typical food anticipation, not accidental pairings. The phenomenon is real; Skinner's explanation is only part of it. - People build rituals the same way in the laboratory and outside it. Accidental reinforcement is intermittent by nature, which is why a superstition survives so many failures and why it takes a deliberate test to drop one. ## Skinner's 1948 experiment The procedure was almost nothing. Skinner reduced pigeons to about three-quarters of their free-feeding weight so that food would work as a [reinforcer](https://operantconditioning.com/reinforcement/), put each bird alone in a chamber, and arranged for the food hopper to swing into reach for a few seconds once every fifteen seconds, with no reference to what the bird was doing. In the language of [schedules](https://operantconditioning.com/schedules-of-reinforcement/) this is a fixed-time schedule; unlike a [fixed-interval schedule](https://operantconditioning.com/fixed-interval-schedule/), it requires no response at all.[1] The process was not nothing. Six of the eight birds developed a clearly defined response that two observers agreed on. One turned counterclockwise around the cage, two or three turns between deliveries. Another repeatedly thrust its head into an upper corner. A third developed a "tossing" motion, as though lifting an invisible bar with its head. Two swung head and body from side to side like a pendulum, and one made incomplete pecking or brushing movements toward the floor. None had been seen before the food started arriving, and none had anything to do with producing it.[1] Skinner's explanation was that reinforcement acts on whatever behavior is in progress when it arrives. At the first delivery the bird was doing *something*; that something was strengthened, so it was more likely to be under way at the next delivery, when it was strengthened again. A feedback loop turned a random act into a ritual, and the shorter the interval, Skinner noted, the more readily the loop closed. His summary is the line quoted ever since: "The bird behaves as if there were a causal relation between its behavior and the presentation of food, although such a relation is lacking."[1] When Skinner lengthened the interval to a minute for one bird, its response became more energetic, a hop from foot to foot; when he then stopped the food altogether, the hopping went through an [extinction](https://operantconditioning.com/extinction/) curve of more than ten thousand responses before it faded, and a single reinstated delivery brought it back. He compared the birds to a bowler who keeps twisting his arm after the ball has left his hand.[1] In *Science and Human Behavior* he generalized: a few accidental pairings between a ritual and a favorable result are enough to set the ritual up and carry it through many unrewarded repetitions.[2] Richard Herrnstein put the point in its strongest form in 1966. Superstition is not an anomaly; it is a corollary of the principles of operant conditioning. If a reinforcer strengthens whatever precedes it, it must sometimes strengthen behavior that had nothing to do with producing it, and the same process runs inside every ordinary contingency: when a rat's press is reinforced, its force, posture, and accompanying head-turn are strengthened along with the feature the contingency required.[3] A click delivered a second late reinforces the spin after the sit, which is why marker timing matters so much in [shaping](https://operantconditioning.com/shaping/) and [dog training](https://operantconditioning.com/dog-training/). ## The reexamination: what the pigeons were actually doing Skinner watched his birds and described what he saw. In 1971 John Staddon and Virginia Simmelhag repeated the experiment with one change that altered its meaning: they catalogued every behavior each pigeon emitted, second by second, through the whole interval, on fixed-time schedules like Skinner's and on matched schedules where a peck was actually required. Two classes of behavior fell out of the records. **Terminal responses** occupied the last part of each interval, just before food was due, and in every bird they were much the same: pecking at the wall where the hopper appeared. **Interim activities** filled the early part of the interval, when food was least likely: the turning, pacing, wing-flapping, and preening that look so much like Skinner's rituals.[4] Neither class fit the accidental-reinforcement story. Terminal pecking emerged in birds that had not been pecking when earlier deliveries arrived, so chance pairings could not have selected it, and it looked the same whether or not a peck was required. Interim activities were almost never adjacent to food, yet they persisted. Periodic food, Staddon and Simmelhag concluded, induces a species-typical pattern: food-directed behavior when food is imminent, other activities when it is not. Behavior, they proposed, is generated by *principles of variation*, including such induced responses, and then selected by *principles of reinforcement*, an explicit analogy to evolution.[4] Reinforcement still selects; it does not select from a blank slate. William Timberlake and Gary Lucas went further in 1985, setting three explanations against one another: Skinner's chance contingency, Pavlovian stimulus substitution (the bird treating the chamber as if it were food, as in [classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/)), and their own account in terms of appetitive behavior systems. Under periodic food their pigeons' behavior was strikingly uniform, almost entirely directed at the wall where food appeared, and resembled the pigeon's natural foraging and social repertoire rather than randomly selected acts or misdirected eating. The rituals, on this view, are pieces of an evolved feeding system that free food switches on.[5] A third study asked whether an animal can even tell that it caused an outcome. Peter Killeen trained pigeons to report whether a stimulus change had been produced by their own peck or by the apparatus, and they discriminated with high accuracy; the errors that remained moved with the payoffs for each answer. Superstition, Killeen argued, is a matter of bias rather than detectability: organisms can usually tell what they did from what merely happened, and what tilts them toward claiming credit is the relative cost of being wrong.[6] Where this leaves Skinner's account Accidental reinforcement is real, and it has been produced cleanly in humans. But the rituals of the 1948 pigeons were mostly food-induced behavior arranged in time by the schedule. Skinner found a real phenomenon, described it accurately, and explained only part of it. ## Superstitious behavior vs. its neighbors Several things can happen when reinforcers arrive without depending on behavior, and they are easy to confuse. | Concept | What the reinforcer depends on | What happens to behavior | Example | | --- | --- | --- | --- | | Contingent operant | The behavior; no response, no reinforcer | Strengthened; comes under [stimulus control](https://operantconditioning.com/stimulus-control/); extinguishes when the contingency ends | A rat presses a lever for food | | Superstitious behavior | Nothing; it arrives on a timer and happens to follow the behavior | Whatever is in progress at delivery is strengthened and self-perpetuates[1] | Skinner's turning pigeon; a batter's glove adjustment | | Induced (interim and terminal) behavior | Nothing; the timing of the reinforcer calls up species-typical responses | Uniform across individuals; predictable from species and schedule, not from chance pairings[4] | Pecking at the hopper wall as food comes due | | Free reinforcers added to a contingency | Partly the behavior, partly nothing | The contingent response *weakens* as free deliveries dilute the contingency[7] | A rat given free pellets presses less | | [Learned helplessness](https://operantconditioning.com/learned-helplessness/) | Nothing; an aversive event starts and stops regardless of behavior | Responding stops; in people, mainly when the procedure leads them to stop trying[8] | Giving up on an unsolvable task | The fourth row is the one most often forgotten. Noncontingent reinforcement does not generally build behavior; added on top of a real contingency, free reinforcers make the contingent response decline, because the behavior no longer makes much difference to how often the reinforcer arrives.[7] Superstition is what accidental reinforcement does to behavior that had no contingency to begin with. ## Superstition in the human laboratory The cleanest human demonstration of accidental reinforcement is one of the earliest. Catania and Cutts gave college students two buttons: presses on one were reinforced on a variable-interval schedule, presses on the other never. Students went on pressing the useless button anyway, because a reinforcer earned by the first sometimes arrived just after a press on the second. A changeover delay, a required pause after switching buttons before any reinforcer could be collected, eliminated the superstitious pressing.[9] Koichi Ono's 1987 study is the human version of Skinner's. Twenty university students sat alone in a booth with three levers and a counter, told only that they need not do anything in particular, that doing something might produce points, and that they should try to get as many as possible. Points arrived on fixed-time and variable-time schedules regardless of what they did. Most developed some pattern of lever-pulling; a few developed elaborate, persistent rituals. The best-known is a woman who, after a point happened to follow a jump, climbed onto the table, jumped to touch the ceiling with her slipper, and went on jumping until she was too tired to continue. Superstition does develop in people under noncontingent reinforcement, Ono concluded, but not uniformly and often only for a while.[10] Children do it too. Wagner and Morris put preschoolers in a room with a mechanical clown that dispensed marbles on a fixed-time schedule, exchangeable afterward for a toy; seven of the twelve developed a recurring behavior between deliveries, such as touching or kissing the clown's face or swinging their bodies.[11] A later computer-based study found superstitious responding whether the noncontingent event was the delivery of points or the removal of an aversive condition, so accidental [negative reinforcement](https://operantconditioning.com/negative-reinforcement/) builds rituals as readily as accidental positive reinforcement.[12] That connects superstition to its apparent opposite. Human learned-helplessness experiments expose people to bursts of loud noise they cannot stop, then test whether they give up on a later, solvable task. Helena Matute noticed that the noise always stops eventually, so whatever the person did last is negatively reinforced by its ending. In her 1994 experiments, students told to try to stop the noise kept pressing keys in patterns, reported that they had found a way to control it, and showed no helplessness afterward; helplessness appeared only when the procedure led people to stop responding.[8] A 1995 follow-up confirmed it: as long as people keep trying, uncontrollable outcomes tend to produce superstition and an illusion of control rather than passivity.[13] ## Superstitions in real life ### Baseball The anthropologist George Gmelch, a former minor-league first baseman, borrowed an observation of Bronislaw Malinowski's: Trobriand Islanders used elaborate magic for open-sea fishing, where catches were uncertain and danger real, and almost none for fishing in the lagoon, where results were reliable. Baseball follows the same rule. Rituals, taboos, and lucky objects cluster in hitting and pitching, where even an excellent player fails much of the time, and are nearly absent in fielding, where players handle the overwhelming majority of chances cleanly. The rituals sit exactly where skill accounts for the least of the outcome.[14] ### Lucky charms Do the rituals do anything? A widely reported set of experiments by Damisch, Stoberock, and Mussweiler in 2010 said yes. Participants who putted with a golf ball described as lucky, who heard "fingers crossed" before a dexterity task, or who kept their own lucky charm with them during a memory game or an anagram task outperformed controls, and the effect ran through higher confidence in their own ability.[15] In 2014 Calin-Jageman and Caldwell ran two direct replications of the lucky-ball putting experiment, with larger samples than the original, and found no benefit.[16] The performance benefit of a lucky charm is unconfirmed. ### The cultural angle Skinner's account explains how an *individual* acquires a ritual from its own history of accidents. Stuart Vyse, whose *Believing in Magic* is the standard survey, points out that most human superstitions were not invented that way. Nobody discovers for themselves that thirteen is unlucky; the rituals come ready-made from parents, teammates, and the surrounding culture, and personal accidental reinforcement then keeps them going. Vyse also documents where superstition concentrates: among students before exams, athletes before competition, and gamblers, where outcomes matter and control is incomplete. It is not a sign of low intelligence or of any disorder, and part of what maintains it may be that the ritual reduces anxiety, which is real negative reinforcement even when the effect on the outcome is imaginary.[17] | Laboratory finding | Human analogue | What actually maintains it | | --- | --- | --- | | A pigeon turns counterclockwise between timed food deliveries[1] | A batter adjusts both gloves before every pitch | Accidental reinforcement plus culture: the ritual is cheap, and some hits happen to follow it[14][17] | | Every bird pecks at the hopper wall as food comes due[4] | Leaning toward the screen as the result loads | Anticipation induced by the timing of the reinforcer; the same in everyone[4][5] | | Students keep pressing a never-reinforced button beside a reinforced one[9] | Wearing the lucky socks while also training hard | Concurrent reinforcement: the real behavior pays and the ritual rides along[9] | | People keep pressing keys during noise that stops on its own, and report control[8] | Pressing the crosswalk button five times; blowing on the dice | Accidental negative reinforcement: the aversive event ends anyway, and whatever came last gets the credit[8][12] | | A "lucky" ball improves putting in one laboratory and not in another[15][16] | The charm in the pocket at the exam | Social transmission and anxiety relief; any effect on the outcome is unconfirmed[16][17] | ## Why superstition is a rational-looking mistake Nothing about a superstition requires stupidity. Four features of ordinary learning are enough to build one and hold it in place. **Reinforcement is blind to causation.** A consequence strengthens the behavior it follows whether or not the behavior produced it; that is the whole content of Skinner's demonstration and Herrnstein's corollary.[1][3] The mechanism that lets a rat find the lever is the one that attaches a ritual to a home run, and a learner that ignored every pairing it could not verify as causal would learn almost nothing. **Accidental reinforcement is intermittent, and intermittent reinforcement is durable.** The good outcome follows the ritual sometimes and not other times, which puts the ritual on a lean, irregular schedule from the start. Behavior maintained that way is far more resistant to extinction than behavior reinforced every time, the [partial reinforcement extinction effect](https://operantconditioning.com/glossary/#partial-reinforcement-extinction-effect) described on the page comparing [continuous and intermittent reinforcement](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/). Skinner's pigeon, still hopping ten thousand responses after the food stopped, is the laboratory version of a fan in the same shirt through eleven losing seasons.[1] **Extinction is never run.** To extinguish a superstition you would have to skip the ritual, repeatedly, and watch the outcomes; but the ritual is cheap and the outcome feels important, so the test is never made. A [contingency](https://operantconditioning.com/abc-model/) is a comparison between two rates, good outcomes with the ritual and without it, and someone who never skips the ritual has only one of the numbers. Killeen's pigeons show this is a bias set by payoffs rather than a failure of perception: when a wrong "I didn't cause it" costs more than a wrong "I did," organisms claim the credit.[6] **Superstition grows where control runs out.** Gmelch's ballplayers were not superstitious about fielding. Rituals attach to outcomes that matter and cannot be fully controlled, which is where accidental pairings are most frequent and real contingencies hardest to find.[14][17] Matute adds the last piece: in that situation, continuing to act produces superstition and stopping produces helplessness, and of the two, superstition is arguably the better deal.[8][13] ## How to test your own superstitions The pigeon could not run the test. You can. The question is the one behind any contingency: does the outcome depend on the behavior, or does it just sometimes follow it? Catania and Cutts' students stopped pressing the useless button as soon as the accidental pairings were broken, and you can break them deliberately.[9] 1. **Name the ritual and the outcome.** Be specific on both sides: "wearing the gray hoodie" and "scoring above 80 on the quiz," not "feeling ready." If you cannot say what outcome the ritual is supposed to produce, it is probably maintained by how it feels rather than by what it does. 2. **Record a baseline.** For a dozen occasions, note whether you performed the ritual and how the outcome went. Most people find they have never once skipped the ritual before an outcome that mattered, so they have no evidence either way. 3. **Withhold the ritual on alternate occasions.** Decide in advance, ideally by coin flip, which occasions get the ritual, so that you are not quietly saving it for the hard days. This is extinction arranged as an experiment. It will feel wrong at first; the discomfort is the negative reinforcement that has been helping to maintain the ritual, and it fades. 4. **Compare the two rates.** Count good outcomes with the ritual and without it. A contingency exists only if the first rate is clearly higher. One bad day without the hoodie proves nothing; you have had bad days with it. 5. **Decide what to keep.** If the outcome does not depend on the ritual, drop it, or keep it with clear eyes: a routine that steadies you may be worth its thirty seconds, provided it does not stand in for preparation and you can do without it when the hoodie is in the wash. ## What the evidence does not show - **It does not show that Skinner's account is the whole story.** The 1948 paper was a demonstration, not a controlled experiment: eight birds, no comparison group, behavior described by observers rather than recorded through the interval, and the measured response chosen after it appeared. Recorded continuously, most of the behavior turned out to be induced, species-typical activity organized by the schedule.[1][4][5] Accidental reinforcement is real, but its clearest demonstrations are the experiments where a real contingency runs alongside.[3][9] - **It does not show that noncontingent reinforcement generally builds behavior.** Added to an existing contingency, free reinforcers weaken the contingent response.[7] Even with no competing contingency the human results are uneven: Ono's students mostly developed something, but few developed anything lasting.[10] - **It does not show that superstitious organisms cannot detect contingency.** Killeen's pigeons could. The bias toward claiming credit is a decision under uncertainty, and it moves with the payoffs.[6] - **It does not show that lucky charms improve performance.** The 2010 result did not survive direct replication.[15][16] - **It does not show that everyday superstitions arise from personal accidents.** Most are learned from other people and then maintained by accidental reinforcement, anxiety relief, and social approval; the pigeon model explains the maintenance better than the origin.[17] - **It does not show that superstition and helplessness are opposite kinds of people.** In Matute's experiments they were the same people under different instructions; the difference was whether they were still responding when the uncontrollable outcome arrived.[8][13] ## Key takeaways - Superstitious behavior is behavior maintained by accidental reinforcement, a reinforcer that happened to follow the behavior without depending on it. Skinner produced it in 1948 by feeding pigeons every fifteen seconds regardless of what they did, and six of eight developed a ritual. - Skinner explained the rituals as reinforcement of whatever the bird happened to be doing when food arrived, and Herrnstein showed that the same process strengthens the incidental features of every reinforced response. Staddon and Simmelhag then found that the pigeons' behavior was mostly species-typical food anticipation organized by the schedule. - People develop rituals under noncontingent points, marbles, and noise that stops on its own, and when they keep trying to control an uncontrollable outcome they become superstitious rather than helpless. Real-world superstitions cluster where outcomes matter and control is incomplete. - Superstitions persist because accidental reinforcement is intermittent, the extinction test is never run, and claiming credit is a cheap bet. To test one, record a baseline, withhold the ritual on randomly chosen occasions, and compare the two rates. - The lucky-charm performance benefit did not replicate, most human superstitions are learned socially rather than built from personal accidents, and the 1948 study was a demonstration rather than a controlled experiment. ### Check yourself **A dog spins in a circle before every meal. Its owner puts the bowl down at the same time each evening, whatever the dog is doing. Is the spinning superstitious behavior in Skinner's sense, and how could you tell?** It could be. The bowl arrives on a fixed-time schedule, so whatever the dog was doing just before it, including a spin that once happened by chance, is accidentally reinforced and becomes more likely to be in progress at the next delivery. But Staddon and Simmelhag would ask whether the spin is species-typical excitement induced by food coming due. The test is contingency: put the bowl down at varied times, unrelated to spinning, and see whether the spin stays tied to the moment of delivery or drifts. **A student wears the same gray hoodie to every exam and has done well in most of them. She says the evidence supports the hoodie. What is missing from her evidence?** The other rate. A contingency is a comparison between the outcome rate with the behavior and the rate without it, and she has never taken an exam without the hoodie. Her results are consistent with the hoodie mattering and with it mattering not at all; only occasions without the ritual, chosen in advance rather than saved for easy exams, can tell those apart. Meanwhile the hoodie is cheap and skipping it feels risky, which is exactly Killeen's bias. **In a replication of Skinner's study, every pigeon pecks at the wall near the hopper just before each food delivery. Is this the accidental reinforcement Skinner described?** No. Accidental reinforcement should produce different rituals in different birds, selected from whatever each happened to be doing. A response that is the same in every bird, appears in birds that were not pecking when earlier deliveries arrived, and is directed at the food source is a terminal response induced by the schedule: species-typical food anticipation, not chance selection. ## Frequently asked questions **What is superstitious behavior in simple terms?** A behavior that keeps happening because a good outcome once followed it by chance, not because the behavior causes the outcome. Skinner fed pigeons on a timer and they developed rituals such as turning and head-tossing; a batter who adjusts his gloves before every pitch is doing the same thing. The reward arrives on its own schedule, and whatever came just before it gets the credit. **What did Skinner's superstition experiment show?** In 1948 Skinner delivered food to hungry pigeons every fifteen seconds regardless of their behavior. Six of eight birds developed a distinctive, repeated response, such as turning counterclockwise or swinging like a pendulum. He concluded that reinforcement strengthens whatever behavior precedes it, whether or not the behavior produced the reinforcer, and that human superstitions form the same way. **Was Skinner's superstition experiment wrong?** Not wrong, but incomplete. When Staddon and Simmelhag recorded pigeons' behavior throughout each interval in 1971, they found that most of it was species-typical food anticipation, the same across birds and not selected by chance pairings. Accidental reinforcement is real and has been shown in other experiments, but the specific rituals Skinner saw were largely induced by the schedule rather than accidentally reinforced. **What is an example of superstitious behavior?** A basketball player who bounces the ball exactly three times before a free throw, a student who uses the same pen for every exam, a gambler who blows on the dice, or a person who presses the crosswalk button repeatedly. In each case the outcome arrives on its own schedule, sometimes just after the ritual, and those coincidences keep the ritual going. **Superstitious behavior vs. learned helplessness: what is the difference?** Both arise when outcomes do not depend on behavior. In learned helplessness the organism stops responding and later fails to act even when control is available. In superstition it keeps responding and credits whatever it did last. Matute's experiments showed that people exposed to uncontrollable noise became superstitious if they kept trying to stop it and helpless mainly when the procedure led them to stop. **Do lucky charms actually work?** The evidence does not support it. A 2010 study reported that a lucky ball or a personal charm improved putting, dexterity, memory, and anagram performance by raising confidence, but two direct replications in 2014 with larger samples found no benefit. A charm may reduce anxiety, which is a real effect on the person, but there is no confirmed effect on the outcome. **Why are athletes so superstitious?** Because their outcomes matter and are only partly under their control. Gmelch found that baseball rituals cluster in hitting and pitching, where even excellent players fail often and luck decides much, and are nearly absent in fielding, where players succeed almost every time. Uncertain, high-stakes activities produce the most accidental pairings between a ritual and a good result. **How do you get rid of a superstition?** Run the extinction test the pigeon could not. Record whether you perform the ritual and how the outcome goes, then withhold the ritual on occasions chosen in advance by a coin flip, and compare the rate of good outcomes with and without it. If the rates are the same, the ritual can go. Expect it to feel uncomfortable at first; the discomfort fades. ## References 1. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 2. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 3. Herrnstein, R. J. (1966). Superstition: A corollary of the principles of operant conditioning. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 33–51). Appleton-Century-Crofts. 4. Staddon, J. E. R., & Simmelhag, V. L. (1971). The "superstition" experiment: A reexamination of its implications for the principles of adaptive behavior. *Psychological Review, 78*(1), 3–43. 5. Timberlake, W., & Lucas, G. A. (1985). The basis of superstitious behavior: Chance contingency, stimulus substitution, or appetitive behavior? *Journal of the Experimental Analysis of Behavior, 44*(3), 279–299. 6. Killeen, P. R. (1978). Superstition: A matter of bias, not detectability. *Science, 199*(4324), 88–90. 7. Hammond, L. J. (1980). The effect of contingency upon the appetitive conditioning of free-operant behavior. *Journal of the Experimental Analysis of Behavior, 34*(3), 297–304. 8. Matute, H. (1994). Learned helplessness and superstitious behavior as opposite effects of uncontrollable reinforcement in humans. *Learning and Motivation, 25*(2), 216–232. 9. Catania, A. C., & Cutts, D. (1963). Experimental control of superstitious responding in humans. *Journal of the Experimental Analysis of Behavior, 6*(2), 203–208. 10. Ono, K. (1987). Superstitious behavior in humans. *Journal of the Experimental Analysis of Behavior, 47*(3), 261–271. 11. Wagner, G. A., & Morris, E. K. (1987). "Superstitious" behavior in children. *The Psychological Record, 37*(4), 471–488. 12. Bloom, C. M., Venard, J., Harden, M., & Seetharaman, S. (2007). Non-contingent positive and negative reinforcement schedules of superstitious behaviors. *Behavioural Processes, 75*(1), 8–13. 13. Matute, H. (1995). Human reactions to uncontrollable outcomes: Further evidence for superstitions rather than helplessness. *Quarterly Journal of Experimental Psychology, 48B*(2), 142–157. 14. Gmelch, G. (1971). Baseball magic. *Trans-Action, 8*(8), 39–41. 15. Damisch, L., Stoberock, B., & Mussweiler, T. (2010). Keep your fingers crossed! How superstition improves performance. *Psychological Science, 21*(7), 1014–1020. 16. Calin-Jageman, R. J., & Caldwell, T. L. (2014). Replication of the superstition and performance study by Damisch, Stoberock, and Mussweiler (2010). *Social Psychology, 45*(3), 239–245. 17. Vyse, S. A. (2014). *Believing in Magic: The Psychology of Superstition* (updated ed.). Oxford University Press. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Continuous vs. intermittent reinforcement](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/): Why behavior reinforced only sometimes is the hardest to extinguish. - [Extinction](https://operantconditioning.com/extinction/): What happens when the reinforcer stops, and why it is never as simple as ignoring. - [Learned helplessness](https://operantconditioning.com/learned-helplessness/): The other thing uncontrollable outcomes can do, and when they do it. # AP Psychology: Operant Conditioning Study Guide & Practice > Operant conditioning as the AP Psychology exam tests it: every term in one line, the studies to name, the confusions that cost points, and a practice set. - Source: https://operantconditioning.com/ap-psychology/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Study guide · AP Psychology* Learning is one of the most reliably tested ideas in AP Psychology, and operant conditioning is the part students most often get backwards under time pressure. This is the exam-facing version of the site: the terms in one line each, the studies by name and year, the confusions the test writers count on, a method for scenario questions, and a practice set with an answer key. **In brief** - In the revised AP Psychology course, effective fall 2024, learning sits within Unit 3, Development and Learning; the Course and Exam Description lists the terms you are responsible for.[1] - Every operant scenario is two questions: was something added or removed, and did the behavior go up or down. Positive means added, negative means removed, and negative reinforcement is not punishment. - Schedules are two more questions — count or time, fixed or variable — and each has a signature pattern the exam expects you to recognize. ## Where operant conditioning sits on the AP exam The College Board revised AP Psychology for the 2024–25 school year. In the revised Course and Exam Description (CED), learning is no longer a unit of its own: classical conditioning, operant conditioning, and social and cognitive factors in learning sit within Unit 3, "Development and Learning."[1] The exam has a multiple-choice section and a free-response section; in the revised course the free-response questions ask you to analyze a description of a study and to argue from several sources of evidence rather than to define terms.[1] Question counts, timing, and weighting are in the CED and can change between editions; check the current one rather than any summary, including this one. You will rarely be asked to define negative reinforcement. You will be given a scenario — a parent, a rat, a paragraph from a study — and asked what kind of learning it shows or whether the design supports the conclusion. The CED is free and lists every term the exam can draw on; read it before you study and again after. Anything not in it is not required, and the last section here sorts out which pages on this site go past it.[1] ## The vocabulary, one line each Definitions are functional, as the exam uses them: a reinforcer or punisher is defined by its effect on behavior, not by intent or feel. Each links to the page that treats it fully. ### Consequences and the four quadrants - **[Law of effect](https://operantconditioning.com/law-of-effect/).** Responses followed by satisfaction are strengthened, responses followed by discomfort weakened; Thorndike's wording is [in the library](https://operantconditioning.com/library/thorndike-animal-intelligence/#chapter-vi-laws-and-hypotheses-for-behavior).[2][3] - **Operant conditioning.** A voluntary behavior changes in frequency because of its consequences; Skinner's term.[4] - **[Reinforcement](https://operantconditioning.com/reinforcement/).** Any consequence that makes the behavior it follows more frequent. - **[Punishment](https://operantconditioning.com/punishment/).** Any consequence that makes the behavior it follows less frequent. - **Positive and negative.** Added and removed. Not good and bad. - **[Positive reinforcement](https://operantconditioning.com/positive-reinforcement/).** Stimulus added, behavior up: a treat after a sit. - **[Negative reinforcement](https://operantconditioning.com/negative-reinforcement/).** Stimulus removed, behavior up: buckling up silences the chime. - **[Positive punishment](https://operantconditioning.com/positive-punishment/).** Stimulus added, behavior down: a scolding after jumping on the couch. - **[Negative punishment](https://operantconditioning.com/negative-punishment/).** Stimulus removed, behavior down: losing the car keys after missing curfew. - **[Extinction](https://operantconditioning.com/extinction/).** The reinforcer stops following the behavior and the behavior declines. Not a quadrant: nothing is added or removed. - **[Extinction burst](https://operantconditioning.com/glossary/#extinction-burst).** The brief rise in rate or intensity when reinforcement is first withheld. - **Spontaneous recovery.** An extinguished response returns after a rest. - **Generalization.** Responding to stimuli that resemble the training stimulus. - **Discrimination.** Responding to the training stimulus but not to similar ones; the signal that reinforcement is available is a [discriminative stimulus](https://operantconditioning.com/discriminative-stimulus/). ### Reinforcers - **[Primary reinforcer](https://operantconditioning.com/primary-and-secondary-reinforcers/).** Reinforcing without learning: food, water, warmth. - **Secondary (conditioned) reinforcer.** Reinforcing through association with primary reinforcers: money, grades, praise. - **[Token economy](https://operantconditioning.com/token-economy/).** Tokens earned for target behaviors are exchanged later for backup reinforcers. - **[Premack principle](https://operantconditioning.com/premack-principle/).** A more probable behavior reinforces a less probable one when made contingent on it: first homework, then the game.[5] ### Building and maintaining behavior - **[Shaping](https://operantconditioning.com/shaping/).** Reinforcing successive approximations to a target behavior. - **[Chaining](https://operantconditioning.com/chaining/).** Linking behaviors so each step cues the next and the last produces the reinforcer. - **[Skinner box](https://operantconditioning.com/skinner-box/).** An operant chamber with a lever or key, a reinforcer dispenser, and a cumulative recorder.[4] - **[Continuous reinforcement](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/).** Every response reinforced: fastest acquisition, fastest extinction. - **Partial (intermittent) reinforcement.** Some responses reinforced: slower acquisition, far greater resistance to extinction. AP materials say partial; behavior analysts say intermittent. - **[Fixed ratio (FR)](https://operantconditioning.com/fixed-ratio-schedule/).** Reinforcement after a set number of responses. Pattern: a pause after each reinforcer, then a fast run.[6] - **[Variable ratio (VR)](https://operantconditioning.com/variable-ratio-schedule/).** After an unpredictable number of responses. Pattern: high and steady, hardest to extinguish.[6] - **[Fixed interval (FI)](https://operantconditioning.com/fixed-interval-schedule/).** The first response after a set time. Pattern: the scallop, accelerating as the interval ends.[6] - **[Variable interval (VI)](https://operantconditioning.com/variable-interval-schedule/).** The first response after an unpredictable time. Pattern: moderate and steady.[6] ### Named phenomena - **[Superstitious behavior](https://operantconditioning.com/superstitious-behavior/).** Strengthened by accidental, non-contingent reinforcement; Skinner's pigeons, fed every fifteen seconds whatever they did, developed stereotyped turning.[7] - **[Learned helplessness](https://operantconditioning.com/learned-helplessness/).** After uncontrollable aversive events, an organism fails to escape when it can; Seligman and Maier's dogs.[8] - **Instinctive drift.** Trained behavior drifts toward species-typical behavior; the Brelands' raccoons "washed" the coins they were trained to deposit.[9] - **Biological preparedness.** Some associations are learned easily and others hardly at all; rats link taste with illness and light or sound with shock, not the reverse.[10][11] ### Contrast cases: cognitive and social learning - **Latent learning.** Learning without reinforcement that appears when there is reason to use it; Tolman and Honzik's rats had a **cognitive map** before food was offered.[12] - **Insight learning.** A sudden, complete solution rather than gradual trial and error; Köhler's chimpanzees.[13] - **Observational learning (modeling).** Learning by watching a model; Bandura's Bobo doll. **Vicarious reinforcement** and **punishment** are consequences the observer sees the model receive.[14] ## Operant vs. classical conditioning, as the exam frames it The exam treats these as the two kinds of associative learning and tests whether you can tell them apart. In classical conditioning two stimuli are paired and a reflex transfers from one to the other: Pavlov's dogs salivated to the signal that preceded food, and Watson and Rayner's infant cried at a white rat that had been paired with a loud noise.[15][16] In operant conditioning a voluntary behavior is followed by a consequence and changes in frequency. Classical responses are *elicited* by what comes before; operant responses are *emitted* and changed by what comes after. | Aspect | Classical (Pavlovian) | Operant (instrumental) | | --- | --- | --- | | What is learned | One stimulus predicts another (CS predicts US) | A behavior produces a consequence | | Kind of response | Reflexive, involuntary: salivation, fear, nausea | Voluntary: pressing, studying, sitting | | Order of events | Key stimulus comes *before* the response | Key event comes *after* the response | | Outcome depends on behavior? | No; food arrives whether or not the dog salivates | Yes; food arrives only if the rat presses | Extinction, spontaneous recovery, generalization, and discrimination occur in both, so a question can ask about any of them in either frame. Most real scenes contain both — a dog excited at the treat pouch, then sitting for the treat. The [operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/) page has the decision checklist and eight worked examples. ## The four confusions the test writers count on Most wrong answers come from one of four confusions, and each has a distractor written for it. ### 1. Negative reinforcement is not punishment *Reinforcement* or *punishment* says what happened to the behavior. *Positive* or *negative* says what happened to the environment. Negative reinforcement is "behavior up, something removed." Nagging that stops when the room is clean negatively reinforces cleaning; nagging that starts at the sight of a mess positively punishes messiness. The same stimulus does opposite jobs. | Change | Behavior increases | Behavior decreases | | --- | --- | --- | | Stimulus added | Positive reinforcement | Positive punishment | | Stimulus removed | Negative reinforcement | Negative punishment | ### 2. Positive means added, even when what is added is unpleasant "Positive punishment" is a contradiction only if positive means good. It means added: a shock, a scolding, an extra chore. And a consequence is a punisher only if the behavior decreases; a detention that does not reduce talking is not punishment, whatever it was called. Items exploit this with a "punishment" followed by more of the behavior. ### 3. Ratio vs. interval, and what "variable" does not tell you Students learn that variable schedules resist extinction and then call anything unpredictable "variable." The first question is ratio or interval. If reinforcement depends on *how many* responses occur, it is ratio and responding faster pays sooner; if it depends on *time*, it is interval and responding faster changes nothing. A slot machine paying after an unpredictable number of pulls is variable ratio; a pool where fish bite after unpredictable stretches of time is variable interval. Only the gambler's rate matters, which is why the gambler pulls fast and the angler casts at a moderate pace.[6] ### 4. Extinction is not punishment Both reduce behavior. In punishment something happens after the response: an aversive is added or a reinforcer removed. In extinction nothing happens; the usual reinforcer simply stops. Tantrums once maintained by attention and now ignored are on extinction; a toy removed after each tantrum is negative punishment. Extinction also has two signatures punishment lacks: a [burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) at the start and spontaneous recovery after a rest. ## The studies you should be able to name Questions name the researcher and expect the term, or describe the procedure and expect the researcher. | Study | Procedure | Finding | Term | | --- | --- | --- | --- | | Thorndike (1898) | Cats worked a latch to escape a puzzle box for food. | Escape times fell gradually, with no sudden insight.[2] | Law of effect | | Watson and Rayner (1920) | A loud noise paired with a white rat for an infant. | Fear of the rat alone, generalizing to a rabbit, a dog, a fur coat.[16] | Conditioned fear; generalization | | Köhler (1925) | Chimpanzees, bananas out of reach, boxes and sticks. | After no progress, sudden and complete solutions.[13] | Insight learning | | Pavlov (1927) | A signal repeatedly preceded food for dogs. | Salivation to the signal; extinction, recovery, generalization, discrimination described.[15] | Classical conditioning | | Tolman and Honzik (1930) | Rats ran a maze daily; one group was unfed until day eleven. | Errors then dropped abruptly to the rewarded group's level.[12] | Latent learning; cognitive map | | Skinner (1938) | Rats pressed a lever for food in an operant chamber. | Rate of response as the measure; reinforcement defined by its effect.[4] | Operant conditioning; Skinner box | | Skinner (1948) | Pigeons fed every fifteen seconds regardless of behavior. | Six of eight developed stereotyped turning and head-tossing.[7] | Superstitious behavior | | Ferster and Skinner (1957) | Pigeons and rats on many schedules, recorded cumulatively. | Each schedule produces a characteristic pattern.[6] | Schedules of reinforcement | | Premack (1959) | Children's candy eating and pinball made contingent on each other. | The preferred activity reinforced the other, not the reverse.[5] | Premack principle | | Breland and Breland (1961) | Raccoons and pigs trained with food to deposit coins. | Drift toward washing and rooting, though it delayed food.[9] | Instinctive drift | | Bandura, Ross, and Ross (1961) | Children watched an adult attack a Bobo doll, a calm adult, or no model. | Those who saw aggression reproduced its specific acts, unreinforced.[14] | Observational learning | | Garcia and Koelling (1966) | Rats drank sweet, "bright, noisy" water, then were made ill or shocked. | Illness attached to the taste; shock to the light and sound. Seligman (1970) named the general principle preparedness.[10][11] | Taste aversion; preparedness | | Seligman and Maier (1967) | Dogs given escapable, inescapable, or no shock, then tested where escape was possible. | Dogs that had had no control mostly failed to escape.[8] | Learned helplessness | The Bobo doll, latent learning, and insight rows are contrast cases: learning occurred without reinforcement, though reinforcement governs whether it is *performed*. ## How to work a scenario question Whether the scenario is a multiple-choice stem or a paragraph from a study, the same six moves name the process. 1. **Find the behavior.** What the organism does, as a verb. If two people are in the scene, decide whose behavior is being asked about. 2. **Find the consequence.** What happened immediately after, for the one behaving. 3. **Added or removed?** Added is positive; removed is negative. 4. **Up or down?** "More likely," "keeps doing it," and "continues" mean up; "stops" and "less often" mean down. 5. **Name the quadrant.** Added and up: positive reinforcement. Removed and up: negative reinforcement. Added and down: positive punishment. Removed and down: negative punishment. 6. **Check that it is a quadrant at all.** If the usual reinforcer simply stopped, it is extinction. If a reflex was elicited by a signal, it is classical. If the learner watched someone else, it is observational learning. A worked case. Maya's brother screams whenever she takes the tablet; she gives it back to stop the noise, and he screams sooner next time. For the brother, the tablet is returned (added) and screaming goes up: positive reinforcement. For Maya, the screaming stops (removed) and handing it back becomes more likely: negative reinforcement. The [A-B-C model](https://operantconditioning.com/abc-model/) page works this decomposition in detail. ### Naming the schedule from the pattern Two questions. Does reinforcement depend on a *count* of responses ("every tenth," "an unpredictable number of") or on *time* ("after a set time," "at varying intervals")? Count is ratio, time is interval. Is the requirement the *same* each time or does it *vary*? Then confirm with the pattern — pause-and-run, high and steady, the scallop, moderate and steady — because a question that describes pacing is handing you the answer. The [schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) page has a simulator that draws each cumulative record live; once you have watched the scallop form you will not confuse it with the ratio pause. When free-response material describes a study, read the method a sentence at a time and name each procedure as you meet it — the reinforcer, the schedule, the dependent variable — because a question about whether the conclusion follows usually turns on a procedure named correctly. ## Practice set: ten items with answers Work each item with the six moves before opening the answer. The [examples page](https://operantconditioning.com/examples/) has fifty more, and the [quiz](https://operantconditioning.com/quiz/) gives instant explanations. 1. A student gets a sticker every time she turns in homework, and submissions increase. Quadrant and schedule? Answer Positive reinforcement (sticker added, behavior up) on continuous reinforcement, so expect fast extinction if the stickers stop. 2. A driver buckles up to silence the chime and, over time, buckles up faster. Quadrant? Answer Negative reinforcement: the chime is removed and the behavior increases. Not punishment, because nothing decreased. 3. A teenager comes home after curfew and loses driving privileges for a week; late arrivals become less frequent. Quadrant? Answer Negative punishment: a reinforcer is removed and the behavior decreases. Had late arrivals continued, it would not have been punishment, whatever the parents called it. 4. A rat gets a pellet after every tenth press; its cumulative record shows a flat stretch after each pellet, then a steep run. Schedule and pattern? Answer Fixed ratio 10. The flat stretch is the post-reinforcement pause and the steep run the high-rate burst: break-and-run. 5. A gambler keeps pulling a machine that pays after an unpredictable number of pulls; an angler keeps casting where fish strike after unpredictable stretches of time. Name each schedule and the phrase that decides it. Answer Variable ratio: "number of pulls" makes reinforcement depend on a count. Variable interval: "stretches of time" makes it depend on time. Both are unpredictable, so "variable" alone does not decide it. 6. A child's tantrums were maintained by attention. The parents now withhold it; tantrums briefly get louder, then decline. Name the process and the brief increase. Answer Extinction, not punishment: the reinforcer stopped, and nothing is added or removed after each tantrum. The increase is the extinction burst; removing a toy after each tantrum would be negative punishment. 7. A pigeon fed every fifteen seconds regardless of what it does is soon turning in circles between feedings. Phenomenon and researcher? Answer Superstitious behavior, Skinner (1948). Whatever the bird was doing when food arrived was accidentally reinforced, and the accident repeated. 8. Rats run a maze daily with no food and seem to wander. On day eleven food appears and their errors drop at once to the level of rats fed all along. Phenomenon and researchers? Answer Latent learning, Tolman and Honzik (1930). The rats had a cognitive map from the unrewarded days; reinforcement gave them a reason to show it. 9. Raccoons trained with food to drop coins in a container increasingly rub and dip the coins instead, delaying the food. Phenomenon and researchers? Answer Instinctive drift, Breland and Breland (1961). The trained response drifted toward innate food-washing at the cost of reinforcement: consequences shape behavior only within a species' evolved repertoire. 10. A dog wags and drools when it hears the treat pouch open, then sits on cue and gets the treat. Which parts are classical and which operant? Answer Both. Wagging and drooling at the pouch sound are a conditioned response to a stimulus that predicts food: classical. The sit is a voluntary behavior followed by a treat that depends on it: operant, positive reinforcement. ## A study plan using this site The [teachers page](https://operantconditioning.com/for-teachers/) gives a two-session route through the front page for a first pass. This is the exam-oriented version: quadrants and schedules first, a term check last. Spread it over a week. 1. **The front page, first half.** The [one-paragraph version](https://operantconditioning.com/#operant-conditioning-in-one-paragraph), the [A-B-C](https://operantconditioning.com/#how-operant-conditioning-works), and the [four quadrants](https://operantconditioning.com/#the-four-quadrants-reinforcement-and-punishment) with the interactive checker. 2. **Negative reinforcement.** The [negative reinforcement](https://operantconditioning.com/negative-reinforcement/) page clears up the confusion the test writers use most. 3. **Schedules, with the simulator.** The [schedules page](https://operantconditioning.com/schedules-of-reinforcement/); run all four until you can sketch each record. If one is shaky, its own page ([FR](https://operantconditioning.com/fixed-ratio-schedule/), [VR](https://operantconditioning.com/variable-ratio-schedule/), [FI](https://operantconditioning.com/fixed-interval-schedule/), [VI](https://operantconditioning.com/variable-interval-schedule/)) has more examples. 4. **Extinction and shaping.** The [extinction](https://operantconditioning.com/extinction/) page for the burst, spontaneous recovery, and the partial reinforcement effect; the [shaping](https://operantconditioning.com/shaping/) page; then ten minutes in [the lab](https://operantconditioning.com/lab/) shaping a lever press. 5. **Operant vs. classical.** The [comparison page](https://operantconditioning.com/operant-vs-classical-conditioning/) and its checklist. The page to reread the night before. 6. **Names and dates.** The [history page](https://operantconditioning.com/history/), with the studies table above. 7. **Practice.** The [examples page](https://operantconditioning.com/examples/), the [quiz](https://operantconditioning.com/quiz/), then the set above. For each wrong answer, find which of the four confusions produced it. 8. **Term check.** Go down the learning terms in the CED and give a one-line definition of each; the [glossary](https://operantconditioning.com/glossary/) has every one.[1] ## Common mistakes, and what you can skip ### Common mistakes - **Classifying by how the consequence felt.** A punisher is a consequence that reduced behavior; a reinforcer is one that increased it. Nothing else counts. - **Naming the schedule from the reinforcer.** Money, praise, and food can be delivered on any schedule. Ask count or time, fixed or variable. - **Calling extinction forgetting, or punishment.** Extinction is new learning that the response no longer pays, which is why it returns after a rest. - **Confusing a signal with a consequence.** A stimulus before a reflex is a conditioned stimulus; a stimulus before a voluntary behavior that signals it will pay is a discriminative stimulus; a stimulus after the behavior is a consequence. - **Treating the contrast cases as refutations.** Latent learning, insight, and observational learning show that learning also occurs without direct reinforcement, not that reinforcement fails to change behavior. ### What the CED does not require This site goes further than the AP course. Unless the current CED says otherwise, you do not need the following: the [matching law](https://operantconditioning.com/matching-law/); behavioral momentum; the response deprivation hypothesis behind the Premack principle; motivating operations; functional analysis; the equations of the Rescorla–Wagner model; autoshaping, sign-tracking, and contrafreeloading; and the debate over whether the positive/negative distinction should be kept at all.[1] Reading them makes the required material easier to hold, and the [neuroscience page](https://operantconditioning.com/neuroscience/) connects reinforcement to the dopamine system covered elsewhere in the course, but none of it is where the points are. ## Key takeaways - In the revised AP Psychology course, learning is taught within Unit 3, Development and Learning, and the Course and Exam Description is the authority on which terms and question formats are tested. - Every operant scenario is two questions — added or removed, and behavior up or down — and the answers name the quadrant; "positive" and "negative" describe the environment, not the outcome. - Negative reinforcement is not punishment, and extinction is not a quadrant: punishment adds an aversive or removes a reinforcer after the response; in extinction the reinforcer simply stops arriving. - Schedules are named by count-or-time and fixed-or-variable: pause-and-run on fixed ratio, a high steady rate on variable ratio, the scallop on fixed interval, a moderate steady rate on variable interval. - Learn the classic studies as procedures with findings, because the exam describes what was done and asks what it showed; latent learning, insight, and the Bobo doll are the contrast cases, and instinctive drift and preparedness mark the biological limits. ### Check yourself **A teacher says she "negatively reinforced" a student by assigning detention, and the student's talking in class decreased. What did she actually do?** Positive punishment. Something was added (detention) and the behavior went down, so the consequence was a punisher, and because it was added rather than removed it was positive. Negative reinforcement would require the behavior to increase because something was taken away. **One rat gets a pellet after every twentieth press; another gets a pellet for the first press after twenty seconds have passed. Sketch the cumulative record you expect from each.** The first rat is on a fixed ratio 20 and should show break-and-run: a pause after each pellet, then a steep run of twenty presses. The second is on a fixed interval 20 seconds and should show a scallop: few presses just after each pellet, then an accelerating rate as the twenty seconds run out. The ratio schedule produces the higher overall rate, because responding faster brings the pellet sooner; on the interval schedule it does not. **Children who watched an adult hit a Bobo doll later hit it themselves. A classmate calls this operant conditioning because "aggression was reinforced." Is that right?** No. In the 1961 study the children were not reinforced for imitating; they reproduced the specific acts they had watched. That is observational learning, and it is on the exam precisely as the contrast with learning through one's own consequences. Operant conditioning would require the children's own hitting to have been followed by a consequence that changed its frequency. ## Frequently asked questions **What is operant conditioning in simple terms for AP Psychology?** Operant conditioning is learning in which a voluntary behavior becomes more or less frequent because of the consequence that follows it. Reinforcement makes the behavior more frequent and punishment makes it less frequent; positive means a stimulus was added and negative means one was removed. Skinner named it and studied it with rats and pigeons in an operant chamber. **Is learning still tested on the AP Psychology exam after the 2024 revision?** Yes. In the revised course, classical conditioning, operant conditioning, and the social and cognitive factors in learning are taught within Unit 3, Development and Learning, rather than as a unit of their own. The Course and Exam Description lists the required terms and describes the question formats; check the current edition for question counts and weighting. **What is the difference between negative reinforcement and punishment on the AP exam?** Negative reinforcement increases a behavior by removing an unpleasant stimulus, as when buckling a seat belt stops the chime. Punishment decreases a behavior, either by adding an unpleasant stimulus (positive punishment) or by removing a pleasant one (negative punishment). The test is always the effect on behavior: if it went up, it was reinforcement, whatever it felt like. **What is an example of a variable ratio schedule for AP Psychology?** A slot machine. It pays after an unpredictable number of pulls, so reinforcement depends on how many responses are made, and the player responds at a high, steady rate that is very hard to extinguish. Fishing, by contrast, is usually variable interval, because a bite depends on unpredictable stretches of time rather than on the number of casts. **What is the difference between fixed interval and variable interval schedules?** In both, reinforcement depends on time rather than on the number of responses. On a fixed interval the wait is the same every time and behavior shows a scallop, with little responding early and a rising rate as the interval ends. On a variable interval the wait is unpredictable and behavior settles to a moderate, steady rate, because there is no point in the interval at which reinforcement is especially likely. **Which studies do I need to know for operant conditioning on the AP exam?** Thorndike's puzzle boxes and the law of effect (1898); Skinner's operant chamber (1938) and superstitious pigeons (1948); Ferster and Skinner on schedules (1957); Breland and Breland on instinctive drift (1961); Garcia and Koelling on taste aversion (1966); Seligman and Maier on learned helplessness (1967); and, as contrast cases, Tolman and Honzik on latent learning (1930), Köhler on insight (1925), and Bandura's Bobo doll (1961). Learn each as a procedure and a finding. **Is the Bobo doll experiment operant conditioning?** No. In Bandura, Ross, and Ross (1961), children imitated an adult's aggressive acts toward an inflatable doll without ever being reinforced for doing so. That is observational learning, and it appears in the learning topics as a contrast with conditioning through one's own consequences. Later work by Bandura added consequences to the model to study vicarious reinforcement and punishment. **How should I answer an operant conditioning scenario question?** Find the behavior and decide whose it is, find what happened right after it, ask whether something was added or removed, ask whether the behavior went up or down, and combine the two answers to name the quadrant. Then check that it is a quadrant at all: if the reinforcer simply stopped, it is extinction; if a reflex was elicited by a signal, it is classical conditioning; if the learner watched someone else, it is observational learning. ## References 1. College Board. (2024). *AP Psychology Course and Exam Description* (effective Fall 2024). College Board. 2. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. 3. Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. 4. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century. 5. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. *Psychological Review, 66*(4), 219–233. 6. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 7. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172. 8. Seligman, M. E. P., & Maier, S. F. (1967). Failure to escape traumatic shock. *Journal of Experimental Psychology, 74*(1), 1–9. 9. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684. 10. Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. *Psychonomic Science, 4*(1), 123–124. 11. Seligman, M. E. P. (1970). On the generality of the laws of learning. *Psychological Review, 77*(5), 406–418. 12. Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. *University of California Publications in Psychology, 4*, 257–275. 13. Köhler, W. (1925). *The Mentality of Apes*. Harcourt, Brace. 14. Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. *Journal of Abnormal and Social Psychology, 63*(3), 575–582. 15. Pavlov, I. P. (1927). *Conditioned Reflexes: An Investigation of the Physiological Activity of the Cerebral Cortex* (G. V. Anrep, Trans.). Oxford University Press. 16. Watson, J. B., & Rayner, R. (1920). Conditioned emotional reactions. *Journal of Experimental Psychology, 3*(1), 1–14. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Operant vs. classical conditioning](https://operantconditioning.com/operant-vs-classical-conditioning/): The decision checklist and eight worked examples the exam questions are built on. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): All four schedules with a live cumulative-record simulator. - [The 20-question quiz](https://operantconditioning.com/quiz/): Mixed classical and operant items with instant explanations and a printable key. # Operant Conditioning in the Workplace: What Actually Works > How operant conditioning runs the workplace: the OBM evidence, pay as a schedule, feedback, Kerr's folly, Wells Fargo, intrinsic motivation, incentive design. - Source: https://operantconditioning.com/workplace/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-13 · Updated: 2026-09-13 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Applications · Work* A paycheck, a deadline, a sales target, a word from the boss: work is the largest system of consequences most adults live inside. Organizational behavior management applies the operant analysis to it, and forty years of studies are clear about what works, what fails, and how a well-built incentive goes wrong. **In brief** - Organizational behavior management applies the three-term contingency to work: pinpoint a behavior, measure it, analyze its consequences, intervene, and evaluate. A meta-analysis of such programs from 1975 to 1995 found an average 17% improvement in task performance. - Feedback works when it is graphed, frequent, and combined with goals and reinforcement, and not reliably otherwise. Money, social recognition, and feedback each raise performance, and the three together beat any one. - An incentive strengthens exactly the behavior it is contingent on, which is Kerr's folly of rewarding A while hoping for B. Wells Fargo's sales goals produced unauthorized accounts because the contingency worked as designed. ## What organizational behavior management is Every workplace already runs on operant conditioning. Pay arrives on a schedule, supervisors deliver attention and correction, customers respond to what an employee does, and each of these consequences shapes behavior whether or not anyone designed it to. **Organizational behavior management** (OBM) takes this seriously: it treats performance as behavior, behavior as a function of its antecedents and consequences, and management as the job of arranging both. Its three branches are performance management, behavioral systems analysis, and behavior-based safety.[1] Two traditions built the field. Fred Luthans and Robert Kreitner's *Organizational Behavior Modification*, first published in 1975, reduced the manager's task to five steps: identify performance-related behaviors, measure their baseline frequency, analyze their [antecedents and consequences](https://operantconditioning.com/abc-model/), intervene, and evaluate the effect on measured performance.[2] Aubrey Daniels, working with industrial clients, called the same approach *performance management* and gave it a vocabulary: **pinpointing**, which means stating a result and then the observable behaviors that produce it, and the **PIC/NIC analysis**, which classifies every consequence a behavior receives as positive or negative, immediate or future, certain or uncertain. Positive, immediate, and certain consequences control behavior; negative, future, and uncertain ones barely register.[3] What separates this from ordinary management is what it refuses to say. It does not explain poor performance by attitude or motivation, since neither can be observed or changed directly; it asks what the employee's behavior produces now, and what it would have to produce for the desired behavior to occur more often. Skinner's chapter on economic control in *Science and Human Behavior* made the founding observation: wages, prices, and supervision are contingencies, so the question is never whether a workplace conditions its people but what it conditions them to do.[4] ## What the evidence shows The largest test of the five-step model is Stajkovic and Luthans's meta-analysis of O.B. Mod. studies published between 1975 and 1995: an average 17% improvement in task performance, with larger effects in manufacturing than in service organizations.[5] Their 2003 meta-analysis separated the three reinforcers organizations use most. Money, social recognition, and performance feedback each raised performance on their own, money the most and feedback the least, and the three together produced a larger improvement than any one of them.[6] Feedback has the longest evidence trail because it is the cheapest intervention. Two reviews of the published applications, Balcazar, Hopkins, and Suarez's through the mid-1980s and Alvero, Bucklin, and Austin's for 1985 to 1998, reached the same conclusion: feedback does not reliably improve performance by itself. It was most consistent when graphed rather than written or spoken, delivered daily or weekly, and combined with goal setting or with reinforcement such as praise or tangible rewards.[7][8] In operant terms, feedback is a conditioned reinforcer only once it has been paired with consequences that matter. A graph nobody acts on is just a graph. Behavior-based safety is the field's clearest demonstration. Komaki, Barwick, and Scott pinpointed specific safe behaviors in two departments of a wholesale bakery, observed them several times a week, gave a short training session, and then posted a graph of each department's safe-performance score. Safe performance rose from about 70% to 96% in one department and from 78% to 99% in the other, and when the feedback was withdrawn it fell back toward baseline, which is what shows the graph, not the training, was the cause.[9] Grindle, Dickinson, and Boettcher's review of behavioral safety studies in manufacturing found the same across sites: feedback, goals, and reinforcement reliably increased safe behavior, though fewer studies tracked injuries long enough to show the change reached the outcome.[10] Bucklin and Dickinson's review of individual monetary incentives found that pay contingent on performance reliably improved it, and that the details managers argue about most, the size of the incentive relative to base pay and whether the pay function was linear or accelerating, made surprisingly little difference. Whether pay was contingent mattered more than how.[11] ## Pay as a schedule of reinforcement Skinner read the wage system as a set of schedules, and the reading holds. **Piece rates** pay per unit produced, which is a [fixed-ratio schedule](https://operantconditioning.com/fixed-ratio-schedule/). Ratio schedules produce high, steady rates of work and, at high requirements, the pausing and breakdown called [ratio strain](https://operantconditioning.com/glossary/#ratio-strain); Skinner noted that the exhausting rates piecework generates are one reason organized labor has resisted it.[4][12] Edward Lazear's study of Safelite Glass, which moved its windshield installers from hourly pay to piece rates with a guaranteed floor in the mid-1990s, is the cleanest field test. Output per worker rose about 44%, roughly half from existing installers working faster and half from sorting: productive workers stayed and joined, less productive ones left. Lazear reports no sign that quality fell, noting that an installer who broke a windshield had to replace it on his own time; the study's outcome measure, though, was output.[13] **Commission** is a ratio schedule with a variable requirement. A salesperson is paid per sale, but each sale takes an unpredictable number of calls, so the calls are reinforced on something close to a [variable-ratio schedule](https://operantconditioning.com/variable-ratio-schedule/), which produces the highest and most persistent responding of any schedule. That is why a dry month does not stop a good salesperson from dialing. **Salary** is not a schedule of reinforcement for output at all. It arrives on a calendar, but the pay does not depend on any response after the interval elapses, so it is not a [fixed-interval schedule](https://operantconditioning.com/fixed-interval-schedule/) in the technical sense; it is closer to noncontingent delivery with attendance as the only requirement. Skinner's observation was that the weekly wage therefore reinforces little directly, and that working is held in place largely by aversive control: supervision and the standing threat of dismissal.[4] Most workplaces are [negative-reinforcement](https://operantconditioning.com/negative-reinforcement/) systems with a salary attached. The **annual bonus** is the weakest arrangement of all. It arrives months after the behavior, once a year, and depends on outcomes shaped by markets, colleagues, and luck: an extremely thin schedule with a long delay, and in Daniels's terms positive, future, and uncertain, the profile that controls behavior least.[3] [How the four schedules differ, with a simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) ## Rewarding A while hoping for B In 1975 Steven Kerr published "On the folly of rewarding A, while hoping for B," an operant paper in everything but vocabulary. Organizations, he argued, routinely reinforce one behavior while hoping for another, then act surprised when they get the one they paid for. His examples have not aged. Universities reward research and hope for teaching. Orphanages funded per child housed hope for adoptions that would empty their beds. Physicians are punished far more for pronouncing a sick patient well than a well patient sick, so they overdiagnose. In the Second World War soldiers went home when the war was won; in Vietnam they went home after a fixed tour regardless, and the army hoped for victory anyway.[14] Kerr traced the folly to four causes: fascination with an "objective" criterion, so that whatever is easy to count gets rewarded; overemphasis on highly visible behaviors; hypocrisy; and a preference for what looks fair over what works.[14] A behavior analyst would add a fifth: a reward contingent on a result can be earned by many behaviors, some cheaper than the ones management had in mind. Wells Fargo is the folly at national scale. Until the goals were eliminated in late 2016, its Community Bank division set aggressive cross-sell goals for branch employees and tied incentive pay, close tracking, and, as employees understood it, their jobs to meeting them. The 2017 investigation commissioned by the board's independent directors found that the root cause of the sales-practice failures was the distortion of the division's sales culture and performance-management system which, combined with aggressive sales management, pressured employees to sell products customers did not want or need and, in some cases, to open accounts customers had not authorized. For years the bank had treated the problem as individual misconduct, dismissing thousands of employees for sales-practice violations while the goals that produced the behavior stayed in place.[15] The contingency did not fail The incentive did what every contingency does: it strengthened the behavior that produced the reinforcer by the fastest available route. "Accounts opened" was the pinpoint, so accounts were opened. Punishing the shortcut without changing the pinpoint leaves the contingency intact and teaches employees to hide the shortcut. The first question to ask of any incentive is not "what do we want?" but "what is the cheapest behavior that earns this?" ## Do rewards undermine intrinsic motivation? The standard objection is that paying or praising people for work they would do anyway makes them do it less once the reward stops. The evidence comes from experiments in which people rewarded for an interesting activity, such as solving puzzles, later spent less free time on it than people never rewarded. Deci, Koestner, and Ryan's meta-analysis of 128 experiments found that expected tangible rewards, contingent on engaging in, completing, or performing well at the task, reduced free-choice persistence, especially in children, while unexpected rewards had no effect and verbal praise increased it.[16] Cameron and Pierce, analyzing much of the same literature, concluded that rewards did not decrease intrinsic motivation overall, that praise increased it, and that the only reliable undermining came from tangible rewards promised simply for doing a task, regardless of how well.[17] Gerhart and Fang asked whether that narrow finding reaches into work, and answered: not far on current evidence. The undermining experiments used interesting tasks, one-time rewards that were then withdrawn, and participants, often children, with no expectation of being paid; jobs differ on every count. In workplace studies, performance-contingent pay is associated with higher performance and, where measured, with intrinsic motivation no lower and sometimes higher, and part of pay's effect works through sorting, who takes and keeps the job, rather than effort alone.[18] Operant theory draws the same line: "intrinsic" names behavior maintained by its natural consequences, the finished design or the solved problem, and a reward for merely showing up can bring the behavior under the reward's control instead. The finding to respect is narrow: do not pay people for engagement in work they already find reinforcing, and do not remove a reward abruptly, because that is [extinction](https://operantconditioning.com/extinction/) and it looks like lost motivation. The finding to doubt is the general one, that recognition or contingent pay poisons work. ## Goal setting as an antecedent A goal is an antecedent: a verbal statement of a contingency, "if you produce this by Friday, that will follow," which functions as a [discriminative stimulus](https://operantconditioning.com/discriminative-stimulus/) only insofar as consequences actually follow it. Locke and Latham's goal-setting theory, built from hundreds of laboratory and field studies, found that specific, difficult goals produce higher performance than easy goals, vague goals, or instructions to do your best, and that the effect depends on commitment to the goal, the ability to reach it, and feedback on progress. Goals without feedback and feedback without goals are each much weaker than the pair.[19] That is the feedback reviews' conclusion from the other side: the goal sets the occasion, the behavior is performance, and feedback plus reinforcement is the consequence. Goals work, in Locke and Latham's account, by directing attention, raising effort and persistence, and prompting the search for strategies.[19] The operant analysis adds two warnings. A goal never followed by reinforcement loses its function: people stop responding to targets that produce nothing. And a goal followed only by consequences for missing it becomes an aversive stimulus, and behavior under aversive control takes the shortest route to relief. Wells Fargo's goals were specific, difficult, and closely tracked; on goal-setting terms they were well built, which is exactly why the missing analysis was of what behavior they would reinforce. ## Common workplace practices in operant terms The table classifies familiar practices by function. Each row is a hypothesis until measured: a consequence is a reinforcer or a punisher only if the behavior it follows goes up or down. | Practice | Operant term | What it actually does | | --- | --- | --- | | Performance feedback | Conditioned reinforcer or punisher; antecedent for the next response | Works when graphed, frequent, and paired with consequences | | Specific praise or recognition | [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) with a social reinforcer | Strengthens the named behavior when contingent, specific, and credible | | Piece rate | Fixed-ratio schedule | High steady output; ratio strain if set too high; quality unreinforced unless paid for | | Commission | Ratio schedule with a variable requirement | Persistent responding through dry spells; reinforces closing, not honesty | | Salary | Noncontingent with respect to output | Maintains attendance; the work itself is held in place by supervision | | Annual bonus | Delayed, thin, uncertain reinforcement | Little control over daily behavior; a burst of effort before it is decided | | Annual performance review | Delayed, usually aversive consequence | Evokes escape: the polished self-assessment and the defensive meeting | | Deadline | Avoidance contingency with a fixed-interval pattern | Effort accelerates as the date nears; Congress passes its bills in a burst before adjournment[20] | | Micromanagement | Aversive control | Compliance and looking busy are negatively reinforced when the manager leaves; initiative is punished by correction | | Performance improvement plan | Avoidance under threat of dismissal | The minimum behavior that removes the threat, plus job searching | | Employee of the month | Competitive, thin reinforcement | Reinforces one person; for everyone else the month was extinction | | "The beatings will continue until morale improves" | [Positive punishment](https://operantconditioning.com/positive-punishment/) aimed at a non-behavior | Morale is not a pinpoint; punishment suppresses everything nearby and teaches escape | Most of what managers call accountability is aversive control. It works, in that coercion reliably produces compliance, and it has the side effects Sidman catalogued: escape and avoidance, countercontrol, aggression, and the disappearance of any behavior not strictly required.[21] People under it do what removes the pressure, which is not always what the organization needs. ## How to design a contingency at work 1. **Pinpoint the behavior.** Name the result, then the observable behaviors that produce it. "Customer satisfaction" is a result; "calls the customer back within one business day" is a behavior. If two observers cannot agree on whether it happened, it is not yet a pinpoint.[3] 2. **Measure a baseline.** Count before you change anything; Komaki's baseline is what later proved the feedback worked.[9] 3. **Analyze the current contingencies.** Ask what the desired behavior produces now (often nothing, or more work) and what the competing behavior produces (often relief). The cheapest fix is often an antecedent: a checklist or a visible cue. 4. **Arrange immediate feedback.** Graph it, post it, deliver it daily or weekly, and have a supervisor rather than a system deliver it where possible.[7][8] 5. **Reinforce the behavior, not only the outcome.** Outcomes lag, depend on other people, and can be gamed; behavior can be reinforced the day it occurs. Reinforce the pinpointed behaviors with attention, recognition, and small contingent rewards, chosen by observing what people do when free, the [Premack principle](https://operantconditioning.com/premack-principle/), rather than by guessing. 6. **Thin the schedule.** Reinforce every occurrence while the behavior is being established, then move to an [intermittent schedule](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/). Behavior reinforced every time extinguishes fast when the program ends, as the reversal phases in the safety studies show.[9] 7. **Check for Kerr's folly.** Ask what the cheapest behavior that earns the reward is; if it is not the one you pinpointed, redesign.[14] 8. **Evaluate and keep the data.** Compare against baseline, and withdraw and reinstate the consequence if you can. If performance did not change, the consequence was not a reinforcer, whatever it cost. ### What managers get wrong The recurring errors follow from skipping steps. Managers reward outcomes whose behavior they never saw, which reinforces whatever produced the number, including luck. They save consequences for the annual review, months late, once, and aversive enough to evoke escape rather than change. They build punishment-heavy cultures because punishment works fast and its side effects arrive slowly: what survives is [avoidance](https://operantconditioning.com/avoidance-learning/), the behavior least sensitive to whether the threat is still real. And they pick reinforcers by assumption, the pizza party and the plaque, instead of by observation. ## Ethics, and what the evidence does not show ### Ethics The manager holds the reinforcers and the employee's livelihood, so a contingency at work is never between equals, and the ethical practice follows from that asymmetry. Use positive reinforcement first and aversive control last, and never manufacture an aversive so that you can remove it. Make the contingency explicit: a hidden contingency is manipulation and a stated one is a deal. Pinpoint behaviors the employee would endorse if asked, and design so that hitting the target serves the customer as well as the firm; Wells Fargo is what happens when that test is skipped. Skinner's case against punishment was not that it fails but that it works quickly and its costs arrive later.[4] ### What the evidence does not show The OBM literature is mostly single-organization studies with reversal or multiple-baseline designs, lasting weeks or months, and published when they worked; the meta-analyses summarize task performance, not innovation, retention, or wellbeing.[5][6] Lazear's 44% is one firm doing a simple, countable job, and the gain includes who left as well as who worked faster.[13] Feedback effects fade when feedback stops, which is proof of function, not of durability.[9] The safety reviews show behavior change more clearly than injury reduction.[10] The undermining effect is real in the laboratory and, on present evidence, small or absent for performance-contingent pay in jobs, but the field literature is still short of good experiments.[16][17][18] And nothing here shows that incentives substitute for the ability to do the job: goals and reinforcement raise the performance of people who already know how, and for people who do not, the tools are training and [shaping](https://operantconditioning.com/shaping/), not a bigger bonus. ## Key takeaways - OBM applies the ABC model to work: pinpoint a behavior, measure it, analyze its consequences, intervene, and evaluate. The meta-analytic record shows an average 17% improvement in task performance. - Feedback is a conditioned reinforcer only once it has been paired with consequences, which is why it works when graphed, frequent, and combined with goals and reinforcement, and not reliably otherwise. - Pay is a schedule: piece rates are fixed-ratio, commission is a ratio schedule with a variable requirement, salary is noncontingent on output and maintained by supervision, and the annual bonus is delayed, thin, and uncertain. - An incentive strengthens exactly the behavior it is contingent on, by the cheapest route available. That is Kerr's folly, and Wells Fargo's unauthorized accounts were the contingency working as designed. - The undermining of intrinsic motivation is real for tangible rewards given for merely engaging in an interesting task and unproven for performance-contingent pay at work. Reinforce performance, not attendance, and never withdraw a reward abruptly. - Most of what organizations call accountability is aversive control, which produces compliance along with escape, avoidance, and countercontrol. Reinforce behavior you can see, and deliver consequences within days rather than at the annual review. ### Check yourself **A call center pays a bonus for keeping average call length under four minutes. Call length drops within a week, and customers start calling back two and three times about the same problem. What happened?** Kerr's folly. The pinpoint was a proxy, call length, and the cheapest behavior that earns a shorter call is ending it before the problem is solved. The contingency worked exactly as designed. The fix is to pinpoint the behavior actually wanted, such as resolving the issue on the first call, and to reinforce that, not to punish the agents who found the shortcut. **A manager emails her team a spreadsheet of last week's error counts every Monday. After three months, error rates are unchanged. She concludes that feedback does not work. Is she right?** No. The reviews found that feedback is not reliably effective on its own; it worked most consistently when graphed rather than tabulated, delivered often, and paired with goals and reinforcement. Nothing follows her spreadsheet, so it has not become a conditioned reinforcer or punisher. Whether a stimulus is feedback in the operant sense is decided by its effect on behavior, and this one has none. **A firm switches from hourly pay to piece rates and output per worker rises 40%. An executive concludes the workers had been lazy. What does the Safelite evidence suggest instead?** That about half of such a gain typically comes from sorting rather than effort: more productive workers stay and join, and less productive ones leave. The other half is the incentive effect on the people who were already there, which is a change in the contingency, not in character. It is also worth checking what happened to quality, which piece rates leave unreinforced unless it is explicitly paid for. ## Frequently asked questions **What is operant conditioning in the workplace in simple terms?** It is the fact that what employees do is shaped by what follows it. Pay, praise, correction, deadlines, and a manager's attention are all consequences, and behavior that produces good consequences quickly and reliably becomes more frequent. Organizational behavior management uses this deliberately: it specifies a behavior, measures it, gives frequent feedback, and reinforces it, instead of appealing to attitude or motivation. **What is organizational behavior management (OBM)?** The application of behavior analysis to work. Its five-step model, from Luthans and Kreitner, is to identify performance-related behaviors, measure them, analyze their antecedents and consequences, intervene, and evaluate. Its branches are performance management, behavioral systems analysis, and behavior-based safety. A meta-analysis of programs from 1975 to 1995 found an average 17% improvement in task performance. **What is an example of operant conditioning at work?** In a bakery, researchers defined specific safe behaviors, observed them several times a week, and posted a graph of each department's safe-performance score. Safe performance rose from about 70% to over 95% and fell again when the graph was withdrawn, which showed the graph was the working part. The posted score was a conditioned reinforcer for safe behavior. Commission, piece rates, and specific praise are everyday examples. **Positive reinforcement vs. punishment at work: which works better?** Punishment and threats produce compliance quickly, which is why organizations use them, but they also produce escape, avoidance, resentment, and the disappearance of any behavior not strictly required. Positive reinforcement builds behavior that persists and generalizes, and it has no such side effects. The evidence from OBM favors reinforcement plus feedback, with aversive control as a last resort for behavior that must stop now. **Does pay for performance work?** Contingent pay reliably raises measured performance on countable tasks. At Safelite Glass, moving installers from hourly pay to piece rates raised output per worker about 44%, half from effort and half from sorting of workers. Reviews find that whether pay is contingent matters more than how large the incentive is. The risks are that pay reinforces only what is measured, and that quality and honesty go unreinforced unless they are paid for too. **Do rewards undermine intrinsic motivation at work?** In the laboratory, expected tangible rewards for merely engaging in an interesting task reduce later free-choice persistence, especially in children; praise increases it. Whether that reaches into work is contested. A review of workplace studies found performance-contingent pay associated with higher performance and no drop in intrinsic motivation. The safe rule is to reinforce performance, not attendance, and never to withdraw a reward abruptly. **Why do annual performance reviews fail to change behavior?** Because a consequence delivered months after the behavior, once a year, and in an aversive setting fails on every dimension that gives consequences their power: it is neither immediate nor certain, and it is often not positive. What it reliably produces is escape behavior, such as polished self-assessments and defensive meetings. Frequent feedback from a supervisor who knows what to look for does more at lower cost. **What did the Wells Fargo scandal show about incentives?** The 2017 board investigation found that aggressive cross-sell goals, backed by incentives and sales pressure, led employees to sell products customers did not need and to open accounts customers had not authorized. The contingency reinforced the pinpointed behavior, accounts opened, by the cheapest route. It is the clearest modern case of rewarding A while hoping for B, and the bank's early response of firing employees left the goals untouched. ## References 1. Wilder, D. A., Austin, J., & Casella, S. (2009). Applying behavior analysis in organizations: Organizational behavior management. *Psychological Services, 6*(3), 202–211. 2. Luthans, F., & Kreitner, R. (1985). *Organizational Behavior Modification and Beyond*. Scott, Foresman. 3. Daniels, A. C., & Bailey, J. S. (2014). *Performance Management: Changing Behavior That Drives Organizational Effectiveness* (5th ed.). Performance Management Publications. 4. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan. 5. Stajkovic, A. D., & Luthans, F. (1997). A meta-analysis of the effects of organizational behavior modification on task performance, 1975–95. *Academy of Management Journal, 40*(5), 1122–1149. 6. Stajkovic, A. D., & Luthans, F. (2003). Behavioral management and task performance in organizations: Conceptual background, meta-analysis, and test of alternative models. *Personnel Psychology, 56*(1), 155–194. 7. Balcazar, F., Hopkins, B. L., & Suarez, Y. (1985). A critical, objective review of performance feedback. *Journal of Organizational Behavior Management, 7*(3–4), 65–89. 8. Alvero, A. M., Bucklin, B. R., & Austin, J. (2001). An objective review of the effectiveness and essential characteristics of performance feedback in organizational settings (1985–1998). *Journal of Organizational Behavior Management, 21*(1), 3–29. 9. Komaki, J., Barwick, K. D., & Scott, L. R. (1978). A behavioral approach to occupational safety: Pinpointing and reinforcing safe performance in a food manufacturing plant. *Journal of Applied Psychology, 63*(4), 434–445. 10. Grindle, A. C., Dickinson, A. M., & Boettcher, W. (2000). Behavioral safety research in manufacturing settings: A review of the literature. *Journal of Organizational Behavior Management, 20*(1), 29–68. 11. Bucklin, B. R., & Dickinson, A. M. (2001). Individual monetary incentives: A review of different types of arrangements between performance and pay. *Journal of Organizational Behavior Management, 21*(3), 45–137. 12. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts. 13. Lazear, E. P. (2000). Performance pay and productivity. *American Economic Review, 90*(5), 1346–1361. 14. Kerr, S. (1975). On the folly of rewarding A, while hoping for B. *Academy of Management Journal, 18*(4), 769–783. 15. Independent Directors of the Board of Wells Fargo & Company. (2017). *Sales Practices Investigation Report*. Wells Fargo & Company. 16. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668. 17. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423. 18. Gerhart, B., & Fang, M. (2015). Pay, intrinsic motivation, extrinsic motivation, performance, and creativity in the workplace: Revisiting long-held beliefs. *Annual Review of Organizational Psychology and Organizational Behavior, 2*, 489–521. 19. Locke, E. A., & Latham, G. P. (2002). Building a practically useful theory of goal setting and task motivation: A 35-year odyssey. *American Psychologist, 57*(9), 705–717. 20. Weisberg, P., & Waldrop, P. B. (1972). Fixed-interval work habits of Congress. *Journal of Applied Behavior Analysis, 5*(1), 93–97. 21. Sidman, M. (1989). *Coercion and Its Fallout*. Authors Cooperative. ## About the author Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards ## Related - [Applications of operant conditioning](https://operantconditioning.com/applications/): Every field where the contingency is used, with the evidence rated. - [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Ratio and interval, fixed and variable: the patterns behind pay. - [The ABC model](https://operantconditioning.com/abc-model/): Antecedent, behavior, consequence: the method behind pinpointing. # About OperantConditioning.com > Who publishes operantconditioning.com, how the content is written and checked, how the site is funded, and how to reach us. - Source: https://operantconditioning.com/about/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-12 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *About* OperantConditioning.com is a free, evidence-based reference on operant conditioning — the science of how consequences shape behavior — written for students, parents, teachers, trainers, clinicians, and anyone trying to change their own habits. ## Who publishes it The site is published by **Operant Conditioning Inc.** and written and maintained by **Ryan Martinson**, the company's founder. The company also makes [Operant](https://operantconditioning.com/app/), a habit app for iPhone and Apple Watch built on Skinner's three-term contingency; building it required reading the primary literature closely, and this site is the reference we wished had existed while doing so. ## Editorial standards - **Primary sources.** Claims about experiments, dates, people, and effect sizes are cited to the original papers and books — Thorndike, Skinner, Ferster, Azrin, Herrnstein, and the applied literature that followed — rather than to secondary summaries. The public-domain books behind the science are republished in full in the [library](https://operantconditioning.com/library/). - **Functional definitions.** We use the technical vocabulary of behavior analysis consistently: reinforcers and punishers are defined by their effect on behavior; "positive" and "negative" mean added and removed. - **Honesty about limits.** Every major page includes what the evidence does *not* show, where researchers disagree, and where the theory has been superseded. - **Corrections.** If you find an error, email us. Corrections are made promptly and the page's "updated" date is changed. ## Expert review We are actively seeking a credentialed reviewer — a Board Certified Behavior Analyst (BCBA-D) or a PhD in behavior analysis or learning — to review the site's content. If that is you and you are interested, please get in touch. ## How the site is funded The site is paid for by Operant Conditioning Inc. It carries no tracking and no paid placements, and everything on it — the guide, the glossary, the quiz, the lab — can be read, assigned, and linked without an account. On iPhone and iPad, Safari may show Apple's standard App Store banner for the company's Operant app at the top of the page; that banner, the App link in the navigation, and the app's [own page](https://operantconditioning.com/app/) are the only promotion on the site. ## Contact Email: [hello@operantconditioning.com](mailto:hello@operantconditioning.com) Publisher: Operant Conditioning Inc. --- # Privacy Policy > How operantconditioning.com handles personal data. In short — this site does not collect it. - Source: https://operantconditioning.com/privacy/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-07 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Legal* Last updated September 7, 2026. ## What we collect OperantConditioning.com is a static website. It does not use cookies, does not require an account, does not include third-party tracking scripts, and does not collect personal information from visitors. Interactive tools on the site (the quadrant finder, the virtual Skinner box, the schedule simulator, the quiz, and the contingency builder) run entirely in your browser and do not send data anywhere. The contingency builder saves what you type in your own browser's local storage so it is still there when you come back; it never leaves your device, and the tool's Clear button erases it. ## Server logs Like virtually all websites, the hosting provider that serves this site may keep standard server logs (IP address, browser type, pages requested, and timestamps) for security and operational purposes. These are retained only as long as the provider's policy requires and are not used to identify individual visitors. ## Links to other sites Every page carries a standard `apple-itunes-app` meta tag, which lets Safari on iPhone and iPad show Apple's own App Store banner for the Operant app. No script or image is loaded for it; the banner is drawn by Safari, and tapping it takes you to the App Store under [Apple's privacy policy](https://www.apple.com/legal/privacy/). The homepage embeds three videos from YouTube (a TED-Ed lesson and two television clips); nothing is requested from YouTube until you press play, after which [Google's privacy policy](https://policies.google.com/privacy) applies to the player (it is loaded from youtube-nocookie.com, YouTube's privacy-enhanced domain). ## Email If you email us, we keep your message and address only as long as needed to respond. ## Changes If we ever add analytics or other data collection, this page will be updated first. ## Contact [hello@operantconditioning.com](mailto:hello@operantconditioning.com) --- # Terms of Use > Terms of use for operantconditioning.com, including the educational nature of the content and copyright. - Source: https://operantconditioning.com/terms/ - Author: Ryan Martinson (https://operantconditioning.com/about/) - Publisher: Operant Conditioning Inc. - Published: 2026-09-07 · Updated: 2026-09-10 - License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching) --- *Legal* Last updated September 7, 2026. ## Educational content The content on OperantConditioning.com is provided for general educational purposes. It is not medical, psychological, veterinary, legal, or professional advice, and it is not a substitute for assessment and treatment by a qualified professional. If you are dealing with dangerous behavior — in a child, an animal, or yourself — consult a licensed professional such as a Board Certified Behavior Analyst, psychologist, physician, or veterinary behaviorist. ## Accuracy We work hard to make the content accurate and cite primary sources, but science moves and mistakes happen. Content is provided "as is" without warranties of any kind. We welcome corrections at [hello@operantconditioning.com](mailto:hello@operantconditioning.com). ## Copyright Text, diagrams, and interactive tools on this site are © Operant Conditioning Inc. You may quote brief excerpts with attribution and a link. Teachers may use pages for non-commercial classroom instruction. Republishing whole pages requires written permission. The diagrams collected on the [diagram library page](https://operantconditioning.com/diagrams/) are additionally released under the [Creative Commons Attribution 4.0 license](https://creativecommons.org/licenses/by/4.0/): you may copy, adapt, and reuse them, including commercially, with attribution to operantconditioning.com. ## Trademarks Operant™ is a trademark of Operant Conditioning Inc. Apple, the Apple logo, iPhone, Apple Watch, and App Store are trademarks of Apple Inc., registered in the U.S. and other countries. This site is not affiliated with or endorsed by Apple Inc. ## Links The site embeds videos from YouTube (a TED-Ed lesson and short clips from television programs, which remain the property of their rights holders) and links to external references. We are not responsible for the content of external sites. ---