# Operant Conditioning: The Complete Guide

> Operant conditioning is learning through consequences. The definition, Skinner's four quadrants, schedules of reinforcement, examples, and how to apply it.

- Source: https://operantconditioning.com/
- Author: Ryan Martinson (https://operantconditioning.com/about/)
- Publisher: Operant Conditioning Inc.
- Published: 2026-09-07 · Updated: 2026-09-12
- License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching)

---

*The complete guide*

Operant conditioning is learning from consequences: a behavior becomes more likely if it’s followed by reinforcement and less likely if it’s followed by punishment.

- **350 BC** — [Aristotle](https://operantconditioning.com/history/#aristotle-on-habit): "Pleasure and pain are also the standards by which we all, in a greater or less degree, regulate our actions."
- **1898** — [Thorndike](https://operantconditioning.com/history/)'s law of effect
- **1937** — [Skinner](https://operantconditioning.com/bf-skinner/) coins the term "operant conditioning"
- **Now** — [Reinforcement learning](https://operantconditioning.com/applications/#economics) is used to improve AI

## What is operant conditioning?

> **Definition**
>
> **Operant conditioning** is a form of learning in which the future frequency of a behavior is changed by the consequences that follow it. Behaviors followed by [reinforcement](https://operantconditioning.com/positive-reinforcement/) become more likely; behaviors followed by [punishment](https://operantconditioning.com/positive-punishment/) become less likely.
>
> The term was introduced by American psychologist [B. F. Skinner](https://operantconditioning.com/bf-skinner/) in 1937, building on Edward Thorndike's law of effect (1898). It is also called *instrumental conditioning*, and it is the foundation of applied behavior analysis, modern animal training, and most evidence-based habit-change methods.[1][2]

**Pick a door.**

- [I can't put my phone down.](https://operantconditioning.com/#the-machine-you-are-in) — Checking and scrolling run on two different schedules. One of them is the strongest of the five.
- [I want a habit that actually sticks.](https://operantconditioning.com/#how-to-use-operant-conditioning-on-yourself) — How motivated you feel is the part you cannot arrange. Three other things you can.
- [My kid keeps doing the thing.](https://operantconditioning.com/parenting/) — What the evidence says about time-out, rewards, and why consistency beats severity.
- [My dog is ignoring me.](https://operantconditioning.com/dog-training/) — Marker timing, clicker mechanics, and the method that replaced dominance theory.
- [Just explain it properly.](https://operantconditioning.com/#operant-conditioning-in-one-paragraph) — The definition, the four quadrants, the evidence, the five schedules, and where the theory runs out.

Teaching this, or studying it? The quiz, glossary, diagrams, citation formats and discussion questions all live on [one page](https://operantconditioning.com/for-teachers/).

## The machine you're already in

Sometime in the last hour you picked up your phone without deciding to. You weren't looking for anything in particular. You just checked. Then, probably, you kept scrolling.

Those are two different mechanisms, and behavior analysts have a name for each. **Checking** pays off on a time basis — either something arrived since you last looked or it didn't, and pressing the button twice as often will not make messages arrive twice as fast. That is a [variable-interval schedule](https://operantconditioning.com/schedules-of-reinforcement/), and it produces a moderate, stubbornly steady rate of responding. It is why you check again four minutes later.

**Scrolling** is the other kind. Each swipe is a response, and whether it pays depends on how many swipes you make rather than on how long you wait. That is a [variable-ratio schedule](https://operantconditioning.com/schedules-of-reinforcement/), and of the five basic schedules — catalogued, with dozens of combinations, in a 700-page book Ferster and Skinner published in 1957 — it is the one that produces the highest and steadiest rate of responding, and the one that keeps behavior going longest after the payoffs stop. It is what a slot machine is. It is what a loot box is. It is what an infinite feed is.

None of which is a character flaw, and in most cases it is not addiction in any clinical sense. It is a schedule doing what seventy years of data say a schedule will do.

Which is the genuinely useful part. A mechanism strong enough to keep you swiping against your own stated wishes is strong enough to aim somewhere else, and aiming it is a skill rather than a personality trait. The rest of this page is that mechanism from the ground up: what a consequence has to do to count as one, the four things it can be, the five ways to time it, and what happens when it stops.

> **One honest caveat before you go further.** Schedules explain a great deal about behavior and not all of it. People also learn by watching someone else, by building a map of a situation before any reward arrives, and — most awkwardly for the theory — through language. Where this runs out, the page says so rather than papering over it.

## Watch: operant vs. classical conditioning in four minutes

Start here if you are new to the topic. This TED-Ed lesson by Peggy Andover walks through Pavlov's dogs, then shows how Skinner's operant conditioning differs — behavior first, consequence second — and how reinforcement and punishment change what an animal (or a person) does next.

*Figure: "The difference between classical and operant conditioning," a TED-Ed lesson by Peggy Andover (2013). Embedded from YouTube's privacy-enhanced player.*

> **What to watch for**
>
> Two things the video makes vivid: in classical conditioning the animal is *passive* — the bell and the food arrive whether or not it does anything — while in operant conditioning the animal's own action is what produces the consequence. And "negative" never means "bad": it means something was *taken away*. The rest of this page builds on both ideas.

## Operant conditioning in one paragraph

Every organism that can learn is constantly running the same experiment: *do something, notice what happens next, adjust.* A rat presses a lever and a food pellet drops, so it presses again. A toddler says "please" and gets the cookie, so "please" becomes a habit. You check your phone, a notification rewards you, and the checking becomes automatic. In each case the behavior **operates** on the environment (hence "operant") and the environment answers back with a consequence. Operant conditioning is the study of how those consequences select which behaviors survive and which fade away.[3]

Three ideas do most of the work:

- **Consequences are defined by their effect, not their appearance.** A "reward" that doesn't increase behavior is not a [reinforcer](https://operantconditioning.com/glossary/#reinforcer). A "punishment" that doesn't decrease behavior is not a punisher. This is a functional definition, and it is the single most important thing to understand about the whole field.
- **Behavior is selected over time, the way evolution selects traits.** Skinner called this "selection by consequences." Reinforcement doesn't teach a rule; it shifts probabilities.[4]
- **The unit of analysis is the three-term contingency:** an antecedent sets the occasion, a behavior occurs, a consequence follows. Learn to see the A-B-C pattern and you can read almost any behavior.

## How operant conditioning works

Skinner's central insight was that behavior is not just triggered by what comes *before* it (as in Pavlov's reflexes), it is shaped by what comes *after* it. He formalized this as the **[three-term contingency](https://operantconditioning.com/glossary/#three-term-contingency)**, often written A → B → C:[3]

Two more variables determine how much a consequence matters:

- **[Contiguity](https://operantconditioning.com/glossary/#contiguity) (timing).** Consequences that follow within seconds are far more effective than delayed ones. In laboratory studies the strength of learning drops off steeply as the delay grows to even a few seconds — one reason a paycheck at the end of the month is a poor reinforcer for any specific behavior on a Tuesday morning.[5]
- **[Contingency](https://operantconditioning.com/glossary/#contingency) (dependability).** The consequence must actually depend on the behavior. If food arrives whether or not the rat presses the lever, lever-pressing doesn't get learned — and, as Skinner showed in his famous "superstition" experiment, pigeons fed on a timer developed odd ritual behaviors that happened to precede the food by chance.[6]

> **Operant vs. respondent behavior**
>
> Skinner distinguished **[respondent](https://operantconditioning.com/glossary/#respondent-behavior)** behavior (reflexes elicited by a prior stimulus — salivation, the eye-blink, the startle response) from **[operant](https://operantconditioning.com/glossary/#operant)** behavior (actions emitted by the organism and controlled by their consequences). Pavlov studied the first; Skinner studied the second. Most of what people mean by "behavior" — walking, talking, working, scrolling — is operant. [Full comparison ›](https://operantconditioning.com/operant-vs-classical-conditioning/)

## Inside the Skinner box

The apparatus that made all of this measurable was Skinner's **operant conditioning chamber** — the "Skinner box." Thorndike had timed cats escaping from puzzle boxes one trial at a time; Skinner's innovation was a box the animal never had to leave, so it could respond whenever it liked and the *rate* of responding could be recorded continuously.[2] [The Skinner box in full: every part, and why it was built that way ›](https://operantconditioning.com/skinner-box/)

![Schematic of an operant conditioning chamber](https://operantconditioning.com/assets/diagrams/schematic-of-an-operant-conditioning-chamber.svg)

*An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a food pellet on any schedule; the light signals when pressing will pay off; the recorder draws responses over time.*

A typical experiment runs in four steps. The rat is kept mildly hungry so food works as a reinforcer. It first learns that the click of the food dispenser means a pellet has arrived (a [conditioned reinforcer](https://operantconditioning.com/glossary/#conditioned-reinforcer)). Then the experimenter [shapes](https://operantconditioning.com/shaping/) lever-pressing, reinforcing closer and closer approximations until the rat presses on its own. Finally the schedule is thinned — every press, then every fifth, then an unpredictable number — and the recorder shows how the pattern of responding changes. Pigeons peck a lit key instead of pressing a lever; the logic is identical. [More on Skinner and the box ›](https://operantconditioning.com/bf-skinner/)

## Run your own Skinner box

At the top of this page you were the rat. Here you hold the pellet button. Reading about shaping is one thing; doing it is another. This lab puts a hungry, untrained rat in a chamber and hands you the pellet button. Magazine-train it, shape a lever press one approximation at a time, put the press on a schedule, then take the food away and watch extinction — all in about three minutes. The counters under the chamber track every pellet you deliver and every press the rat makes, from the first step on.

The rat is a stylized model — its tendencies shift with what you reinforce, drift back when you don't, and follow the schedule patterns Ferster and Skinner documented — not a replay of real data. Real shaping takes longer and real rats are more surprising.

Want the lab on its own page to link, share, or assign? [Open the virtual Skinner box ›](https://operantconditioning.com/lab/)

## The four quadrants: reinforcement and punishment

Every consequence can be sorted along two questions. **Did the behavior increase or decrease?** (That tells you whether it was reinforcement or punishment.) **Was a stimulus added or removed?** (That tells you whether it was "positive" or "negative.") Crucially, in this vocabulary *positive* and *negative* mean plus and minus — added and removed — not good and bad.

- **Positive reinforcement** — A pleasant stimulus is added after the behavior. The dog sits; the dog gets a treat. You finish a task; you feel a hit of satisfaction. (https://operantconditioning.com/positive-reinforcement/)
- **Negative reinforcement** — An aversive stimulus is removed after the behavior. You buckle up; the seat-belt chime stops. You take an aspirin; the headache goes away. (https://operantconditioning.com/negative-reinforcement/)
- **Positive punishment** — An aversive stimulus is added after the behavior. You touch a hot pan; it burns. You speed; you get a ticket. (https://operantconditioning.com/positive-punishment/)
- **Negative punishment** — A pleasant stimulus is removed after the behavior. A teenager breaks curfew; the car keys are gone for a week. A player fouls; they sit out. (https://operantconditioning.com/negative-punishment/)

> **The two-question test**
>
> Ask, in this order: **(1) Did the behavior become more or less likely?** More likely means reinforcement; less likely means punishment. **(2) Was something added or removed?** Added means positive; removed means negative. Those two answers name the quadrant every time.

### The four types of operant conditioning at a glance

| Type | What happens after the behavior | Effect on the behavior | Example |
| --- | --- | --- | --- |
| [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/) | A stimulus is added | Increases | A dog sits and gets a treat; sitting becomes more frequent |
| [Negative reinforcement](https://operantconditioning.com/negative-reinforcement/) | A stimulus is removed | Increases | You buckle up and the seat-belt chime stops; buckling up becomes faster and more reliable |
| [Positive punishment](https://operantconditioning.com/positive-punishment/) | A stimulus is added | Decreases | You touch a hot pan and get burned; touching hot pans becomes rarer |
| [Negative punishment](https://operantconditioning.com/negative-punishment/) | A stimulus is removed | Decreases | A teenager breaks curfew and loses the car keys; breaking curfew becomes rarer |

A fifth process, [extinction](https://operantconditioning.com/extinction/), is not a quadrant: the reinforcer that used to follow the behavior simply stops arriving, and the behavior fades.

### Which quadrant is it? An interactive check

Almost everyone gets one pair backwards the first time, and it is nearly always *negative reinforcement* mistaken for *punishment*. Answer the two questions about any scenario and the tool will classify it.

### Five mistakes almost everyone makes

1. **Treating "negative reinforcement" as a polite word for punishment.** It is the opposite: negative reinforcement makes a behavior *more* likely by taking something unpleasant away. Taking an aspirin to end a headache is negative reinforcement of aspirin-taking.
2. **Reading "positive" as good and "negative" as bad.** They mean added and removed. A spray of water in a dog's face is *positive* punishment.
3. **Calling something a reinforcer because it seems nice.** A reinforcer is anything that increases the behavior it follows — it is defined by its effect. Scolding that a bored child finds attention-worthy is a reinforcer; a sticker a teenager finds embarrassing is not.
4. **Mixing up operant and classical conditioning.** If the key event comes *before* the response and the response is a reflex (salivating, flinching), it is classical. If the key event comes *after* a voluntary behavior, it is operant.
5. **Confusing extinction with punishment.** Extinction means the reinforcer simply stops arriving; nothing is added or taken away as a consequence. The behavior fades — usually after a brief burst — rather than being suppressed.

### Try three

## Reinforcement versus punishment: what the evidence says

Skinner believed punishment was a poor way to change behavior, in part because an early experiment by his student W. K. Estes suggested that punishment only temporarily suppressed responding.[7] Later research complicated that picture: punishment *can* produce lasting decreases when it is immediate, consistent, and sufficiently intense from the outset.[8] But those same studies documented why practitioners still prefer reinforcement:

- Punishment teaches what *not* to do without teaching what to do instead. Reinforcement builds a replacement.
- Punishment tends to produce escape and avoidance — of the punisher as much as the behavior. (The child learns not to get caught.)
- It can elicit aggression and emotional side effects, and it models the use of aversive control.
- It works best at intensities and consistencies that are ethically or practically unavailable in most human settings.

In parenting specifically, a large body of research links corporal punishment to worse, not better, long-term behavioral outcomes.[9] Modern applied behavior analysis therefore treats reinforcement-based procedures as the default and reserves punishment for narrow, supervised cases where reinforcement alone has failed and the behavior is dangerous.[10]

[Reinforcement: the two types, kinds of reinforcers, and what makes it work ›](https://operantconditioning.com/reinforcement/) [Punishment: the two types, the side effects, and the alternatives ›](https://operantconditioning.com/punishment/)

## Schedules of reinforcement

Once a behavior is learned, *how often* it gets reinforced changes both how fast the organism responds and how long the behavior persists when reinforcement stops. Ferster and Skinner catalogued these patterns in a 700-page 1957 volume, and the basic findings have held up for seventy years.[11]

| Schedule | Reinforcer delivered… | Typical response pattern | Everyday example |
| --- | --- | --- | --- |
| **Continuous (CRF)** | after every response | Fast learning; fast extinction | A vending machine |
| **Fixed ratio (FR)** | after a set number of responses | High rate with a pause after each reinforcer | Paid per piece; "buy 10, get 1 free" |
| **Variable ratio (VR)** | after an unpredictable number of responses | Highest, steadiest rate; most resistant to extinction | Slot machines; social-media feeds |
| **Fixed interval (FI)** | for the first response after a set time | "Scallop": slow after a reinforcer, accelerating as the interval ends | Checking the oven as the timer nears zero |
| **Variable interval (VI)** | for the first response after an unpredictable time | Moderate, steady rate | Checking email |

The practical rule: **use continuous reinforcement to build a behavior, then thin to an [intermittent schedule](https://operantconditioning.com/glossary/#intermittent-reinforcement) to make it durable.** The variable-ratio schedule is why gambling and infinite-scroll apps are so hard to quit — and why a behavior you reinforce only sometimes can end up stronger than one you reinforce every time. [Run the interactive schedule simulator ›](https://operantconditioning.com/schedules-of-reinforcement/) · [How organisms choose between schedules: the matching law ›](https://operantconditioning.com/matching-law/)

## Extinction, shaping, and stimulus control

### Extinction

When a previously reinforced behavior stops producing reinforcement, it gradually declines. But not immediately: there is usually an **[extinction burst](https://operantconditioning.com/glossary/#extinction-burst)** — a temporary spike in the frequency, intensity, and variability of the behavior — before it fades. (Push the elevator button; nothing happens; you push it harder and faster before giving up.) Behavior that has been extinguished can also show **[spontaneous recovery](https://operantconditioning.com/glossary/#spontaneous-recovery)** after a rest period. [More on extinction ›](https://operantconditioning.com/extinction/)

### Shaping

Complex behavior is rarely emitted fully formed, so it can't simply be reinforced. **Shaping** solves this by reinforcing *[successive approximations](https://operantconditioning.com/glossary/#successive-approximations)* — first any movement toward the lever, then touching it, then pressing it. Skinner used shaping to teach pigeons to play ping-pong; trainers use it to teach dolphins to jump through hoops; speech therapists use it to build words from sounds. [How shaping works ›](https://operantconditioning.com/shaping/)

### Stimulus control and the antecedent

A behavior reinforced in one context and not in another comes under **[stimulus control](https://operantconditioning.com/glossary/#stimulus-control)**: it appears when the "discriminative stimulus" (S^D) is present and not otherwise. The rat presses only when the light is on. You swear with friends and not with your grandmother. This is the "A" in A-B-C, and it is the most under-used lever in self-improvement — changing the cue is often easier than willing a new response. [Stimulus control in depth ›](https://operantconditioning.com/stimulus-control/) · [The ABC model ›](https://operantconditioning.com/abc-model/)

## Operant vs. classical conditioning

The two great forms of associative learning are often confused. The clean distinction is *what gets associated with what*:

| Aspect | Classical (Pavlovian) conditioning | Operant (instrumental) conditioning |
| --- | --- | --- |
| **Association** | Stimulus ↔ stimulus (bell → food) | Behavior ↔ consequence (press → food) |
| **Behavior type** | Involuntary, reflexive (salivation, fear, nausea) | Voluntary, "emitted" (pressing, speaking, working) |
| **Organism's role** | Passive; the stimulus is presented regardless | Active; the consequence depends on what it does |
| **Timing of key event** | Stimulus comes *before* the response | Consequence comes *after* the response |
| **Founders** | Ivan Pavlov (1890s–1927) | Edward Thorndike (1898); B. F. Skinner (1937–38) |

In real life the two run together. The sound of the treat bag classically conditions excitement in a dog *and* operantly reinforces running to the kitchen. [Full comparison with examples ›](https://operantconditioning.com/operant-vs-classical-conditioning/)

## Can you spot it?

Eight scenarios, instant explanations, no score kept. When you can call all eight without hesitating, you have it — and there is a [longer set of twenty](https://operantconditioning.com/quiz/).

## What happens in the brain

Reinforcement has a physical address. In 1953 James Olds and Peter Milner found that a rat would press a lever thousands of times an hour for a pulse of electricity to its own brain, and in 1997 Wolfram Schultz and colleagues showed what the relevant neurons are doing: midbrain **dopamine** cells fire when a reinforcer is *better than expected*, fall silent when it is exactly as expected, and dip when an expected one fails to arrive.[21][16] That **reward prediction error** is the brain's teaching signal, and it explains why unpredictable reinforcers hold behavior so well and why a fully predictable one stops teaching. It is not a pleasure signal: Kent Berridge and Terry Robinson showed that animals without dopamine still *like* sugar but no longer *want* it.[22] [The neuroscience of operant conditioning, in depth ›](https://operantconditioning.com/neuroscience/)

## A brief history

- **1898** — Thorndike's puzzle boxes Edward Thorndike times cats escaping from latched boxes. Escapes get faster, and he proposes the **[law of effect](https://operantconditioning.com/glossary/#law-of-effect)**: responses followed by satisfaction are "stamped in"; those followed by discomfort are "stamped out."[1]
- **1930–1938** — Skinner's operant chamber At Harvard, as a graduate student and then a junior fellow, Skinner builds the apparatus later nicknamed the "Skinner box" and invents the cumulative recorder; at Minnesota he coins "operant" (1937) and lays out the science of operant behavior in *The Behavior of Organisms* (1938).[12][2]
- **1948–1957** — Superstition, schedules, and Verbal Behavior The pigeon "superstition" study (1948), *Science and Human Behavior* (1953), *Schedules of Reinforcement* with Ferster (1957), and *Verbal Behavior* (1957) extend the analysis to society and language.
- **1968** — Applied behavior analysis is born Baer, Wolf, and Risley publish the founding paper of ABA in the first issue of the *Journal of Applied Behavior Analysis*.[15]
- **1997** — The dopamine connection Schultz, Dayan, and Montague show that midbrain dopamine neurons encode a *reward prediction error* — a biological implementation of the learning signal operant theory had assumed.[16]

[The full history of operant conditioning ›](https://operantconditioning.com/history/) · [B. F. Skinner: life, work, and the Skinner box ›](https://operantconditioning.com/bf-skinner/) · [The original books, full text, in the library ›](https://operantconditioning.com/library/)

## Examples of operant conditioning in everyday life

Once you know the pattern you see it everywhere:

- **Positive reinforcement:** A barista is thanked warmly for a latte-art heart and starts making them on every cup. A student's essay earns praise and she writes more. Your phone lights up with a like.
- **Negative reinforcement:** You clean the kitchen to end your partner's nagging. A student finishes homework early to escape the anxiety of a deadline. A driver takes the side street to avoid the traffic jam.
- **Positive punishment:** A dog gets sprayed with water for jumping on the couch. A late invoice draws a fee. You bite into a moldy strawberry.
- **Negative punishment:** A child loses screen time for hitting a sibling. A driver loses points from their license. A team member loses a project after missing deadlines.
- **Schedules in the wild:** Slot machines (variable ratio). A bakery loyalty card (fixed ratio). Watching the clock in the last minutes of a shift (fixed interval). Fishing (variable interval).

[50+ examples, sorted by quadrant and setting ›](https://operantconditioning.com/examples/)

## Operant conditioning in pop culture

Two scenes worth watching a second time. Both are embedded from YouTube and load only when you press play.

*Figure: **What to look for:** — every time Penny does something Sheldon likes, a chocolate appears — — [positive reinforcement](https://operantconditioning.com/positive-reinforcement/) — , delivered immediately, on a continuous schedule. When Leonard objects, Sheldon reaches for a spray bottle — that would be — [positive punishment](https://operantconditioning.com/positive-punishment/) — . Sheldon even names the procedure.*

*Figure: **What almost everyone gets wrong here:** — Jim calls it Pavlov, but is it? Dwight's hand reaching out is a voluntary behavior that has been reinforced with a mint whenever the chime sounds — the chime is working as a — [discriminative stimulus](https://operantconditioning.com/glossary/#discriminative-stimulus) — , which makes the reaching — *operant* — . The dry mouth he notices is the — [classical](https://operantconditioning.com/operant-vs-classical-conditioning/) — part. Most real learning is both at once.*

Clips are uploaded by third parties and may disappear; NBC hosts [the official Office clip](https://www.nbc.com/the-office/video/jims-pavlovian-prank-on-dwight-the-office/4141507).

## Key terms to know

The twelve terms that do most of the work, defined the way behavior analysts define them. Hover or tap any underlined term anywhere on this page for its definition; the [full glossary](https://operantconditioning.com/glossary/) has 128 entries.

- **Operant**: A class of behavior defined by its effect on the environment (what it accomplishes), not by its exact form. Lever-pressing with the left paw or the right paw is the same operant.
- **Reinforcer**: Any consequence that increases the future frequency of the behavior it follows. *Positive* reinforcers are added; *negative* reinforcers are removed.
- **Punisher**: Any consequence that decreases the future frequency of the behavior it follows.
- **Primary vs. secondary reinforcer**: Primary reinforcers work without learning (food, water, warmth). Secondary (conditioned) reinforcers acquire their power by being paired with primary ones — money, praise, grades, a clicker.
- **Three-term contingency**: Antecedent → Behavior → Consequence: the basic unit of analysis. [The ABC model ›](https://operantconditioning.com/abc-model/)
- **Discriminative stimulus (S^D)**: A cue that signals a behavior will be reinforced. When a behavior reliably occurs in its presence and not otherwise, the behavior is under **stimulus control**.
- **Schedule of reinforcement**: The rule for which responses get reinforced: continuous, fixed ratio, variable ratio, fixed interval, or variable interval. [Schedules ›](https://operantconditioning.com/schedules-of-reinforcement/)
- **Extinction**: Withholding the reinforcer that maintained a behavior, so the behavior declines — often after an **extinction burst**, a temporary increase. [Extinction ›](https://operantconditioning.com/extinction/)
- **Shaping**: Building a new behavior by reinforcing successive approximations of it. [Shaping ›](https://operantconditioning.com/shaping/)
- **Motivating operation**: A condition such as deprivation or satiation that changes how effective a reinforcer is. Food reinforces a hungry rat, not a full one.
- **Premack principle**: A more probable behavior can reinforce a less probable one — "finish your homework, then you can play." [Premack ›](https://operantconditioning.com/premack-principle/)
- **Law of effect**: Thorndike's 1898 principle that responses followed by satisfying consequences are strengthened and those followed by discomfort are weakened — the ancestor of operant conditioning.

*In practice*

## Where operant conditioning is used.

The same three-term contingency runs a therapy session, a classroom, a family dinner, a factory floor, and a habit tracker.

- [Applied behavior analysis](https://operantconditioning.com/applications/#aba): The clinical discipline built on operant principles, used in autism intervention, developmental disability support, and behavioral medicine.
- [Education](https://operantconditioning.com/classroom/): Token economies, positive behavioral supports, immediate feedback, and the teaching machines Skinner pioneered in the 1950s.
- [Parenting](https://operantconditioning.com/parenting/): Catch them being good, planned ignoring, time-out done correctly, and why consistency beats severity.
- [Animal training](https://operantconditioning.com/dog-training/): Clicker training, marker signals, and the reinforcement-based methods that replaced dominance theory.
- [Workplace & management](https://operantconditioning.com/applications/#workplace): Organizational behavior management, safety programs, and why annual reviews fail to change daily behavior.
- [Habits & self-management](https://operantconditioning.com/habits/): Using antecedents, tiny behaviors, and immediate consequences to build habits that stick — on yourself.

## Criticisms and limitations

An honest account includes what operant conditioning does *not* explain well.

- **Biological constraints.** Organisms are not blank slates. The Brelands found that raccoons trained to deposit coins would instead "wash" them — an instinctive food-handling behavior that intruded even though it delayed reinforcement, a drift toward species-typical behavior they called [instinctive drift](https://operantconditioning.com/glossary/#instinctive-drift).[14] Reinforcement works with an animal's evolved tendencies, not against them.
- **Cognition and language.** Chomsky argued that reinforcement cannot explain how children acquire grammar from limited input;[13] Tolman's rats learned mazes without obvious reinforcement, and Bandura's children learned by watching.[17] Operant learning is one powerful process among several, not a complete theory of mind.
- **Rewards and intrinsic motivation.** Expected, tangible rewards for an activity someone already enjoys can reduce their interest once the rewards stop — the overjustification effect. A large meta-analysis found it; a rival meta-analysis found it small and narrow.[18][19] The fair reading: praise and feedback rarely undermine motivation; paying people for things they already love sometimes does.
- **Ethics of control.** Skinner's *Beyond Freedom and Dignity* (1971) argued that since behavior is always controlled by its environment, we should design that environment deliberately. Critics saw a road to manipulation; the ethics codes of behavior analysis answer with consent, least-restrictive procedures, and the client's own goals.[10]

What survived every critique is the core: the law of effect, schedule effects, extinction bursts, stimulus control, and shaping replicate across species and remain the working toolkit of clinicians, teachers, and trainers. Reinforcement learning — the branch of AI behind game-playing systems and the tuning of language models — is a mathematical descendant of the same ideas.[20]

## How to use operant conditioning on yourself

The same contingency that trains a pigeon can be turned inward, and most self-improvement advice ignores two-thirds of it: it obsesses over motivation, which is not a term in the equation, and neglects antecedents and consequences, which are. The protocol is short. Attach the new behavior to a cue that already happens every day; shrink the behavior until it is almost embarrassing; deliver a small reinforcer within seconds, every time at first; then thin the schedule so the habit survives missed days. [The full seven-step protocol, with the evidence on how long habits take ›](https://operantconditioning.com/habits/)

### Design your own contingency

Fill in the three terms for a habit you actually want. The tool checks each against the science — is the cue stable, is the behavior really a behavior, is the consequence immediate — and gives you a card to print. Nothing you type leaves your browser. [Full guide to building habits ›](https://operantconditioning.com/habits/)

> **The app: Operant runs this loop for you.** A habit app for iPhone and Apple Watch from the publisher of this site. Each habit is set up as an antecedent, a behavior and a consequence — the three terms, not just the middle one. Free to download and try; a subscription unlocks the full app. [About the app](https://operantconditioning.com/app/) · [Download on the App Store](https://apps.apple.com/us/app/operant-behavior-change-app/id6802081776)

## Before you leave: can you answer these without opening them?

**What is operant conditioning in simple terms?**

Operant conditioning is learning from consequences. When a behavior is followed by something good (or the removal of something bad), it happens more often. When it is followed by something bad (or the loss of something good), it happens less often. The organism "operates" on its environment and the results shape what it does next.

**Who discovered operant conditioning?**

The underlying principle — the law of effect — was discovered by Edward Thorndike in 1898 through his puzzle-box experiments with cats. B. F. Skinner named it "operant" conditioning in 1937, developed the experimental methods to study it, and built the field around it beginning with *The Behavior of Organisms* in 1938.

**What are the four types of operant conditioning?**

Positive reinforcement (add something, behavior increases), negative reinforcement (remove something, behavior increases), positive punishment (add something, behavior decreases), and negative punishment (remove something, behavior decreases). "Positive" and "negative" refer to adding and removing a stimulus, not to whether the outcome is good or bad.

**What is Skinner's theory of operant conditioning?**

Skinner's theory is that behavior is selected by its consequences, much as species are selected by their environments. A behavior that is followed by reinforcement becomes more frequent; one followed by punishment or by no reinforcement at all becomes less frequent. He distinguished this *operant* behavior, which acts on the environment, from *respondent* behavior, the reflexes studied by Pavlov, and he showed that the schedule on which reinforcement arrives controls how fast and how persistently an organism responds. The theory deliberately explains behavior by its history of consequences rather than by inner states such as wants or intentions, which Skinner treated as behavior to be explained rather than as causes. [B. F. Skinner: the theory, the experiments, and the critiques ›](https://operantconditioning.com/bf-skinner/)

**Why is it called "operant" conditioning?**

Skinner chose the word in 1937 because the behavior *operates* on the environment to produce a consequence: the rat's press operates the lever, the lever delivers food, and the food changes future pressing. He contrasted operant behavior with respondent behavior, which is elicited by a stimulus that comes before it, as a puff of air elicits a blink. The older name, *instrumental* conditioning, makes the same point from the other side: the behavior is instrumental in producing the outcome.

**What are the three components of operant conditioning?**

The antecedent, the behavior, and the consequence, usually written A-B-C and called the three-term contingency. The antecedent is the situation or cue that sets the occasion for the behavior; the behavior is what the organism does; the consequence is what follows, and it is the consequence that changes how likely the behavior is next time. Every example on this site can be broken into those three parts. [The ABC model in depth ›](https://operantconditioning.com/abc-model/)

**Is negative reinforcement the same as punishment?**

No — this is the most common mistake in the whole subject. Negative reinforcement *increases* a behavior by removing something unpleasant (taking a painkiller to end a headache). Punishment *decreases* a behavior. "Negative" only means something was taken away. The two forms of negative reinforcement, escape and avoidance, have [their own page](https://operantconditioning.com/avoidance-learning/).

**Which schedule of reinforcement is most resistant to extinction?**

The variable-ratio schedule, in which reinforcement follows an unpredictable number of responses. Because the organism can never tell whether the next response will pay off, responding persists long after reinforcement has stopped. This is why gambling and social-media checking are hard to extinguish.

## References

1. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. *Psychological Review Monograph Supplement, 2*(4), 1–109. See also Thorndike, E. L. (1911). *Animal Intelligence: Experimental Studies*. Macmillan. Read the 1898 monograph as Chapter II of the 1911 book in the library: https://operantconditioning.com/library/thorndike-animal-intelligence/#chapter-ii-animal-intelligence-an-experimental-study-of-the-associative-processes-in-animals
2. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century.
3. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan.
4. Skinner, B. F. (1981). Selection by consequences. *Science, 213*(4507), 501–504.
5. Grice, G. R. (1948). The relative effects of delay of reinforcement in the discrimination learning of rats. *Journal of Experimental Psychology, 38*(1), 1–16. See also Lattal, K. A. (2010). Delayed reinforcement of operant behavior. *Journal of the Experimental Analysis of Behavior, 93*(1), 129–139.
6. Skinner, B. F. (1948). 'Superstition' in the pigeon. *Journal of Experimental Psychology, 38*(2), 168–172.
7. Estes, W. K. (1944). An experimental study of punishment. *Psychological Monographs, 57*(3), i–40.
8. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), *Operant Behavior: Areas of Research and Application* (pp. 380–447). Appleton-Century-Crofts.
9. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. *Journal of Family Psychology, 30*(4), 453–469.
10. Behavior Analyst Certification Board. (2020). *Ethics Code for Behavior Analysts*. BACB. See also Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson.
11. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts.
12. Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. *Journal of General Psychology, 16*, 272–279.
13. Chomsky, N. (1959). A review of B. F. Skinner's *Verbal Behavior*. *Language, 35*(1), 26–58.
14. Breland, K., & Breland, M. (1961). The misbehavior of organisms. *American Psychologist, 16*(11), 681–684.
15. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. *Journal of Applied Behavior Analysis, 1*(1), 91–97.
16. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. *Science, 275*(5306), 1593–1599.
17. Tolman, E. C. (1948). Cognitive maps in rats and men. *Psychological Review, 55*(4), 189–208; Bandura, A. (1977). *Social Learning Theory*. Prentice-Hall.
18. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. *Psychological Bulletin, 125*(6), 627–668.
19. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. *Review of Educational Research, 64*(3), 363–423.
20. Sutton, R. S., & Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press.
21. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. *Journal of Comparative and Physiological Psychology, 47*(6), 419–427; Olds, J. (1958). Self-stimulation of the brain. *Science, 127*(3294), 315–324.
22. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? *Brain Research Reviews, 28*(3), 309–369.

*Keep going*

## Go deeper.

- [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The most-used tool in the kit — and the most misunderstood word in it.
- [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): Fixed, variable, ratio, interval — with a live cumulative-record simulator.
- [B. F. Skinner](https://operantconditioning.com/bf-skinner/): The man, the box, the pigeons, the controversies.
- [50+ examples](https://operantconditioning.com/examples/): Everyday, classroom, workplace, and animal examples for every quadrant.
- [Glossary](https://operantconditioning.com/glossary/): Every term from "abolishing operation" to "variable interval," defined.
- [Can you spot it?](https://operantconditioning.com/quiz/): A 20-question quiz with instant explanations. Can you tell the quadrants apart?
- [The library](https://operantconditioning.com/library/): Thorndike, Morgan, James, Yerkes, Darwin: the public-domain sources in full text, for readers and for AI.
- [The Operant app](https://operantconditioning.com/app/): A habit app for iPhone and Apple Watch that runs the A-B-C loop for each habit.

## About the author

Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards
