Study guide · AP Psychology
Operant Conditioning for AP Psychology: Terms, Studies, and Practice
Learning is one of the most reliably tested ideas in AP Psychology, and operant conditioning is the part students most often get backwards under time pressure. This is the exam-facing version of the site: the terms in one line each, the studies by name and year, the confusions the test writers count on, a method for scenario questions, and a practice set with an answer key.
In brief
- In the revised AP Psychology course, effective fall 2024, learning sits within Unit 3, Development and Learning; the Course and Exam Description lists the terms you are responsible for.[1]
- Every operant scenario is two questions: was something added or removed, and did the behavior go up or down. Positive means added, negative means removed, and negative reinforcement is not punishment.
- Schedules are two more questions — count or time, fixed or variable — and each has a signature pattern the exam expects you to recognize.
Where operant conditioning sits on the AP exam
The College Board revised AP Psychology for the 2024–25 school year. In the revised Course and Exam Description (CED), learning is no longer a unit of its own: classical conditioning, operant conditioning, and social and cognitive factors in learning sit within Unit 3, "Development and Learning."[1] The exam has a multiple-choice section and a free-response section; in the revised course the free-response questions ask you to analyze a description of a study and to argue from several sources of evidence rather than to define terms.[1] Question counts, timing, and weighting are in the CED and can change between editions; check the current one rather than any summary, including this one.
You will rarely be asked to define negative reinforcement. You will be given a scenario — a parent, a rat, a paragraph from a study — and asked what kind of learning it shows or whether the design supports the conclusion. The CED is free and lists every term the exam can draw on; read it before you study and again after. Anything not in it is not required, and the last section here sorts out which pages on this site go past it.[1]
The vocabulary, one line each
Definitions are functional, as the exam uses them: a reinforcer or punisher is defined by its effect on behavior, not by intent or feel. Each links to the page that treats it fully.
Consequences and the four quadrants
- Law of effect. Responses followed by satisfaction are strengthened, responses followed by discomfort weakened; Thorndike's wording is in the library.[2][3]
- Operant conditioning. A voluntary behavior changes in frequency because of its consequences; Skinner's term.[4]
- Reinforcement. Any consequence that makes the behavior it follows more frequent.
- Punishment. Any consequence that makes the behavior it follows less frequent.
- Positive and negative. Added and removed. Not good and bad.
- Positive reinforcement. Stimulus added, behavior up: a treat after a sit.
- Negative reinforcement. Stimulus removed, behavior up: buckling up silences the chime.
- Positive punishment. Stimulus added, behavior down: a scolding after jumping on the couch.
- Negative punishment. Stimulus removed, behavior down: losing the car keys after missing curfew.
- Extinction. The reinforcer stops following the behavior and the behavior declines. Not a quadrant: nothing is added or removed.
- Extinction burst. The brief rise in rate or intensity when reinforcement is first withheld.
- Spontaneous recovery. An extinguished response returns after a rest.
- Generalization. Responding to stimuli that resemble the training stimulus.
- Discrimination. Responding to the training stimulus but not to similar ones; the signal that reinforcement is available is a discriminative stimulus.
Reinforcers
- Primary reinforcer. Reinforcing without learning: food, water, warmth.
- Secondary (conditioned) reinforcer. Reinforcing through association with primary reinforcers: money, grades, praise.
- Token economy. Tokens earned for target behaviors are exchanged later for backup reinforcers.
- Premack principle. A more probable behavior reinforces a less probable one when made contingent on it: first homework, then the game.[5]
Building and maintaining behavior
- Shaping. Reinforcing successive approximations to a target behavior.
- Chaining. Linking behaviors so each step cues the next and the last produces the reinforcer.
- Skinner box. An operant chamber with a lever or key, a reinforcer dispenser, and a cumulative recorder.[4]
- Continuous reinforcement. Every response reinforced: fastest acquisition, fastest extinction.
- Partial (intermittent) reinforcement. Some responses reinforced: slower acquisition, far greater resistance to extinction. AP materials say partial; behavior analysts say intermittent.
- Fixed ratio (FR). Reinforcement after a set number of responses. Pattern: a pause after each reinforcer, then a fast run.[6]
- Variable ratio (VR). After an unpredictable number of responses. Pattern: high and steady, hardest to extinguish.[6]
- Fixed interval (FI). The first response after a set time. Pattern: the scallop, accelerating as the interval ends.[6]
- Variable interval (VI). The first response after an unpredictable time. Pattern: moderate and steady.[6]
Named phenomena
- Superstitious behavior. Strengthened by accidental, non-contingent reinforcement; Skinner's pigeons, fed every fifteen seconds whatever they did, developed stereotyped turning.[7]
- Learned helplessness. After uncontrollable aversive events, an organism fails to escape when it can; Seligman and Maier's dogs.[8]
- Instinctive drift. Trained behavior drifts toward species-typical behavior; the Brelands' raccoons "washed" the coins they were trained to deposit.[9]
- Biological preparedness. Some associations are learned easily and others hardly at all; rats link taste with illness and light or sound with shock, not the reverse.[10][11]
Contrast cases: cognitive and social learning
- Latent learning. Learning without reinforcement that appears when there is reason to use it; Tolman and Honzik's rats had a cognitive map before food was offered.[12]
- Insight learning. A sudden, complete solution rather than gradual trial and error; Köhler's chimpanzees.[13]
- Observational learning (modeling). Learning by watching a model; Bandura's Bobo doll. Vicarious reinforcement and punishment are consequences the observer sees the model receive.[14]
Operant vs. classical conditioning, as the exam frames it
The exam treats these as the two kinds of associative learning and tests whether you can tell them apart. In classical conditioning two stimuli are paired and a reflex transfers from one to the other: Pavlov's dogs salivated to the signal that preceded food, and Watson and Rayner's infant cried at a white rat that had been paired with a loud noise.[15][16] In operant conditioning a voluntary behavior is followed by a consequence and changes in frequency. Classical responses are elicited by what comes before; operant responses are emitted and changed by what comes after.
| Aspect | Classical (Pavlovian) | Operant (instrumental) |
|---|---|---|
| What is learned | One stimulus predicts another (CS predicts US) | A behavior produces a consequence |
| Kind of response | Reflexive, involuntary: salivation, fear, nausea | Voluntary: pressing, studying, sitting |
| Order of events | Key stimulus comes before the response | Key event comes after the response |
| Outcome depends on behavior? | No; food arrives whether or not the dog salivates | Yes; food arrives only if the rat presses |
Extinction, spontaneous recovery, generalization, and discrimination occur in both, so a question can ask about any of them in either frame. Most real scenes contain both — a dog excited at the treat pouch, then sitting for the treat. The operant vs. classical conditioning page has the decision checklist and eight worked examples.
The four confusions the test writers count on
Most wrong answers come from one of four confusions, and each has a distractor written for it.
1. Negative reinforcement is not punishment
Reinforcement or punishment says what happened to the behavior. Positive or negative says what happened to the environment. Negative reinforcement is "behavior up, something removed." Nagging that stops when the room is clean negatively reinforces cleaning; nagging that starts at the sight of a mess positively punishes messiness. The same stimulus does opposite jobs.
| Change | Behavior increases | Behavior decreases |
|---|---|---|
| Stimulus added | Positive reinforcement | Positive punishment |
| Stimulus removed | Negative reinforcement | Negative punishment |
2. Positive means added, even when what is added is unpleasant
"Positive punishment" is a contradiction only if positive means good. It means added: a shock, a scolding, an extra chore. And a consequence is a punisher only if the behavior decreases; a detention that does not reduce talking is not punishment, whatever it was called. Items exploit this with a "punishment" followed by more of the behavior.
3. Ratio vs. interval, and what "variable" does not tell you
Students learn that variable schedules resist extinction and then call anything unpredictable "variable." The first question is ratio or interval. If reinforcement depends on how many responses occur, it is ratio and responding faster pays sooner; if it depends on time, it is interval and responding faster changes nothing. A slot machine paying after an unpredictable number of pulls is variable ratio; a pool where fish bite after unpredictable stretches of time is variable interval. Only the gambler's rate matters, which is why the gambler pulls fast and the angler casts at a moderate pace.[6]
4. Extinction is not punishment
Both reduce behavior. In punishment something happens after the response: an aversive is added or a reinforcer removed. In extinction nothing happens; the usual reinforcer simply stops. Tantrums once maintained by attention and now ignored are on extinction; a toy removed after each tantrum is negative punishment. Extinction also has two signatures punishment lacks: a burst at the start and spontaneous recovery after a rest.
The studies you should be able to name
Questions name the researcher and expect the term, or describe the procedure and expect the researcher.
| Study | Procedure | Finding | Term |
|---|---|---|---|
| Thorndike (1898) | Cats worked a latch to escape a puzzle box for food. | Escape times fell gradually, with no sudden insight.[2] | Law of effect |
| Watson and Rayner (1920) | A loud noise paired with a white rat for an infant. | Fear of the rat alone, generalizing to a rabbit, a dog, a fur coat.[16] | Conditioned fear; generalization |
| Köhler (1925) | Chimpanzees, bananas out of reach, boxes and sticks. | After no progress, sudden and complete solutions.[13] | Insight learning |
| Pavlov (1927) | A signal repeatedly preceded food for dogs. | Salivation to the signal; extinction, recovery, generalization, discrimination described.[15] | Classical conditioning |
| Tolman and Honzik (1930) | Rats ran a maze daily; one group was unfed until day eleven. | Errors then dropped abruptly to the rewarded group's level.[12] | Latent learning; cognitive map |
| Skinner (1938) | Rats pressed a lever for food in an operant chamber. | Rate of response as the measure; reinforcement defined by its effect.[4] | Operant conditioning; Skinner box |
| Skinner (1948) | Pigeons fed every fifteen seconds regardless of behavior. | Six of eight developed stereotyped turning and head-tossing.[7] | Superstitious behavior |
| Ferster and Skinner (1957) | Pigeons and rats on many schedules, recorded cumulatively. | Each schedule produces a characteristic pattern.[6] | Schedules of reinforcement |
| Premack (1959) | Children's candy eating and pinball made contingent on each other. | The preferred activity reinforced the other, not the reverse.[5] | Premack principle |
| Breland and Breland (1961) | Raccoons and pigs trained with food to deposit coins. | Drift toward washing and rooting, though it delayed food.[9] | Instinctive drift |
| Bandura, Ross, and Ross (1961) | Children watched an adult attack a Bobo doll, a calm adult, or no model. | Those who saw aggression reproduced its specific acts, unreinforced.[14] | Observational learning |
| Garcia and Koelling (1966) | Rats drank sweet, "bright, noisy" water, then were made ill or shocked. | Illness attached to the taste; shock to the light and sound. Seligman (1970) named the general principle preparedness.[10][11] | Taste aversion; preparedness |
| Seligman and Maier (1967) | Dogs given escapable, inescapable, or no shock, then tested where escape was possible. | Dogs that had had no control mostly failed to escape.[8] | Learned helplessness |
The Bobo doll, latent learning, and insight rows are contrast cases: learning occurred without reinforcement, though reinforcement governs whether it is performed.
How to work a scenario question
Whether the scenario is a multiple-choice stem or a paragraph from a study, the same six moves name the process.
- Find the behavior. What the organism does, as a verb. If two people are in the scene, decide whose behavior is being asked about.
- Find the consequence. What happened immediately after, for the one behaving.
- Added or removed? Added is positive; removed is negative.
- Up or down? "More likely," "keeps doing it," and "continues" mean up; "stops" and "less often" mean down.
- Name the quadrant. Added and up: positive reinforcement. Removed and up: negative reinforcement. Added and down: positive punishment. Removed and down: negative punishment.
- Check that it is a quadrant at all. If the usual reinforcer simply stopped, it is extinction. If a reflex was elicited by a signal, it is classical. If the learner watched someone else, it is observational learning.
A worked case. Maya's brother screams whenever she takes the tablet; she gives it back to stop the noise, and he screams sooner next time. For the brother, the tablet is returned (added) and screaming goes up: positive reinforcement. For Maya, the screaming stops (removed) and handing it back becomes more likely: negative reinforcement. The A-B-C model page works this decomposition in detail.
Naming the schedule from the pattern
Two questions. Does reinforcement depend on a count of responses ("every tenth," "an unpredictable number of") or on time ("after a set time," "at varying intervals")? Count is ratio, time is interval. Is the requirement the same each time or does it vary? Then confirm with the pattern — pause-and-run, high and steady, the scallop, moderate and steady — because a question that describes pacing is handing you the answer. The schedules of reinforcement page has a simulator that draws each cumulative record live; once you have watched the scallop form you will not confuse it with the ratio pause.
When free-response material describes a study, read the method a sentence at a time and name each procedure as you meet it — the reinforcer, the schedule, the dependent variable — because a question about whether the conclusion follows usually turns on a procedure named correctly.
Practice set: ten items with answers
Work each item with the six moves before opening the answer. The examples page has fifty more, and the quiz gives instant explanations.
A student gets a sticker every time she turns in homework, and submissions increase. Quadrant and schedule?
Answer
Positive reinforcement (sticker added, behavior up) on continuous reinforcement, so expect fast extinction if the stickers stop.
A driver buckles up to silence the chime and, over time, buckles up faster. Quadrant?
Answer
Negative reinforcement: the chime is removed and the behavior increases. Not punishment, because nothing decreased.
A teenager comes home after curfew and loses driving privileges for a week; late arrivals become less frequent. Quadrant?
Answer
Negative punishment: a reinforcer is removed and the behavior decreases. Had late arrivals continued, it would not have been punishment, whatever the parents called it.
A rat gets a pellet after every tenth press; its cumulative record shows a flat stretch after each pellet, then a steep run. Schedule and pattern?
Answer
Fixed ratio 10. The flat stretch is the post-reinforcement pause and the steep run the high-rate burst: break-and-run.
A gambler keeps pulling a machine that pays after an unpredictable number of pulls; an angler keeps casting where fish strike after unpredictable stretches of time. Name each schedule and the phrase that decides it.
Answer
Variable ratio: "number of pulls" makes reinforcement depend on a count. Variable interval: "stretches of time" makes it depend on time. Both are unpredictable, so "variable" alone does not decide it.
A child's tantrums were maintained by attention. The parents now withhold it; tantrums briefly get louder, then decline. Name the process and the brief increase.
Answer
Extinction, not punishment: the reinforcer stopped, and nothing is added or removed after each tantrum. The increase is the extinction burst; removing a toy after each tantrum would be negative punishment.
A pigeon fed every fifteen seconds regardless of what it does is soon turning in circles between feedings. Phenomenon and researcher?
Answer
Superstitious behavior, Skinner (1948). Whatever the bird was doing when food arrived was accidentally reinforced, and the accident repeated.
Rats run a maze daily with no food and seem to wander. On day eleven food appears and their errors drop at once to the level of rats fed all along. Phenomenon and researchers?
Answer
Latent learning, Tolman and Honzik (1930). The rats had a cognitive map from the unrewarded days; reinforcement gave them a reason to show it.
Raccoons trained with food to drop coins in a container increasingly rub and dip the coins instead, delaying the food. Phenomenon and researchers?
Answer
Instinctive drift, Breland and Breland (1961). The trained response drifted toward innate food-washing at the cost of reinforcement: consequences shape behavior only within a species' evolved repertoire.
A dog wags and drools when it hears the treat pouch open, then sits on cue and gets the treat. Which parts are classical and which operant?
Answer
Both. Wagging and drooling at the pouch sound are a conditioned response to a stimulus that predicts food: classical. The sit is a voluntary behavior followed by a treat that depends on it: operant, positive reinforcement.
A study plan using this site
The teachers page gives a two-session route through the front page for a first pass. This is the exam-oriented version: quadrants and schedules first, a term check last. Spread it over a week.
- The front page, first half. The one-paragraph version, the A-B-C, and the four quadrants with the interactive checker.
- Negative reinforcement. The negative reinforcement page clears up the confusion the test writers use most.
- Schedules, with the simulator. The schedules page; run all four until you can sketch each record. If one is shaky, its own page (FR, VR, FI, VI) has more examples.
- Extinction and shaping. The extinction page for the burst, spontaneous recovery, and the partial reinforcement effect; the shaping page; then ten minutes in the lab shaping a lever press.
- Operant vs. classical. The comparison page and its checklist. The page to reread the night before.
- Names and dates. The history page, with the studies table above.
- Practice. The examples page, the quiz, then the set above. For each wrong answer, find which of the four confusions produced it.
- Term check. Go down the learning terms in the CED and give a one-line definition of each; the glossary has every one.[1]
Common mistakes, and what you can skip
Common mistakes
- Classifying by how the consequence felt. A punisher is a consequence that reduced behavior; a reinforcer is one that increased it. Nothing else counts.
- Naming the schedule from the reinforcer. Money, praise, and food can be delivered on any schedule. Ask count or time, fixed or variable.
- Calling extinction forgetting, or punishment. Extinction is new learning that the response no longer pays, which is why it returns after a rest.
- Confusing a signal with a consequence. A stimulus before a reflex is a conditioned stimulus; a stimulus before a voluntary behavior that signals it will pay is a discriminative stimulus; a stimulus after the behavior is a consequence.
- Treating the contrast cases as refutations. Latent learning, insight, and observational learning show that learning also occurs without direct reinforcement, not that reinforcement fails to change behavior.
What the CED does not require
This site goes further than the AP course. Unless the current CED says otherwise, you do not need the following: the matching law; behavioral momentum; the response deprivation hypothesis behind the Premack principle; motivating operations; functional analysis; the equations of the Rescorla–Wagner model; autoshaping, sign-tracking, and contrafreeloading; and the debate over whether the positive/negative distinction should be kept at all.[1] Reading them makes the required material easier to hold, and the neuroscience page connects reinforcement to the dopamine system covered elsewhere in the course, but none of it is where the points are.
Key takeaways
- In the revised AP Psychology course, learning is taught within Unit 3, Development and Learning, and the Course and Exam Description is the authority on which terms and question formats are tested.
- Every operant scenario is two questions — added or removed, and behavior up or down — and the answers name the quadrant; "positive" and "negative" describe the environment, not the outcome.
- Negative reinforcement is not punishment, and extinction is not a quadrant: punishment adds an aversive or removes a reinforcer after the response; in extinction the reinforcer simply stops arriving.
- Schedules are named by count-or-time and fixed-or-variable: pause-and-run on fixed ratio, a high steady rate on variable ratio, the scallop on fixed interval, a moderate steady rate on variable interval.
- Learn the classic studies as procedures with findings, because the exam describes what was done and asks what it showed; latent learning, insight, and the Bobo doll are the contrast cases, and instinctive drift and preparedness mark the biological limits.
Check yourself
A teacher says she "negatively reinforced" a student by assigning detention, and the student's talking in class decreased. What did she actually do?
Positive punishment. Something was added (detention) and the behavior went down, so the consequence was a punisher, and because it was added rather than removed it was positive. Negative reinforcement would require the behavior to increase because something was taken away.
One rat gets a pellet after every twentieth press; another gets a pellet for the first press after twenty seconds have passed. Sketch the cumulative record you expect from each.
The first rat is on a fixed ratio 20 and should show break-and-run: a pause after each pellet, then a steep run of twenty presses. The second is on a fixed interval 20 seconds and should show a scallop: few presses just after each pellet, then an accelerating rate as the twenty seconds run out. The ratio schedule produces the higher overall rate, because responding faster brings the pellet sooner; on the interval schedule it does not.
Children who watched an adult hit a Bobo doll later hit it themselves. A classmate calls this operant conditioning because "aggression was reinforced." Is that right?
No. In the 1961 study the children were not reinforced for imitating; they reproduced the specific acts they had watched. That is observational learning, and it is on the exam precisely as the contrast with learning through one's own consequences. Operant conditioning would require the children's own hitting to have been followed by a consequence that changed its frequency.
Want more? The 20-question quiz covers every page on the site with instant explanations.
Frequently asked questions
What is operant conditioning in simple terms for AP Psychology?
Operant conditioning is learning in which a voluntary behavior becomes more or less frequent because of the consequence that follows it. Reinforcement makes the behavior more frequent and punishment makes it less frequent; positive means a stimulus was added and negative means one was removed. Skinner named it and studied it with rats and pigeons in an operant chamber.
Is learning still tested on the AP Psychology exam after the 2024 revision?
Yes. In the revised course, classical conditioning, operant conditioning, and the social and cognitive factors in learning are taught within Unit 3, Development and Learning, rather than as a unit of their own. The Course and Exam Description lists the required terms and describes the question formats; check the current edition for question counts and weighting.
What is the difference between negative reinforcement and punishment on the AP exam?
Negative reinforcement increases a behavior by removing an unpleasant stimulus, as when buckling a seat belt stops the chime. Punishment decreases a behavior, either by adding an unpleasant stimulus (positive punishment) or by removing a pleasant one (negative punishment). The test is always the effect on behavior: if it went up, it was reinforcement, whatever it felt like.
What is an example of a variable ratio schedule for AP Psychology?
A slot machine. It pays after an unpredictable number of pulls, so reinforcement depends on how many responses are made, and the player responds at a high, steady rate that is very hard to extinguish. Fishing, by contrast, is usually variable interval, because a bite depends on unpredictable stretches of time rather than on the number of casts.
What is the difference between fixed interval and variable interval schedules?
In both, reinforcement depends on time rather than on the number of responses. On a fixed interval the wait is the same every time and behavior shows a scallop, with little responding early and a rising rate as the interval ends. On a variable interval the wait is unpredictable and behavior settles to a moderate, steady rate, because there is no point in the interval at which reinforcement is especially likely.
Which studies do I need to know for operant conditioning on the AP exam?
Thorndike's puzzle boxes and the law of effect (1898); Skinner's operant chamber (1938) and superstitious pigeons (1948); Ferster and Skinner on schedules (1957); Breland and Breland on instinctive drift (1961); Garcia and Koelling on taste aversion (1966); Seligman and Maier on learned helplessness (1967); and, as contrast cases, Tolman and Honzik on latent learning (1930), Köhler on insight (1925), and Bandura's Bobo doll (1961). Learn each as a procedure and a finding.
Is the Bobo doll experiment operant conditioning?
No. In Bandura, Ross, and Ross (1961), children imitated an adult's aggressive acts toward an inflatable doll without ever being reinforced for doing so. That is observational learning, and it appears in the learning topics as a contrast with conditioning through one's own consequences. Later work by Bandura added consequences to the model to study vicarious reinforcement and punishment.
How should I answer an operant conditioning scenario question?
Find the behavior and decide whose it is, find what happened right after it, ask whether something was added or removed, ask whether the behavior went up or down, and combine the two answers to name the quadrant. Then check that it is a quadrant at all: if the reinforcer simply stopped, it is extinction; if a reflex was elicited by a signal, it is classical conditioning; if the learner watched someone else, it is observational learning.
References
- College Board. (2024). AP Psychology Course and Exam Description (effective Fall 2024). College Board.
- Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. Psychological Review Monograph Supplement, 2(4), 1–109.
- Thorndike, E. L. (1911). Animal Intelligence: Experimental Studies. Macmillan.
- Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
- Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. Psychological Review, 66(4), 219–233.
- Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
- Skinner, B. F. (1948). 'Superstition' in the pigeon. Journal of Experimental Psychology, 38(2), 168–172.
- Seligman, M. E. P., & Maier, S. F. (1967). Failure to escape traumatic shock. Journal of Experimental Psychology, 74(1), 1–9.
- Breland, K., & Breland, M. (1961). The misbehavior of organisms. American Psychologist, 16(11), 681–684.
- Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. Psychonomic Science, 4(1), 123–124.
- Seligman, M. E. P. (1970). On the generality of the laws of learning. Psychological Review, 77(5), 406–418.
- Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. University of California Publications in Psychology, 4, 257–275.
- Köhler, W. (1925). The Mentality of Apes. Harcourt, Brace.
- Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575–582.
- Pavlov, I. P. (1927). Conditioned Reflexes: An Investigation of the Physiological Activity of the Cerebral Cortex (G. V. Anrep, Trans.). Oxford University Press.
- Watson, J. B., & Rayner, R. (1920). Conditioned emotional reactions. Journal of Experimental Psychology, 3(1), 1–14.