Schedules · Fundamentals

Continuous vs. Intermittent Reinforcement: Why Less Reinforcement Produces More Persistence

Reinforce every response and a behavior is learned quickly and abandoned quickly. Reinforce only some and it is learned slowly and kept for a very long time. The choice between the two is the most practical decision in operant conditioning, and it hides one of the field's oldest puzzles.

Updated 17 min read

Definition

Continuous reinforcement (CRF) is a schedule on which every occurrence of a behavior is followed by the reinforcer. Intermittent reinforcement, also called partial reinforcement, is any schedule on which only some occurrences are. Ferster and Skinner treated continuous reinforcement as the baseline from which every other schedule departs.[1]

The two schedules do different jobs. Continuous reinforcement establishes a new behavior fastest, because every response contacts its consequence. Intermittent reinforcement makes an established behavior more durable: it delays satiation, it costs fewer reinforcers, and above all it persists far longer when reinforcement stops — the partial reinforcement extinction effect.[2]

In brief

  • On a continuous schedule every response is reinforced; on an intermittent schedule only some are. Continuous reinforcement teaches a new behavior fastest, and intermittent reinforcement keeps an established one going.
  • Behavior built on intermittent reinforcement persists far longer once reinforcement stops, even though it received fewer reinforcers: the partial reinforcement extinction effect, first demonstrated by Humphreys in 1939.
  • The working rule is continuous first, then thin in small steps toward a variable schedule. Thin too fast and responding breaks down (ratio strain); never thin and the behavior collapses the first time the reinforcer fails to arrive.

How continuous and intermittent reinforcement differ

Every schedule answers one question: which responses will be reinforced? On continuous reinforcement the answer is all of them (FR 1 is the same thing under another name). A rat whose every lever press produces a pellet, a toddler whose every "please" produces the cookie, a light switch that works every time — each is on CRF, and the learner cannot fail to notice the relation.[1]

On an intermittent schedule some responses go unreinforced by design, under a rule that counts responses (fixed and variable ratio) or clock time (fixed and variable interval); the schedules page covers those four patterns. Skinner argued that almost all reinforcement outside the laboratory is intermittent — the angler does not catch a fish on every cast — whereas continuous reinforcement is mostly something people arrange on purpose: a trainer with a pouch of treats, a vending machine.[3]

Two schedules, two jobs

During acquisition — while a behavior is new, weak, or being shaped — reinforce every occurrence; an unreinforced response early on is a lost opportunity. During maintenance — once the behavior is reliable — move to an intermittent schedule, which holds the behavior with fewer reinforcers and makes it robust when reinforcement is occasionally missed.[4] Ferster and Skinner followed the same rule with pigeons: establish the response on continuous reinforcement, then introduce the schedule.[1] Three differences drive it.

The comparison at a glance

FeatureContinuous reinforcement (CRF)Intermittent (partial) reinforcement
Which responses are reinforcedEvery oneOnly some, by a ratio or interval rule
Speed of acquisitionFastestSlower; may never be acquired on a lean schedule
Response rate once establishedModerate and steadyHigher on ratio schedules
SatiationReached quicklyDelayed
Resistance to extinctionLow; extinction is rapidHigh; highest on variable ratio
Best used forTeaching a new behavior; shapingMaintaining an established behavior

Examples of continuous and intermittent reinforcement

In each row the behavior is the same; only the rule for reinforcing it differs.

SettingBehaviorContinuousIntermittent
MachinesPressing a buttonVending machine: every purchase delivers the snackSlot machine: a payout after an unpredictable number of plays
Dog trainingSitting on cueWeek one: a treat for every sitLater: a treat for some sits, praise for all of them
ClassroomRaising a hand before speakingNew routine: the teacher calls on the student every timeEstablished routine: called on some of the time, never for calling out
ParentingAsking politelyEvery polite request granted while the phrase is being learnedPolite requests granted often but not always
FishingCastingA stocked pond, brieflySome casts catch a fish; the angler keeps casting for hours

The dog-training row shows the usual sequence: continuous reinforcement while the sit is being taught, then intermittent treats with praise on every trial as a conditioned reinforcer. The slot-machine row shows a designer who wants persistence rather than learning: no one needs to be taught to press a button, so the machine can start lean and stay lean — Skinner's gambler, held through long droughts by a variable ratio.[3]

The partial reinforcement extinction effect

When reinforcement is withdrawn, behavior that was reinforced only some of the time keeps going longer than behavior that was reinforced every time. This is the partial reinforcement extinction effect (PRE), and it is paradoxical on its face: the intermittently reinforced behavior received fewer reinforcers and yet is harder to get rid of.

Skinner came to intermittent reinforcement by accident, when a shortage of hand-made pellets led him to reinforce his rats only once a minute rather than for every press.[5] In The Behavior of Organisms he reported that rats reinforced on this "periodic" schedule produced far larger extinction curves than rats reinforced for every press.[6] The first experiment designed to isolate the effect was Lloyd Humphreys's in 1939, and it was not operant at all. Humphreys paired a light with a puff of air to the eye until his human subjects blinked at the light; one group received the puff on every trial, another on a random half. When the puff was discontinued, the half-reinforced group went on blinking for far longer.[2]

Mowrer and Jones extended the finding to operant behavior in 1945. Rats pressed a bar for food on schedules that paid every press, every second, third, or fourth press, or an irregular mixture; the leaner the schedule, the more presses the rats made in extinction. They also noticed that if each rat's "response" was counted as the run of presses its schedule required, the groups extinguished after roughly the same number of runs — the response-unit hypothesis.[7] The effect has since been reproduced in runways and Skinner boxes, in rats, pigeons, and people, and Mackintosh's 1974 review treated it as one of the most reliable phenomena in the study of extinction.[8]

Why less reinforcement produces more persistence

The effect took decades to explain, and each explanation captures something the others miss.

The discrimination hypothesis

The oldest account, from Mowrer and Jones, is that extinction has to be noticed before it can take hold. After continuous reinforcement the first unreinforced response is unmistakable evidence that conditions have changed; after intermittent reinforcement a run of unreinforced responses looks like more of the same.[7] It is correct as far as it goes. But in two 1962 experiments, animals given partial reinforcement, then a block of continuous reinforcement, and only then extinction still showed the effect, although the shift to extinction should have been just as detectable for them as for animals trained on continuous reinforcement alone.[9][10] Something learned during partial reinforcement survived the intervening block.

Amsel's frustration theory

Abram Amsel argued that nonreward is not a neutral event. When an organism expects a reinforcer and does not get it, the omission produces frustrative nonreward, an aversive state that disrupts behavior and drives the organism away. On a continuous schedule, frustration is first met in extinction, where it does exactly that. On an intermittent schedule the organism meets frustration during training, keeps responding, and is reinforced, so the cues of anticipated frustration become signals for continuing rather than quitting.[11] This explains why the effect survives a block of continuous reinforcement.

Capaldi's sequential theory

E. J. Capaldi's account is about memory rather than emotion. On each trial the organism carries a trace of what happened on the last one, and that trace is part of the situation in which the next response occurs. When a nonrewarded trial is followed by a rewarded one, responding is reinforced in the presence of the memory of nonreward. In extinction that memory is the only kind on offer; a subject whose responding has been reinforced in its presence keeps going, and a subject on continuous reinforcement, which never has, stops.[12] The theory predicts effects of the sequence of trials, not just their proportion, and it handles effects that appear after only a handful of trials, where frustration has little time to develop.

Nevin's behavioral momentum

John Nevin came at the problem from the other direction. In his research on behavioral momentum, resistance to disruption is measured as the proportional decline from baseline, in the same subject, across situations that differ in rate of reinforcement. Measured this way, richer reinforcement produces more resistance to extinction, not less. Nevin argued that the classic group comparison counts absolute responses without controlling for the different baseline rates the schedules produce, and that what remains of the PRE reflects the size of the change from training to extinction: from continuous reinforcement to none is a large change, from lean reinforcement to none a small one.[13] In the fuller theory, persistence depends on the rate of reinforcement in a context, and the speed of extinction also on how detectable the change is.[14]

What the theories agree on

Intermittent reinforcement teaches something continuous reinforcement never can: that unreinforced responses are ordinary and that responding through them pays. Whether that lesson is called a discrimination, a counterconditioned frustration, a memory, or a small stimulus change, the prescription is the same. An organism that has never met an unreinforced response has no defense against the first one.

How to thin reinforcement from continuous to intermittent

The transition is called schedule thinning, and it is where most applied failures happen: the goal is to reinforce fewer responses without the behavior faltering.[4][15]

  1. Establish the behavior on continuous reinforcement. Stay on CRF until the behavior occurs promptly and reliably whenever the occasion arises: the dog sits on the first cue, the student raises a hand without a reminder.
  2. Take the smallest step. From every response, go to two of every three, then every other (FR 2), then two of every five — steps small enough that the learner hardly notices. Big jumps are the usual cause of collapse.
  3. Make it variable as soon as it is intermittent. Past every-other, vary the requirement around the average rather than fixing it. Variable schedules produce steadier responding with fewer pauses, and they are what the world will impose anyway.[1]
  4. Move only when performance is stable. Raise the requirement after the behavior has held steady at the current step, not on a timetable.
  5. Watch for ratio strain and back off. Pausing, slowing, errors, and emotional behavior after a step up are ratio strain: the schedule has been stretched faster than the behavior can bear. Return to the last requirement that worked and stretch again more gradually.[4]
  6. Keep a conditioned reinforcer on every response. The tangible reinforcer becomes intermittent; praise, a click, or a check mark can stay continuous. In a token economy, the tokens thin and the exchange for backup reinforcers is spaced out separately.[15]
  7. Hand the behavior to natural reinforcers. The endpoint is no contrived schedule at all: the sit that produces the walk, the hand-raise that produces the turn to speak; the habits page covers this handoff for your own behavior.

Why intermittently reinforced problem behavior is so hard to extinguish

The partial reinforcement extinction effect has a dark side. Problem behavior in homes, classrooms, and institutions is almost never reinforced continuously: a tantrum is given in to on some occasions and not others; a self-injurious act draws attention from one staff member and is ignored by the next. By the time anyone tries extinction, the behavior has a long history on a lean, variable schedule — the condition that produces the greatest persistence.

Lerman and Iwata's 1996 review of basic and applied extinction research made the point directly: laboratory findings predict that intermittently reinforced behavior will resist extinction, the natural environment supplies intermittent reinforcement almost by default, and applied studies that varied the schedule before extinction were scarce. Their practical conclusions still hold: expect extinction to be slow, expect an extinction burst, and combine extinction with differential reinforcement of an alternative behavior rather than relying on withholding alone.[16]

Behavioral momentum adds a warning: reinforcing an alternative behavior in the same context enriches it, and Nevin and Grace's theory predicts that a richer context makes all behavior in it more resistant to change, the problem behavior included.[14] Practitioners therefore deliver the alternative reinforcement in a clearly different situation where they can, and above all they stay consistent: a single give-in during extinction is a reinforcer delivered on the leanest schedule yet.

"Intermittent reinforcement" in relationships

The term has escaped the laboratory. In articles about dating and abuse, "intermittent reinforcement" describes a partner who is warm one day and cold or cruel the next, and the argument is that the unpredictability itself keeps the other person attached, often under the heading of "trauma bonding." It is worth being precise about what the schedule concept supports.

What it supports is a claim about behavior. If someone's texting, apologizing, or effort to please has been reinforced by affection on an unpredictable schedule, that behavior will persist longer when the affection stops than it would have under reliable affection. People are not exempt from the effect; Humphreys's original subjects were people.[2]

What it does not support is any claim about the bond. The laboratory effect is measured in lever presses and eyeblinks after reinforcement stops; it says nothing about whether unpredictable affection produces stronger attachment than reliable affection, and no schedule experiment has tested that. The "trauma bonding" claims come from clinical and popular writing, not schedule research, and the evidence for them as operant findings is thin. Real relationships also contain much that a schedule analysis leaves out: relief when tension breaks (negative reinforcement), fear, money, children, and the practical difficulty of leaving. What the term fairly says is that behavior reinforced unpredictably will be hard to stop, and will get worse before it fades.

Common mistakes

What the evidence does not show

Key takeaways

Check yourself

A new owner starts teaching a puppy to sit and gives a treat for roughly every third sit "so the behavior will be strong." Two weeks later the puppy still does not sit on cue. What went wrong?

The schedule was thinned before the behavior existed. On an intermittent schedule most sits go unreinforced, so the puppy rarely contacts the relation between sitting and the treat and the behavior never gets established. Intermittent reinforcement maintains a behavior that has already been acquired; it does not build one. Start on continuous reinforcement, and thin only once the sit is prompt and reliable.

A teacher has kept a class working quietly for a month by praising them after every good five-minute block. She decides that is too much praise and switches to praising once per lesson. Within a week the class is noisy again. What happened, and what should she do?

She stretched the schedule far too fast, from every block to roughly one in ten, and the behavior showed ratio strain and then broke down. The class did not lose motivation; the schedule did not hold. She should return to a requirement close to the one that worked, hold there until quiet work is stable, then thin in small, variable steps, keeping a brief acknowledgment on most blocks as a conditioned reinforcer while the tangible reinforcer thins.

A friend says her partner's unpredictable affection has "intermittently reinforced" her into staying. Is the schedule concept being used correctly?

Partly. The concept applies to behavior: if her efforts to reach out have been reinforced by affection only unpredictably, those efforts will persist longer than they would have under reliable affection, and there is no reason to think people are exempt. But the laboratory effect says nothing about attachment or the strength of a bond, and staying in a relationship involves negative reinforcement, fear, and practical constraints that a schedule analysis leaves out. The fair conclusion is only that the behavior will be hard to stop and will get worse before it fades.

Want more? The 20-question quiz covers every page on the site with instant explanations.

Frequently asked questions

What is the difference between continuous and intermittent reinforcement in simple terms?

With continuous reinforcement, every time the behavior happens it is reinforced: every sit earns a treat, every coin in the vending machine delivers a snack. With intermittent reinforcement, only some occurrences are reinforced: some sits earn a treat, some pulls on a slot machine pay out. Continuous reinforcement teaches a behavior fastest; intermittent reinforcement keeps a learned behavior going and makes it much harder to stop.

Continuous vs. intermittent reinforcement: which is better?

It depends on the stage. Continuous reinforcement is better for teaching a new behavior, because the learner cannot miss the connection between the behavior and its consequence. Intermittent reinforcement is better for maintaining a behavior that is already reliable: it needs fewer reinforcers, delays satiation, and produces behavior that survives when reinforcement is occasionally missed. The standard sequence is continuous first, then a gradual shift to intermittent.

What is an example of continuous reinforcement, and what is an example of intermittent reinforcement?

A light switch is continuous reinforcement: flip it and the light comes on every time. A treat for every sit during a puppy's first week of training is another. A slot machine is intermittent reinforcement: it pays after an unpredictable number of plays. So is fishing, where only some casts catch a fish, and a parent who gives in to whining some of the time.

What is the partial reinforcement extinction effect?

The finding that behavior reinforced only some of the time persists longer after reinforcement stops than behavior reinforced every time. Lloyd Humphreys first demonstrated it in 1939 with human eyeblink conditioning, and Mowrer and Jones extended it to rats pressing a bar in 1945. It is paradoxical because the intermittently reinforced behavior received fewer reinforcers yet is harder to extinguish.

Why does intermittent reinforcement make behavior harder to extinguish?

Several explanations are in use. The change to extinction is harder to detect, because unreinforced responses were always common. Amsel argued that the organism learns to keep responding through the frustration of nonreward; Capaldi that responding becomes conditioned to the memory of nonreward; Nevin that the shift from a lean schedule to none is a smaller change than from continuous reinforcement to none. All agree that an organism that has never met an unreinforced response has no defense against it.

How do you switch from continuous to intermittent reinforcement?

Gradually, and only after the behavior is reliable. Go from every response to two of every three, then every other, then vary the requirement around a slowly growing average. Move to the next step only when performance is stable at the current one. If responding slows, pauses, or falls apart, that is ratio strain: drop back to the last step that worked. Keep praise or another conditioned reinforcer on every response while the tangible reinforcer thins.

Is intermittent reinforcement the same as being inconsistent?

No. Intermittent reinforcement is a rule applied deliberately to a behavior that has already been learned, to maintain it economically. Inconsistency is an accidental schedule applied to whatever behavior happens to be occurring, which is often the behavior nobody wanted. A parent who gives in to a tantrum one time in five has put the tantrum on a lean variable schedule, which is the most persistent kind of behavior there is.

What does "intermittent reinforcement" mean in relationships?

In popular writing it describes a partner whose affection is unpredictable and the claim that the unpredictability itself keeps the other person attached. The operant concept supports only part of this: behavior that has been reinforced unpredictably, such as repeatedly reaching out, will persist longer when the affection stops. It says nothing about the strength of a bond, and the "trauma bonding" claims are not established findings from schedule research.

References

  1. Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
  2. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. Journal of Experimental Psychology, 25(2), 141–158.
  3. Skinner, B. F. (1953). Science and Human Behavior. Macmillan.
  4. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.
  5. Skinner, B. F. (1956). A case history in scientific method. American Psychologist, 11(5), 221–233.
  6. Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
  7. Mowrer, O. H., & Jones, H. (1945). Habit strength as a function of the pattern of reinforcement. Journal of Experimental Psychology, 35(4), 293–311.
  8. Mackintosh, N. J. (1974). The Psychology of Animal Learning. Academic Press.
  9. Jenkins, H. M. (1962). Resistance to extinction when partial reinforcement is followed by regular reinforcement. Journal of Experimental Psychology, 64(5), 441–450.
  10. Theios, J. (1962). The partial reinforcement effect sustained through blocks of continuous reinforcement. Journal of Experimental Psychology, 64(1), 1–6.
  11. Amsel, A. (1958). The role of frustrative nonreward in noncontinuous reward situations. Psychological Bulletin, 55(2), 102–119.
  12. Capaldi, E. J. (1966). Partial reinforcement: A hypothesis of sequential effects. Psychological Review, 73(5), 459–477.
  13. Nevin, J. A. (1988). Behavioral momentum and the partial reinforcement effect. Psychological Bulletin, 103(1), 44–56.
  14. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. Behavioral and Brain Sciences, 23(1), 73–90.
  15. Kazdin, A. E. (2013). Behavior Modification in Applied Settings (7th ed.). Waveland Press.
  16. Lerman, D. C., & Iwata, B. A. (1996). Developing a technology for the use of operant extinction in clinical settings: An examination of basic and applied research. Journal of Applied Behavior Analysis, 29(3), 345–382.