Reinforcement · Timing and rate

Schedules of Reinforcement

A behavior's strength depends not just on whether it is reinforced but on how often and on what rule. Those rules are schedules, and they explain everything from slot machines to why you keep checking your phone.

Updated 14 min read

Definition

A schedule of reinforcement is the rule that determines which occurrences of a behavior are followed by a reinforcer. A continuous schedule reinforces every response; an intermittent (partial) schedule reinforces only some, based on the number of responses (ratio schedules) or the passage of time (interval schedules), on a fixed or variable basis.

Each schedule produces a characteristic pattern of responding and a characteristic resistance to extinction. The systematic study of schedules was the work of Charles Ferster and B. F. Skinner, whose 1957 book Schedules of Reinforcement reported hundreds of cumulative records from pigeons and rats.[1]

In brief

  • A schedule of reinforcement is the rule deciding which responses earn a reinforcer: every response (continuous) or only some (intermittent).
  • Intermittent schedules are based on a count of responses (ratio) or on time elapsed (interval), with a fixed or variable requirement.
  • Continuous reinforcement teaches a behavior fastest; variable ratio maintains the highest, steadiest rate and is the hardest to extinguish.

Go deeper on each schedule. Every basic schedule now has its own page, with the laboratory evidence, everyday examples, and how to use it.

Try it: the schedule simulator

Below is a cumulative record — the same kind of graph Skinner's recorder drew on a rolling strip of paper. Time runs left to right; every response steps the line up by one; a small green tick marks each reinforcer. Pick a schedule and let the simulated subject run to watch the classic patterns appear — or switch the mode to "I'll respond" and press Respond yourself.

responses per reinforcer
Responses 0
Reinforcers 0
Elapsed 0s
Rate 0/min

The simulated subject is a stylized model of the response patterns Ferster and Skinner documented — post-reinforcement pauses on fixed ratio, steady high rates on variable ratio, the fixed-interval "scallop," and a moderate steady rate on variable interval — not a replay of real data. Try FR 20 to see a longer post-reinforcement pause, or FI 15 to see the scallop.

Continuous versus intermittent reinforcement

On a continuous reinforcement schedule (CRF), every response is reinforced. It is the fastest way to establish a new behavior: the relationship between behavior and consequence is unmistakable. It is also the fastest schedule to extinguish, because the first unreinforced response is immediately informative — something has changed.

On an intermittent schedule, only some responses are reinforced. Learning is slower, but the behavior becomes far more durable. This is the partial reinforcement extinction effect: behavior maintained on an intermittent schedule persists much longer once reinforcement stops than behavior maintained on CRF.[2] A vending machine that fails once loses you immediately; a slot machine that fails a hundred times in a row is doing exactly what it always does.

The practical rule

Reinforce continuously while a behavior is being learned. Once it is reliable, thin gradually to an intermittent schedule — ideally a variable-ratio schedule — to make it resistant to extinction. Thin too fast and you get ratio strain: pausing, erratic responding, and eventually extinction.

The four basic intermittent schedules

Two questions define the basic schedules. Is reinforcement based on the number of responses (ratio) or on time elapsed (interval)? And is the requirement fixed or variable?

Idealized cumulative records for the four basic schedules Four panels. Fixed ratio: steep runs separated by flat post-reinforcement pauses. Variable ratio: a steady steep line. Fixed interval: repeated scallops that start flat and curve upward before each reinforcer. Variable interval: a moderate, steady slope. Fixed ratio (FR) Pause, then a burst ("break and run") Time → Variable ratio (VR) High, steady rate; hardest to extinguish Time → Fixed interval (FI) The "scallop": slow after each reinforcer, accelerating as the interval ends Time → Variable interval (VI) Moderate, steady rate Time →
Idealized cumulative records. Each upward step is a response; green ticks mark reinforcers. Patterns after Ferster & Skinner (1957).

Fixed ratio (FR)

Reinforcement follows every nth response. FR 1 is continuous reinforcement; FR 10 pays after every tenth response. Fixed-ratio schedules produce a high rate of responding with a distinctive post-reinforcement pause after each reinforcer — the larger the ratio, the longer the pause — followed by a rapid, steady run to the next one. Piece-rate pay, "buy ten coffees, get one free," and a set of ten push-ups before a rest are fixed ratios.

Variable ratio (VR)

Reinforcement follows an unpredictable number of responses that averages n. VR 10 might pay after 3, then 15, then 8, then 14 responses. Because the very next response might be the one that pays, there is little pausing; the result is the highest, steadiest response rate of any basic schedule and the greatest resistance to extinction. Slot machines are the textbook example. So are fishing casts, sales calls, refreshing a social feed, and opening loot boxes.[1]

Fixed interval (FI)

The first response after a fixed time has elapsed is reinforced; responses during the interval earn nothing. FI 60 s pays the first response after a minute. Animals on FI schedules produce the famous scallop: almost no responding right after a reinforcer, then an accelerating rate as the end of the interval approaches. Checking the oven as the timer runs down, studying more as the exam nears, and glancing at the mailbox around delivery time follow this pattern. (A salary is not a good example — it is paid on a time basis regardless of the number of responses, which makes it closer to a fixed-time schedule than a fixed-interval one.)

Variable interval (VI)

The first response after an unpredictable interval that averages t is reinforced. VI schedules produce a moderate, remarkably steady rate — the organism keeps checking, because the reinforcer could become available at any moment, but doesn't race, because responding faster doesn't make it come sooner. Checking email, waiting for a reply to a text, and a manager's random walk-throughs of the shop floor all maintain behavior on variable intervals. Because VI produces such stable baselines, it is the workhorse schedule in laboratory research on choice and drugs.

ScheduleReinforcer deliveredResponse patternResistance to extinctionEveryday examples
Continuous (CRF)After every responseSteady while satiation is far offLowLight switch, vending machine, a new skill being taught
Fixed ratio (FR)After a fixed number of responsesHigh rate; post-reinforcement pauseModerate–highPiece-rate pay, loyalty punch cards, sets of reps
Variable ratio (VR)After a variable number of responsesVery high, steady rate; little pausingHighestSlot machines, sales calls, social-media feeds, fishing
Fixed interval (FI)First response after a fixed timeScallop: pause then accelerationModerateWatching the oven timer, cramming before a scheduled exam
Variable interval (VI)First response after a variable timeModerate, steady rateHighChecking email or texts, surprise inspections

Which schedule is it? A two-question rule

Students confuse the four schedules almost as often as they confuse the quadrants, and the same trick works: two questions, in order. (1) Does the reinforcer depend on how many responses were made, or on how much time has passed? Count means ratio; time means interval. (2) Is the requirement always the same, or does it vary around an average? Same means fixed; varies means variable. (If every single response is reinforced, it is continuous reinforcement, and if none are, it is extinction.) A detail that catches people: on an interval schedule, responding faster does not bring the reinforcer sooner — only the first response after the interval counts.

BasisRequirement fixedRequirement varies
Depends on a count of responsesFixed ratio (FR)Variable ratio (VR)
Depends on time elapsedFixed interval (FI)Variable interval (VI)

Practice: name the schedule

Decide, then open the answer.

A factory pays a worker $2 for every 50 shirts sewn.

Fixed ratio (FR 50). The reinforcer depends on a count, and the count is always the same. Expect a brief pause after each payment, then a run.

A slot machine pays out after an unpredictable number of pulls, averaging about one in twenty.

Variable ratio (VR 20). Count-based, requirement varies around an average. High, steady responding; the hardest schedule to extinguish.

Your paycheck arrives every other Friday, provided you have worked that fortnight.

Fixed interval (FI two weeks) — approximately. Time-based and fixed. Real paychecks are a poor example of pure FI because the response requirement is not a single response after the interval, but the classic textbook classification is FI, and the "scallop" of end-of-period effort is real.

You check your phone for messages that arrive at unpredictable times; the first check after a message arrives finds it.

Variable interval. A message becomes available after a varying time, and only a check after that point is reinforced. Checking faster does not make messages arrive sooner — which is why VI produces a moderate, steady rate rather than a frantic one.

A dog gets a treat for every single sit during the first week of training.

Continuous reinforcement (CRF, or FR 1). Every response pays. Fast learning, fast extinction; the right schedule for building a behavior, the wrong one for keeping it.

A quality inspector checks a machine that jams at random intervals; she can only fix a jam once it has happened.

Variable interval. The opportunity to be reinforced (finding and fixing a jam) becomes available after a varying time; checks before a jam cannot be reinforced. Time-based, variable.

Why variable ratio is so powerful — and so dangerous

Three features make VR the schedule of choice for anyone who wants behavior to persist: no signal ever tells the organism that the next response is pointless; the average payout can be made arbitrarily lean once the behavior is established; and the occasional early payout after a long dry run is itself a powerful reinforcer of persistence. Casinos, mobile games, and social platforms have converged on variable-ratio designs for exactly these reasons.[3] The same mechanism that makes a behavior hard to quit makes it hard to build deliberately — you cannot start on VR. The order is always CRF, then a gradual stretch.

Beyond the basic four

ScheduleRuleWhat it's used for
Fixed / variable time (FT, VT)Reinforcer delivered after time passes, regardless of behaviorNoncontingent reinforcement; the procedure behind Skinner's "superstition" study[4] and a treatment for attention-maintained problem behavior
Differential reinforcement of low rate (DRL)Reinforced only if at least t seconds have passed since the last responseSlowing behavior that is fine at a low rate (talking in class, eating speed)
Differential reinforcement of high rate (DRH)Reinforced only if n responses occur within t secondsSpeeding up fluent skills (math facts, typing)
Differential reinforcement of other behavior (DRO)Reinforced if the target behavior does not occur for an intervalReducing problem behavior without punishment; see alternatives to punishment
Progressive ratio (PR)Ratio increases after each reinforcer until the organism stops (the "breakpoint")Measuring how hard an organism will work for a reinforcer — a standard index of reinforcer value and drug abuse liability[5]
Concurrent schedulesTwo or more schedules available at once on different responsesStudying choice; produced Herrnstein's matching law: relative response rate matches relative reinforcement rate[6]
Chained schedulesCompleting one schedule produces the stimulus for the next; only the last delivers the primary reinforcerAnalyzing sequences of behavior; conditioned reinforcement
Multiple and mixed schedulesSchedules alternate, with (multiple) or without (mixed) a signalStudying stimulus control and behavioral contrast

Do humans follow these schedules?

Mostly, with an important caveat. Human performance on simple schedules is often shaped by verbal rules and self-instruction as much as by the schedule itself. Adults told (or who figure out) that reinforcement is time-based may respond at a very low rate or a very high steady rate on FI rather than producing the clean scallop seen in pigeons, and instructions can override contingencies for a surprisingly long time.[7] Infants and young children, who have less verbal behavior to lean on, look more like the animal data. The lesson for anyone applying schedules to people: the description of the contingency is itself a variable, and rules that don't match the real schedule eventually lose to it.

Using schedules deliberately

  1. Start continuous. While a behavior is new — a child's first attempts at a chore, a dog's first sits, your first week of a habit — reinforce every occurrence.
  2. Stretch the ratio slowly. Move to reinforcing two out of three, then every other, then unpredictably around one in three. Watch for ratio strain and back off if the behavior falters.
  3. Prefer variable over fixed. Variable schedules produce steadier behavior with fewer pauses and greater persistence.
  4. Prefer ratio over interval when you want rate. If you want more of a behavior, tie reinforcement to responses. If you want steady monitoring, an interval schedule is fine.
  5. Let natural reinforcers take over. The goal of any contrived schedule is to hand the behavior off to the consequences the world already provides — the clean kitchen, the finished chapter, the dog that comes when called. How to do this for your own habits ›

Schedules and problem behavior

The same principles explain why unwanted behavior is so persistent. A parent who gives in to a tantrum "only occasionally" has put tantrums on a lean variable-ratio schedule — the most extinction-resistant schedule there is. A manager who answers after-hours emails "only when urgent" has done the same for after-hours emailing. When you are trying to reduce a behavior, the first question is always: what schedule is currently maintaining it, and who is delivering the reinforcer? See extinction and the extinction burst ›

Key takeaways

Explain it to a friend. Explain why a slot machine keeps people playing long after a broken vending machine would have lost them, without using the words schedule, ratio, or interval.

Frequently asked questions

What are the four schedules of reinforcement?

Fixed ratio (reinforcement after a set number of responses), variable ratio (after an unpredictable number of responses), fixed interval (first response after a set time), and variable interval (first response after an unpredictable time). Continuous reinforcement — every response reinforced — is the fifth, simplest case.

Which schedule of reinforcement is most effective?

It depends on the goal. Continuous reinforcement is best for teaching a new behavior. Variable ratio produces the highest, steadiest rate and the greatest resistance to extinction, so it is best for maintaining an established behavior. Variable interval produces the most stable moderate rate.

Which schedule is most resistant to extinction?

Variable ratio. Because reinforcement has always been unpredictable, a long run without it is not a signal that anything has changed, so responding persists. This is why gambling and social-media checking are so hard to stop.

Is a salary a fixed interval schedule?

Not really. Fixed interval requires a response after the interval; a salary is paid on a time basis regardless of the amount of behavior, which makes it more like a fixed-time (noncontingent) schedule. Piece-rate pay and commissions are ratio schedules. This is one reason salaries do little to reinforce specific daily behaviors.

What is the fixed-interval scallop?

The characteristic pattern on a cumulative record under a fixed-interval schedule: little or no responding right after a reinforcer, then an accelerating rate as the end of the interval approaches, producing a curve that looks like a series of scallops.

What is a post-reinforcement pause?

The pause in responding that follows each reinforcer on a fixed-ratio (and fixed-interval) schedule. On fixed ratio, the pause grows with the size of the ratio — the "break" in break-and-run responding.

What is the partial reinforcement extinction effect?

The finding that behavior reinforced only some of the time persists longer during extinction than behavior that was reinforced every time. Intermittent reinforcement makes the transition to no reinforcement harder to detect, so the behavior keeps going.

What is ratio strain?

The breakdown in responding — long pauses, erratic bursts, eventually extinction — that occurs when a ratio schedule is thinned too quickly or set too high for the value of the reinforcer. The cure is to drop back to a richer schedule and stretch more gradually.

References

  1. Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
  2. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. Journal of Experimental Psychology, 25(2), 141–158. See also Mowrer, O. H., & Jones, H. (1945). Habit strength as a function of the pattern of reinforcement. Journal of Experimental Psychology, 35(4), 293–311.
  3. Schüll, N. D. (2012). Addiction by Design: Machine Gambling in Las Vegas. Princeton University Press.
  4. Skinner, B. F. (1948). 'Superstition' in the pigeon. Journal of Experimental Psychology, 38(2), 168–172.
  5. Hodos, W. (1961). Progressive ratio as a measure of reward strength. Science, 134(3483), 943–944.
  6. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267–272.
  7. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour (pp. 159–192). Wiley.