# Variable Ratio Schedule: Definition, Examples, and Evidence

> A variable ratio schedule reinforces after an unpredictable number of responses. VR notation, random ratio vs. VR, slot machines, thinning, and ratio strain.

- Source: https://operantconditioning.com/variable-ratio-schedule/
- Author: Ryan Martinson (https://operantconditioning.com/about/)
- Publisher: Operant Conditioning Inc.
- Published: 2026-09-13 · Updated: 2026-09-13
- License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching)

---

*Schedules · Ratio*

A slot machine, a fishing rod, and a well-trained dog have one thing in common: the payoff comes after an unpredictable number of tries. That arrangement is the variable ratio schedule, and it produces the fastest, steadiest, and most stubborn behavior in operant conditioning. This page explains how it is built, what it does to behavior, and where the evidence stops.

> **Definition**
>
> A variable ratio schedule (VR) is a schedule of reinforcement in which a reinforcer follows an unpredictable number of responses that varies around an average. On a VR 10 schedule a reinforcer might come after 4 responses, then 17, then 9, then 10; over many reinforcers the count averages 10. The requirement depends only on responses, never on time, and nothing in the situation tells the organism which response will be the one that pays.[1]
>
> Variable ratio schedules produce the highest and steadiest rates of any basic schedule and the greatest persistence once reinforcement stops. Ferster and Skinner mapped them in *Schedules of Reinforcement* (1957), and Skinner had already identified them as the schedule behind gambling.[1][2]

**In brief**

- A variable ratio schedule reinforces after a number of responses that varies unpredictably around an average; VR 10 means ten responses per reinforcer on average, and VR 1 is continuous reinforcement.
- Because the very next response could be the one that pays, VR schedules produce a high, steady rate with little or no [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause), and behavior built on them is the hardest to [extinguish](https://operantconditioning.com/extinction/).
- The way to reach a VR schedule is to reinforce every response first and then thin gradually and irregularly; thin too fast and you get [ratio strain](https://operantconditioning.com/glossary/#ratio-strain) instead of persistence.

## How a variable ratio schedule works

Every [schedule of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) answers one question: which responses earn the reinforcer? A ratio schedule answers it with a count. On a [fixed ratio](https://operantconditioning.com/fixed-ratio-schedule/) the count is always the same. On a variable ratio the count changes from one reinforcer to the next, and only its average is fixed. The notation gives that average: VR 5, VR 20, VR 100. A VR 1 reinforces every response and is [continuous reinforcement](https://operantconditioning.com/continuous-vs-intermittent-reinforcement/) under another name.[1]

Two features follow. First, the reinforcer depends on responding and on nothing else: waiting accomplishes nothing, and responding faster brings the next reinforcer sooner. Second, no signal marks the response that will pay. After a reinforcer on a fixed ratio of 20, the organism has nineteen unpaid responses ahead of it and pauses before starting the next run. After a reinforcer on a VR 20, the next reinforcer might be one response away.[1][3]

Procedure and process

"VR 10" names a procedure: the rule by which reinforcers are arranged. The high, steady rate is a process: what the rule does to behavior once the consequence is actually reinforcing. You can arrange a flawless VR 10 and get nothing if the consequence is not a reinforcer for that organism at that moment.

## Two ways to build one: arranged variable ratio and random ratio

A variable ratio is built in one of two ways, and the difference matters.

The first is the method Ferster and Skinner used. The experimenter writes a list of ratios whose mean is the value wanted — for a VR 10, something like 1, 3, 5, 8, 10, 12, 14, 17, and 20 — and the apparatus works through it in an irregular order. This is an **arranged variable ratio**. It has a smallest ratio and a largest one, so the longest unpaid run the subject will ever meet is known in advance.[1]

The second is a **random ratio** (RR). Each response is reinforced with a fixed probability, independently of every other response. On an RR 10, every response has a one-in-ten chance of paying, whatever happened on the previous nine or ninety. The mean is still 10, but the ratios are unbounded: runs of fifty or a hundred unpaid responses are rare but certain to occur eventually. A random ratio has no memory: a long losing run does not make the next response any more likely to pay, which is what the gambler who feels a win is "due" gets wrong.

Modern slot machines are random-ratio devices: each spin is an independent draw at a fixed probability set by the program. Haw has argued that calling them variable-ratio schedules blurs a real distinction, because the unbounded losing runs and occasional early wins of a random ratio are part of what the player experiences.[4] The two produce similar behavior, and this page uses "variable ratio" for both except where the difference matters.

## The behavior it produces

### A high, steady rate

The signature of a VR schedule is a high rate of responding that runs on without breaks. The [post-reinforcement pause](https://operantconditioning.com/glossary/#post-reinforcement-pause) that punctuates fixed-ratio performance is short or absent. Schlinger, Derenne, and Baron's review of fifty years of research on pausing explains why: the pause on ratio schedules is better understood as a pre-ratio pause, controlled by the size of the ratio ahead rather than by the reinforcer just received. On a fixed ratio the ratio ahead is always the full count, and the pause grows with it. On a variable ratio the ratio ahead might be one, so the pause shrinks, lengthening again only when the average ratio becomes large.[5]

Ratio schedules also produce higher rates than interval schedules at the same rate of reinforcement. Zeiler's review of the controlling variables and Baum's comparison of pigeons on ratio and interval schedules both trace this to the **feedback function**, the relation between response rate and reinforcement rate. On a ratio schedule it is a straight line: double the rate and you double the reinforcers. On an interval schedule it flattens; past a modest rate, responding faster earns almost nothing extra.[6][7]

### The greatest resistance to extinction

When reinforcement stops, behavior that was reinforced intermittently persists far longer than behavior that was reinforced every time. This is the [partial reinforcement extinction effect](https://operantconditioning.com/glossary/#partial-reinforcement-extinction-effect), and among the basic schedules the variable ratio produces the most of it.[1][3] The effect was first shown in classical conditioning: Humphreys found in 1939 that eyeblink responses conditioned with the air puff on only half the trials extinguished more slowly than responses conditioned with the puff on every trial.[8] Mowrer and Jones then trained rats to press a lever for food on ratios of one to four and on an irregular pattern; the more presses had been required per reinforcer in training, the more presses the rats made in extinction.[9]

The usual explanation is discrimination. On continuous reinforcement, the first unreinforced response is unlike anything in the organism's history. On a VR 50, a run of eighty unreinforced responses is within normal experience, so nothing has visibly changed, and behavior continues until the run grows long enough to be discriminated from training.[3]

### Ratio strain

The high rate has a limit. If the ratio is stretched too far, or too fast, responding breaks down: long pauses, bursts separated by inactivity, and eventually extinction even though reinforcement is still available. Ferster and Skinner saw this in pigeons when ratios were pushed too high, and trainers see it whenever reinforcement is thinned faster than the behavior can bear.[1][3] [Ratio strain](https://operantconditioning.com/glossary/#ratio-strain) is the schedule failing to maintain the behavior, not the organism being stubborn, and the remedy is a richer schedule and a slower stretch.

## What Ferster and Skinner's cumulative records show

The evidence for all of this is a stack of cumulative records. Skinner's cumulative recorder, described in *The Behavior of Organisms*, made rate visible: a pen steps up once for every response as the paper rolls, so a steep line means a high rate, a flat line means no responding, and a small diagonal tick marks each reinforcer.[10] Ferster and Skinner's 1957 book is hundreds of such records, mostly from pigeons pecking a lighted key.[1]

A fixed-ratio record is a staircase: a flat pause after each reinforcer, then a steep run to the next. A fixed-interval record shows the scallop, a curve that accelerates toward each reinforcer. A variable-interval record is a straight line at a moderate slope. A variable-ratio record is a straight line at a steep slope, with the reinforcer ticks scattered irregularly along it and almost no visible reaction to any of them. When the mean ratio was pushed high enough, the record broke into runs separated by flat stretches, which is what ratio strain looks like on paper.[1] The site's [virtual Skinner box](https://operantconditioning.com/lab/) produces a stylized VR record, and the [Skinner box page](https://operantconditioning.com/skinner-box/) explains how to read one.

## Variable ratio vs. fixed ratio, variable interval, and fixed interval

The [schedules hub](https://operantconditioning.com/schedules-of-reinforcement/) gives the two-question rule for telling the basic schedules apart: count or time, fixed or variable? Here the point is narrower — what each of the other three does that a variable ratio does not.

| Schedule | Reinforcer follows | Pattern on the record | Resistance to extinction | How it differs from VR |
| --- | --- | --- | --- | --- |
| **Variable ratio (VR)** | An unpredictable number of responses, averaging *n* | Steep, straight line; little or no pause | Highest | — |
| **Fixed ratio (FR)** | Every *n*th response | Staircase: pause, then run; pause grows with *n* | High | The count ahead is known, so the organism pauses before each ratio |
| **Variable interval (VI)** | First response after an unpredictable time, averaging *t* | Straight line, moderate slope | High | Responding faster does not bring the reinforcer sooner, so the rate is moderate |
| **Fixed interval (FI)** | First response after a fixed time | Scallop: pause, then acceleration | Moderate | Time, not effort, sets up the reinforcer; early responses are wasted |

Two contrasts matter most. Against the fixed ratio, the variable ratio trades predictability for steadiness: the same reinforcers per hundred responses, delivered on an irregular count, removes the pauses and produces a smoother and usually higher rate.[1][5] Against the [variable interval](https://operantconditioning.com/variable-interval-schedule/), the variable ratio rewards speed: on VI, a pigeon that pecks twice as fast earns barely any more food and settles into a moderate rate, while on VR twice the pecks means twice the food.[7] The [fixed interval](https://operantconditioning.com/fixed-interval-schedule/) is the schedule least like VR: the clock sets up the reinforcer, and the organism does little until the time is nearly up.

## Where variable ratio schedules show up

In each row below, a response is followed by a reinforcer after an unpredictable count, and more responses mean more reinforcers. The schedule column is honest about how close the fit is.

| Setting | Response | Reinforcer | Schedule | Note |
| --- | --- | --- | --- | --- |
| Casino | A spin of a slot machine | A payout | Random ratio, set by the program | Each spin independent; a near miss is a loss |
| Fishing | A cast | A fish on the line | Approximately VR | More casts, more fish, on an unpredictable count |
| Sales | A cold call | An appointment or a sale | Approximately VR | "A numbers game" describes a ratio schedule |
| Dog training | Sit on cue | A treat | VR 3 to VR 5 after thinning | The standard way to maintain a trained behavior |
| Phone | Pulling to refresh, scrolling | A message, a like, a good post | Resembles a variable schedule | Ratio and interval features mixed; the contingency is not public |

### The slot machine

Skinner made the connection himself. In *Science and Human Behavior* he observed that the effectiveness of variable-ratio schedules in generating high rates had long been known to the proprietors of gambling establishments, and he described the compulsive gambler as the result of such a schedule.[2] Schüll's ethnography of machine gambling in Las Vegas shows the industry's side: designers working explicitly to extend "time on device," with fast play, small frequent payouts, and features tuned to keep players seated.[11]

### The near miss

One feature of slot machines is not a schedule effect at all. A **near miss** is a loss that looks almost like a win — two matching symbols on the payline and the third just above or below it. Reid argued in 1986 that near misses encourage continued play even though they pay nothing, and that a game can be built to produce them.[12] In one laboratory test, people playing a simulated slot machine whose losing spins were near misses about 30 percent of the time kept playing longer after the machine stopped paying than people who saw near misses more rarely or more often.[13] On a pure random ratio every loss is the same event. The near miss is a stimulus added on top of the schedule, a reminder that the schedule is only one of the variables at work.

### Feeds and notifications

Refreshing a social feed is now the textbook example after the slot machine, and it deserves more care. The resemblance is real: each pull or scroll is a response, most produce nothing, and the count between payoffs is unpredictable. But new posts arrive with time whether or not you refresh, which is an interval feature, and no one outside the companies knows the actual contingency. The careful statement is that feed-checking *resembles* behavior on a variable schedule; the studies behind this page used pigeons, rats, and slot machines, not feeds.

## How to use a variable ratio schedule

A variable ratio is a maintenance schedule, not a teaching schedule: a learner who has never been reinforced for a behavior will not persist through nine unpaid attempts to reach the tenth. The sequence, in [dog training](https://operantconditioning.com/dog-training/), classrooms, and self-management alike, is continuous reinforcement first and then a gradual stretch.[3]

1. **Establish the behavior on continuous reinforcement.** Reinforce every correct response until it is fluent — a dog that sits on cue nine times in ten, a child who starts homework when asked. If you are still shaping the behavior, you are not ready to thin.
2. **Stretch to a small, variable ratio.** Reinforce about two responses in three, then one in two, irregularly rather than in a pattern. Write the sequence down in advance or draw it from a hat; a "variable" schedule that always pays on the third response is a fixed ratio with extra steps.
3. **Keep some short ratios in the mix.** Include ratios of one and two, so that a reinforcer is sometimes followed almost immediately by another. If a reinforcer reliably means the next one is far off, you have rebuilt the fixed-ratio pause.
4. **Stretch slowly, and watch the behavior rather than the calendar.** Raise the mean only when responding at the current ratio is steady. Pauses, refusals, and a drop in quality are the early signs of ratio strain; when they appear, drop back to the last ratio that worked and hold it longer.[3]
5. **Keep the reinforcer worth working for.** Thinning changes how often the reinforcer arrives, not how much it has to matter. A full dog or a token that buys nothing will defeat any schedule.
6. **Hand the behavior off to natural consequences.** A contrived VR keeps the behavior alive until the world's own reinforcers take over — the recall that ends in a game, the homework habit that pays off at school. The [habits page](https://operantconditioning.com/habits/) covers the same sequence for your own behavior.

## Common mistakes

- **Thinning too fast.** The commonest error. Jumping from continuous reinforcement to "every fifth time or so" produces ratio strain, and the trainer concludes that intermittent reinforcement does not work. It does; the stretch was too steep.
- **Starting on a variable ratio.** A behavior that has never been reinforced cannot be maintained by a schedule that pays one attempt in ten. Continuous reinforcement comes first.
- **Being predictable.** Every third response, or a strict alternation, is a fixed schedule, and the learner will find the pattern. Randomize.
- **Calling an interval schedule a ratio schedule.** If the reinforcer becomes available with time and responding faster does not bring it sooner, it is an interval schedule. Waiting for a reply is variable interval; the slot machine and the cold call are variable ratio.
- **Arranging one by accident.** Giving in "just this once," now and then, is the most effective way known to make an unwanted behavior permanent. Before trying to extinguish anything, ask what schedule it is on and who is delivering the reinforcer, and expect an [extinction burst](https://operantconditioning.com/extinction/#what-is-an-extinction-burst) when the schedule ends.

## What the evidence does not show

**It does not show that variable ratio schedules are addictive.** A schedule is a description of when reinforcers arrive. It explains why a behavior persists through unpaid runs; it does not, by itself, make a consequence reinforcing or a behavior harmful. Most people who fish, sell, or play an occasional slot machine develop no problem, and Schüll's account makes clear that the schedule is embedded in design choices — speed of play, credits in place of coins, near misses, the ergonomics of the seat — and in the circumstances players bring with them, none of which is a schedule.[11] "Variable ratio" is one variable in that analysis, not a diagnosis.

**It does not show that the partial reinforcement extinction effect is simple.** Nevin argued in 1988 that the effect, though reliably found when groups trained on different schedules are compared, sits awkwardly beside another well-established finding: within a single organism, behavior maintained by more frequent reinforcement is more resistant to disruption. How the effect comes out depends on how extinction is measured and on how easily the change from training to extinction can be discriminated.[14] Mowrer and Jones suggested something similar: the effective unit after ratio training may be the whole run ending in food, so counting single responses in extinction overstates the persistence.[9]

**It does not show that adults behave like pigeons.** Lowe found that adults on simple schedules often produce patterns unlike the animal records, and argued that the reason is verbal: people describe the schedule to themselves and then follow the description, which may be wrong. His evidence came mostly from fixed-interval schedules, but the implication for ratio schedules is direct.[15] A person who has decided that a machine is "due" is responding to a rule, not to the random ratio in front of them, and a person who has correctly concluded that the odds are fixed can stop in a way no pigeon can.

## Key takeaways

- A variable ratio schedule reinforces after a number of responses that varies unpredictably around a mean; VR 10 averages ten responses per reinforcer, and VR 1 is continuous reinforcement.
- Because the ratio ahead might be one, a variable ratio produces a high, steady rate with little or no post-reinforcement pause, and higher rates than interval schedules at the same rate of reinforcement.
- Behavior on a variable ratio is the most resistant to extinction of any basic schedule: a long unpaid run looks like ordinary experience, so nothing signals that the contingency has changed.
- An arranged variable ratio draws from a bounded list; a random ratio pays each response with a fixed probability and has no memory. Slot machines are random-ratio devices, and a win is never "due."
- Reinforce every response first, then thin gradually and irregularly, keeping some short ratios in the mix; thinning too fast produces ratio strain rather than persistence.
- The schedule is a description, not a diagnosis: it does not make a consequence reinforcing or gambling harmful by itself, and adult humans layer rules over it that pigeons do not.

### Check yourself

**A trainer moves a dog straight from a treat for every sit to a treat for roughly every fifth sit. Within one session the dog starts wandering off between cues. What happened, and what should the trainer do?**

Ratio strain. The schedule was thinned faster than the behavior could bear, so responding broke down: pausing, wandering, and, if it continues, extinction. The trainer should drop back to a richer schedule — every sit, or every other sit — hold it until sitting is steady again, then stretch in smaller steps, keeping some ratios of one and two in the mix.

**A friend says that checking for text replies is a variable ratio schedule, "like a slot machine." Is it?**

No. A reply becomes available with the passage of time, not with the number of checks, and checking faster does not make it arrive sooner; only the first check after it arrives is reinforced. That is a variable interval schedule, which produces a moderate, steady rate. The slot machine is a ratio schedule because every spin is a chance to win and more spins mean more wins.

**After forty losing spins someone says the machine "must be about to pay." On a random ratio schedule, is the next spin more likely to win?**

No. On a random ratio each response is reinforced with the same fixed probability, independently of everything before it, so forty losses change nothing about the forty-first spin. The belief that a win is due is a rule the person has formed, not a property of the schedule, and it is a good example of how human persistence on ratio schedules can be governed by self-talk that misdescribes the contingency.

## Frequently asked questions

**What is a variable ratio schedule in simple terms?**

A variable ratio schedule pays off after a number of responses that changes unpredictably, averaging some value. On a VR 10, a reinforcer might come after 3 responses, then 15, then 12, averaging 10 over time. Because the next response could always be the one that pays, the behavior stays fast and steady, and it keeps going for a long time after the payoffs stop.

**What is an example of a variable ratio schedule?**

A slot machine: each spin is a response, most spins pay nothing, and a payout arrives after an unpredictable number of them. Other cases are casting a fishing line, making sales calls, submitting job applications, and a trained dog that gets a treat for an unpredictable one sit in four. In every case, more responses mean more reinforcers, but the count between reinforcers varies.

**What does VR 10 mean?**

VR stands for variable ratio, and the number is the average ratio of responses to reinforcers. VR 10 means that, averaged over many reinforcers, ten responses are required for each one; the actual counts vary, perhaps from one to twenty. VR 1 is continuous reinforcement. A random ratio 10 (RR 10) is the related schedule in which every response has a one-in-ten chance of being reinforced.

**What is the difference between a variable ratio and a variable interval schedule?**

Both are unpredictable, but a variable ratio depends on how many responses you make, while a variable interval depends on how much time has passed. On VR, responding faster earns reinforcers faster, so the rate is high. On VI, only the first response after the interval counts, so responding faster gains nothing and the rate is moderate. A slot machine is VR; checking for a text reply is VI.

**What is the difference between a variable ratio and a fixed ratio schedule?**

A fixed ratio reinforces after the same number of responses every time; a variable ratio reinforces after a number that varies around an average. Fixed ratios produce a pause after each reinforcer followed by a fast run, because the organism knows the full count lies ahead. Variable ratios produce a steady rate with little pausing, because the next reinforcer might be one response away.

**Why is the variable ratio schedule the most resistant to extinction?**

Because a long run without reinforcement looks like ordinary experience. On continuous reinforcement, the first unreinforced response is a clear signal that something has changed. On a lean variable ratio, dozens of unreinforced responses are normal, so the organism has no way to tell that reinforcement has stopped and keeps responding. This is the partial reinforcement extinction effect, and variable ratio schedules produce the most of it.

**Is a slot machine a variable ratio or a random ratio schedule?**

Strictly a random ratio. Each spin is an independent draw with a fixed probability of paying, so the number of spins between wins is unbounded and a win is never due. A laboratory variable ratio draws from a fixed list of ratios. The behavior the two produce is similar, a high and steady rate, but the random ratio's long losing runs and occasional back-to-back wins are part of what machine gambling is.

**Is social media a variable ratio schedule?**

It resembles one. Pulling to refresh or scrolling is a response, most produce nothing, and something interesting turns up after an unpredictable number. But new posts and messages arrive with time regardless of your checking, which is an interval feature, and the actual contingency inside any app is not public. The careful statement is that feed-checking resembles behavior on a variable schedule, not that it has been shown to be one.

## References

1. Ferster, C. B., & Skinner, B. F. (1957). *Schedules of Reinforcement*. Appleton-Century-Crofts.
2. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan.
3. Mazur, J. E. (2017). *Learning and Behavior* (8th ed.). Routledge.
4. Haw, J. (2008). Random-ratio schedules of reinforcement: The role of early wins and unreinforced trials. *Journal of Gambling Issues, 21*, 56–67.
5. Schlinger, H. D., Derenne, A., & Baron, A. (2008). What 50 years of research tell us about pausing under ratio schedules of reinforcement. *The Behavior Analyst, 31*(1), 39–60.
6. Zeiler, M. D. (1977). Schedules of reinforcement: The controlling variables. In W. K. Honig & J. E. R. Staddon (Eds.), *Handbook of Operant Behavior* (pp. 201–232). Prentice-Hall.
7. Baum, W. M. (1993). Performances on ratio and interval schedules of reinforcement: Data and theory. *Journal of the Experimental Analysis of Behavior, 59*(2), 245–264.
8. Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. *Journal of Experimental Psychology, 25*(2), 141–158.
9. Mowrer, O. H., & Jones, H. (1945). Habit strength as a function of the pattern of reinforcement. *Journal of Experimental Psychology, 35*(4), 293–311.
10. Skinner, B. F. (1938). *The Behavior of Organisms: An Experimental Analysis*. Appleton-Century.
11. Schüll, N. D. (2012). *Addiction by Design: Machine Gambling in Las Vegas*. Princeton University Press.
12. Reid, R. L. (1986). The psychology of the near miss. *Journal of Gambling Behavior, 2*(1), 32–39.
13. Kassinove, J. I., & Schare, M. L. (2001). Effects of the "near miss" and the "big win" on persistence at slot machine gambling. *Psychology of Addictive Behaviors, 15*(2), 155–158.
14. Nevin, J. A. (1988). Behavioral momentum and the partial reinforcement effect. *Psychological Bulletin, 103*(1), 44–56.
15. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), *Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour* (pp. 159–192). Wiley.

## About the author

Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards

## Related

- [Schedules of reinforcement](https://operantconditioning.com/schedules-of-reinforcement/): All five basic schedules, the two-question rule, and the interactive simulator.
- [Fixed ratio schedule](https://operantconditioning.com/fixed-ratio-schedule/): The pause-and-run pattern, piece-rate pay, and why the pause grows with the ratio.
- [Extinction](https://operantconditioning.com/extinction/): What happens when reinforcement stops: bursts, recovery, and resurgence.
