Choice · How time gets divided

The Matching Law

Every behavior is a choice among alternatives, and organisms divide their behavior among alternatives in proportion to what each one pays. That single regularity, found in pigeons in 1961, now predicts basketball shot selection, classroom disruption, and why we buy things we do not need.

Updated 15 min read

Definition

The matching law states that when two or more responses are available, the relative rate of each response equals — matches — the relative rate of reinforcement it produces. For two alternatives:

B1B1+B2=R1R1+R2

where B is the rate of each behavior and R the rate of reinforcement it earns. An option that delivers 70% of the available reinforcement gets about 70% of the behavior.[1]

A worked example. A pigeon can peck either of two keys. Key A pays off about 30 times an hour, key B about 10 times an hour, so A delivers 30 ÷ (30 + 10) = 75% of the reinforcement. The matching law predicts that the pigeon will make about 75% of its pecks on A and 25% on B — not all of them on A, even though A is plainly better. That the pigeon does not simply pick the better key every time, and that the poorer option keeps a share proportional to what it pays, is the surprise of the law: on schedules like these, behavior is spread in proportion to payoff rather than piled onto the best option.

In brief

  • Organisms allocate behavior in proportion to reinforcement: an option that delivers 70% of the available reinforcement gets about 70% of the behavior.
  • A behavior's strength depends not only on what it earns but on what everything else earns, so enriching the alternatives reduces it without punishment.
  • Because a reinforcer's value falls hyperbolically with delay, preference flips toward a small, soon reward as it approaches; commitment means choosing before the flip.

Herrnstein's experiment

Richard Herrnstein put pigeons in a chamber with two response keys, each paying grain on its own variable-interval schedule — a concurrent VI VI schedule. A bird could switch between keys whenever it liked. He varied how the reinforcement was divided between the keys while holding the total constant (VI 3-minute against VI 3-minute, VI 2.25 against VI 4.5, and so on), and to stop the birds simply alternating, he added a changeover delay: a brief period after each switch during which no reinforcement could be collected.[1]

The result was startlingly orderly. Plot the proportion of pecks on the left key against the proportion of reinforcers earned on the left key and the points fall on the diagonal. Birds did not go exclusively to the richer key, and they did not split their pecks evenly; they matched.

The matching relation A scatter plot with proportion of reinforcers on the horizontal axis and proportion of responses on the vertical axis. Data points lie close to the diagonal line where the two proportions are equal. Proportion of reinforcers from the left key Proportion of responses on the left key 011 Perfect matching (dashed)
Schematic of the matching relation. Real data cluster around the diagonal; systematic departures from it are captured by the generalized matching law below.

From matching to a new law of effect

In 1970 Herrnstein drew the larger conclusion. If behavior is allocated in proportion to reinforcement, then even a single response in a Skinner box is a choice — between pressing the lever and everything else the rat could do (grooming, exploring, resting), all of which produce some reinforcement of their own. Writing the "everything else" reinforcement as Re, the matching law for one measured response becomes a hyperbola:

B=kRR+Re

Response rate rises with reinforcement rate but with diminishing returns, leveling off at a maximum k. This fit decades of single-schedule data that the original law of effect had described only qualitatively, and Herrnstein proposed it as the quantitative law of effect.[2] It also carried a practical implication: a behavior's strength depends not just on what it earns but on what everything else earns. Enrich the alternatives and the behavior declines without any punishment at all.

The generalized matching law

Real organisms do not match perfectly. William Baum showed in 1974 that the departures are systematic and can be captured by two parameters. Taking logarithms of the ratio form:

log(B1B2)=alog(R1R2)+logb

The slope a is sensitivity. When a = 1, matching is perfect; the usual finding is undermatching, a slope around 0.8, meaning organisms are somewhat less extreme in their preference than the reinforcement ratio warrants. The intercept b is bias: a constant preference for one alternative unrelated to reinforcement — a key that is easier to reach, a side the animal favors.[3][4] The generalized form has been fitted to hundreds of data sets across species, and its parameters turn out to be sensitive to procedural details in interpretable ways: shorter changeover delays produce more undermatching, for instance, because switching itself is reinforced.

Reinforcement rate is not the only thing organisms match to. Reinforcer magnitude, delay, and quality all enter the equation, which is what makes the framework a general account of choice rather than a fact about grain.[4]

Matching beyond the pigeon

Sports

A basketball player choosing between a two-point and a three-point attempt is on a concurrent schedule. Vollmer and Bourret analyzed a season of college basketball and found that the proportion of three-point shots taken by teams and by individual players matched the proportion of points those shots produced.[5] Reed, Critchfield, and Martens applied the generalized matching law to NFL play-calling and found that the ratio of passing to rushing plays tracked the ratio of yards each type gained, with the undermatching and bias the generalized law allows for.[6] No coach was computing logarithms; the law describes what allocation looks like when consequences are doing the selecting.

Classrooms and problem behavior

Martens and Houk observed a student whose disruptive and on-task behavior each drew teacher attention at different rates, and found the two behaviors allocated in proportion to the attention each earned.[7] This is the theoretical spine of differential reinforcement of alternative behavior: to reduce a problem behavior maintained by attention, you do not need to punish it. You need the alternative to pay better — more attention, more reliably, sooner. A large applied literature has since evaluated problem behavior as choice, with concurrent-schedule arrangements that make the appropriate response the richer option.[8]

Everyday human behavior

McDowell argued in 1988 that Herrnstein's hyperbola predicts a common frustration: adding a little reinforcement for a desired behavior has a large effect in an environment that is otherwise barren and almost none in an environment that is already rich. The same praise that transforms a child's behavior in a bleak classroom does nothing in one full of competing reinforcers.[9] Humans, it must be said, match less cleanly than pigeons; people given instructions or forming their own rules about a schedule often follow the rule rather than the contingency, a theme that runs through all human operant research.

Melioration and the mechanism of matching

The matching law is a description, not a mechanism. Two candidate mechanisms competed. Maximizing accounts, borrowed from economics, hold that organisms distribute behavior so as to obtain the most total reinforcement, and on concurrent VI VI schedules matching happens to be nearly optimal.[10] Herrnstein and Vaughan's melioration holds instead that organisms shift behavior toward whichever alternative currently has the higher local rate of return until the local rates are equal — a myopic rule that produces matching without any computation of totals, and that predicts the systematically suboptimal choices people make when a locally better option worsens the long-run outcome.[11] Melioration is one reason the matching law connects so naturally to impulsiveness.

Self-control: when the alternatives differ in time

The most consequential extension of matching is to choices between a smaller reinforcer available sooner and a larger one available later. Rachlin and Green showed in 1972 that pigeons facing that choice directly took the small immediate grain, but that if the choice was made well in advance, the same pigeons committed themselves to the larger, later reward — the first laboratory demonstration of a commitment device.[12] George Ainslie explained why in 1975: if the value of a reinforcer falls with delay along a hyperbola rather than an exponential curve, the curves for a small-soon and a large-late reward cross, so preference reverses as the small reward approaches.[13] James Mazur's adjusting-delay procedure confirmed the hyperbolic shape precisely.[14]

Steep delay discounting — a fast drop in value with delay — has since been documented in people with substance-use disorders, in problem gamblers, and in smokers, and is studied as a process that cuts across many conditions.[15] The everyday translation: the environment that makes you impulsive is one in which the small reward is near and the large reward is far, and the fix is to move the choice point earlier, when the curves have not yet crossed.

Organisms are not uniformly impulsive, though. Cole found that rats on a schedule in which retrieving food pellets from the tray started a one-minute period without further pellets learned to let pellets accumulate and collect them in batches — operant hoarding, a form of self-control the impulsivity findings would not have predicted, and a reminder that the details of the contingency matter.[16]

Behavioral economics: demand, price, and elasticity

Once behavior is allocation, the tools of economics apply. Steven Hursh proposed in 1980 that a schedule requirement is a price (responses per reinforcer), that consumption plotted against price gives a demand curve, and that the slope of that curve — elasticity — measures how essential a reinforcer is.[17] Food in a closed economy, where the animal earns all of its food in the chamber, is inelastic: raise the price and the animal works harder to keep consumption up. Sweetened water in an open economy is elastic: raise the price and consumption collapses. Whether reinforcers are substitutes (one replaces another) or complements (consumed together) can be measured the same way.[18] Kagel, Battalio, and Green showed, in a research program running from the 1970s onward, that rats and pigeons obey demand theory in detail, including some of its odder predictions, such as Giffen goods.[19]

The approach has direct policy uses. Demand curves for cigarettes, alcohol, and drugs measured in the laboratory predict how consumption responds to taxation, and "essential value" derived from demand analysis compares the reinforcing efficacy of drugs on a common scale.[20] It also explains a stubborn feature of behavior change: a reinforcer you offer competes in a market, and if the problem behavior is a cheap, inelastic, non-substitutable good, small incentives for the alternative will not move it.

Limits and criticisms

Within those limits, the matching law is the closest thing behavior analysis has to a physical law. It made choice measurable, connected the laboratory to economics, and gave clinicians a simple instruction that holds up: to change what someone does, change what the alternatives pay.

Key takeaways

Check yourself

A pigeon can peck key A, which pays about 30 times an hour, or key B, which pays about 10. A classmate predicts the bird will peck A almost exclusively, since A is plainly better. What does the matching law predict?

About 75% of pecks on A and 25% on B, because A delivers 30 out of every 40 reinforcers. On concurrent variable-interval schedules behavior is spread in proportion to payoff rather than piled onto the best option; the poorer key keeps a share proportional to what it pays.

A praise program that transformed a student's behavior in one classroom does nothing in another. Was the praise too weak?

Not necessarily. In Herrnstein's hyperbola a behavior's strength depends on its reinforcement relative to the reinforcement for everything else, so adding a little reinforcement has a large effect in a barren environment and almost none in one already full of competing reinforcers. The praise is the same; the alternatives are not.

A student's disruptive behavior earns teacher attention more reliably than on-task behavior does. Without punishing anything, how does the matching law say to reduce the disruption?

Make the alternative pay better: attend to on-task behavior more, more reliably, and sooner, so that it earns the larger share of attention and therefore draws the larger share of behavior. This is differential reinforcement of alternative behavior, and the matching law is its theoretical spine.

Pigeons choosing between a small immediate grain and a larger delayed one take the small one, yet when the same choice is offered well in advance they commit to the larger one. Why does preference reverse?

Because value falls with delay along a hyperbola rather than an exponential curve, the value curves for the small-soon and large-late rewards cross. Far from both rewards the larger one is worth more; as the small one becomes imminent it overtakes, so choosing early, before the curves cross, is a commitment device.

Explain it to a friend. Explain the matching law to someone who dislikes math, using either the basketball or the classroom example and no numbers at all.

Frequently asked questions

What is the matching law in simple terms?

Organisms spread their behavior across options in proportion to how much reinforcement each option provides. If one option delivers twice as much reinforcement as another, it gets about twice as much behavior. Richard Herrnstein discovered it in pigeons in 1961, and it holds, with some systematic deviations, across species and settings.

What is the generalized matching law?

Baum's 1974 extension, which adds two parameters: sensitivity (how strongly behavior tracks reinforcement; usually a little less than 1, called undermatching) and bias (a constant preference for one option unrelated to reinforcement). It is written as a straight line in logarithmic ratios and fits most choice data.

How does the matching law explain problem behavior?

Problem behavior and appropriate behavior are alternatives on a concurrent schedule. If misbehavior earns attention more reliably than good behavior, the matching law predicts a lot of misbehavior. The treatment is to make the appropriate alternative pay more — differential reinforcement of alternative behavior — rather than to punish the problem.

What is the difference between the matching law and the law of effect?

Thorndike's law of effect says responses followed by satisfying consequences are strengthened. Herrnstein's matching law quantifies it: response strength is relative, depending on the reinforcement for a behavior compared with the reinforcement for everything else. Herrnstein proposed the hyperbolic form of matching as the quantitative law of effect.

Does the matching law apply to humans?

Yes, though less cleanly. Sports play-calling, classroom behavior, and conversation have been shown to match. Human deviations mostly come from rules and instructions: people who can describe a schedule often follow their description rather than the contingency.

What does the matching law have to do with self-control?

Choices between a smaller, sooner reward and a larger, later one are matching choices in which delay reduces value. Because value falls hyperbolically with delay, preference flips toward the small reward as it becomes imminent. Making the choice early, before the flip, is what commitment devices do.

References

  1. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267–272.
  2. Herrnstein, R. J. (1970). On the law of effect. Journal of the Experimental Analysis of Behavior, 13(2), 243–266.
  3. Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. Journal of the Experimental Analysis of Behavior, 22(1), 231–242.
  4. Baum, W. M. (1979). Matching, undermatching, and overmatching in studies of choice. Journal of the Experimental Analysis of Behavior, 32(2), 269–281.
  5. Vollmer, T. R., & Bourret, J. (2000). An application of the matching law to evaluate the allocation of two- and three-point shots by college basketball players. Journal of Applied Behavior Analysis, 33(2), 137–150.
  6. Reed, D. D., Critchfield, T. S., & Martens, B. K. (2006). The generalized matching law in elite sport competition: Football play calling as operant choice. Journal of Applied Behavior Analysis, 39(3), 281–297.
  7. Martens, B. K., & Houk, J. L. (1989). The application of Herrnstein's law of effect to disruptive and on-task behavior of a retarded adolescent girl. Journal of the Experimental Analysis of Behavior, 51(1), 17–27.
  8. Fisher, W. W., & Mazur, J. E. (1997). Basic and applied research on choice responding. Journal of Applied Behavior Analysis, 30(3), 387–410.
  9. McDowell, J. J. (1988). Matching theory in natural human environments. The Behavior Analyst, 11(2), 95–109.
  10. Rachlin, H., Green, L., Kagel, J. H., & Battalio, R. C. (1976). Economic demand theory and psychological studies of choice. In G. H. Bower (Ed.), The Psychology of Learning and Motivation (Vol. 10). Academic Press.
  11. Herrnstein, R. J., & Vaughan, W. (1980). Melioration and behavioral allocation. In J. E. R. Staddon (Ed.), Limits to Action: The Allocation of Individual Behavior. Academic Press.
  12. Rachlin, H., & Green, L. (1972). Commitment, choice and self-control. Journal of the Experimental Analysis of Behavior, 17(1), 15–22.
  13. Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. Psychological Bulletin, 82(4), 463–496.
  14. Mazur, J. E. (1987). An adjusting procedure for studying delayed reinforcement. In M. L. Commons, J. E. Mazur, J. A. Nevin, & H. Rachlin (Eds.), Quantitative Analyses of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value (pp. 55–73). Erlbaum.
  15. Bickel, W. K., & Marsch, L. A. (2001). Toward a behavioral economic understanding of drug dependence: Delay discounting processes. Addiction, 96(1), 73–86.
  16. Cole, M. R. (1990). Operant hoarding: A new paradigm for the study of self-control. Journal of the Experimental Analysis of Behavior, 53(2), 247–262.
  17. Hursh, S. R. (1980). Economic concepts for the analysis of behavior. Journal of the Experimental Analysis of Behavior, 34(2), 219–238.
  18. Hursh, S. R. (1984). Behavioral economics. Journal of the Experimental Analysis of Behavior, 42(3), 435–452.
  19. Kagel, J. H., Battalio, R. C., & Green, L. (1995). Economic Choice Theory: An Experimental Analysis of Animal Behavior. Cambridge University Press.
  20. Hursh, S. R., & Silberberg, A. (2008). Economic demand and essential value. Psychological Review, 115(1), 186–198.
  21. Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour (pp. 159–192). Wiley.