Choice · How time gets divided
The Matching Law
Every behavior is a choice among alternatives, and organisms divide their behavior among alternatives in proportion to what each one pays. That single regularity, found in pigeons in 1961, now predicts basketball shot selection, classroom disruption, and why we buy things we do not need.
Definition
The matching law states that when two or more responses are available, the relative rate of each response equals — matches — the relative rate of reinforcement it produces. For two alternatives:
where B is the rate of each behavior and R the rate of reinforcement it earns. An option that delivers 70% of the available reinforcement gets about 70% of the behavior.[1]
A worked example. A pigeon can peck either of two keys. Key A pays off about 30 times an hour, key B about 10 times an hour, so A delivers 30 ÷ (30 + 10) = 75% of the reinforcement. The matching law predicts that the pigeon will make about 75% of its pecks on A and 25% on B — not all of them on A, even though A is plainly better. That the pigeon does not simply pick the better key every time, and that the poorer option keeps a share proportional to what it pays, is the surprise of the law: on schedules like these, behavior is spread in proportion to payoff rather than piled onto the best option.
In brief
- Organisms allocate behavior in proportion to reinforcement: an option that delivers 70% of the available reinforcement gets about 70% of the behavior.
- A behavior's strength depends not only on what it earns but on what everything else earns, so enriching the alternatives reduces it without punishment.
- Because a reinforcer's value falls hyperbolically with delay, preference flips toward a small, soon reward as it approaches; commitment means choosing before the flip.
Herrnstein's experiment
Richard Herrnstein put pigeons in a chamber with two response keys, each paying grain on its own variable-interval schedule — a concurrent VI VI schedule. A bird could switch between keys whenever it liked. He varied how the reinforcement was divided between the keys while holding the total constant (VI 3-minute against VI 3-minute, VI 2.25 against VI 4.5, and so on), and to stop the birds simply alternating, he added a changeover delay: a brief period after each switch during which no reinforcement could be collected.[1]
The result was startlingly orderly. Plot the proportion of pecks on the left key against the proportion of reinforcers earned on the left key and the points fall on the diagonal. Birds did not go exclusively to the richer key, and they did not split their pecks evenly; they matched.
From matching to a new law of effect
In 1970 Herrnstein drew the larger conclusion. If behavior is allocated in proportion to reinforcement, then even a single response in a Skinner box is a choice — between pressing the lever and everything else the rat could do (grooming, exploring, resting), all of which produce some reinforcement of their own. Writing the "everything else" reinforcement as Re, the matching law for one measured response becomes a hyperbola:
Response rate rises with reinforcement rate but with diminishing returns, leveling off at a maximum k. This fit decades of single-schedule data that the original law of effect had described only qualitatively, and Herrnstein proposed it as the quantitative law of effect.[2] It also carried a practical implication: a behavior's strength depends not just on what it earns but on what everything else earns. Enrich the alternatives and the behavior declines without any punishment at all.
The generalized matching law
Real organisms do not match perfectly. William Baum showed in 1974 that the departures are systematic and can be captured by two parameters. Taking logarithms of the ratio form:
The slope a is sensitivity. When a = 1, matching is perfect; the usual finding is undermatching, a slope around 0.8, meaning organisms are somewhat less extreme in their preference than the reinforcement ratio warrants. The intercept b is bias: a constant preference for one alternative unrelated to reinforcement — a key that is easier to reach, a side the animal favors.[3][4] The generalized form has been fitted to hundreds of data sets across species, and its parameters turn out to be sensitive to procedural details in interpretable ways: shorter changeover delays produce more undermatching, for instance, because switching itself is reinforced.
Reinforcement rate is not the only thing organisms match to. Reinforcer magnitude, delay, and quality all enter the equation, which is what makes the framework a general account of choice rather than a fact about grain.[4]
Matching beyond the pigeon
Sports
A basketball player choosing between a two-point and a three-point attempt is on a concurrent schedule. Vollmer and Bourret analyzed a season of college basketball and found that the proportion of three-point shots taken by teams and by individual players matched the proportion of points those shots produced.[5] Reed, Critchfield, and Martens applied the generalized matching law to NFL play-calling and found that the ratio of passing to rushing plays tracked the ratio of yards each type gained, with the undermatching and bias the generalized law allows for.[6] No coach was computing logarithms; the law describes what allocation looks like when consequences are doing the selecting.
Classrooms and problem behavior
Martens and Houk observed a student whose disruptive and on-task behavior each drew teacher attention at different rates, and found the two behaviors allocated in proportion to the attention each earned.[7] This is the theoretical spine of differential reinforcement of alternative behavior: to reduce a problem behavior maintained by attention, you do not need to punish it. You need the alternative to pay better — more attention, more reliably, sooner. A large applied literature has since evaluated problem behavior as choice, with concurrent-schedule arrangements that make the appropriate response the richer option.[8]
Everyday human behavior
McDowell argued in 1988 that Herrnstein's hyperbola predicts a common frustration: adding a little reinforcement for a desired behavior has a large effect in an environment that is otherwise barren and almost none in an environment that is already rich. The same praise that transforms a child's behavior in a bleak classroom does nothing in one full of competing reinforcers.[9] Humans, it must be said, match less cleanly than pigeons; people given instructions or forming their own rules about a schedule often follow the rule rather than the contingency, a theme that runs through all human operant research.
Melioration and the mechanism of matching
The matching law is a description, not a mechanism. Two candidate mechanisms competed. Maximizing accounts, borrowed from economics, hold that organisms distribute behavior so as to obtain the most total reinforcement, and on concurrent VI VI schedules matching happens to be nearly optimal.[10] Herrnstein and Vaughan's melioration holds instead that organisms shift behavior toward whichever alternative currently has the higher local rate of return until the local rates are equal — a myopic rule that produces matching without any computation of totals, and that predicts the systematically suboptimal choices people make when a locally better option worsens the long-run outcome.[11] Melioration is one reason the matching law connects so naturally to impulsiveness.
Self-control: when the alternatives differ in time
The most consequential extension of matching is to choices between a smaller reinforcer available sooner and a larger one available later. Rachlin and Green showed in 1972 that pigeons facing that choice directly took the small immediate grain, but that if the choice was made well in advance, the same pigeons committed themselves to the larger, later reward — the first laboratory demonstration of a commitment device.[12] George Ainslie explained why in 1975: if the value of a reinforcer falls with delay along a hyperbola rather than an exponential curve, the curves for a small-soon and a large-late reward cross, so preference reverses as the small reward approaches.[13] James Mazur's adjusting-delay procedure confirmed the hyperbolic shape precisely.[14]
Steep delay discounting — a fast drop in value with delay — has since been documented in people with substance-use disorders, in problem gamblers, and in smokers, and is studied as a process that cuts across many conditions.[15] The everyday translation: the environment that makes you impulsive is one in which the small reward is near and the large reward is far, and the fix is to move the choice point earlier, when the curves have not yet crossed.
Organisms are not uniformly impulsive, though. Cole found that rats on a schedule in which retrieving food pellets from the tray started a one-minute period without further pellets learned to let pellets accumulate and collect them in batches — operant hoarding, a form of self-control the impulsivity findings would not have predicted, and a reminder that the details of the contingency matter.[16]
Behavioral economics: demand, price, and elasticity
Once behavior is allocation, the tools of economics apply. Steven Hursh proposed in 1980 that a schedule requirement is a price (responses per reinforcer), that consumption plotted against price gives a demand curve, and that the slope of that curve — elasticity — measures how essential a reinforcer is.[17] Food in a closed economy, where the animal earns all of its food in the chamber, is inelastic: raise the price and the animal works harder to keep consumption up. Sweetened water in an open economy is elastic: raise the price and consumption collapses. Whether reinforcers are substitutes (one replaces another) or complements (consumed together) can be measured the same way.[18] Kagel, Battalio, and Green showed, in a research program running from the 1970s onward, that rats and pigeons obey demand theory in detail, including some of its odder predictions, such as Giffen goods.[19]
The approach has direct policy uses. Demand curves for cigarettes, alcohol, and drugs measured in the laboratory predict how consumption responds to taxation, and "essential value" derived from demand analysis compares the reinforcing efficacy of drugs on a common scale.[20] It also explains a stubborn feature of behavior change: a reinforcer you offer competes in a market, and if the problem behavior is a cheap, inelastic, non-substitutable good, small incentives for the alternative will not move it.
Limits and criticisms
- Matching is descriptive. It tells you the outcome of allocation, not how the organism gets there; melioration, maximizing, and momentary-maximizing accounts all reproduce it under various conditions and are hard to separate.
- Humans often follow rules instead. Verbal instructions and self-generated rules can override contingencies, so human matching is weaker and more variable than animal matching unless the schedule is hard to describe.[21]
- Ratio schedules break the pattern. On concurrent ratio schedules, exclusive preference for the better option is the optimal strategy and is what animals do, so matching in its simple form applies mainly to interval schedules, where spreading behavior across options pays.
- Parameters need estimating. The generalized law fits almost anything with a free slope and intercept; its value lies in the parameters being stable and interpretable, which they generally are, not in the fit alone.
Within those limits, the matching law is the closest thing behavior analysis has to a physical law. It made choice measurable, connected the laboratory to economics, and gave clinicians a simple instruction that holds up: to change what someone does, change what the alternatives pay.
Key takeaways
- On concurrent variable-interval schedules, organisms do not pile behavior onto the best option; they spread it in proportion to payoff. Herrnstein's pigeons matched the proportion of pecks on each key to the proportion of reinforcers it delivered.
- Even a single response is a choice against everything else the organism could do. Herrnstein's hyperbola makes response rate rise with reinforcement at diminishing returns, which is why the same praise transforms behavior in a barren environment and does almost nothing in a rich one.
- Real organisms deviate systematically. The generalized matching law adds sensitivity (usually undermatching, a slope around 0.8) and bias (a constant preference unrelated to reinforcement), and reinforcer magnitude, delay, and quality enter the equation alongside rate.
- Matching is a description, not a mechanism; melioration, which shifts behavior toward the locally richer option, is one candidate. Humans match less cleanly because rules can override contingencies, and on concurrent ratio schedules exclusive preference for the better option is optimal and is what animals do.
- The practical instruction: to change what someone does, change what the alternatives pay. Problem behavior maintained by attention is reduced by making the appropriate alternative pay better, and impulsive choices are avoided by moving the choice point earlier, before the value curves cross.
Check yourself
A pigeon can peck key A, which pays about 30 times an hour, or key B, which pays about 10. A classmate predicts the bird will peck A almost exclusively, since A is plainly better. What does the matching law predict?
About 75% of pecks on A and 25% on B, because A delivers 30 out of every 40 reinforcers. On concurrent variable-interval schedules behavior is spread in proportion to payoff rather than piled onto the best option; the poorer key keeps a share proportional to what it pays.
A praise program that transformed a student's behavior in one classroom does nothing in another. Was the praise too weak?
Not necessarily. In Herrnstein's hyperbola a behavior's strength depends on its reinforcement relative to the reinforcement for everything else, so adding a little reinforcement has a large effect in a barren environment and almost none in one already full of competing reinforcers. The praise is the same; the alternatives are not.
A student's disruptive behavior earns teacher attention more reliably than on-task behavior does. Without punishing anything, how does the matching law say to reduce the disruption?
Make the alternative pay better: attend to on-task behavior more, more reliably, and sooner, so that it earns the larger share of attention and therefore draws the larger share of behavior. This is differential reinforcement of alternative behavior, and the matching law is its theoretical spine.
Pigeons choosing between a small immediate grain and a larger delayed one take the small one, yet when the same choice is offered well in advance they commit to the larger one. Why does preference reverse?
Because value falls with delay along a hyperbola rather than an exponential curve, the value curves for the small-soon and large-late rewards cross. Far from both rewards the larger one is worth more; as the small one becomes imminent it overtakes, so choosing early, before the curves cross, is a commitment device.
Explain it to a friend. Explain the matching law to someone who dislikes math, using either the basketball or the classroom example and no numbers at all.
Frequently asked questions
What is the matching law in simple terms?
Organisms spread their behavior across options in proportion to how much reinforcement each option provides. If one option delivers twice as much reinforcement as another, it gets about twice as much behavior. Richard Herrnstein discovered it in pigeons in 1961, and it holds, with some systematic deviations, across species and settings.
What is the generalized matching law?
Baum's 1974 extension, which adds two parameters: sensitivity (how strongly behavior tracks reinforcement; usually a little less than 1, called undermatching) and bias (a constant preference for one option unrelated to reinforcement). It is written as a straight line in logarithmic ratios and fits most choice data.
How does the matching law explain problem behavior?
Problem behavior and appropriate behavior are alternatives on a concurrent schedule. If misbehavior earns attention more reliably than good behavior, the matching law predicts a lot of misbehavior. The treatment is to make the appropriate alternative pay more — differential reinforcement of alternative behavior — rather than to punish the problem.
What is the difference between the matching law and the law of effect?
Thorndike's law of effect says responses followed by satisfying consequences are strengthened. Herrnstein's matching law quantifies it: response strength is relative, depending on the reinforcement for a behavior compared with the reinforcement for everything else. Herrnstein proposed the hyperbolic form of matching as the quantitative law of effect.
Does the matching law apply to humans?
Yes, though less cleanly. Sports play-calling, classroom behavior, and conversation have been shown to match. Human deviations mostly come from rules and instructions: people who can describe a schedule often follow their description rather than the contingency.
What does the matching law have to do with self-control?
Choices between a smaller, sooner reward and a larger, later one are matching choices in which delay reduces value. Because value falls hyperbolically with delay, preference flips toward the small reward as it becomes imminent. Making the choice early, before the flip, is what commitment devices do.
References
- Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267–272.
- Herrnstein, R. J. (1970). On the law of effect. Journal of the Experimental Analysis of Behavior, 13(2), 243–266.
- Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. Journal of the Experimental Analysis of Behavior, 22(1), 231–242.
- Baum, W. M. (1979). Matching, undermatching, and overmatching in studies of choice. Journal of the Experimental Analysis of Behavior, 32(2), 269–281.
- Vollmer, T. R., & Bourret, J. (2000). An application of the matching law to evaluate the allocation of two- and three-point shots by college basketball players. Journal of Applied Behavior Analysis, 33(2), 137–150.
- Reed, D. D., Critchfield, T. S., & Martens, B. K. (2006). The generalized matching law in elite sport competition: Football play calling as operant choice. Journal of Applied Behavior Analysis, 39(3), 281–297.
- Martens, B. K., & Houk, J. L. (1989). The application of Herrnstein's law of effect to disruptive and on-task behavior of a retarded adolescent girl. Journal of the Experimental Analysis of Behavior, 51(1), 17–27.
- Fisher, W. W., & Mazur, J. E. (1997). Basic and applied research on choice responding. Journal of Applied Behavior Analysis, 30(3), 387–410.
- McDowell, J. J. (1988). Matching theory in natural human environments. The Behavior Analyst, 11(2), 95–109.
- Rachlin, H., Green, L., Kagel, J. H., & Battalio, R. C. (1976). Economic demand theory and psychological studies of choice. In G. H. Bower (Ed.), The Psychology of Learning and Motivation (Vol. 10). Academic Press.
- Herrnstein, R. J., & Vaughan, W. (1980). Melioration and behavioral allocation. In J. E. R. Staddon (Ed.), Limits to Action: The Allocation of Individual Behavior. Academic Press.
- Rachlin, H., & Green, L. (1972). Commitment, choice and self-control. Journal of the Experimental Analysis of Behavior, 17(1), 15–22.
- Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. Psychological Bulletin, 82(4), 463–496.
- Mazur, J. E. (1987). An adjusting procedure for studying delayed reinforcement. In M. L. Commons, J. E. Mazur, J. A. Nevin, & H. Rachlin (Eds.), Quantitative Analyses of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value (pp. 55–73). Erlbaum.
- Bickel, W. K., & Marsch, L. A. (2001). Toward a behavioral economic understanding of drug dependence: Delay discounting processes. Addiction, 96(1), 73–86.
- Cole, M. R. (1990). Operant hoarding: A new paradigm for the study of self-control. Journal of the Experimental Analysis of Behavior, 53(2), 247–262.
- Hursh, S. R. (1980). Economic concepts for the analysis of behavior. Journal of the Experimental Analysis of Behavior, 34(2), 219–238.
- Hursh, S. R. (1984). Behavioral economics. Journal of the Experimental Analysis of Behavior, 42(3), 435–452.
- Kagel, J. H., Battalio, R. C., & Green, L. (1995). Economic Choice Theory: An Experimental Analysis of Animal Behavior. Cambridge University Press.
- Hursh, S. R., & Silberberg, A. (2008). Economic demand and essential value. Psychological Review, 115(1), 186–198.
- Lowe, C. F. (1979). Determinants of human operant behaviour. In M. D. Zeiler & P. Harzem (Eds.), Advances in Analysis of Behaviour: Vol. 1. Reinforcement and the Organization of Behaviour (pp. 159–192). Wiley.