Reinforcement · Types of reinforcers

Primary and Secondary Reinforcers: What Is Built In, What Is Learned, and Why Both Fail

Food reinforces a hungry rat because of what a rat is. A click reinforces the same rat because of what the click has come to predict. That difference, built in versus learned, sorts every reinforcer into one of two kinds, and it explains why praise fails with some children, why poker chips can teach a chimpanzee, and why a clicker in a trainer's hand can mark a behavior faster than any treat.

Updated 13 min read

Definition

A primary reinforcer (also called an unconditioned reinforcer) is a stimulus that strengthens the behavior it follows without any prior learning: food, water, warmth, sleep, sexual contact, and escape from pain are the standard examples. A secondary reinforcer (also called a conditioned reinforcer) is a stimulus that has acquired the power to strengthen behavior through its relation to other reinforcers: a magazine click, praise, a grade, money, a token. Both are reinforcers only if they do the job: the test is the effect on behavior, not the biology or the history.[1]

"Secondary" does not mean weaker. Most of what reinforces adult human behavior is conditioned, and money keeps working even when nothing behind it is currently wanted.

In brief

  • A primary reinforcer works without a learning history, but not without conditions: its power rises with deprivation and falls with satiation, so "unlearned" never means "always effective."
  • A secondary reinforcer earns its function by predicting another reinforcer, keeps it only while the prediction holds, and loses it by extinction when the pairing stops.
  • Generalized reinforcers such as money, tokens, and attention are backed by many reinforcers at once, which is why they work almost regardless of the person's state, and why a token economy collapses when the tokens stop buying anything.

Primary reinforcers: unlearned, but not unconditional

A primary reinforcer works the first time. A rat that has never seen a lever will eat the pellet that arrives after it presses, and pressing will increase; no pairing is needed, because the reinforcing effect of food is part of what a rat is. The standard list is short: food, water, warmth when cold and cooling when hot, sleep, sexual contact, and the removal of pain or other aversive stimulation, which is the negative reinforcement side of the same list.[1] For some species physical contact belongs on it too.

Unlearned does not mean always effective. Food does not reinforce a rat that has just eaten. Skinner treated the rat's hunger as an experimental variable in its own right, set by hours of food deprivation, and reported that the rate of lever pressing rose and fell with it.[2] A primary reinforcer's power is a property of the stimulus and the organism's current state together, and a trainer who works after dinner with kibble has changed the second without noticing.

The modern name for that state is the motivating operation. Jack Michael defined an establishing operation as an event that does two things at once: it raises the effectiveness of some stimulus as a reinforcer, and it raises the frequency of whatever behavior has produced that stimulus in the past.[3] Deprivation is the establishing operation for food; satiation is its opposite, an abolishing operation; the two were later grouped under the term motivating operation.[4]

A list of effects, not of pleasures

The list of primary reinforcers is a list of things shown to strengthen behavior without a learning history, not of things that feel good. Escape from pain reinforces without anyone enjoying the pain, and Olds and Milner's rats, below, pressed for a pulse of electricity that answers no known need.

Secondary reinforcers: how a click comes to count

A secondary reinforcer starts as a stimulus that does nothing. The sound of the food magazine in a Skinner box means nothing to a naive rat. In The Behavior of Organisms Skinner described what happens once the sound has been followed by food a number of times: the rat comes to the tray at the click, and the click itself, delivered after a lever press with no food behind it, is enough to strengthen pressing, for a while. As clicks keep arriving without food, the effect fades.[2] That fading is the signature of a conditioned reinforcer: its function is borrowed, and it is withdrawn when the relation that produced it is broken.

The procedure here is pairing: a neutral stimulus is followed by a reinforcer. The process is the change in what the stimulus does, and it is the only evidence that a conditioned reinforcer exists. Trainers who "charge" a clicker are running the procedure; the dog's head snapping toward them at the sound is the first sign of the process. Magazine training is step one in the virtual Skinner box.

Paul Bersh measured how much pairing it takes. Rats received a light followed by food, with the number of pairings and the light–food interval varied across groups; the light was then tested by making it the only consequence of lever pressing. More pairings produced a stronger conditioned reinforcer, up to a limit, and a light that came on shortly before the food worked far better than one that came on long before it.[5] The order is the one classical conditioning requires. How the two kinds of conditioning combine ›

Kelleher and Gollub's review sorted the evidence by method: maintaining an old response in extinction with the stimulus still following it, or teaching a new response with the stimulus as its only consequence. Both methods demonstrated conditioned reinforcement, and both showed it to be short-lived when the stimulus was never paired again, because the test itself extinguishes the pairing.[6] Zimmerman's answer was to establish the stimulus intermittently: a buzzer followed by water only some of the time, for thirsty rats, went on to reinforce a new bar-pressing response that persisted far longer than earlier methods had achieved.[7] The reason is familiar from schedules of reinforcement: a relation that was never reliable is slow to be detected as broken.

Pairing is not enough: prediction, information, and dopamine

Pairing is the procedure, but not the variable that matters. Egger and Miller gave rats a food pellet after a sequence of two stimuli. For one group the first stimulus reliably predicted the food, so the second added nothing; for another the first stimulus was unreliable, so the second carried the news. The second stimulus had been paired with food equally often in both groups, and it was a far stronger reinforcer when it was the informative one.[8] A stimulus becomes a conditioned reinforcer to the extent that it predicts a reinforcer, not merely accompanies one.

Edmund Fantino's delay-reduction hypothesis makes the point quantitative: a stimulus reinforces to the extent that its onset signals a reduction in the time to the next primary reinforcer, relative to the average wait in that situation. A stimulus that means "food soon" when food is usually far off is a strong conditioned reinforcer; one that means "no food" is not a reinforcer at all, however informative. In the observing-response experiments he reviewed, pigeons work to produce a stimulus that tells them which schedule is in effect only when it can announce the better outcome.[9] Good news reinforces; news does not. Ben Williams's review reached the same place: conditioned reinforcement is a sound and necessary concept, but it is a statement about predictive relations.[10]

The brain agrees, up to a point

Schultz, Dayan, and Montague recorded midbrain dopamine neurons in monkeys while a cue came to predict a squirt of juice. Early in training the neurons fired at the juice; once the cue reliably predicted it, the burst moved to the cue, and the fully predicted juice evoked nothing.[11] A predictive cue inherits the signal, which is close to a physiological description of a conditioned reinforcer. Olds and Milner's rats, which pressed a lever for nothing but a brief pulse of electrical stimulation to the septal area, mark the other boundary: a reinforcer that needs no learning history and serves no known biological need.[12] The full account is on the neuroscience page.

Generalized reinforcers: money, tokens, and attention

A conditioned reinforcer backed by one primary reinforcer inherits that reinforcer's motivating operations: a click backed only by food is worth little to a full dog. Skinner's explanation of why money, attention, and approval work almost all the time was the generalized reinforcer, a conditioned reinforcer paired with many different reinforcers, so that at any moment at least one of the deprivations behind it is likely to be in force. Money buys food, warmth, shelter, and company; attention is the precondition for everything another person might provide. Skinner added that a generalized reinforcer can eventually work even when none of the reinforcers behind it is currently relevant, which is as close as behavior analysis comes to explaining the miser.[1]

The laboratory version is the token. John Wolfe taught chimpanzees to operate a weighted lever to earn poker chips, which they could insert into a vending machine, the "Chimp-O-Mat," for grapes. The chimps worked for chips, preferred a chip that bought two grapes to one that bought one, learned to ignore a chip that bought nothing, and kept working when the chips could not be cashed in until later, though longer delays weakened the effect.[13] John Cowles then showed that tokens alone, exchanged for food only after a run of trials, could sustain the learning of new discriminations.[14] A token is a conditioned reinforcer that can be carried, counted, and saved, which is most of what money is.

Timothy Hackenberg's review draws the modern picture: a token system is three schedules at once, one for earning tokens, one for reaching the exchange period, and one for exchanging, and the tokens' power depends on all three, above all on how reliably and how soon exchange happens.[15] A token economy in a classroom or a ward is the same machinery, and it fails the same way: when the exchange stops, the tokens become paper, and the behavior they were maintaining goes with them.

Three cases that strain the categories

The primary/secondary distinction is clean in the laboratory and leaky everywhere else. Three cases show the seams.

Social reinforcers and Harlow's caution

Praise, attention, a smile, being listened to: social reinforcers are usually filed under "conditioned," on the theory that a caregiver was paired with food and warmth from birth. Harry Harlow's surrogate-mother experiments showed that this cannot be the whole story. Infant rhesus monkeys raised with two artificial mothers, one of bare wire that held the milk bottle and one covered in soft cloth that gave no milk, spent most of their time clinging to the cloth mother and ran to her when frightened, whichever mother had fed them.[16] Contact comfort reinforced on its own, not as a by-product of feeding. The category of a reinforcer is an empirical question: some social reinforcers are primary for some species, some are conditioned, and some are not reinforcers at all for a particular person.

The clicker

A clicker is a conditioned reinforcer built on purpose. Karen Pryor, who learned the method training dolphins with a whistle, describes charging the clicker by pairing it with food and then using the click to mark the exact instant of the behavior, with the food following a moment later.[17] Feng, Howell, and Bennett compared the three explanations on offer, that the click reinforces, that it marks, and that it bridges, and concluded that the published studies cannot yet tell them apart; direct comparisons of clicker-plus-food against food alone are few and mixed.[18] What is not in doubt is how the click acquires its function. Marker training in practice ›

Activity reinforcers

David Premack's work sits across the distinction. In his account a reinforcer is not a stimulus but a behavior, eating rather than food, and any behavior can reinforce a less probable one.[19] Running needs no pairing to reinforce drinking in a rat that has been kept from running, so it behaves like a primary reinforcer without being a biological necessity. The Premack principle page covers the experiments; here it is a reminder that the primary/secondary sorting was built for stimuli, and a good deal of what reinforces is not a stimulus.

Types of reinforcers and how each fails

When a reinforcer stops working, its type tells you where to look; the last column is the practical one.

TypeWhere the function comes fromExamplesHow it fails
Primary (unconditioned)Biology; no learning history neededFood, water, warmth, sleep, sexual contact, escape from pain, contact comfort in infant primatesSatiation, or the wrong motivating operation: kibble after dinner, a blanket in July
Secondary (conditioned)Pairing with, and prediction of, another reinforcerThe magazine click, a clicker, a marker word, praise, a grade, a check markExtinction when the pairing stops; satiation on the reinforcer behind it; redundancy when it predicts nothing new
GeneralizedPairing with many reinforcersMoney, tokens, points, attention, approvalLoss of backing: a token that never buys anything, points with an empty menu, attention that is free anyway
SocialDelivered by another person; primary or conditioned depending on the reinforcer and the speciesA smile, thanks, being listened to, physical contactNot a reinforcer for this person; delivered whether or not the behavior occurs; paired with criticism until it becomes a warning
ActivityThe opportunity to perform a more probable behaviorPlay after homework, the walk after the sitThe requirement deprives nobody; the activity has become less probable through satiation

"Social" and "activity" describe what form a reinforcer takes; "primary," "secondary," and "generalized" describe where its function comes from. A smile can be any of the latter three.

How to build a conditioned reinforcer

The procedure is the same for a rat, a dog, a first-grader, or you, and the same variables govern it: order, interval, number of pairings, and prediction.

  1. Choose a stimulus that is brief, distinct, and otherwise meaningless. A click, a short word you do not use in conversation, a specific check mark, a chip; not a word the learner already hears forty times a day for free.
  2. Confirm that the backing reinforcer works right now. A conditioned reinforcer can only be as strong as what stands behind it at the moment of pairing. Pair before the meal, not after it, and use something the learner would actually choose.
  3. Stimulus first, then reinforcer, within about a second. A light that came on shortly before food became a stronger conditioned reinforcer than one that came on long before, and more pairings made a stronger one up to a limit.[5] A few dozen pairings across two or three short sessions is a reasonable start; it has worked when the learner orients to the stimulus before the reinforcer appears.
  4. Make it predictive, not merely frequent. No backing reinforcer without the stimulus, no stimulus without the backing reinforcer. A redundant stimulus, one whose reinforcer was already predicted by something else, acquires little function no matter how often it is paired.[8]
  5. Use it to mark, then pay. The stimulus goes at the instant the behavior occurs; the backing reinforcer follows. A conditioned reinforcer can be delivered with a precision the primary reinforcer cannot match, which is what makes shaping possible.
  6. Keep it backed. Every click earns a treat while a behavior is being taught; tokens are exchanged on a schedule the learner can count on. Intermittent pairing makes a conditioned reinforcer more durable, but nothing makes it permanent.[7] If the learner stops orienting to the stimulus, recharge it.

Common mistakes

What the evidence does not show

The literature on conditioned reinforcement is large, and several claims made in its name go beyond it.

Key takeaways

Check yourself

A teacher hands out points every day for on-task behavior. For the first month, points buy ten minutes of free time on Friday; then the free-time period is dropped but the points continue. On-task behavior falls back to where it started. What happened to the points?

The points were a conditioned reinforcer backed by free time. When the exchange stopped, the pairing stopped, and the points' function extinguished; they went on being delivered but predicted nothing. This is the standard failure mode of a token system, and the remedy is to restore the backing, not to hand out more points.

A trainer charges a clicker with treats for two sessions, then starts using it, but also keeps tossing the dog treats "to keep him motivated" whether or not a click has sounded. Within a week the dog barely reacts to the click. Why?

The click stopped predicting anything. Treats that arrive without a click make the click redundant, and Egger and Miller showed that a redundant stimulus acquires little reinforcing function even when it has been paired with food just as often as an informative one. Restore the rule that treats follow clicks and only clicks, and recharge.

Is a warm blanket a primary or a secondary reinforcer?

Warmth is a primary reinforcer: it strengthens behavior in a cold organism with no learning history. But whether the blanket reinforces anything right now depends on the motivating operation, and in a warm room it will not. "Primary" describes where the function comes from, not whether it is currently in force.

Want more? The 20-question quiz covers every page on the site with instant explanations.

Frequently asked questions

What are primary and secondary reinforcers in simple terms?

A primary reinforcer works without being learned: food, water, warmth, sleep, escape from pain. A secondary (conditioned) reinforcer has to be learned; it works because it has come to predict something else that reinforces, the way a clicker predicts a treat or money predicts everything money buys. Both are defined by their effect: if the behavior does not increase, neither one is a reinforcer.

What is the difference between a primary and a secondary reinforcer?

History. A primary reinforcer strengthens behavior the first time it is delivered, because of the organism's biology. A secondary reinforcer starts out neutral and acquires its power by being paired with, and coming to predict, another reinforcer. Both depend on conditions: a primary reinforcer needs the right deprivation to be in force, and a secondary reinforcer needs the reinforcer behind it to keep arriving.

What is an example of a secondary reinforcer?

The click of a clicker in dog training. It means nothing to an untrained dog; after being followed by treats a few dozen times it comes to strengthen whatever behavior it follows, and the trainer can use it to mark a sit at the exact instant it happens. Praise, grades, money, tokens, and a check mark in a habit app are secondary reinforcers in the same way.

Is money a primary or secondary reinforcer?

Secondary, and specifically a generalized conditioned reinforcer. Money has no biological value of its own, but it has been paired with nearly everything a person needs and wants, so it works almost regardless of which need is currently active. Skinner used it as the standard example of a generalized reinforcer, and the chimpanzee token studies of the 1930s showed the same thing in animals working for poker chips.

Is praise a primary or secondary reinforcer?

Usually secondary: praise acquires its power from what has followed it in a person's history, which is why it reinforces some people strongly, others weakly, and some not at all. Whether social contact in general is primary is less clear. Harlow's monkeys preferred a cloth mother that never fed them, so some social reinforcement, at least contact comfort in infant primates, appears to be unlearned.

Why do secondary reinforcers stop working?

Two reasons. First, extinction: a conditioned reinforcer keeps its function only while it goes on predicting the reinforcer behind it, so a token that no longer buys anything, or a click no longer followed by food, gradually loses its power. Second, satiation of the backing reinforcer: a food-backed clicker is weak after a large meal. Generalized reinforcers resist the second problem, not the first.

What is a generalized reinforcer?

A conditioned reinforcer that has been paired with many different reinforcers rather than one. Money, tokens, attention, and approval are the standard examples. Because several deprivations are likely to be in force at any moment, a generalized reinforcer works almost regardless of the person's state, which is why a token economy can run all day in a classroom without anyone being hungry.

Is a conditioned reinforcer the same as a discriminative stimulus?

No, but the same stimulus often does both jobs. A discriminative stimulus comes before a behavior and signals that the behavior will be reinforced; a conditioned reinforcer comes after a behavior and strengthens it. In a behavior chain, each link's stimulus reinforces the response that produced it and sets the occasion for the next one, which is why chaining works.

References

  1. Skinner, B. F. (1953). Science and Human Behavior. Macmillan.
  2. Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
  3. Michael, J. (1993). Establishing operations. The Behavior Analyst, 16(2), 191–206.
  4. Laraway, S., Snycerski, S., Michael, J., & Poling, A. (2003). Motivating operations and terms to describe them: Some further refinements. Journal of Applied Behavior Analysis, 36(3), 407–414.
  5. Bersh, P. J. (1951). The influence of two variables upon the establishment of a secondary reinforcer for operant responses. Journal of Experimental Psychology, 41(1), 62–73.
  6. Kelleher, R. T., & Gollub, L. R. (1962). A review of positive conditioned reinforcement. Journal of the Experimental Analysis of Behavior, 5(S4), 543–597.
  7. Zimmerman, D. W. (1957). Durable secondary reinforcement: Method and theory. Psychological Review, 64(6, Pt. 1), 373–383.
  8. Egger, M. D., & Miller, N. E. (1962). Secondary reinforcement in rats as a function of information value and reliability of the stimulus. Journal of Experimental Psychology, 64(2), 97–104.
  9. Fantino, E. (1977). Conditioned reinforcement: Choice and information. In W. K. Honig & J. E. R. Staddon (Eds.), Handbook of Operant Behavior (pp. 313–339). Prentice-Hall.
  10. Williams, B. A. (1994). Conditioned reinforcement: Experimental and theoretical issues. The Behavior Analyst, 17(2), 261–285.
  11. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593–1599.
  12. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology, 47(6), 419–427.
  13. Wolfe, J. B. (1936). Effectiveness of token-rewards for chimpanzees. Comparative Psychology Monographs, 12(5), 1–72.
  14. Cowles, J. T. (1937). Food-tokens as incentives for learning by chimpanzees. Comparative Psychology Monographs, 14(5), 1–96.
  15. Hackenberg, T. D. (2009). Token reinforcement: A review and analysis. Journal of the Experimental Analysis of Behavior, 91(2), 257–286.
  16. Harlow, H. F. (1958). The nature of love. American Psychologist, 13(12), 673–685.
  17. Pryor, K. (1999). Don't Shoot the Dog! The New Art of Teaching and Training (rev. ed.). Bantam.
  18. Feng, L. C., Howell, T. J., & Bennett, P. C. (2016). How clicker training works: Comparing reinforcing, marking, and bridging hypotheses. Applied Animal Behaviour Science, 181, 34–40.
  19. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. Psychological Review, 66(4), 219–233.