Procedures · Reinforcement-based reduction

Differential Reinforcement: DRA, DRI, DRO, DRL, and DRH Explained

Most problem behavior is not a failure of discipline; it is a behavior that works. Differential reinforcement is the family of procedures that makes something else work better — an alternative behavior, or simply the absence of the problem — while the problem behavior stops paying. It is the reason modern behavior analysis reaches for reinforcement before punishment.

Updated 13 min read

Definition

Differential reinforcement is the procedure of reinforcing one response, or one class of responses, while withholding reinforcement from another. The reinforced class becomes more frequent; the unreinforced class is placed on extinction and declines. Every use of the term has both halves — reinforcement of something and extinction of something else — running at the same time.[1]

It is the mechanism inside shaping and discrimination training, and in applied work it names a family of procedures — DRA, DRI, DRO, DRL, and DRH — that reduce a problem behavior by paying for something other than it, rather than by punishing it.[2]

In brief

  • Differential reinforcement always has two halves: one response class is reinforced and another is placed on extinction. Drop either half and the procedure is not differential.
  • The five applied procedures differ in what earns the reinforcer: an alternative behavior (DRA), an incompatible behavior (DRI), the absence of the behavior for an interval (DRO), a low rate (DRL), or a high rate (DRH).
  • It works when a functional assessment has identified what reinforces the problem behavior, so that the same reinforcer can be withheld for the problem and delivered for the alternative.

How differential reinforcement works

Reinforcement strengthens whatever it follows, and left to itself it is not selective: if the pellet comes after hard presses and soft presses alike, both persist. Differential reinforcement adds a criterion. Responses that meet it are reinforced, responses that do not are not, and because the two classes now have different consequences their frequencies diverge.

Skinner described the procedure in The Behavior of Organisms under two headings. In differentiation the criterion is a property of the response itself: reinforce only lever presses above a certain force and the whole distribution of forces shifts upward. In discrimination the criterion is the stimulus present when the response occurs: reinforce presses when a light is on and never when it is off, and the rat comes to press only in the light.[1] Both are the same operation — pay for one class, not the other — applied to different dimensions of behavior.

Three central procedures are differential reinforcement under other names. Shaping is differential reinforcement of successive approximations, with the criterion moved step by step. Stimulus control is built by differential reinforcement with respect to a discriminative stimulus. And the reduction procedures below are differential reinforcement with respect to a problem behavior: something else is reinforced, and the problem behavior is placed on extinction.[2]

Procedure and process

"Differential reinforcement" names what you do: deliver a consequence after one class of responses and withhold it after another. What happens to behavior — one class rising, the other declining — is the process, and the only test that the procedure worked. If the "reinforced" behavior does not increase, the consequence was not a reinforcer for that person; if what you withheld was not the reinforcer maintaining the problem behavior, nothing was put on extinction.

The five procedures: DRA, DRI, DRO, DRL, and DRH

Applied behavior analysis uses differential reinforcement mainly to reduce behavior without punishment. The five procedures differ in what has to happen for the reinforcer to arrive.[2]

ProcedureWhat is reinforcedWhat is withheldExampleBest for
DRA — alternative behaviorA specific appropriate behavior that produces what the problem behavior producedThe reinforcer for the problem behaviorA child who screams for help is taught to tap the adult's arm; tapping is answered, screaming is notBehavior with a clear function that an acceptable behavior can serve
DRI — incompatible behaviorAn alternative that physically cannot occur at the same time as the problem behaviorThe reinforcer for the problem behaviorA dog that jumps on guests is greeted only while sittingBehavior with an obvious physical opposite
DRO — other behaviorThe absence of the problem behavior for a set intervalThe reinforcer for the problem behavior; a response usually resets the intervalA point for every five minutes without calling outBehavior with no obvious alternative
DRL — low ratesResponding at or below a limit, or spaced far enough apartReinforcement when responding is too frequentA class earns free time if there are five or fewer talk-outs in a periodBehavior that is fine in moderation
DRH — high ratesResponding at or above a set rateReinforcement when responding is too slowA break for finishing twenty math facts in a minuteFluency in a skill that is accurate but slow

DRA and DRI

DRA is the workhorse. The alternative can be anything the person can do, or can be taught, that gets them what the problem behavior got them: asking instead of grabbing, raising a hand instead of shouting, requesting a break instead of shoving the worksheet away. DRI adds one constraint — the alternative cannot coexist with the problem behavior — which makes the extinction half easier to keep, since a sitting dog is not jumping. Not every behavior has a useful opposite.[2]

DRO

DRO is the odd one out, because nothing in particular is reinforced. The reinforcer is delivered when an interval passes without the target behavior, whatever else the person was doing. Reynolds coined the term in a 1961 pigeon experiment on behavioral contrast, in which one schedule delivered food only when the bird had refrained from pecking for a set time.[3] In interval DRO the behavior must be absent for the whole interval; in momentary DRO only at the instant the interval ends. The whole-interval version is the stronger treatment; momentary DRO is easier to run and useful for maintenance. A response usually resets the clock.[2]

DRL and DRH

DRL and DRH act on rate rather than on which behavior occurs. DRL reinforces responding only when it is infrequent: either the total for a session is at or below a limit (full-session DRL), or each response is reinforced only if enough time has passed since the last (spaced-responding DRL). It is the tool for behavior that should be reduced, not eliminated. Deitz and Repp's 1973 classroom study is the standard example: a boy in a special-education class earned candy when talk-outs in a period stayed at five or fewer, a whole class earned it as a group on the same terms, and high-school students earned a free period for keeping off-topic remarks under a limit lowered in stages toward zero. In each case the behavior fell to the criterion.[4] DRH is the mirror image, reinforcing only when responding is fast enough, and belongs to fluency training rather than behavior reduction.[2] Both are schedules before they are treatments.

Why the function of the behavior matters

The extinction half only works if the reinforcer you withhold is the one maintaining the behavior, and that is not something you can see by looking. The same tantrum can be maintained by attention, by escape from a demand, by access to an item, or by the sensation it produces, and each calls for a different thing to be withheld and paid.[5]

Iwata and colleagues' functional analysis, first published in 1982, made the function testable. Children who injured themselves were observed under a series of conditions — an adult who responded to self-injury with attention, an adult who withdrew task demands when it occurred, a room with nothing to do, and a play condition as a control — and for most of them the behavior was reliably higher in one condition than the others.[5] Vollmer and Iwata's 1992 review drew the consequence for treatment: differential reinforcement should use the functional reinforcer, the one identified by the analysis, rather than an arbitrary one that merely seems appealing. Withholding it for the problem behavior is extinction; delivering it for the alternative gives the alternative the job the problem behavior used to do.[6]

FunctionWhat is withheldWhat the alternative earnsExample alternative
AttentionReactions to the problem behaviorAttention, promptlyTapping an arm; a raised hand
Escape from demandsRemoval of the taskA break, help, or an easier step"Break, please"; "help"
Access to items or activitiesThe itemThe itemAsking; pointing; a picture card
Automatic (sensory)Difficult; the behavior produces its own reinforcerA matched sensory alternativeChewing a safe object instead of a sleeve

A 1993 study by Vollmer, Iwata, and colleagues shows the logic at full strength. Three women whose self-injury had been shown by functional analysis to be attention-maintained were treated with DRO using attention as the reinforcer, and with noncontingent attention delivered on a time schedule regardless of behavior. Both reduced self-injury. The authors noted an advantage of the noncontingent schedule worth remembering when designing a DRO: attention arrived densely from the start, whereas DRO begins with stretches in which the person earns nothing, which is the condition that produces bursts.[7] The escape case is the one that catches parents and teachers: ignoring a tantrum maintained by getting out of a task is not extinction, because the task still went away. Escape-maintained behavior ›

Functional communication training: DRA with a request

The most studied form of DRA teaches the person to ask for what the problem behavior produced. Carr and Durand's 1985 study established the method. Four children with developmental disabilities were observed while task difficulty and adult attention were varied; for some, problem behavior rose when attention was scarce, for others when tasks were hard. Each child was then taught a phrase — "Am I doing good work?" for attention, "I don't understand" for help — and the phrase was answered whenever it was used. Problem behavior fell when the phrase matched the child's function and did not fall when the child was taught the other one.[8] The reinforcer, not the words, was the active ingredient.

Tiger, Hanley, and Bruzek's practical guide lays out the modern package: a functional analysis; a communicative response chosen for the learner — vocal, signed, a card, a switch — and easy enough to beat the problem behavior; teaching by prompting and reinforcing it every time; and then, once the problem behavior is low, thinning the schedule with signals for when requests will and will not be honored, or with gradually longer delays.[9] Thinning is where FCT most often fails: a request refused too often stops being worth making, and the problem behavior returns.

Whether extinction is required has been tested directly. In a summary of 21 inpatient cases, functional communication training without extinction did not reduce problem behavior; adding extinction produced large reductions in many cases but not all, and the remainder required punishment components before the behavior came down.[10] Teaching the request is necessary; it is not sufficient while the scream still works.

What the research shows

DRA has the strongest evidence base of the family. Petscher, Rey, and Bailey's 2009 review concluded that DRA has substantial empirical support as a treatment for problem behavior in people with developmental disabilities, and observed that most studies combined it with extinction, so the effect of reinforcing an alternative without withholding reinforcement for the problem behavior is much less well established.[11] A methodological review of the adult literature reached a similar verdict with a caution: the procedures worked in most studies, but the studies were often small and short, so durability in adults is less well documented than the initial reductions.[12]

DRO's evidence is broad but its mechanism is unsettled. Jessel and Ingvarsson's 2016 summary of recent DRO research notes that the procedure may reduce behavior less by reinforcing "other behavior" — which is not measured, and need not increase in any specific form — than through the extinction and the response-contingent postponement of reinforcement built into it.[13] For practice this matters little: DRO reduces behavior. For understanding, it means the name is partly a misnomer.

Two laboratory findings are cautions. Reynolds' pigeons showed behavioral contrast: cutting reinforcement for pecking in one component of a multiple schedule raised the rate of pecking in the other, where reinforcement was unchanged.[3] Put a behavior on extinction in one setting and it may rise in another. And research on behavioral momentum shows that adding reinforcement to a situation — which DRA does — makes all behavior there more resistant to change, including the problem behavior when its reinforcement is later withheld.[14] Neither is a reason not to use the procedure; both are reasons to run it everywhere and expect persistence.

Examples across settings

In every row the two halves are named: what the alternative earns, and what the problem behavior no longer earns. The function column is a guess; in real life, check it first.

SettingProblem behaviorLikely functionProcedure
ClassroomCalling outTeacher attentionDRA: raised hands are answered promptly; called-out answers get no response
ParentingWhining for snacksAccess to the itemDRA: a plain request in a normal voice gets the snack, or a clear answer; whining never does
ParentingScreaming during homeworkEscape from the taskFCT: "break, please" earns a two-minute break; screaming does not end the homework
Dog trainingJumping on guestsAttention from the guestDRI: sitting is greeted and petted; jumping is turned away from
Self-managementChecking your phone mid-taskEscape from boredom or difficultyDRA: a short walk or a glass of water for the same relief, with notifications off so checking pays less

How to run differential reinforcement

  1. Define the problem behavior and count it. Observable, measurable, with a baseline rate. A DRO interval or a DRL limit depends on how often the behavior occurs now.
  2. Find the function. A functional assessment, formal or informal: what arrives right after the behavior, and what goes away when it starts? The answer tells you what to withhold and what to pay with.[5][6]
  3. Choose the alternative. Something the person can do or can be taught quickly, that produces the same reinforcer, and that takes less effort than the problem behavior. If nothing fits, use DRO; if the behavior is acceptable in moderation, use DRL.
  4. Start the schedule dense. Reinforce the alternative every time at first. For DRO, set the first interval a little shorter than the average gap between problem behaviors at baseline, so the person contacts the reinforcer early and often.[2]
  5. Withhold the maintaining reinforcer for the problem behavior. This is extinction, and it has to be as consistent as the reinforcement. If the function is escape, the demand stays; if attention, the reaction stops; if an item, the item is not delivered.
  6. Plan for the burst. Expect the problem behavior to increase briefly when it stops working. The extinction burst is less likely when extinction is combined with reinforcement of an alternative, but everyone involved should know it may come and agree not to give in.[15]
  7. Thin the schedule, then generalize. Once the problem behavior is low, lengthen DRO intervals, tighten DRL limits, and move the alternative onto intermittent reinforcement with clear signals for when it will pay.[9] Then run the procedure everywhere the behavior occurs, with everyone it occurs with.

Common mistakes

What the evidence does not show

Most of the evidence comes from single-case experimental designs with people with developmental disabilities, in clinics and schools with trained staff. The reductions are large and replicated. What the literature does not establish is how well the same procedures work when run by parents and teachers without support, how often the effects hold months later, and how much of the effect is due to reinforcing the alternative rather than to the extinction that almost always accompanies it.[11][12]

Nor does the evidence show that differential reinforcement is free of side effects. Extinction bursts, contrast in other settings, and the momentum problem — a richer context making the problem behavior more persistent once its reinforcement is withheld — are all documented.[15][3][14] They are milder than the side effects of punishment, and manageable when expected. Differential reinforcement is the default not because it is perfect but because it builds something while it removes something, and because its failures are usually failures of function, schedule, or consistency that better design can fix.

Key takeaways

Check yourself

A teacher decides to reduce a student's calling out by praising him whenever he raises his hand, but keeps answering his called-out questions so he does not fall behind. Two weeks later calling out is unchanged. What is missing?

The extinction half. Both responses still produce the same reinforcer, and the older, easier one is better practiced, so it persists. DRA requires that the reinforcer be withheld for the problem behavior: called-out questions go unanswered, and raised hands are answered promptly and every time.

A child's screaming during homework has been shown to be escape-maintained. A parent sets up a DRO in which five minutes without screaming earns a sticker, but screaming continues. Why might the DRO be failing?

Two likely reasons. The sticker is an arbitrary reinforcer competing with escape, which is the functional one, and the parent may still be ending homework when screaming occurs, so the problem behavior is not on extinction. The better design uses the functional reinforcer: a request for a break earns one, and screaming no longer does. And check the interval; if screaming occurs every two minutes at baseline, five minutes is never reached and no reinforcement is delivered.

A trainer uses DRI for a dog that jumps on people: at home the dog is greeted only while sitting, and jumping at home stops. At the park, where the trainer does not run the procedure, jumping increases. What is happening?

Jumping is still reinforced at the park, and behavioral contrast means that reducing reinforcement in one setting can raise the behavior in another where reinforcement is unchanged. Generalization has to be programmed: run the procedure across settings and with the other people the dog meets.

Want more? The 20-question quiz covers every page on the site with instant explanations.

Frequently asked questions

What is differential reinforcement in simple terms?

Reinforcing one behavior while no longer reinforcing another. The reinforced behavior becomes more common and the other fades. It is how shaping works, how animals learn to respond only to certain cues, and, in applied settings, the main way to reduce a problem behavior without punishment: pay for something else, and stop paying for the problem.

What is the difference between DRA and DRO?

DRA reinforces a specific alternative behavior that serves the same purpose as the problem behavior, such as asking instead of grabbing. DRO reinforces the absence of the problem behavior for a set interval, whatever else the person does. DRA teaches a replacement; DRO does not, which is why DRA is preferred when a suitable alternative exists.

What is an example of differential reinforcement?

A child who whines for snacks is given the snack, or a clear answer, whenever she asks in a normal voice, and never when she whines. Asking increases and whining declines. In dog training, a dog that jumps on guests is greeted only while sitting, so sitting replaces jumping.

Is differential reinforcement the same as extinction?

No, but extinction is half of it. Extinction alone withholds the reinforcer for a behavior and leaves a gap where the behavior was. Differential reinforcement withholds that reinforcer and at the same time delivers it for something else, so the person still gets what they were after, through the behavior you chose. The combination produces smaller bursts and more durable change.

What is DRL, and when is it used?

Differential reinforcement of low rates reinforces a behavior only when it occurs infrequently, either below a limit for the session or with enough time between responses. It is used for behavior that is acceptable in moderation, such as asking questions in class or eating quickly, where the goal is fewer, not none. Deitz and Repp reduced classroom talk-outs this way in 1973.

What is functional communication training?

A form of DRA in which the person is taught a communicative response, such as a word, sign, or card, that produces the same reinforcer the problem behavior produced, while the problem behavior no longer produces it. Introduced by Carr and Durand in 1985, it is the most studied function-based treatment for problem behavior.

Does differential reinforcement work without extinction?

Rarely, in the published research. Reviews of DRA find that nearly all successful studies also placed the problem behavior on extinction, and a series of clinical cases found that functional communication training without extinction did not reduce problem behavior. If the problem behavior still works, teaching an alternative gives the person two routes to the same reinforcer.

Is differential reinforcement a type of punishment?

No. Punishment adds or removes a stimulus after a behavior to make it less likely. Differential reinforcement reduces a behavior by reinforcing something else and withholding the reinforcer for the problem behavior, which is extinction, not punishment. It is the standard alternative to punishment in applied behavior analysis.

References

  1. Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
  2. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.
  3. Reynolds, G. S. (1961). Behavioral contrast. Journal of the Experimental Analysis of Behavior, 4(1), 57–71.
  4. Deitz, S. M., & Repp, A. C. (1973). Decreasing classroom misbehavior through the use of DRL schedules of reinforcement. Journal of Applied Behavior Analysis, 6(3), 457–463.
  5. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. Journal of Applied Behavior Analysis, 27(2), 197–209. (Reprinted from Analysis and Intervention in Developmental Disabilities, 2, 3–20, 1982.)
  6. Vollmer, T. R., & Iwata, B. A. (1992). Differential reinforcement as treatment for behavior disorders: Procedural and functional variations. Research in Developmental Disabilities, 13(4), 393–417.
  7. Vollmer, T. R., Iwata, B. A., Zarcone, J. R., Smith, R. G., & Mazaleski, J. L. (1993). The role of attention in the treatment of attention-maintained self-injurious behavior: Noncontingent reinforcement and differential reinforcement of other behavior. Journal of Applied Behavior Analysis, 26(1), 9–21.
  8. Carr, E. G., & Durand, V. M. (1985). Reducing behavior problems through functional communication training. Journal of Applied Behavior Analysis, 18(2), 111–126.
  9. Tiger, J. H., Hanley, G. P., & Bruzek, J. (2008). Functional communication training: A review and practical guide. Behavior Analysis in Practice, 1(1), 16–23.
  10. Hagopian, L. P., Fisher, W. W., Sullivan, M. T., Acquisto, J., & LeBlanc, L. A. (1998). Effectiveness of functional communication training with and without extinction and punishment: A summary of 21 inpatient cases. Journal of Applied Behavior Analysis, 31(2), 211–235.
  11. Petscher, E. S., Rey, C., & Bailey, J. S. (2009). A review of empirical support for differential reinforcement of alternative behavior. Research in Developmental Disabilities, 30(3), 409–425.
  12. Chowdhury, M., & Benson, B. A. (2011). Use of differential reinforcement to reduce behavior problems in adults with intellectual disabilities: A methodological review. Research in Developmental Disabilities, 32(2), 383–394.
  13. Jessel, J., & Ingvarsson, E. T. (2016). Recent advances in applied research on DRO procedures. Journal of Applied Behavior Analysis, 49(4), 991–995.
  14. Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. Behavioral and Brain Sciences, 23(1), 73–90.
  15. Lerman, D. C., & Iwata, B. A. (1996). Developing a technology for the use of operant extinction in clinical settings: An examination of basic and applied research. Journal of Applied Behavior Analysis, 29(3), 345–382.