Procedures · Reinforcement-based reduction
Differential Reinforcement: DRA, DRI, DRO, DRL, and DRH Explained
Most problem behavior is not a failure of discipline; it is a behavior that works. Differential reinforcement is the family of procedures that makes something else work better — an alternative behavior, or simply the absence of the problem — while the problem behavior stops paying. It is the reason modern behavior analysis reaches for reinforcement before punishment.
Definition
Differential reinforcement is the procedure of reinforcing one response, or one class of responses, while withholding reinforcement from another. The reinforced class becomes more frequent; the unreinforced class is placed on extinction and declines. Every use of the term has both halves — reinforcement of something and extinction of something else — running at the same time.[1]
It is the mechanism inside shaping and discrimination training, and in applied work it names a family of procedures — DRA, DRI, DRO, DRL, and DRH — that reduce a problem behavior by paying for something other than it, rather than by punishing it.[2]
In brief
- Differential reinforcement always has two halves: one response class is reinforced and another is placed on extinction. Drop either half and the procedure is not differential.
- The five applied procedures differ in what earns the reinforcer: an alternative behavior (DRA), an incompatible behavior (DRI), the absence of the behavior for an interval (DRO), a low rate (DRL), or a high rate (DRH).
- It works when a functional assessment has identified what reinforces the problem behavior, so that the same reinforcer can be withheld for the problem and delivered for the alternative.
How differential reinforcement works
Reinforcement strengthens whatever it follows, and left to itself it is not selective: if the pellet comes after hard presses and soft presses alike, both persist. Differential reinforcement adds a criterion. Responses that meet it are reinforced, responses that do not are not, and because the two classes now have different consequences their frequencies diverge.
Skinner described the procedure in The Behavior of Organisms under two headings. In differentiation the criterion is a property of the response itself: reinforce only lever presses above a certain force and the whole distribution of forces shifts upward. In discrimination the criterion is the stimulus present when the response occurs: reinforce presses when a light is on and never when it is off, and the rat comes to press only in the light.[1] Both are the same operation — pay for one class, not the other — applied to different dimensions of behavior.
Three central procedures are differential reinforcement under other names. Shaping is differential reinforcement of successive approximations, with the criterion moved step by step. Stimulus control is built by differential reinforcement with respect to a discriminative stimulus. And the reduction procedures below are differential reinforcement with respect to a problem behavior: something else is reinforced, and the problem behavior is placed on extinction.[2]
Procedure and process
"Differential reinforcement" names what you do: deliver a consequence after one class of responses and withhold it after another. What happens to behavior — one class rising, the other declining — is the process, and the only test that the procedure worked. If the "reinforced" behavior does not increase, the consequence was not a reinforcer for that person; if what you withheld was not the reinforcer maintaining the problem behavior, nothing was put on extinction.
The five procedures: DRA, DRI, DRO, DRL, and DRH
Applied behavior analysis uses differential reinforcement mainly to reduce behavior without punishment. The five procedures differ in what has to happen for the reinforcer to arrive.[2]
| Procedure | What is reinforced | What is withheld | Example | Best for |
|---|---|---|---|---|
| DRA — alternative behavior | A specific appropriate behavior that produces what the problem behavior produced | The reinforcer for the problem behavior | A child who screams for help is taught to tap the adult's arm; tapping is answered, screaming is not | Behavior with a clear function that an acceptable behavior can serve |
| DRI — incompatible behavior | An alternative that physically cannot occur at the same time as the problem behavior | The reinforcer for the problem behavior | A dog that jumps on guests is greeted only while sitting | Behavior with an obvious physical opposite |
| DRO — other behavior | The absence of the problem behavior for a set interval | The reinforcer for the problem behavior; a response usually resets the interval | A point for every five minutes without calling out | Behavior with no obvious alternative |
| DRL — low rates | Responding at or below a limit, or spaced far enough apart | Reinforcement when responding is too frequent | A class earns free time if there are five or fewer talk-outs in a period | Behavior that is fine in moderation |
| DRH — high rates | Responding at or above a set rate | Reinforcement when responding is too slow | A break for finishing twenty math facts in a minute | Fluency in a skill that is accurate but slow |
DRA and DRI
DRA is the workhorse. The alternative can be anything the person can do, or can be taught, that gets them what the problem behavior got them: asking instead of grabbing, raising a hand instead of shouting, requesting a break instead of shoving the worksheet away. DRI adds one constraint — the alternative cannot coexist with the problem behavior — which makes the extinction half easier to keep, since a sitting dog is not jumping. Not every behavior has a useful opposite.[2]
DRO
DRO is the odd one out, because nothing in particular is reinforced. The reinforcer is delivered when an interval passes without the target behavior, whatever else the person was doing. Reynolds coined the term in a 1961 pigeon experiment on behavioral contrast, in which one schedule delivered food only when the bird had refrained from pecking for a set time.[3] In interval DRO the behavior must be absent for the whole interval; in momentary DRO only at the instant the interval ends. The whole-interval version is the stronger treatment; momentary DRO is easier to run and useful for maintenance. A response usually resets the clock.[2]
DRL and DRH
DRL and DRH act on rate rather than on which behavior occurs. DRL reinforces responding only when it is infrequent: either the total for a session is at or below a limit (full-session DRL), or each response is reinforced only if enough time has passed since the last (spaced-responding DRL). It is the tool for behavior that should be reduced, not eliminated. Deitz and Repp's 1973 classroom study is the standard example: a boy in a special-education class earned candy when talk-outs in a period stayed at five or fewer, a whole class earned it as a group on the same terms, and high-school students earned a free period for keeping off-topic remarks under a limit lowered in stages toward zero. In each case the behavior fell to the criterion.[4] DRH is the mirror image, reinforcing only when responding is fast enough, and belongs to fluency training rather than behavior reduction.[2] Both are schedules before they are treatments.
Why the function of the behavior matters
The extinction half only works if the reinforcer you withhold is the one maintaining the behavior, and that is not something you can see by looking. The same tantrum can be maintained by attention, by escape from a demand, by access to an item, or by the sensation it produces, and each calls for a different thing to be withheld and paid.[5]
Iwata and colleagues' functional analysis, first published in 1982, made the function testable. Children who injured themselves were observed under a series of conditions — an adult who responded to self-injury with attention, an adult who withdrew task demands when it occurred, a room with nothing to do, and a play condition as a control — and for most of them the behavior was reliably higher in one condition than the others.[5] Vollmer and Iwata's 1992 review drew the consequence for treatment: differential reinforcement should use the functional reinforcer, the one identified by the analysis, rather than an arbitrary one that merely seems appealing. Withholding it for the problem behavior is extinction; delivering it for the alternative gives the alternative the job the problem behavior used to do.[6]
| Function | What is withheld | What the alternative earns | Example alternative |
|---|---|---|---|
| Attention | Reactions to the problem behavior | Attention, promptly | Tapping an arm; a raised hand |
| Escape from demands | Removal of the task | A break, help, or an easier step | "Break, please"; "help" |
| Access to items or activities | The item | The item | Asking; pointing; a picture card |
| Automatic (sensory) | Difficult; the behavior produces its own reinforcer | A matched sensory alternative | Chewing a safe object instead of a sleeve |
A 1993 study by Vollmer, Iwata, and colleagues shows the logic at full strength. Three women whose self-injury had been shown by functional analysis to be attention-maintained were treated with DRO using attention as the reinforcer, and with noncontingent attention delivered on a time schedule regardless of behavior. Both reduced self-injury. The authors noted an advantage of the noncontingent schedule worth remembering when designing a DRO: attention arrived densely from the start, whereas DRO begins with stretches in which the person earns nothing, which is the condition that produces bursts.[7] The escape case is the one that catches parents and teachers: ignoring a tantrum maintained by getting out of a task is not extinction, because the task still went away. Escape-maintained behavior ›
Functional communication training: DRA with a request
The most studied form of DRA teaches the person to ask for what the problem behavior produced. Carr and Durand's 1985 study established the method. Four children with developmental disabilities were observed while task difficulty and adult attention were varied; for some, problem behavior rose when attention was scarce, for others when tasks were hard. Each child was then taught a phrase — "Am I doing good work?" for attention, "I don't understand" for help — and the phrase was answered whenever it was used. Problem behavior fell when the phrase matched the child's function and did not fall when the child was taught the other one.[8] The reinforcer, not the words, was the active ingredient.
Tiger, Hanley, and Bruzek's practical guide lays out the modern package: a functional analysis; a communicative response chosen for the learner — vocal, signed, a card, a switch — and easy enough to beat the problem behavior; teaching by prompting and reinforcing it every time; and then, once the problem behavior is low, thinning the schedule with signals for when requests will and will not be honored, or with gradually longer delays.[9] Thinning is where FCT most often fails: a request refused too often stops being worth making, and the problem behavior returns.
Whether extinction is required has been tested directly. In a summary of 21 inpatient cases, functional communication training without extinction did not reduce problem behavior; adding extinction produced large reductions in many cases but not all, and the remainder required punishment components before the behavior came down.[10] Teaching the request is necessary; it is not sufficient while the scream still works.
What the research shows
DRA has the strongest evidence base of the family. Petscher, Rey, and Bailey's 2009 review concluded that DRA has substantial empirical support as a treatment for problem behavior in people with developmental disabilities, and observed that most studies combined it with extinction, so the effect of reinforcing an alternative without withholding reinforcement for the problem behavior is much less well established.[11] A methodological review of the adult literature reached a similar verdict with a caution: the procedures worked in most studies, but the studies were often small and short, so durability in adults is less well documented than the initial reductions.[12]
DRO's evidence is broad but its mechanism is unsettled. Jessel and Ingvarsson's 2016 summary of recent DRO research notes that the procedure may reduce behavior less by reinforcing "other behavior" — which is not measured, and need not increase in any specific form — than through the extinction and the response-contingent postponement of reinforcement built into it.[13] For practice this matters little: DRO reduces behavior. For understanding, it means the name is partly a misnomer.
Two laboratory findings are cautions. Reynolds' pigeons showed behavioral contrast: cutting reinforcement for pecking in one component of a multiple schedule raised the rate of pecking in the other, where reinforcement was unchanged.[3] Put a behavior on extinction in one setting and it may rise in another. And research on behavioral momentum shows that adding reinforcement to a situation — which DRA does — makes all behavior there more resistant to change, including the problem behavior when its reinforcement is later withheld.[14] Neither is a reason not to use the procedure; both are reasons to run it everywhere and expect persistence.
Examples across settings
In every row the two halves are named: what the alternative earns, and what the problem behavior no longer earns. The function column is a guess; in real life, check it first.
| Setting | Problem behavior | Likely function | Procedure |
|---|---|---|---|
| Classroom | Calling out | Teacher attention | DRA: raised hands are answered promptly; called-out answers get no response |
| Parenting | Whining for snacks | Access to the item | DRA: a plain request in a normal voice gets the snack, or a clear answer; whining never does |
| Parenting | Screaming during homework | Escape from the task | FCT: "break, please" earns a two-minute break; screaming does not end the homework |
| Dog training | Jumping on guests | Attention from the guest | DRI: sitting is greeted and petted; jumping is turned away from |
| Self-management | Checking your phone mid-task | Escape from boredom or difficulty | DRA: a short walk or a glass of water for the same relief, with notifications off so checking pays less |
How to run differential reinforcement
- Define the problem behavior and count it. Observable, measurable, with a baseline rate. A DRO interval or a DRL limit depends on how often the behavior occurs now.
- Find the function. A functional assessment, formal or informal: what arrives right after the behavior, and what goes away when it starts? The answer tells you what to withhold and what to pay with.[5][6]
- Choose the alternative. Something the person can do or can be taught quickly, that produces the same reinforcer, and that takes less effort than the problem behavior. If nothing fits, use DRO; if the behavior is acceptable in moderation, use DRL.
- Start the schedule dense. Reinforce the alternative every time at first. For DRO, set the first interval a little shorter than the average gap between problem behaviors at baseline, so the person contacts the reinforcer early and often.[2]
- Withhold the maintaining reinforcer for the problem behavior. This is extinction, and it has to be as consistent as the reinforcement. If the function is escape, the demand stays; if attention, the reaction stops; if an item, the item is not delivered.
- Plan for the burst. Expect the problem behavior to increase briefly when it stops working. The extinction burst is less likely when extinction is combined with reinforcement of an alternative, but everyone involved should know it may come and agree not to give in.[15]
- Thin the schedule, then generalize. Once the problem behavior is low, lengthen DRO intervals, tighten DRL limits, and move the alternative onto intermittent reinforcement with clear signals for when it will pay.[9] Then run the procedure everywhere the behavior occurs, with everyone it occurs with.
Common mistakes
- Paying the alternative less than the problem behavior was paid. If screaming brought a parent to the room in three seconds and asking gets "just a minute," screaming wins. The matching law predicts it: behavior goes where reinforcement is richer, faster, and more reliable.[9]
- DRO intervals that are too long. An interval the person never completes delivers no reinforcement, and a procedure that delivers no reinforcement is extinction alone. Start below the baseline gap and lengthen from there.[2]
- Forgetting the extinction half. Reinforcing the alternative while the problem behavior still works gives the person two ways to earn the same thing, and the old one is better practiced. The evidence for DRA is largely evidence for DRA with extinction.[10][11]
- Withholding the wrong reinforcer. Ignoring escape-maintained behavior is not extinction; it delivers the escape. Only the reinforcer that matches the function can be withheld.[6]
- An alternative that is harder than the problem, or thinned too fast. A five-word sentence competes badly with a scream, and moving from every time to occasionally in one step puts the alternative on extinction. Start with the easiest response that does the job, and thin in small steps.[9]
- Running it in one setting only. Behavioral contrast and plain failure to generalize both mean the behavior can rise wherever the procedure is not running.[3]
What the evidence does not show
Most of the evidence comes from single-case experimental designs with people with developmental disabilities, in clinics and schools with trained staff. The reductions are large and replicated. What the literature does not establish is how well the same procedures work when run by parents and teachers without support, how often the effects hold months later, and how much of the effect is due to reinforcing the alternative rather than to the extinction that almost always accompanies it.[11][12]
Nor does the evidence show that differential reinforcement is free of side effects. Extinction bursts, contrast in other settings, and the momentum problem — a richer context making the problem behavior more persistent once its reinforcement is withheld — are all documented.[15][3][14] They are milder than the side effects of punishment, and manageable when expected. Differential reinforcement is the default not because it is perfect but because it builds something while it removes something, and because its failures are usually failures of function, schedule, or consistency that better design can fix.
Key takeaways
- Differential reinforcement is two procedures at once: one response class is reinforced and another is placed on extinction. Shaping, discrimination training, and the DR family of treatments are the same operation on different dimensions of behavior.
- DRA reinforces a specific alternative, DRI an incompatible one, DRO the absence of the behavior for an interval, DRL a low rate, and DRH a high rate. DRA, especially as functional communication training, has the strongest evidence.
- The reinforcer you withhold must be the one maintaining the behavior, and the reinforcer you deliver should be the same one. A functional assessment tells you which it is; guessing is how "ignoring" ends up delivering the escape.
- Start dense and thin slowly: reinforce the alternative every time, set DRO intervals shorter than the baseline gap, and expect a burst when the problem behavior stops working.
- The side effects — bursts, contrast in other settings, and momentum — are documented, but they are milder and more manageable than those of punishment, which is why reinforcement-based reduction comes first.
Check yourself
A teacher decides to reduce a student's calling out by praising him whenever he raises his hand, but keeps answering his called-out questions so he does not fall behind. Two weeks later calling out is unchanged. What is missing?
The extinction half. Both responses still produce the same reinforcer, and the older, easier one is better practiced, so it persists. DRA requires that the reinforcer be withheld for the problem behavior: called-out questions go unanswered, and raised hands are answered promptly and every time.
A child's screaming during homework has been shown to be escape-maintained. A parent sets up a DRO in which five minutes without screaming earns a sticker, but screaming continues. Why might the DRO be failing?
Two likely reasons. The sticker is an arbitrary reinforcer competing with escape, which is the functional one, and the parent may still be ending homework when screaming occurs, so the problem behavior is not on extinction. The better design uses the functional reinforcer: a request for a break earns one, and screaming no longer does. And check the interval; if screaming occurs every two minutes at baseline, five minutes is never reached and no reinforcement is delivered.
A trainer uses DRI for a dog that jumps on people: at home the dog is greeted only while sitting, and jumping at home stops. At the park, where the trainer does not run the procedure, jumping increases. What is happening?
Jumping is still reinforced at the park, and behavioral contrast means that reducing reinforcement in one setting can raise the behavior in another where reinforcement is unchanged. Generalization has to be programmed: run the procedure across settings and with the other people the dog meets.
Want more? The 20-question quiz covers every page on the site with instant explanations.
Frequently asked questions
What is differential reinforcement in simple terms?
Reinforcing one behavior while no longer reinforcing another. The reinforced behavior becomes more common and the other fades. It is how shaping works, how animals learn to respond only to certain cues, and, in applied settings, the main way to reduce a problem behavior without punishment: pay for something else, and stop paying for the problem.
What is the difference between DRA and DRO?
DRA reinforces a specific alternative behavior that serves the same purpose as the problem behavior, such as asking instead of grabbing. DRO reinforces the absence of the problem behavior for a set interval, whatever else the person does. DRA teaches a replacement; DRO does not, which is why DRA is preferred when a suitable alternative exists.
What is an example of differential reinforcement?
A child who whines for snacks is given the snack, or a clear answer, whenever she asks in a normal voice, and never when she whines. Asking increases and whining declines. In dog training, a dog that jumps on guests is greeted only while sitting, so sitting replaces jumping.
Is differential reinforcement the same as extinction?
No, but extinction is half of it. Extinction alone withholds the reinforcer for a behavior and leaves a gap where the behavior was. Differential reinforcement withholds that reinforcer and at the same time delivers it for something else, so the person still gets what they were after, through the behavior you chose. The combination produces smaller bursts and more durable change.
What is DRL, and when is it used?
Differential reinforcement of low rates reinforces a behavior only when it occurs infrequently, either below a limit for the session or with enough time between responses. It is used for behavior that is acceptable in moderation, such as asking questions in class or eating quickly, where the goal is fewer, not none. Deitz and Repp reduced classroom talk-outs this way in 1973.
What is functional communication training?
A form of DRA in which the person is taught a communicative response, such as a word, sign, or card, that produces the same reinforcer the problem behavior produced, while the problem behavior no longer produces it. Introduced by Carr and Durand in 1985, it is the most studied function-based treatment for problem behavior.
Does differential reinforcement work without extinction?
Rarely, in the published research. Reviews of DRA find that nearly all successful studies also placed the problem behavior on extinction, and a series of clinical cases found that functional communication training without extinction did not reduce problem behavior. If the problem behavior still works, teaching an alternative gives the person two routes to the same reinforcer.
Is differential reinforcement a type of punishment?
No. Punishment adds or removes a stimulus after a behavior to make it less likely. Differential reinforcement reduces a behavior by reinforcing something else and withholding the reinforcer for the problem behavior, which is extinction, not punishment. It is the standard alternative to punishment in applied behavior analysis.
References
- Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
- Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.
- Reynolds, G. S. (1961). Behavioral contrast. Journal of the Experimental Analysis of Behavior, 4(1), 57–71.
- Deitz, S. M., & Repp, A. C. (1973). Decreasing classroom misbehavior through the use of DRL schedules of reinforcement. Journal of Applied Behavior Analysis, 6(3), 457–463.
- Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. Journal of Applied Behavior Analysis, 27(2), 197–209. (Reprinted from Analysis and Intervention in Developmental Disabilities, 2, 3–20, 1982.)
- Vollmer, T. R., & Iwata, B. A. (1992). Differential reinforcement as treatment for behavior disorders: Procedural and functional variations. Research in Developmental Disabilities, 13(4), 393–417.
- Vollmer, T. R., Iwata, B. A., Zarcone, J. R., Smith, R. G., & Mazaleski, J. L. (1993). The role of attention in the treatment of attention-maintained self-injurious behavior: Noncontingent reinforcement and differential reinforcement of other behavior. Journal of Applied Behavior Analysis, 26(1), 9–21.
- Carr, E. G., & Durand, V. M. (1985). Reducing behavior problems through functional communication training. Journal of Applied Behavior Analysis, 18(2), 111–126.
- Tiger, J. H., Hanley, G. P., & Bruzek, J. (2008). Functional communication training: A review and practical guide. Behavior Analysis in Practice, 1(1), 16–23.
- Hagopian, L. P., Fisher, W. W., Sullivan, M. T., Acquisto, J., & LeBlanc, L. A. (1998). Effectiveness of functional communication training with and without extinction and punishment: A summary of 21 inpatient cases. Journal of Applied Behavior Analysis, 31(2), 211–235.
- Petscher, E. S., Rey, C., & Bailey, J. S. (2009). A review of empirical support for differential reinforcement of alternative behavior. Research in Developmental Disabilities, 30(3), 409–425.
- Chowdhury, M., & Benson, B. A. (2011). Use of differential reinforcement to reduce behavior problems in adults with intellectual disabilities: A methodological review. Research in Developmental Disabilities, 32(2), 383–394.
- Jessel, J., & Ingvarsson, E. T. (2016). Recent advances in applied research on DRO procedures. Journal of Applied Behavior Analysis, 49(4), 991–995.
- Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. Behavioral and Brain Sciences, 23(1), 73–90.
- Lerman, D. C., & Iwata, B. A. (1996). Developing a technology for the use of operant extinction in clinical settings: An examination of basic and applied research. Journal of Applied Behavior Analysis, 29(3), 345–382.