Core process

Shaping

You cannot reinforce a behavior that never happens. Shaping is how operant conditioning builds behavior that does not yet exist — one small approximation at a time, in pigeons, children, patients, and yourself.

Updated 16 min read

Definition

Shaping is the differential reinforcement of successive approximations to a target behavior. Responses that come closer to the goal are reinforced; earlier, cruder forms are no longer reinforced; and the criterion for reinforcement is moved step by step until the target behavior appears.

The term is B. F. Skinner's, and so is the analogy: "Operant conditioning shapes behavior as a sculptor shapes a lump of clay." Shaping is the standard method for teaching a new response in applied behavior analysis, animal training, rehabilitation, and skill learning of every kind.[1]

In brief

  • Shaping is differential reinforcement of successive approximations: reinforce responses closer to the goal, extinguish cruder ones, and raise the criterion step by step.
  • Reinforcement acts only on behavior that occurs; shaping builds behavior that does not yet exist by selecting from the learner's natural variation.
  • Mark each approximation immediately, keep the steps small enough that reinforcement stays frequent, and thin the schedule once the terminal behavior is reliable.

Why shaping is needed

Reinforcement can only act on behavior that occurs. A rat in an operant chamber will not press the lever by chance for a very long time; a child who has never said "water" cannot be praised for saying it; a stroke patient who cannot lift an arm cannot be rewarded for lifting it. If you wait for the finished behavior, you wait forever.

Shaping solves the problem by reinforcing whatever the organism can do that resembles the goal, then demanding a little more. It relies on a fact that is easy to miss: no response is ever repeated exactly. Every lever press has a slightly different force, every attempt at a word a slightly different sound. Shaping selects from that natural variation, the way breeding selects from variation in a population.[1]

How shaping works: the mechanism

Shaping is two procedures running together — reinforcement of the current approximation and extinction of everything else — and it cycles through four phases:

Shaping as a rising criterion Successive bell-shaped distributions of a behavior's form move rightward across the page; a criterion line steps upward so that only the upper tail of each distribution is reinforced, dragging the next distribution further toward the target. Form of the behavior (closer to the target →) Step 1 Step 2 Step 3 Step 4 Target reinforced → Each curve: the range of forms the animal produces at that step. Green line: the criterion — only forms to its right are reinforced.
Shaping as a moving criterion. Reinforcing only the upper tail of what the animal currently does shifts the whole distribution; the criterion is then raised again. Variability (the width of each curve) is what makes the next step possible.
  1. Reinforce a starting behavior. The rat turns toward the lever; a pellet drops. Turning toward the lever increases in frequency.
  2. Variability appears. As the behavior is repeated it varies: some turns are closer, some include a step forward, some a raised paw. Extinction increases variability further — when a response stops paying off, organisms do it in new ways.[2][3]
  3. Raise the criterion. Once turns are reliable, only turns that include a step forward are reinforced. Plain turns go on extinction; steps forward increase.
  4. Repeat. Steps toward, touching, pawing, pressing. Each stage is built on the reinforced behavior of the previous one and the extinction-driven variability that behavior produces.

The variability step is the one people forget. Allen Neuringer's research showed that variability is not noise around a "true" response but a dimension of behavior that reinforcement controls: reinforce variety and organisms become more variable; reinforce sameness and they become stereotyped.[2] Good shaping keeps the current form reinforced often enough to persist but not so often that it hardens into a fixed habit.

Differential reinforcement is the engine

"Differential" means some forms of the response are reinforced and others are not. Without the extinction half you are not shaping — you are reinforcing whatever happens, and the behavior settles at the easiest form that still pays. Without the reinforcement half, the behavior disappears. The skill is holding both at once and moving the line between them.

Skinner's discovery: the pigeon that learned to bowl

Skinner dated his understanding of shaping to a single day in 1943. He, Keller Breland, and Norman Guttman were working on Project Pigeon on the top floor of a flour mill in Minneapolis and decided, for amusement, to teach a pigeon to swipe a small wooden ball down a miniature alley with its beak. Waiting for a full swipe went nowhere. So they reinforced any response that resembled it — a look at the ball, a move toward it, a touch — and within minutes the bird was bowling.[4][5] Skinner later wrote that the day changed how he thought about behavior: complex acts could be built rapidly from nothing by hand-delivered reinforcement.

He showed the public how in a 1951 Scientific American article, "How to Teach Animals," which walks a reader through establishing a conditioned reinforcer (a sound paired with food) and using it to shape a dog or pigeon to a new behavior in a single session — the recipe clicker trainers still follow.[6] His pigeons went on to play a version of ping-pong, pecking a ball back and forth across a table.[7]

How to shape a behavior: step-by-step

  1. Define the terminal behavior. Be exact about the finished form: "sits on the mat with all four paws for 10 seconds," "says 'water' clearly," "raises the affected arm to shoulder height." A vague goal cannot be approximated.
  2. Find the starting behavior. Something the learner already does, at least occasionally, that is on the road to the goal. Glancing at the mat. Saying "wa." Moving the arm two inches. If it never occurs, you need an easier start.
  3. Plan the approximations. Write down the steps you expect, but hold the plan loosely — the learner's variability will suggest better ones.
  4. Reinforce immediately and every time. Use a conditioned reinforcer (a click, a "yes," a checkmark) to mark the exact instant the criterion is met, then deliver the real reinforcer. A second's delay reinforces whatever happened in that second instead.
  5. Raise the criterion in small steps. Move on when the current approximation is reliable — Karen Pryor's rule of thumb is when the learner succeeds most of the time — and keep each step small enough that reinforcement stays frequent.[8]
  6. Don't stay too long; don't move too fast. Staying reinforces a plateau until it resists change. Moving too fast puts the learner on extinction: variability spikes, then the behavior collapses. If that happens, drop back a step.
  7. Thin the schedule at the end. Once the terminal behavior is reliable, shift to intermittent reinforcement and bring the behavior under the cue you want it to answer to.

Examples of shaping across settings

SettingTerminal behaviorSuccessive approximations
LaboratoryRat presses a leverFaces lever → approaches → touches → rests paw on it → presses
Speech / early interventionChild says "water"Any vocalization → "wa" → "wa-wa" → "wa-ter" → "water" clearly; the same method Lovaas used to build first words[9]
ParentingToilet trainingSits on potty clothed → sits unclothed → sits at scheduled times → urinates on potty → initiates independently[10]
ClassroomShy student answers in classNods → one-word answer to a direct question → a sentence → volunteers an answer → asks a question
Dog training"Go to mat"Looks at mat → steps toward → one paw on → all four paws → lies down → stays as handler moves away
RehabilitationStroke patient uses affected armSmall movements reinforced, then larger range, then functional tasks — the core of constraint-induced movement therapy[11]
Psychiatric careMute patient speaksEye movement toward gum → lip movement → any sound → a word → answering questions, in a classic 1960 case[12]
Self-managementWriting 500 words a dayOpen the document → one sentence → one paragraph → 100 words → 250 → 500
Business (analogy)A finished productMinimum viable version → customer feedback → iterate; a loose analogy, since the "reinforcer" is market data rather than an immediate consequence for a specific response

Shaping vs. chaining

Shaping builds a new form of a single response. Chaining links existing responses into a sequence, where each step produces the cue for the next and the whole chain ends in a reinforcer. Making a bed, brushing teeth, and a dog's retrieve are chains. The stimulus each step produces does two jobs at once: it is the discriminative stimulus for the next response and a conditioned reinforcer for the one just completed, which is why chains hold together — and why a link that stops paying early in a chain lets everything after it fall apart.[16] The first step in chaining is a task analysis: breaking the sequence into teachable components.[13]

AspectShapingForward chainingBackward chainingTotal-task chaining
What is taughtA new response formA sequence, first step firstA sequence, last step firstWhole sequence every trial
HowReinforce closer approximations; extinguish earlier onesTeach step 1 to mastery, then 1+2, and so onTrainer does all but the last step; learner completes it and is reinforced; then last two…Learner attempts all steps with prompts where needed
ReinforcerAfter each approximationAfter the last mastered stepAlways at the natural end of the chainAt the end, plus prompts faded
Best forBehavior that does not yet occur in any formLearners who can already do the early stepsLearners who benefit from finishing every trial with successLearners who can do most steps already
ExampleTeaching a first wordPutting on a coatZipping a jacket (start with the last inch)Making a sandwich with a picture schedule

In practice the two combine: a step within a chain that the learner cannot yet perform is shaped. Backward chaining has a special advantage — the learner completes the chain and contacts the terminal reinforcer on every trial, and each newly added step is reinforced by the chance to perform the steps already mastered.[13]

Shaping vs. prompting and fading

A prompt is help added before or during the response to make it occur: a verbal instruction, a gesture, a model to imitate, or physical guidance. Prompting produces the behavior now; shaping waits for it to emerge. Both end in the same place — the behavior under the control of its natural cue — but prompting requires fading: removing the help gradually so the learner does not become dependent on it.[13]

The rule of thumb: prompt when the behavior is physically possible but the learner does not know what is wanted; shape when the behavior does not yet exist in the repertoire. Many programs do both — prompt a rough form, fade the prompt, then shape toward fluency. More on cues, prompts, and stimulus control ›

Shaping vs. luring in animal training

Trainers distinguish free shaping — waiting for the animal to offer approximations and marking them — from luring, in which food is used to steer the animal into position (a treat moved over a dog's head produces a sit). Luring is faster for simple positions but has a cost: the dog learns to follow the food, and the lure must itself be faded or the behavior never happens without it. Shaped behavior tends to be more durable and produces animals that actively experiment; luring is a prompt and should be treated like one. Reinforcement-based dog training ›

Common shaping mistakes

Shaping yourself: technology and self-management

Most successful habit-building is shaping in disguise. The fitness watch that suggests a slightly higher daily goal after a week of hitting the old one, the running program that adds a minute each week, the language app that lengthens sessions as you improve — all are raising criteria on a behavior they reinforce immediately. The ones that fail usually fail by lumping: they ask for the terminal behavior on day one.

To shape your own behavior, start with an approximation you will certainly emit — one push-up, one sentence, shoes on by the door — and reinforce it on the spot, even if only by marking it done. When the small behavior is happening on most days, raise the criterion a little. Real-world habit research suggests automaticity takes weeks to months and varies widely between people and behaviors, so raise the bar on the evidence of your own record, not a calendar.[14] A missed step is data: the criterion was raised too far, and the correct response is to split it, not to try harder.

The percentile schedule: shaping as a formula

Because human shapers drift — too generous on good days, too strict on bad ones — researchers formalized the procedure. In a percentile schedule the criterion is set from the learner's own recent performance: a response is reinforced if it beats a fixed percentage of the last several responses (say, half of the last ten). The criterion rises exactly as fast as the learner improves, falls if performance drops, and keeps the rate of reinforcement roughly constant. Gregory Galbicka's 1994 review proposed bringing the method into applied settings; it remains the clearest statement of what a good shaper does intuitively.[15]

Key takeaways

Check yourself

You are shaping yourself toward writing 500 words a day. After a week at 100 words you jump to 400 and miss three days in a row. What went wrong, and what is the fix?

The criterion was raised too far, so reinforcement dropped and the behavior went on extinction. A missed step is data, not a character flaw: the fix is always to split the step, not to try harder.

A trainer shaping a dog to lie on its mat sometimes clicks when only two paws are on the mat, "because it tried." Why is this a problem?

Reinforcing a below-criterion response teaches that the criterion is negotiable. Differential reinforcement is the engine of shaping, with some forms reinforced and others not; without the extinction half, the behavior settles at the easiest form that still pays.

A trainer moves a treat over a dog's head so that it sits, and repeats this for a week. The dog now sits only when a treat is in the hand. Was the sit shaped?

No. That is luring, in which food steers the animal into position; it is a prompt, and it must be faded or the behavior never happens without it. Shaping waits for the animal to offer approximations and marks them, which produces more durable behavior.

A child can put each arm into a coat sleeve, pull the coat up, and zip it, but never does them in order without help. Shaping or chaining?

Chaining. The responses already exist; what is missing is the sequence, in which each step produces the cue for the next and the whole chain ends in a reinforcer. Shaping is for a behavior that does not yet occur in any form.

Explain it to a friend. Explain how shaping builds a behavior that has never happened yet, without using the words "step," "approximation," or "criterion."

Frequently asked questions

What is shaping in psychology?

Shaping is a procedure in operant conditioning for teaching a behavior that does not yet occur. The trainer reinforces successive approximations — responses that come progressively closer to the target — while withholding reinforcement from earlier, cruder forms, until the target behavior is performed.

What are successive approximations?

The series of intermediate behaviors between the learner's starting point and the goal, each a little closer to the goal than the last. Teaching a rat to press a lever, the approximations might be facing the lever, approaching it, touching it, and pressing it.

What is an example of shaping?

A parent teaching a toddler to say "water" praises "wa," then only "wa-wa," then only the full word. A dog trainer teaching "go to your mat" clicks and treats for looking at the mat, then a step toward it, then a paw on it, then all four paws, then lying down.

What is the difference between shaping and chaining?

Shaping creates a new form of a single behavior by reinforcing closer approximations. Chaining links behaviors the learner can already do into a sequence, such as the steps of brushing teeth. Chaining can be done forward, backward, or as a total task, and a step the learner cannot perform is often shaped separately.

Who invented shaping?

B. F. Skinner, with Keller Breland and Norman Guttman, discovered shaping in 1943 while teaching a pigeon to "bowl" during Project Pigeon. Skinner described the method in Science and Human Behavior (1953) and for a general audience in "How to Teach Animals" (Scientific American, 1951).

What is differential reinforcement of successive approximations?

It is the technical definition of shaping. "Differential reinforcement" means some responses are reinforced and others are placed on extinction; "successive approximations" means the reinforced responses are, in sequence, progressively closer to the target behavior.

How is shaping used in ABA therapy?

Behavior analysts use shaping to build first words and speech sounds, motor skills, feeding and self-care behaviors, social responses such as eye contact, and tolerance of medical or dental procedures. It is usually combined with prompting, fading, and chaining within a task-analyzed program.

References

  1. Skinner, B. F. (1953). Science and Human Behavior. Macmillan.
  2. Neuringer, A. (2002). Operant variability: Evidence, functions, and theory. Psychonomic Bulletin & Review, 9(4), 672–705. See also Page, S., & Neuringer, A. (1985). Variability is an operant. Journal of Experimental Psychology: Animal Behavior Processes, 11(3), 429–452.
  3. Antonitis, J. J. (1951). Response variability in the white rat during conditioning, extinction, and reconditioning. Journal of Experimental Psychology, 42(4), 273–281.
  4. Peterson, G. B. (2004). A day of great illumination: B. F. Skinner's discovery of shaping. Journal of the Experimental Analysis of Behavior, 82(3), 317–328.
  5. Skinner, B. F. (1958). Reinforcement today. American Psychologist, 13(3), 94–99.
  6. Skinner, B. F. (1951). How to teach animals. Scientific American, 185(6), 26–29.
  7. Skinner, B. F. (1962). Two "synthetic social relations." Journal of the Experimental Analysis of Behavior, 5(4), 531–533.
  8. Pryor, K. (1984). Don't Shoot the Dog! The New Art of Teaching and Training. Simon & Schuster.
  9. Lovaas, O. I., Berberich, J. P., Perloff, B. F., & Schaeffer, B. (1966). Acquisition of imitative speech by schizophrenic children. Science, 151(3711), 705–707.
  10. Azrin, N. H., & Foxx, R. M. (1971). A rapid method of toilet training the institutionalized retarded. Journal of Applied Behavior Analysis, 4(2), 89–99.
  11. Taub, E., Crago, J. E., Burgio, L. D., Fleming, W. C., Nepomuceno, C. S., Connell, J. S., & Miller, N. E. (1994). An operant approach to rehabilitation medicine: Overcoming learned nonuse by shaping. Journal of the Experimental Analysis of Behavior, 61(2), 281–293.
  12. Isaacs, W., Thomas, J., & Goldiamond, I. (1960). Application of operant conditioning to reinstate verbal behavior in psychotics. Journal of Speech and Hearing Disorders, 25(1), 8–12.
  13. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.
  14. Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. European Journal of Social Psychology, 40(6), 998–1009.
  15. Galbicka, G. (1994). Shaping in the 21st century: Moving percentile schedules into applied settings. Journal of Applied Behavior Analysis, 27(4), 739–760.
  16. Kelleher, R. T., & Gollub, L. R. (1962). A review of positive conditioned reinforcement. Journal of the Experimental Analysis of Behavior, 5(4), 543–597.