# Shaping in Operant Conditioning: Successive Approximations, Step by Step

> Shaping is differential reinforcement of successive approximations to a target behavior. How it works, a step-by-step protocol, examples, shaping vs. chaining.

- Source: https://operantconditioning.com/shaping/
- Author: Ryan Martinson (https://operantconditioning.com/about/)
- Publisher: Operant Conditioning Inc.
- Published: 2026-09-07 · Updated: 2026-09-10
- License: https://operantconditioning.com/terms/#copyright (quote with attribution and a link; free to reproduce for non-commercial teaching)

---

*Core process*

You cannot reinforce a behavior that never happens. Shaping is how operant conditioning builds behavior that does not yet exist — one small approximation at a time, in pigeons, children, patients, and yourself.

> **Definition**
>
> **Shaping** is the **differential reinforcement of successive approximations** to a target behavior. Responses that come closer to the goal are reinforced; earlier, cruder forms are no longer reinforced; and the criterion for reinforcement is moved step by step until the target behavior appears.
>
> The term is B. F. Skinner's, and so is the analogy: "Operant conditioning shapes behavior as a sculptor shapes a lump of clay." Shaping is the standard method for teaching a new response in [applied behavior analysis](https://operantconditioning.com/applications/#aba), animal training, rehabilitation, and skill learning of every kind.[1]

**In brief**

- Shaping is differential reinforcement of successive approximations: reinforce responses closer to the goal, extinguish cruder ones, and raise the criterion step by step.
- Reinforcement acts only on behavior that occurs; shaping builds behavior that does not yet exist by selecting from the learner's natural variation.
- Mark each approximation immediately, keep the steps small enough that reinforcement stays frequent, and [thin the schedule](https://operantconditioning.com/schedules-of-reinforcement/) once the terminal behavior is reliable.

## Why shaping is needed

[Reinforcement](https://operantconditioning.com/positive-reinforcement/) can only act on behavior that occurs. A rat in an operant chamber will not press the lever by chance for a very long time; a child who has never said "water" cannot be praised for saying it; a stroke patient who cannot lift an arm cannot be rewarded for lifting it. If you wait for the finished behavior, you wait forever.

Shaping solves the problem by reinforcing whatever the organism *can* do that resembles the goal, then demanding a little more. It relies on a fact that is easy to miss: no response is ever repeated exactly. Every lever press has a slightly different force, every attempt at a word a slightly different sound. Shaping selects from that natural variation, the way breeding selects from variation in a population.[1]

## How shaping works: the mechanism

Shaping is two procedures running together — reinforcement of the current approximation and [extinction](https://operantconditioning.com/extinction/) of everything else — and it cycles through four phases:

![Shaping as a rising criterion](https://operantconditioning.com/assets/diagrams/shaping-as-a-rising-criterion.svg)

*Shaping as a moving criterion. Reinforcing only the upper tail of what the animal currently does shifts the whole distribution; the criterion is then raised again. Variability (the width of each curve) is what makes the next step possible.*

1. **Reinforce a starting behavior.** The rat turns toward the lever; a pellet drops. Turning toward the lever increases in frequency.
2. **Variability appears.** As the behavior is repeated it varies: some turns are closer, some include a step forward, some a raised paw. Extinction increases variability further — when a response stops paying off, organisms do it in new ways.[2][3]
3. **Raise the criterion.** Once turns are reliable, only turns that include a step forward are reinforced. Plain turns go on extinction; steps forward increase.
4. **Repeat.** Steps toward, touching, pawing, pressing. Each stage is built on the reinforced behavior of the previous one and the extinction-driven variability that behavior produces.

The variability step is the one people forget. Allen Neuringer's research showed that variability is not noise around a "true" response but a dimension of behavior that reinforcement controls: reinforce variety and organisms become more variable; reinforce sameness and they become stereotyped.[2] Good shaping keeps the current form reinforced often enough to persist but not so often that it hardens into a fixed habit.

> **Differential reinforcement is the engine**
>
> "Differential" means some forms of the response are reinforced and others are not. Without the extinction half you are not shaping — you are reinforcing whatever happens, and the behavior settles at the easiest form that still pays. Without the reinforcement half, the behavior disappears. The skill is holding both at once and moving the line between them.

## Skinner's discovery: the pigeon that learned to bowl

Skinner dated his understanding of shaping to a single day in 1943. He, Keller Breland, and Norman Guttman were working on [Project Pigeon](https://operantconditioning.com/bf-skinner/) on the top floor of a flour mill in Minneapolis and decided, for amusement, to teach a pigeon to swipe a small wooden ball down a miniature alley with its beak. Waiting for a full swipe went nowhere. So they reinforced any response that resembled it — a look at the ball, a move toward it, a touch — and within minutes the bird was bowling.[4][5] Skinner later wrote that the day changed how he thought about behavior: complex acts could be built rapidly from nothing by hand-delivered reinforcement.

He showed the public how in a 1951 *Scientific American* article, "How to Teach Animals," which walks a reader through establishing a conditioned reinforcer (a sound paired with food) and using it to shape a dog or pigeon to a new behavior in a single session — the recipe clicker trainers still follow.[6] His pigeons went on to play a version of ping-pong, pecking a ball back and forth across a table.[7]

## How to shape a behavior: step-by-step

1. **Define the terminal behavior.** Be exact about the finished form: "sits on the mat with all four paws for 10 seconds," "says 'water' clearly," "raises the affected arm to shoulder height." A vague goal cannot be approximated.
2. **Find the starting behavior.** Something the learner already does, at least occasionally, that is on the road to the goal. Glancing at the mat. Saying "wa." Moving the arm two inches. If it never occurs, you need an easier start.
3. **Plan the approximations.** Write down the steps you expect, but hold the plan loosely — the learner's variability will suggest better ones.
4. **Reinforce immediately and every time.** Use a conditioned reinforcer (a click, a "yes," a checkmark) to mark the exact instant the criterion is met, then deliver the real reinforcer. A second's delay reinforces whatever happened in that second instead.
5. **Raise the criterion in small steps.** Move on when the current approximation is reliable — Karen Pryor's rule of thumb is when the learner succeeds most of the time — and keep each step small enough that reinforcement stays frequent.[8]
6. **Don't stay too long; don't move too fast.** Staying reinforces a plateau until it resists change. Moving too fast puts the learner on extinction: variability spikes, then the behavior collapses. If that happens, drop back a step.
7. **Thin the schedule at the end.** Once the terminal behavior is reliable, shift to [intermittent reinforcement](https://operantconditioning.com/schedules-of-reinforcement/) and bring the behavior under the cue you want it to answer to.

## Examples of shaping across settings

| Setting | Terminal behavior | Successive approximations |
| --- | --- | --- |
| Laboratory | Rat presses a lever | Faces lever → approaches → touches → rests paw on it → presses |
| Speech / early intervention | Child says "water" | Any vocalization → "wa" → "wa-wa" → "wa-ter" → "water" clearly; the same method Lovaas used to build first words[9] |
| Parenting | Toilet training | Sits on potty clothed → sits unclothed → sits at scheduled times → urinates on potty → initiates independently[10] |
| Classroom | Shy student answers in class | Nods → one-word answer to a direct question → a sentence → volunteers an answer → asks a question |
| Dog training | "Go to mat" | Looks at mat → steps toward → one paw on → all four paws → lies down → stays as handler moves away |
| Rehabilitation | Stroke patient uses affected arm | Small movements reinforced, then larger range, then functional tasks — the core of constraint-induced movement therapy[11] |
| Psychiatric care | Mute patient speaks | Eye movement toward gum → lip movement → any sound → a word → answering questions, in a classic 1960 case[12] |
| Self-management | Writing 500 words a day | Open the document → one sentence → one paragraph → 100 words → 250 → 500 |
| Business (analogy) | A finished product | Minimum viable version → customer feedback → iterate; a loose analogy, since the "reinforcer" is market data rather than an immediate consequence for a specific response |

## Shaping vs. chaining

Shaping builds a *new form* of a single response. **Chaining** links *existing* responses into a sequence, where each step produces the cue for the next and the whole chain ends in a reinforcer. Making a bed, brushing teeth, and a dog's retrieve are chains. The stimulus each step produces does two jobs at once: it is the discriminative stimulus for the next response and a conditioned reinforcer for the one just completed, which is why chains hold together — and why a link that stops paying early in a chain lets everything after it fall apart.[16] The first step in chaining is a **task analysis**: breaking the sequence into teachable components.[13]

| Aspect | Shaping | Forward chaining | Backward chaining | Total-task chaining |
| --- | --- | --- | --- | --- |
| **What is taught** | A new response form | A sequence, first step first | A sequence, last step first | Whole sequence every trial |
| **How** | Reinforce closer approximations; extinguish earlier ones | Teach step 1 to mastery, then 1+2, and so on | Trainer does all but the last step; learner completes it and is reinforced; then last two… | Learner attempts all steps with prompts where needed |
| **Reinforcer** | After each approximation | After the last mastered step | Always at the natural end of the chain | At the end, plus prompts faded |
| **Best for** | Behavior that does not yet occur in any form | Learners who can already do the early steps | Learners who benefit from finishing every trial with success | Learners who can do most steps already |
| **Example** | Teaching a first word | Putting on a coat | Zipping a jacket (start with the last inch) | Making a sandwich with a picture schedule |

In practice the two combine: a step within a chain that the learner cannot yet perform is shaped. Backward chaining has a special advantage — the learner completes the chain and contacts the terminal reinforcer on every trial, and each newly added step is reinforced by the chance to perform the steps already mastered.[13]

## Shaping vs. prompting and fading

A **prompt** is help added before or during the response to make it occur: a verbal instruction, a gesture, a model to imitate, or physical guidance. Prompting produces the behavior now; shaping waits for it to emerge. Both end in the same place — the behavior under the control of its natural cue — but prompting requires **fading**: removing the help gradually so the learner does not become dependent on it.[13]

- **Most-to-least prompting** starts with full physical guidance and fades to lighter prompts; useful for learners who make many errors.
- **Least-to-most prompting** gives the learner a chance to respond alone and adds help only as needed; it lets independent responding happen sooner.
- **Time delay** keeps the prompt but waits progressively longer before giving it, so the learner has room to beat the prompt.

The rule of thumb: prompt when the behavior is physically possible but the learner does not know what is wanted; shape when the behavior does not yet exist in the repertoire. Many programs do both — prompt a rough form, fade the prompt, then shape toward fluency. [More on cues, prompts, and stimulus control ›](https://operantconditioning.com/abc-model/)

## Shaping vs. luring in animal training

Trainers distinguish **free shaping** — waiting for the animal to offer approximations and marking them — from **luring**, in which food is used to steer the animal into position (a treat moved over a dog's head produces a sit). Luring is faster for simple positions but has a cost: the dog learns to follow the food, and the lure must itself be faded or the behavior never happens without it. Shaped behavior tends to be more durable and produces animals that actively experiment; luring is a prompt and should be treated like one. [Reinforcement-based dog training ›](https://operantconditioning.com/dog-training/)

## Common shaping mistakes

- **Steps that are too big.** The learner rarely meets the new criterion, reinforcement drops, and the behavior extinguishes. The fix is always to split the step.
- **Staying too long on one step.** A heavily reinforced approximation becomes stereotyped and the learner stops varying. Move on while variability is still present.
- **Inconsistent criteria.** Reinforcing a below-criterion response "because they tried" teaches that the criterion is negotiable.
- **Late reinforcement.** A marker delivered a second late reinforces whatever followed the target — the head-turn after the sit. This is how [superstitious behaviors](https://operantconditioning.com/glossary/#superstitious-behavior) get built into a shaped response.
- **Lumping instead of splitting.** Karen Pryor's term for demanding two improvements at once — a longer stay *and* a straighter sit. Raise one criterion at a time, and relax the others temporarily when you introduce a new one.[8]
- **Not planning the endgame.** Shaping ends with the terminal behavior on continuous reinforcement, which is fragile. Thin the schedule and add the cue, or the behavior will not survive real life.

## Shaping yourself: technology and self-management

Most successful [habit-building](https://operantconditioning.com/habits/) is shaping in disguise. The fitness watch that suggests a slightly higher daily goal after a week of hitting the old one, the running program that adds a minute each week, the language app that lengthens sessions as you improve — all are raising criteria on a behavior they reinforce immediately. The ones that fail usually fail by lumping: they ask for the terminal behavior on day one.

To shape your own behavior, start with an approximation you will certainly emit — one push-up, one sentence, shoes on by the door — and reinforce it on the spot, even if only by marking it done. When the small behavior is happening on most days, raise the criterion a little. Real-world habit research suggests automaticity takes weeks to months and varies widely between people and behaviors, so raise the bar on the evidence of your own record, not a calendar.[14] A missed step is data: the criterion was raised too far, and the correct response is to split it, not to try harder.

### The percentile schedule: shaping as a formula

Because human shapers drift — too generous on good days, too strict on bad ones — researchers formalized the procedure. In a **percentile schedule** the criterion is set from the learner's own recent performance: a response is reinforced if it beats a fixed percentage of the last several responses (say, half of the last ten). The criterion rises exactly as fast as the learner improves, falls if performance drops, and keeps the rate of reinforcement roughly constant. Gregory Galbicka's 1994 review proposed bringing the method into applied settings; it remains the clearest statement of what a good shaper does intuitively.[15]

## Key takeaways

- Shaping is two procedures running together: reinforcement of the current approximation and extinction of everything else. Without the extinction half you are reinforcing whatever happens; without the reinforcement half the behavior disappears.
- Variability is what makes each step possible. No response is ever repeated exactly, extinction increases variability further, and reinforcement itself can make behavior more variable or more stereotyped.
- Define the terminal behavior exactly, start with something the learner already does, mark the instant the criterion is met with a conditioned reinforcer, and raise the criterion in small steps. Steps that are too big put the learner on extinction; staying too long hardens a plateau.
- Shaping builds a new form of a single response; chaining links existing responses into a sequence; prompting adds help that must then be faded. Shape when the behavior does not yet exist, and prompt when it is possible but the learner does not know what is wanted.
- Shaping ends with the terminal behavior on continuous reinforcement, which is fragile. Thin the schedule and bring the behavior under its cue, or it will not survive real life.

### Check yourself

**You are shaping yourself toward writing 500 words a day. After a week at 100 words you jump to 400 and miss three days in a row. What went wrong, and what is the fix?**

The criterion was raised too far, so reinforcement dropped and the behavior went on extinction. A missed step is data, not a character flaw: the fix is always to split the step, not to try harder.

**A trainer shaping a dog to lie on its mat sometimes clicks when only two paws are on the mat, "because it tried." Why is this a problem?**

Reinforcing a below-criterion response teaches that the criterion is negotiable. Differential reinforcement is the engine of shaping, with some forms reinforced and others not; without the extinction half, the behavior settles at the easiest form that still pays.

**A trainer moves a treat over a dog's head so that it sits, and repeats this for a week. The dog now sits only when a treat is in the hand. Was the sit shaped?**

No. That is luring, in which food steers the animal into position; it is a prompt, and it must be faded or the behavior never happens without it. Shaping waits for the animal to offer approximations and marks them, which produces more durable behavior.

**A child can put each arm into a coat sleeve, pull the coat up, and zip it, but never does them in order without help. Shaping or chaining?**

Chaining. The responses already exist; what is missing is the sequence, in which each step produces the cue for the next and the whole chain ends in a reinforcer. Shaping is for a behavior that does not yet occur in any form.

**Explain it to a friend.** Explain how shaping builds a behavior that has never happened yet, without using the words "step," "approximation," or "criterion."

## Frequently asked questions

**What is shaping in psychology?**

Shaping is a procedure in operant conditioning for teaching a behavior that does not yet occur. The trainer reinforces successive approximations — responses that come progressively closer to the target — while withholding reinforcement from earlier, cruder forms, until the target behavior is performed.

**What are successive approximations?**

The series of intermediate behaviors between the learner's starting point and the goal, each a little closer to the goal than the last. Teaching a rat to press a lever, the approximations might be facing the lever, approaching it, touching it, and pressing it.

**What is an example of shaping?**

A parent teaching a toddler to say "water" praises "wa," then only "wa-wa," then only the full word. A dog trainer teaching "go to your mat" clicks and treats for looking at the mat, then a step toward it, then a paw on it, then all four paws, then lying down.

**What is the difference between shaping and chaining?**

Shaping creates a new form of a single behavior by reinforcing closer approximations. Chaining links behaviors the learner can already do into a sequence, such as the steps of brushing teeth. Chaining can be done forward, backward, or as a total task, and a step the learner cannot perform is often shaped separately.

**Who invented shaping?**

B. F. Skinner, with Keller Breland and Norman Guttman, discovered shaping in 1943 while teaching a pigeon to "bowl" during Project Pigeon. Skinner described the method in *Science and Human Behavior* (1953) and for a general audience in "How to Teach Animals" (*Scientific American*, 1951).

**What is differential reinforcement of successive approximations?**

It is the technical definition of shaping. "Differential reinforcement" means some responses are reinforced and others are placed on extinction; "successive approximations" means the reinforced responses are, in sequence, progressively closer to the target behavior.

**How is shaping used in ABA therapy?**

Behavior analysts use shaping to build first words and speech sounds, motor skills, feeding and self-care behaviors, social responses such as eye contact, and tolerance of medical or dental procedures. It is usually combined with prompting, fading, and chaining within a task-analyzed program.

## References

1. Skinner, B. F. (1953). *Science and Human Behavior*. Macmillan.
2. Neuringer, A. (2002). Operant variability: Evidence, functions, and theory. *Psychonomic Bulletin & Review, 9*(4), 672–705. See also Page, S., & Neuringer, A. (1985). Variability is an operant. *Journal of Experimental Psychology: Animal Behavior Processes, 11*(3), 429–452.
3. Antonitis, J. J. (1951). Response variability in the white rat during conditioning, extinction, and reconditioning. *Journal of Experimental Psychology, 42*(4), 273–281.
4. Peterson, G. B. (2004). A day of great illumination: B. F. Skinner's discovery of shaping. *Journal of the Experimental Analysis of Behavior, 82*(3), 317–328.
5. Skinner, B. F. (1958). Reinforcement today. *American Psychologist, 13*(3), 94–99.
6. Skinner, B. F. (1951). How to teach animals. *Scientific American, 185*(6), 26–29.
7. Skinner, B. F. (1962). Two "synthetic social relations." *Journal of the Experimental Analysis of Behavior, 5*(4), 531–533.
8. Pryor, K. (1984). *Don't Shoot the Dog! The New Art of Teaching and Training*. Simon & Schuster.
9. Lovaas, O. I., Berberich, J. P., Perloff, B. F., & Schaeffer, B. (1966). Acquisition of imitative speech by schizophrenic children. *Science, 151*(3711), 705–707.
10. Azrin, N. H., & Foxx, R. M. (1971). A rapid method of toilet training the institutionalized retarded. *Journal of Applied Behavior Analysis, 4*(2), 89–99.
11. Taub, E., Crago, J. E., Burgio, L. D., Fleming, W. C., Nepomuceno, C. S., Connell, J. S., & Miller, N. E. (1994). An operant approach to rehabilitation medicine: Overcoming learned nonuse by shaping. *Journal of the Experimental Analysis of Behavior, 61*(2), 281–293.
12. Isaacs, W., Thomas, J., & Goldiamond, I. (1960). Application of operant conditioning to reinstate verbal behavior in psychotics. *Journal of Speech and Hearing Disorders, 25*(1), 8–12.
13. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). *Applied Behavior Analysis* (3rd ed.). Pearson.
14. Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. *European Journal of Social Psychology, 40*(6), 998–1009.
15. Galbicka, G. (1994). Shaping in the 21st century: Moving percentile schedules into applied settings. *Journal of Applied Behavior Analysis, 27*(4), 739–760.
16. Kelleher, R. T., & Gollub, L. R. (1962). A review of positive conditioned reinforcement. *Journal of the Experimental Analysis of Behavior, 5*(4), 543–597.


## About the author

Ryan holds a master's degree from UCLA, where he studied animal behavior in Daniel Blumstein's lab and was part of the university's Evolutionary Medicine Program, which applies findings from evolutionary biology and animal behavior to human health. He founded Operant Conditioning Inc. and built the Operant habit app (https://operantconditioning.com/app/). Every page here is written from the primary literature and cites it. How pages are checked: https://operantconditioning.com/about/#editorial-standards

## Related

- [Positive reinforcement](https://operantconditioning.com/positive-reinforcement/): The consequence that does the work in every approximation.
- [Extinction](https://operantconditioning.com/extinction/): The other half of differential reinforcement — bursts, variability, and recovery.
- [Habits](https://operantconditioning.com/habits/): Shaping your own behavior with antecedents, tiny steps, and immediate consequences.
