The complete guide

Operant Conditioning

Operant conditioning is learning from consequences: a behavior becomes more likely if it’s followed by reinforcement and less likely if it’s followed by punishment.

0 / 100
0
pellets
Press the lever. Go on.
350 BC
Aristotle: “Pleasure and pain are also the standards by which we all, in a greater or less degree, regulate our actions.”
1898
Thorndike’s law of effect
1937
Skinner coins the term “operant conditioning”
Now
Reinforcement learning is used to improve AI

What is operant conditioning?

Definition

Operant conditioning is a form of learning in which the future frequency of a behavior is changed by the consequences that follow it. Behaviors followed by reinforcement become more likely; behaviors followed by punishment become less likely.

The term was introduced by American psychologist B. F. Skinner in 1937, building on Edward Thorndike's law of effect (1898). It is also called instrumental conditioning, and it is the foundation of applied behavior analysis, modern animal training, and most evidence-based habit-change methods.[1][2]

The machine you’re already in

Sometime in the last hour you picked up your phone without deciding to. You weren’t looking for anything in particular. You just checked. Then, probably, you kept scrolling.

Those are two different mechanisms, and behavior analysts have a name for each. Checking pays off on a time basis — either something arrived since you last looked or it didn’t, and pressing the button twice as often will not make messages arrive twice as fast. That is a variable-interval schedule, and it produces a moderate, stubbornly steady rate of responding. It is why you check again four minutes later.

Scrolling is the other kind. Each swipe is a response, and whether it pays depends on how many swipes you make rather than on how long you wait. That is a variable-ratio schedule, and of the five basic schedules — catalogued, with dozens of combinations, in a 700-page book Ferster and Skinner published in 1957 — it is the one that produces the highest and steadiest rate of responding, and the one that keeps behavior going longest after the payoffs stop. It is what a slot machine is. It is what a loot box is. It is what an infinite feed is.

None of which is a character flaw, and in most cases it is not addiction in any clinical sense. It is a schedule doing what seventy years of data say a schedule will do.

Which is the genuinely useful part. A mechanism strong enough to keep you swiping against your own stated wishes is strong enough to aim somewhere else, and aiming it is a skill rather than a personality trait. The rest of this page is that mechanism from the ground up: what a consequence has to do to count as one, the four things it can be, the five ways to time it, and what happens when it stops.

One honest caveat before you go further. Schedules explain a great deal about behavior and not all of it. People also learn by watching someone else, by building a map of a situation before any reward arrives, and — most awkwardly for the theory — through language. Where this runs out, the page says so rather than papering over it.

Watch: operant vs. classical conditioning in four minutes

Start here if you are new to the topic. This TED-Ed lesson by Peggy Andover walks through Pavlov's dogs, then shows how Skinner's operant conditioning differs — behavior first, consequence second — and how reinforcement and punishment change what an animal (or a person) does next.

"The difference between classical and operant conditioning," a TED-Ed lesson by Peggy Andover (2013). Embedded from YouTube's privacy-enhanced player.

What to watch for

Two things the video makes vivid: in classical conditioning the animal is passive — the bell and the food arrive whether or not it does anything — while in operant conditioning the animal's own action is what produces the consequence. And "negative" never means "bad": it means something was taken away. The rest of this page builds on both ideas.

Operant conditioning in one paragraph

Every organism that can learn is constantly running the same experiment: do something, notice what happens next, adjust. A rat presses a lever and a food pellet drops, so it presses again. A toddler says "please" and gets the cookie, so "please" becomes a habit. You check your phone, a notification rewards you, and the checking becomes automatic. In each case the behavior operates on the environment (hence "operant") and the environment answers back with a consequence. Operant conditioning is the study of how those consequences select which behaviors survive and which fade away.[3]

Three ideas do most of the work:

How operant conditioning works

Skinner's central insight was that behavior is not just triggered by what comes before it (as in Pavlov's reflexes), it is shaped by what comes after it. He formalized this as the three-term contingency, often written A → B → C:[3]

A

Antecedent

The cue, context, or "discriminative stimulus" that signals a behavior will pay off. The phone buzzes. The dog sees the leash. It's 9 a.m. and you're at your desk.

B

Behavior

The operant: what the organism does. Defined by its function (what it accomplishes), not by exact muscle movements.

C

Consequence

What happens immediately after. Reinforcement strengthens the behavior; punishment weakens it. If the behavior stops producing its reinforcer, extinction follows.

Two more variables determine how much a consequence matters:

Operant vs. respondent behavior

Skinner distinguished respondent behavior (reflexes elicited by a prior stimulus — salivation, the eye-blink, the startle response) from operant behavior (actions emitted by the organism and controlled by their consequences). Pavlov studied the first; Skinner studied the second. Most of what people mean by "behavior" — walking, talking, working, scrolling — is operant. Full comparison ›

Inside the Skinner box

The apparatus that made all of this measurable was Skinner's operant conditioning chamber — the "Skinner box." Thorndike had timed cats escaping from puzzle boxes one trial at a time; Skinner's innovation was a box the animal never had to leave, so it could respond whenever it liked and the rate of responding could be recorded continuously.[2] The Skinner box in full: every part, and why it was built that way ›

Schematic of an operant conditioning chamber A rectangular chamber. On the left wall: a stimulus light near the top, a lever at mid-height, and a food tray at the bottom. A speaker sits high on the right wall. The floor is a grid. Outside the chamber, a cumulative recorder draws a rising stepped line on a moving strip of paper. Stimulus light (SD) Lever — the response Food tray — the reinforcer Speaker Grid floor · the animal is free to move and respond at any time Cumulative recorder Paper moves → · pen steps up per response Slope = rate of responding
An operant conditioning chamber for a rat. A press on the lever can be programmed to deliver a food pellet on any schedule; the light signals when pressing will pay off; the recorder draws responses over time.

A typical experiment runs in four steps. The rat is kept mildly hungry so food works as a reinforcer. It first learns that the click of the food dispenser means a pellet has arrived (a conditioned reinforcer). Then the experimenter shapes lever-pressing, reinforcing closer and closer approximations until the rat presses on its own. Finally the schedule is thinned — every press, then every fifth, then an unpredictable number — and the recorder shows how the pattern of responding changes. Pigeons peck a lit key instead of pressing a lever; the logic is identical. More on Skinner and the box ›

Run your own Skinner box

At the top of this page you were the rat. Here you hold the pellet button. Reading about shaping is one thing; doing it is another. This lab puts a hungry, untrained rat in a chamber and hands you the pellet button. Magazine-train it, shape a lever press one approximation at a time, put the press on a schedule, then take the food away and watch extinction — all in about three minutes. The counters under the chamber track every pellet you deliver and every press the rat makes, from the first step on.

signal light lever food tray

Pellets delivered 0
Lever presses 0
Press rate
Phase Magazine training
  1. Magazine trainingTeach the rat what the click means.
  2. ShapingReinforce successive approximations of a lever press.
  3. SchedulesCompare CRF, ratio, and interval schedules.
  4. ExtinctionStop reinforcing; watch the burst, then recovery.

The rat is a stylized model — its tendencies shift with what you reinforce, drift back when you don't, and follow the schedule patterns Ferster and Skinner documented — not a replay of real data. Real shaping takes longer and real rats are more surprising.

Want the lab on its own page to link, share, or assign? Open the virtual Skinner box ›

The four quadrants: reinforcement and punishment

Every consequence can be sorted along two questions. Did the behavior increase or decrease? (That tells you whether it was reinforcement or punishment.) Was a stimulus added or removed? (That tells you whether it was "positive" or "negative.") Crucially, in this vocabulary positive and negative mean plus and minus — added and removed — not good and bad.

The two-question test

Ask, in this order: (1) Did the behavior become more or less likely? More likely means reinforcement; less likely means punishment. (2) Was something added or removed? Added means positive; removed means negative. Those two answers name the quadrant every time.

The four types of operant conditioning at a glance

TypeWhat happens after the behaviorEffect on the behaviorExample
Positive reinforcementA stimulus is addedIncreasesA dog sits and gets a treat; sitting becomes more frequent
Negative reinforcementA stimulus is removedIncreasesYou buckle up and the seat-belt chime stops; buckling up becomes faster and more reliable
Positive punishmentA stimulus is addedDecreasesYou touch a hot pan and get burned; touching hot pans becomes rarer
Negative punishmentA stimulus is removedDecreasesA teenager breaks curfew and loses the car keys; breaking curfew becomes rarer

A fifth process, extinction, is not a quadrant: the reinforcer that used to follow the behavior simply stops arriving, and the behavior fades.

Which quadrant is it? An interactive check

Almost everyone gets one pair backwards the first time, and it is nearly always negative reinforcement mistaken for punishment. Answer the two questions about any scenario and the tool will classify it.

Five mistakes almost everyone makes

  1. Treating "negative reinforcement" as a polite word for punishment. It is the opposite: negative reinforcement makes a behavior more likely by taking something unpleasant away. Taking an aspirin to end a headache is negative reinforcement of aspirin-taking.
  2. Reading "positive" as good and "negative" as bad. They mean added and removed. A spray of water in a dog's face is positive punishment.
  3. Calling something a reinforcer because it seems nice. A reinforcer is anything that increases the behavior it follows — it is defined by its effect. Scolding that a bored child finds attention-worthy is a reinforcer; a sticker a teenager finds embarrassing is not.
  4. Mixing up operant and classical conditioning. If the key event comes before the response and the response is a reflex (salivating, flinching), it is classical. If the key event comes after a voluntary behavior, it is operant.
  5. Confusing extinction with punishment. Extinction means the reinforcer simply stops arriving; nothing is added or taken away as a consequence. The behavior fades — usually after a brief burst — rather than being suppressed.

Try three

Reinforcement versus punishment: what the evidence says

Skinner believed punishment was a poor way to change behavior, in part because an early experiment by his student W. K. Estes suggested that punishment only temporarily suppressed responding.[7] Later research complicated that picture: punishment can produce lasting decreases when it is immediate, consistent, and sufficiently intense from the outset.[8] But those same studies documented why practitioners still prefer reinforcement:

In parenting specifically, a large body of research links corporal punishment to worse, not better, long-term behavioral outcomes.[9] Modern applied behavior analysis therefore treats reinforcement-based procedures as the default and reserves punishment for narrow, supervised cases where reinforcement alone has failed and the behavior is dangerous.[10]

Reinforcement: the two types, kinds of reinforcers, and what makes it work ›
Punishment: the two types, the side effects, and the alternatives ›

Schedules of reinforcement

Once a behavior is learned, how often it gets reinforced changes both how fast the organism responds and how long the behavior persists when reinforcement stops. Ferster and Skinner catalogued these patterns in a 700-page 1957 volume, and the basic findings have held up for seventy years.[11]

ScheduleReinforcer delivered…Typical response patternEveryday example
Continuous (CRF)after every responseFast learning; fast extinctionA vending machine
Fixed ratio (FR)after a set number of responsesHigh rate with a pause after each reinforcerPaid per piece; "buy 10, get 1 free"
Variable ratio (VR)after an unpredictable number of responsesHighest, steadiest rate; most resistant to extinctionSlot machines; social-media feeds
Fixed interval (FI)for the first response after a set time"Scallop": slow after a reinforcer, accelerating as the interval endsChecking the oven as the timer nears zero
Variable interval (VI)for the first response after an unpredictable timeModerate, steady rateChecking email

The practical rule: use continuous reinforcement to build a behavior, then thin to an intermittent schedule to make it durable. The variable-ratio schedule is why gambling and infinite-scroll apps are so hard to quit — and why a behavior you reinforce only sometimes can end up stronger than one you reinforce every time. Run the interactive schedule simulator › · How organisms choose between schedules: the matching law ›

Extinction, shaping, and stimulus control

Extinction

When a previously reinforced behavior stops producing reinforcement, it gradually declines. But not immediately: there is usually an extinction burst — a temporary spike in the frequency, intensity, and variability of the behavior — before it fades. (Push the elevator button; nothing happens; you push it harder and faster before giving up.) Behavior that has been extinguished can also show spontaneous recovery after a rest period. More on extinction ›

Shaping

Complex behavior is rarely emitted fully formed, so it can't simply be reinforced. Shaping solves this by reinforcing successive approximations — first any movement toward the lever, then touching it, then pressing it. Skinner used shaping to teach pigeons to play ping-pong; trainers use it to teach dolphins to jump through hoops; speech therapists use it to build words from sounds. How shaping works ›

Stimulus control and the antecedent

A behavior reinforced in one context and not in another comes under stimulus control: it appears when the "discriminative stimulus" (SD) is present and not otherwise. The rat presses only when the light is on. You swear with friends and not with your grandmother. This is the "A" in A-B-C, and it is the most under-used lever in self-improvement — changing the cue is often easier than willing a new response. Stimulus control in depth › · The ABC model ›

Operant vs. classical conditioning

The two great forms of associative learning are often confused. The clean distinction is what gets associated with what:

AspectClassical (Pavlovian) conditioningOperant (instrumental) conditioning
AssociationStimulus ↔ stimulus (bell → food)Behavior ↔ consequence (press → food)
Behavior typeInvoluntary, reflexive (salivation, fear, nausea)Voluntary, "emitted" (pressing, speaking, working)
Organism's rolePassive; the stimulus is presented regardlessActive; the consequence depends on what it does
Timing of key eventStimulus comes before the responseConsequence comes after the response
FoundersIvan Pavlov (1890s–1927)Edward Thorndike (1898); B. F. Skinner (1937–38)

In real life the two run together. The sound of the treat bag classically conditions excitement in a dog and operantly reinforces running to the kitchen. Full comparison with examples ›

Can you spot it?

Eight scenarios, instant explanations, no score kept. When you can call all eight without hesitating, you have it — and there is a longer set of twenty.

What happens in the brain

Reinforcement has a physical address. In 1953 James Olds and Peter Milner found that a rat would press a lever thousands of times an hour for a pulse of electricity to its own brain, and in 1997 Wolfram Schultz and colleagues showed what the relevant neurons are doing: midbrain dopamine cells fire when a reinforcer is better than expected, fall silent when it is exactly as expected, and dip when an expected one fails to arrive.[21][16] That reward prediction error is the brain's teaching signal, and it explains why unpredictable reinforcers hold behavior so well and why a fully predictable one stops teaching. It is not a pleasure signal: Kent Berridge and Terry Robinson showed that animals without dopamine still like sugar but no longer want it.[22] The neuroscience of operant conditioning, in depth ›

A brief history

The full history of operant conditioning › · B. F. Skinner: life, work, and the Skinner box › · The original books, full text, in the library ›

Examples of operant conditioning in everyday life

Once you know the pattern you see it everywhere:

50+ examples, sorted by quadrant and setting ›

Operant conditioning in pop culture

Two scenes worth watching a second time. Both are embedded from YouTube and load only when you press play.

What to look for: every time Penny does something Sheldon likes, a chocolate appears — positive reinforcement, delivered immediately, on a continuous schedule. When Leonard objects, Sheldon reaches for a spray bottle — that would be positive punishment. Sheldon even names the procedure.
What almost everyone gets wrong here: Jim calls it Pavlov, but is it? Dwight's hand reaching out is a voluntary behavior that has been reinforced with a mint whenever the chime sounds — the chime is working as a discriminative stimulus, which makes the reaching operant. The dry mouth he notices is the classical part. Most real learning is both at once.

Clips are uploaded by third parties and may disappear; NBC hosts the official Office clip.

Key terms to know

The twelve terms that do most of the work, defined the way behavior analysts define them. Hover or tap any underlined term anywhere on this page for its definition; the full glossary has 128 entries.

Operant
A class of behavior defined by its effect on the environment (what it accomplishes), not by its exact form. Lever-pressing with the left paw or the right paw is the same operant.
Reinforcer
Any consequence that increases the future frequency of the behavior it follows. Positive reinforcers are added; negative reinforcers are removed.
Punisher
Any consequence that decreases the future frequency of the behavior it follows.
Primary vs. secondary reinforcer
Primary reinforcers work without learning (food, water, warmth). Secondary (conditioned) reinforcers acquire their power by being paired with primary ones — money, praise, grades, a clicker.
Three-term contingency
Antecedent → Behavior → Consequence: the basic unit of analysis. The ABC model ›
Discriminative stimulus (SD)
A cue that signals a behavior will be reinforced. When a behavior reliably occurs in its presence and not otherwise, the behavior is under stimulus control.
Schedule of reinforcement
The rule for which responses get reinforced: continuous, fixed ratio, variable ratio, fixed interval, or variable interval. Schedules ›
Extinction
Withholding the reinforcer that maintained a behavior, so the behavior declines — often after an extinction burst, a temporary increase. Extinction ›
Shaping
Building a new behavior by reinforcing successive approximations of it. Shaping ›
Motivating operation
A condition such as deprivation or satiation that changes how effective a reinforcer is. Food reinforces a hungry rat, not a full one.
Premack principle
A more probable behavior can reinforce a less probable one — "finish your homework, then you can play." Premack ›
Law of effect
Thorndike's 1898 principle that responses followed by satisfying consequences are strengthened and those followed by discomfort are weakened — the ancestor of operant conditioning.

Criticisms and limitations

An honest account includes what operant conditioning does not explain well.

What survived every critique is the core: the law of effect, schedule effects, extinction bursts, stimulus control, and shaping replicate across species and remain the working toolkit of clinicians, teachers, and trainers. Reinforcement learning — the branch of AI behind game-playing systems and the tuning of language models — is a mathematical descendant of the same ideas.[20]

How to use operant conditioning on yourself

The same contingency that trains a pigeon can be turned inward, and most self-improvement advice ignores two-thirds of it: it obsesses over motivation, which is not a term in the equation, and neglects antecedents and consequences, which are. The protocol is short. Attach the new behavior to a cue that already happens every day; shrink the behavior until it is almost embarrassing; deliver a small reinforcer within seconds, every time at first; then thin the schedule so the habit survives missed days. The full seven-step protocol, with the evidence on how long habits take ›

Design your own contingency

Fill in the three terms for a habit you actually want. The tool checks each against the science — is the cue stable, is the behavior really a behavior, is the consequence immediate — and gives you a card to print. Nothing you type leaves your browser. Full guide to building habits ›

The app

Operant runs this loop for you

A habit app for iPhone and Apple Watch from the publisher of this site. Each habit is set up as an antecedent, a behavior and a consequence — the three terms, not just the middle one. Free to download and try; a subscription unlocks the full app.

Before you leave: can you answer these without opening them?

What is operant conditioning in simple terms?

Operant conditioning is learning from consequences. When a behavior is followed by something good (or the removal of something bad), it happens more often. When it is followed by something bad (or the loss of something good), it happens less often. The organism "operates" on its environment and the results shape what it does next.

Who discovered operant conditioning?

The underlying principle — the law of effect — was discovered by Edward Thorndike in 1898 through his puzzle-box experiments with cats. B. F. Skinner named it "operant" conditioning in 1937, developed the experimental methods to study it, and built the field around it beginning with The Behavior of Organisms in 1938.

What are the four types of operant conditioning?

Positive reinforcement (add something, behavior increases), negative reinforcement (remove something, behavior increases), positive punishment (add something, behavior decreases), and negative punishment (remove something, behavior decreases). "Positive" and "negative" refer to adding and removing a stimulus, not to whether the outcome is good or bad.

What is Skinner's theory of operant conditioning?

Skinner's theory is that behavior is selected by its consequences, much as species are selected by their environments. A behavior that is followed by reinforcement becomes more frequent; one followed by punishment or by no reinforcement at all becomes less frequent. He distinguished this operant behavior, which acts on the environment, from respondent behavior, the reflexes studied by Pavlov, and he showed that the schedule on which reinforcement arrives controls how fast and how persistently an organism responds. The theory deliberately explains behavior by its history of consequences rather than by inner states such as wants or intentions, which Skinner treated as behavior to be explained rather than as causes. B. F. Skinner: the theory, the experiments, and the critiques ›

Why is it called "operant" conditioning?

Skinner chose the word in 1937 because the behavior operates on the environment to produce a consequence: the rat's press operates the lever, the lever delivers food, and the food changes future pressing. He contrasted operant behavior with respondent behavior, which is elicited by a stimulus that comes before it, as a puff of air elicits a blink. The older name, instrumental conditioning, makes the same point from the other side: the behavior is instrumental in producing the outcome.

What are the three components of operant conditioning?

The antecedent, the behavior, and the consequence, usually written A-B-C and called the three-term contingency. The antecedent is the situation or cue that sets the occasion for the behavior; the behavior is what the organism does; the consequence is what follows, and it is the consequence that changes how likely the behavior is next time. Every example on this site can be broken into those three parts. The ABC model in depth ›

Is negative reinforcement the same as punishment?

No — this is the most common mistake in the whole subject. Negative reinforcement increases a behavior by removing something unpleasant (taking a painkiller to end a headache). Punishment decreases a behavior. "Negative" only means something was taken away. The two forms of negative reinforcement, escape and avoidance, have their own page.

Which schedule of reinforcement is most resistant to extinction?

The variable-ratio schedule, in which reinforcement follows an unpredictable number of responses. Because the organism can never tell whether the next response will pay off, responding persists long after reinforcement has stopped. This is why gambling and social-media checking are hard to extinguish.

References

  1. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. Psychological Review Monograph Supplement, 2(4), 1–109. See also Thorndike, E. L. (1911). Animal Intelligence: Experimental Studies. Macmillan. Read the 1898 monograph as Chapter II of the 1911 book in the library ›
  2. Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
  3. Skinner, B. F. (1953). Science and Human Behavior. Macmillan.
  4. Skinner, B. F. (1981). Selection by consequences. Science, 213(4507), 501–504.
  5. Grice, G. R. (1948). The relative effects of delay of reinforcement in the discrimination learning of rats. Journal of Experimental Psychology, 38(1), 1–16. See also Lattal, K. A. (2010). Delayed reinforcement of operant behavior. Journal of the Experimental Analysis of Behavior, 93(1), 129–139.
  6. Skinner, B. F. (1948). 'Superstition' in the pigeon. Journal of Experimental Psychology, 38(2), 168–172.
  7. Estes, W. K. (1944). An experimental study of punishment. Psychological Monographs, 57(3), i–40.
  8. Azrin, N. H., & Holz, W. C. (1966). Punishment. In W. K. Honig (Ed.), Operant Behavior: Areas of Research and Application (pp. 380–447). Appleton-Century-Crofts.
  9. Gershoff, E. T., & Grogan-Kaylor, A. (2016). Spanking and child outcomes: Old controversies and new meta-analyses. Journal of Family Psychology, 30(4), 453–469.
  10. Behavior Analyst Certification Board. (2020). Ethics Code for Behavior Analysts. BACB. See also Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.
  11. Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
  12. Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. Journal of General Psychology, 16, 272–279.
  13. Chomsky, N. (1959). A review of B. F. Skinner's Verbal Behavior. Language, 35(1), 26–58.
  14. Breland, K., & Breland, M. (1961). The misbehavior of organisms. American Psychologist, 16(11), 681–684.
  15. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. Journal of Applied Behavior Analysis, 1(1), 91–97.
  16. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593–1599.
  17. Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208; Bandura, A. (1977). Social Learning Theory. Prentice-Hall.
  18. Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6), 627–668.
  19. Cameron, J., & Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. Review of Educational Research, 64(3), 363–423.
  20. Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
  21. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology, 47(6), 419–427; Olds, J. (1958). Self-stimulation of the brain. Science, 127(3294), 315–324.
  22. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? Brain Research Reviews, 28(3), 309–369.