The brain · Dopamine and prediction error
The Neuroscience of Operant Conditioning
Skinner deliberately treated the organism as a black box. Seventy years of neuroscience have opened it, and what is inside looks remarkably like the law of effect: a broadcast teaching signal, a window of a few seconds, and two separate systems for wanting and liking.
In one paragraph
Reinforcement has a physical address. When a consequence is better than the brain predicted, midbrain dopamine neurons fire a brief burst that is broadcast to the striatum and frontal cortex, strengthening whichever synapses were active in the preceding seconds. When the consequence is exactly as predicted, they stay quiet; when an expected reinforcer fails to arrive, they dip below baseline. This reward prediction error is the teaching signal of operant conditioning, and it is why immediacy, contingency, and unpredictability matter so much.[4][5]
In brief
- Midbrain dopamine neurons signal reward prediction error: a burst when a consequence is better than predicted, silence when expected, a dip when worse.
- Dopamine drives wanting, not liking: rats without dopamine still enjoy sugar but stop seeking it, and addiction is sensitized wanting.
- Dopamine strengthens only synapses active in the preceding seconds, which is why immediacy matters, and extended training turns goal-directed actions into habits.
Reinforcement has an address: Olds and Milner
In 1953 James Olds and Peter Milner, working at McGill, implanted an electrode into the septal area of a rat's brain and arranged for a lever press to deliver a brief electrical pulse. The rat pressed, and kept pressing; they published the finding the following year. Olds went on to map the sites that supported intracranial self-stimulation: rats would respond thousands of times an hour and cross electrified grids to reach the lever. Routtenberg and Lindy later gave hungry rats a daily hour with both a food lever and a stimulation lever; some spent the hour on stimulation and lost weight.[1][2][3] For the first time, a reinforcer had been produced by acting directly on the nervous system, and the anatomy of the effective sites — the medial forebrain bundle and the pathways it carries — pointed toward a particular chemical system.
Dopamine is a prediction-error signal, not a pleasure signal
The pathways Olds had stimulated carry axons of dopamine neurons from the midbrain (the ventral tegmental area and substantia nigra) to the striatum and prefrontal cortex. Through the 1980s the working assumption was that dopamine was pleasure. Wolfram Schultz's recordings from monkeys overturned that.
Schultz trained monkeys on a simple task in which a light or sound was followed, a second or two later, by a squirt of juice, while recording individual dopamine neurons. Early in training the neurons fired when the juice arrived. Once the cue reliably predicted juice, the burst moved to the cue, and the fully predicted juice produced no response at all. And when the cue was followed by no juice, the neurons paused — activity dropped below baseline at the moment the juice should have come.[4][5] That is exactly the profile of a prediction error: positive for better-than-expected, zero for as-expected, negative for worse-than-expected. In 1996 Montague, Dayan, and Sejnowski had shown that this is the quantity a temporal-difference learning algorithm needs, and the 1997 paper by Schultz, Dayan, and Montague joined the two: the brain appeared to be running the same computation that computer scientists had derived from the law of effect.[6][4][7]
Correlation became causation with optogenetics. Activating dopamine neurons with light at the moment a reward is delivered is enough to make rats learn about a cue that would otherwise be blocked, and phasic stimulation of these neurons is sufficient to produce a conditioned preference for a place.[8][9] The dopamine burst does not merely accompany learning; it drives it.
Why this explains schedules of reinforcement
A reinforcer that is fully predictable generates no prediction error and, eventually, no dopamine response: a continuous schedule teaches fast, then goes flat. A reinforcer that arrives after an unpredictable number of responses cannot be fully predicted, so every one produces a burst. Schultz's group later found that dopamine neurons also show a slow ramp of activity that is largest when reward probability is 50% — maximal uncertainty.[10] Whether that ramp is what makes gambling compelling is still argued, but the variable-ratio schedule has, at minimum, a plausible neural signature.
Wanting is not liking
If dopamine is not pleasure, what is it for? Kent Berridge and Terry Robinson answered with a dissociation that has held up for three decades. Rats whose dopamine neurons were destroyed with the toxin 6-hydroxydopamine stopped eating — they would starve unless tube-fed — yet when sugar was placed on their tongues they showed exactly the lip-licking, paw-licking facial reactions that intact rats show. They still liked sugar. What they had lost was wanting: the motivation to work for it, approach it, and treat cues for it as attractive.[11][12] Berridge and Robinson called this attribution of attractiveness to reinforcers and their cues incentive salience. Liking, meanwhile, turned out to depend on tiny opioid "hedonic hotspots" in the nucleus accumbens and elsewhere, anatomically separate from the dopamine system.[13]
Reinforcement, at the level of neurons, is therefore mostly about wanting. Dopamine also sets how much effort an animal will spend: with dopamine reduced, rats still choose food, but they shift from a lever that pays well and requires many presses to freely available, less preferred chow.[14] The everyday consequence is familiar to anyone who has kept doing something they no longer enjoy.
Addiction as sensitized wanting
The dissociation explains the central puzzle of addiction: people keep wanting drugs that have long since stopped delivering much pleasure. Robinson and Berridge's incentive-sensitization theory proposes that repeated drug exposure sensitizes the dopamine system's response to the drug and to its cues, so that wanting grows even as liking shrinks. The syringe, the bar, the friend, the time of day become cues with enormous incentive salience, which is why relapse is so often triggered by the environment rather than by withdrawal.[15][16] In operant terms: drug taking is positively reinforced by the drug and negatively reinforced by relief from withdrawal, and the cues that precede it become discriminative stimuli and conditioned reinforcers with a sensitized grip.
Learning from the stick: punishment in the brain
Worse-than-expected outcomes produce a dopamine dip, and the brain appears to learn from dips through a different route than from bursts. Michael Frank and colleagues gave a probabilistic learning task to people with Parkinson's disease, a condition of dopamine loss. Off medication, patients were better at learning from negative feedback — which choices to avoid — than from positive feedback. On dopamine-replacing medication, the pattern reversed: they learned from positive feedback and became worse at learning from negative.[17] Frank's model attributes the two to separate striatal pathways, one facilitating action when dopamine is high and one suppressing it when dopamine is low. Separately, neurons in the lateral habenula fire when a reward is omitted or a punishment predicted, and inhibit dopamine neurons — a candidate source of the dip.[18][36] Reinforcement and punishment are not mirror images at the neural level any more than they are at the behavioral one.
The window of a few seconds: why immediacy matters
Every practical guide to reinforcement says the consequence must be immediate. The neural reason is now visible. A dopamine burst cannot strengthen every synapse in the striatum; it strengthens the ones that were recently active, which carry a short-lived molecular "eligibility trace." Yagishita and colleagues, using glutamate uncaging on single dendritic spines, found that dopamine enlarged a spine only if it arrived within roughly 0.3 to 2 seconds after the spine had been stimulated. Earlier or later, nothing happened.[19] That window is the cellular basis of contiguity: a reinforcer that comes thirty seconds after a behavior finds the trace gone and strengthens whatever happened in the last two seconds instead.
Dopamine is not the only teaching signal. Neurons of the nucleus basalis release acetylcholine across the cortex and respond to reinforcers and to the stimuli that predict them; pairing a tone with electrical stimulation of the nucleus basalis is enough to expand the tone's representation in the auditory cortex, without any behavior at all.[20][21] Reinforcement, in other words, reshapes perception as well as action.
From action to habit: two systems in the striatum
Press a lever a few hundred times for food and then make the food unappealing — pair it with a mild poison, or feed the animal to satiety. A rat with moderate training stops pressing; it "knows" what the lever produces and no longer wants it. A rat with extensive training keeps pressing anyway.[22][23] Anthony Dickinson used this reinforcer-devaluation test to distinguish goal-directed actions, which are sensitive to the current value of their outcome, from habits, which are triggered by antecedents and run off regardless.
The two have different homes. Lesions of the dorsomedial striatum leave animals unable to act on outcome value, while lesions of the dorsolateral striatum prevent habits from forming, so that over-trained animals stay sensitive to devaluation.[24][25][26] Training on interval schedules produces habits faster than training on ratio schedules, presumably because on an interval schedule the connection between how much you respond and how much you get is loose.[27] Human imaging finds the same division of labor: prediction errors in the ventral striatum during learning, with the dorsal striatum engaged when the learning must guide action.[28] Everitt and Robbins argued that addiction is this transition run to its end — from action to habit to compulsion, with control migrating from ventral to dorsal striatum as the behavior becomes cue-driven and insensitive to consequences.[29]
This is the neuroscience behind a piece of practical advice on the habits page: a well-formed habit survives the loss of motivation, for good and ill. The cue keeps producing the behavior after the outcome has lost its appeal, which is why habits are hard to break by deciding to and easier to break by changing the antecedent.
Operant conditioning in a single neuron
The principle scales down remarkably far. In 1969 Eberhard Fetz reinforced monkeys with food pellets whenever a single recorded neuron in motor cortex fired faster; within minutes the monkeys raised that neuron's firing rate, with no instruction about what they were doing.[30] In the sea slug Aplysia, Brembs and colleagues reinforced a feeding movement by stimulating a dopaminergic nerve immediately after it, and then reproduced the learning in a single identified neuron in a dish: contingent dopamine applied to neuron B51 changed its excitability the way training changed the whole animal's behavior.[31] Operant conditioning is not a trick of large brains; it is a property of neurons.
The brain as a reinforcement learner
Reinforcement learning, the branch of artificial intelligence in which an agent learns from reward signals, was built on Thorndike and Skinner and on temporal-difference learning, and the dopamine findings turned it into a theory of the brain.[7][32] The current picture has two learners running in parallel: a "model-free" system that caches the value of actions from prediction errors — the habit system — and a "model-based" system that plans using a map of how actions lead to outcomes — the goal-directed system — with control shifting between them according to which is more reliable.[33] The same framework connects to classical conditioning through the Rescorla–Wagner model of 1972, which was itself a prediction-error rule.[34]
What this means in practice
- Immediacy is not a rule of thumb; it is a molecular window. If the real reinforcer must be delayed, deliver a conditioned reinforcer — a click, a word, a checkmark — inside the window and let it bridge the gap.
- Predictable reinforcers stop teaching. Once a behavior is learned, thinning to an intermittent schedule keeps prediction errors, and dopamine, alive. The same mechanism is what makes slot machines and feeds hard to leave.
- Wanting and liking come apart. A behavior can be maintained by cues long after its outcome stopped being enjoyable; treat the cues, not just the outcome.
- Habits outlive motivation. An over-trained behavior is insensitive to devaluation. To change it, change the antecedent or make the response impossible rather than relying on wanting it less.
- Reinforcement and punishment use different circuitry. Which one a person learns from best can depend on the state of their dopamine system — one reason blanket claims that "punishment doesn't work" or "rewards don't work" are both too simple.
What is still unsettled
The prediction-error account is the best-supported theory of dopamine, not the whole story. Dopamine also ramps up as animals approach rewards, participates in movement and vigor, and is released in patterns that a single scalar error signal does not obviously explain; some researchers argue it broadcasts several different messages on different timescales.[35][14] Most of the causal work is in rodents, and human evidence rests on imaging and on patient groups. And none of it changes the functional definitions: a reinforcer is still whatever increases the behavior it follows. The neuroscience explains why the law of effect holds; it does not replace it.
Key takeaways
- Dopamine is a prediction-error signal, not a pleasure signal: midbrain dopamine neurons burst when a consequence is better than predicted, stay quiet when it is as predicted, and dip when it is worse. Because a fully predictable reinforcer produces no error, continuous reinforcement teaches fast and then goes flat, while intermittent schedules keep prediction errors alive.
- Reinforcement and punishment are not mirror images in the brain. Learning from dips runs through a different striatal pathway than learning from bursts, and which one a person learns from best can depend on the state of their dopamine system.
- Wanting and liking are separate systems: dopamine drives wanting (incentive salience) and effort, while liking depends on opioid hotspots. Addiction is sensitized wanting, which is why cues trigger relapse long after the drug stopped being enjoyable.
- A dopamine burst strengthens only synapses that were active within roughly 0.3 to 2 seconds before it, the cellular basis of contiguity. A delayed reinforcer strengthens whatever happened just before it arrived, so bridge any delay with a conditioned reinforcer.
- With extended training, control shifts from a goal-directed system in the dorsomedial striatum, sensitive to outcome value, to a habit system in the dorsolateral striatum, triggered by cues regardless of value. Habits outlive motivation, so they are easier to break by changing the antecedent than by wanting the outcome less.
Check yourself
A monkey has learned that a light predicts juice. When the juice arrives exactly as predicted, do its dopamine neurons fire?
No. Once the cue reliably predicts juice, the burst moves to the cue and the fully predicted juice produces no response, because dopamine signals prediction error rather than pleasure. If the juice is then omitted, the neurons dip below baseline at the moment it should have come.
A rat whose dopamine neurons have been destroyed stops eating and would starve unless tube-fed. Has it lost the ability to enjoy food?
No. When sugar is placed on its tongue it shows the same lip-licking and paw-licking reactions as an intact rat, so liking is intact. What it has lost is wanting: the motivation to seek food, work for it, and treat its cues as attractive, which depends on dopamine while liking depends on separate opioid hotspots.
A trainer gives a treat about thirty seconds after a good sit, once she has found the treat bag. What does the dopamine burst strengthen?
Whatever the dog did in the last couple of seconds before the treat arrived, not the sit. Dopamine enlarges a synapse only if it arrives within roughly 0.3 to 2 seconds of the synapse's activity, and the sit's eligibility trace is long gone. The fix is a conditioned reinforcer, such as a click, delivered inside the window to bridge the gap.
Two rats were trained to press a lever for food, one moderately and one extensively. The food is then made unappealing. Which rat keeps pressing, and why?
The extensively trained rat. Its pressing has become a habit, supported by the dorsolateral striatum and triggered by antecedents regardless of the outcome's current value. The moderately trained rat's pressing is still goal-directed, supported by the dorsomedial striatum, so it stops once the outcome is devalued.
Explain it to a friend. Explain why a reinforcer that arrives every single time eventually stops teaching anything, using the word "surprise" and without using the word "dopamine."
Frequently asked questions
Is dopamine the "pleasure chemical"?
No. Dopamine neurons signal reward prediction error — how much better or worse an outcome was than expected — and drive wanting (motivation and cue attraction). Pleasure, or "liking," depends on separate opioid systems. Animals with dopamine removed still show every sign of enjoying sugar; they simply stop seeking it.
What is a reward prediction error?
The difference between the reward that arrives and the reward that was predicted. A positive error (better than expected) produces a burst of dopamine and strengthens the preceding behavior; zero error (as expected) produces nothing; a negative error (worse than expected, including an omitted reward) produces a dip. It is the neural version of the law of effect and the core of reinforcement-learning algorithms.
Which part of the brain is responsible for operant conditioning?
No single part. Dopamine neurons in the midbrain provide the teaching signal; the ventral striatum learns predictions; the dorsomedial striatum supports goal-directed action; the dorsolateral striatum supports habits; the prefrontal cortex supports planning and rule use; and the amygdala and lateral habenula handle aversive outcomes. Operant learning also occurs in invertebrates with far simpler nervous systems, and even in single neurons.
Why must reinforcement be immediate?
Because the synapses that were active during a behavior stay "eligible" for strengthening only briefly. In mouse striatal neurons, dopamine strengthened a synapse only if it arrived within about 0.3–2 seconds of the synapse's activity. A delayed reinforcer strengthens whatever happened just before it arrived, not the behavior you intended.
Does the brain treat punishment as the opposite of reinforcement?
Not exactly. Omitted rewards and predicted punishments produce dopamine dips, partly driven by the lateral habenula, and learning from them appears to run through a different striatal pathway than learning from rewards. In people with Parkinson's disease, dopamine medication improves learning from positive feedback and worsens learning from negative feedback.
How does this relate to habits?
With extended practice, control of a behavior shifts from a goal-directed system (dorsomedial striatum, sensitive to whether the outcome is still valuable) to a habit system (dorsolateral striatum, triggered by cues regardless of outcome value). That is why long-standing habits continue after their rewards have lost appeal, and why changing cues works better than willpower.
References
- Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology, 47(6), 419–427.
- Olds, J. (1958). Self-stimulation of the brain. Science, 127(3294), 315–324.
- Routtenberg, A., & Lindy, J. (1965). Effects of the availability of rewarding septal and hypothalamic stimulation on bar pressing for food under conditions of deprivation. Journal of Comparative and Physiological Psychology, 60(2), 158–161.
- Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593–1599.
- Schultz, W. (1998). Predictive reward signal of dopamine neurons. Journal of Neurophysiology, 80(1), 1–27.
- Montague, P. R., Dayan, P., & Sejnowski, T. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. Journal of Neuroscience, 16(5), 1936–1947.
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
- Steinberg, E. E., Keiflin, R., Boivin, J. R., Witten, I. B., Deisseroth, K., & Janak, P. H. (2013). A causal link between prediction errors, dopamine neurons and learning. Nature Neuroscience, 16(7), 966–973.
- Tsai, H.-C., Zhang, F., Adamantidis, A., Stuber, G. D., Bonci, A., de Lecea, L., & Deisseroth, K. (2009). Phasic firing in dopaminergic neurons is sufficient for behavioral conditioning. Science, 324(5930), 1080–1084.
- Fiorillo, C. D., Tobler, P. N., & Schultz, W. (2003). Discrete coding of reward probability and uncertainty by dopamine neurons. Science, 299(5614), 1898–1902.
- Berridge, K. C., Venier, I. L., & Robinson, T. E. (1989). Taste reactivity analysis of 6-hydroxydopamine-induced aphagia: Implications for arousal and anhedonia hypotheses of dopamine function. Behavioral Neuroscience, 103(1), 36–45.
- Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? Brain Research Reviews, 28(3), 309–369.
- Peciña, S., & Berridge, K. C. (2005). Hedonic hot spot in nucleus accumbens shell: Where do μ-opioids cause increased hedonic impact of sweetness? Journal of Neuroscience, 25(50), 11777–11786.
- Salamone, J. D., & Correa, M. (2012). The mysterious motivational functions of mesolimbic dopamine. Neuron, 76(3), 470–485.
- Robinson, T. E., & Berridge, K. C. (1993). The neural basis of drug craving: An incentive-sensitization theory of addiction. Brain Research Reviews, 18(3), 247–291.
- Volkow, N. D., Koob, G. F., & McLellan, A. T. (2016). Neurobiologic advances from the brain disease model of addiction. New England Journal of Medicine, 374(4), 363–371.
- Frank, M. J., Seeberger, L. C., & O'Reilly, R. C. (2004). By carrot or by stick: Cognitive reinforcement learning in parkinsonism. Science, 306(5703), 1940–1943.
- Matsumoto, M., & Hikosaka, O. (2007). Lateral habenula as a source of negative reward signals in dopamine neurons. Nature, 447(7148), 1111–1115.
- Yagishita, S., Hayashi-Takagi, A., Ellis-Davies, G. C. R., Urakubo, H., Ishii, S., & Kasai, H. (2014). A critical time window for dopamine actions on the structural plasticity of dendritic spines. Science, 345(6204), 1616–1620.
- Richardson, R. T., & DeLong, M. R. (1990). Context-dependent responses of primate nucleus basalis neurons in a go/no-go task. Journal of Neuroscience, 10(8), 2528–2540.
- Kilgard, M. P., & Merzenich, M. M. (1998). Cortical map reorganization enabled by nucleus basalis activity. Science, 279(5357), 1714–1718.
- Adams, C. D., & Dickinson, A. (1981). Instrumental responding following reinforcer devaluation. Quarterly Journal of Experimental Psychology B, 33(2), 109–121.
- Dickinson, A. (1985). Actions and habits: The development of behavioural autonomy. Philosophical Transactions of the Royal Society B, 308(1135), 67–78.
- Yin, H. H., Knowlton, B. J., & Balleine, B. W. (2004). Lesions of dorsolateral striatum preserve outcome expectancy but disrupt habit formation in instrumental learning. European Journal of Neuroscience, 19(1), 181–189.
- Yin, H. H., Ostlund, S. B., Knowlton, B. J., & Balleine, B. W. (2005). The role of the dorsomedial striatum in instrumental conditioning. European Journal of Neuroscience, 22(2), 513–523.
- Yin, H. H., & Knowlton, B. J. (2006). The role of the basal ganglia in habit formation. Nature Reviews Neuroscience, 7(6), 464–476.
- Dickinson, A., Nicholas, D. J., & Adams, C. D. (1983). The effect of the instrumental training contingency on susceptibility to reinforcer devaluation. Quarterly Journal of Experimental Psychology B, 35(1), 35–51.
- O'Doherty, J., Dayan, P., Schultz, J., Deichmann, R., Friston, K., & Dolan, R. J. (2004). Dissociable roles of ventral and dorsal striatum in instrumental conditioning. Science, 304(5669), 452–454.
- Everitt, B. J., & Robbins, T. W. (2005). Neural systems of reinforcement for drug addiction: From actions to habits to compulsion. Nature Neuroscience, 8(11), 1481–1489.
- Fetz, E. E. (1969). Operant conditioning of cortical unit activity. Science, 163(3870), 955–958.
- Brembs, B., Lorenzetti, F. D., Reyes, F. D., Baxter, D. A., & Byrne, J. H. (2002). Operant reward learning in Aplysia: Neuronal correlates and mechanisms. Science, 296(5573), 1706–1709.
- Dayan, P., & Niv, Y. (2008). Reinforcement learning: The good, the bad and the ugly. Current Opinion in Neurobiology, 18(2), 185–196.
- Daw, N. D., Niv, Y., & Dayan, P. (2005). Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control. Nature Neuroscience, 8(12), 1704–1711.
- Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical Conditioning II: Current Research and Theory (pp. 64–99). Appleton-Century-Crofts.
- Berke, J. D. (2018). What does dopamine mean? Nature Neuroscience, 21(6), 787–793.
- Matsumoto, M., & Hikosaka, O. (2009). Representation of negative motivational value in the primate lateral habenula. Nature Neuroscience, 12(1), 77–84.