History

The History of Operant Conditioning

Who discovered operant conditioning, where behaviorism came from, what the cognitive revolution did and did not overturn, and how a law about cats in boxes became a profession, a finding about dopamine, and the algorithm behind modern AI.

Updated 21 min read

Who discovered operant conditioning?

The principle was discovered by Edward L. Thorndike, whose puzzle-box experiments with cats (1898) produced the law of effect: responses followed by a satisfying consequence are strengthened, and those followed by discomfort are weakened. B. F. Skinner named it operant conditioning in 1937, invented the methods for studying it, and built the science around it beginning with The Behavior of Organisms (1938).[4][11]

So the honest answer is two names: Thorndike found the law; Skinner turned it into a field. Behind both stand a century of earlier thinking about how animals learn from the results of their actions.

In brief

  • Thorndike found the law of effect with his 1898 puzzle-box cats; Skinner named operant conditioning in 1937 and built the science around it.
  • The cognitive revolution bounded the law of effect rather than overturning it: reinforcement still changes behavior, within limits set by evolution.
  • The tradition continues as applied behavior analysis, in the dopamine prediction-error signal, and in reinforcement learning, a formalization of learning from consequences.

Who discovered operant conditioning: Thorndike or Skinner?

Thorndike was the first to show experimentally, with learning curves rather than anecdotes, that the consequences of a response change its future probability. Skinner was the first to treat that fact as the foundation of a science: he defined the operant as a class of behavior controlled by its consequences, distinguished it from Pavlov's reflexes, built the free-operant chamber and cumulative recorder, and discovered the phenomena — schedules, shaping, extinction, stimulus control — that make up the subject today. Thorndike's own term, instrumental learning, survives as "instrumental conditioning," a synonym still used in the discrete-trial tradition.

Timeline: the history of operant conditioning

Read the sources

The books this timeline begins with are in the public domain and republished in full in the library: Thorndike's Animal Intelligence (with the 1898 monograph), Morgan's Animal Behaviour, James's Principles of Psychology, Yerkes's Dancing Mouse, Pfungst's Clever Hans, Romanes and Darwin. Skinner's own books are still under copyright.

From the law of effect to the operant: what Skinner changed

It is tempting to read Skinner as Thorndike with better equipment. The differences are deeper, and they explain why the field is called the experimental analysis of behavior rather than the study of trial-and-error learning.

AspectThorndike (1898–1932)Skinner (1938 onward)
Basic measureTime to escape on each trialRate of response — continuous, sensitive, and the measure that made schedule effects visible[11]
MethodDiscrete trials; experimenter resets the boxFree operant; the animal responds whenever it likes and the cumulative recorder draws rate as a slope
Unit of behaviorA specific movement connected to a situationThe operant: a class of responses defined by their common consequence, whatever their form (left paw, right paw, nose)[34]
Why reinforcement works"Satisfiers" stamp in connections (Hull later: drive reduction)No theory required: a reinforcer is whatever strengthens the behavior it follows; deprivation is an operation you perform and measure[17]
Research designLearning curves across trials, later group comparisonsSingle organisms studied to steady state, with effects shown by reversal in the same animal — codified by Sidman in 1960[35]
ScopeAnimal learning and educationAll behavior, including thinking, language, and culture — via the three-term contingency of antecedent, behavior, consequence

Add the distinction from Pavlovian conditioning and you have the framework every page on this site uses. Operant vs. classical conditioning ›

The cognitive revolution: what it overturned and what it didn't

Between roughly 1956 and 1970 psychology's center of gravity moved from behavior to mind. George Miller's paper on the limits of short-term memory (1956), Chomsky's review of Verbal Behavior (1959), and Ulric Neisser's Cognitive Psychology (1967) are the usual markers.[23][36] Within learning research itself, a series of findings showed that consequences were not the whole story:

What survived

None of these findings overturned the law of effect; they bounded it. Reinforcement still changes behavior, schedules still produce their signature patterns, extinction bursts still occur, and shaping and stimulus control still work — in rats, pigeons, dolphins, children, and adults. What changed is that operant conditioning is now understood as one powerful learning process among several, operating within limits set by evolution, and describable in the language of prediction error. Behavior analysts kept publishing in their own journals throughout; the "death of behaviorism" was mostly a change in who wrote the introductory textbook.

Behavior analysis today

The experimental tradition became a profession in the decades after 1968. Applied behavior analysis now has a certifying body (the BACB, founded in 1998), a credential (the Board Certified Behavior Analyst), licensure in most U.S. states, insurance coverage for autism services in every state, and a research literature spanning developmental disabilities, education, organizational management, addiction treatment, brain-injury rehabilitation, and animal welfare.[39] Its methods — reinforcement-based teaching, functional assessment before treatment, single-subject designs with continuous measurement — descend directly from the laboratory. How ABA is practiced ›

The neurodiversity critique

An honest history has to note that ABA is also contested. Many autistic adults and advocates object to its early history (Lovaas's original 1960s program used aversives, including electric shock, a practice long since abandoned), to goals that aimed at making autistic children "indistinguishable" from peers, to suppressing harmless self-stimulatory behavior, and to intensive programs delivered without the child's assent.[40] Practitioners have responded that modern ABA is reinforcement-based, individualized, and increasingly assent-driven, and the field's own journals now publish on these concerns.[40] The debate is about goals, consent, and history, not about whether reinforcement changes behavior — which everyone in it agrees it does.

Beyond the clinic

Outside ABA, the operant tradition underlies reinforcement-based animal training, contingency management in addiction medicine, school-wide positive behavior supports, organizational behavior management, and the analysis of choice that became behavioral economics. Herrnstein's matching law and George Ainslie's hyperbolic discounting — the finding that we overvalue immediate consequences along a predictable curve — gave economists a behavioral account of impulsiveness years before it was fashionable.[21][41]

The neuroscience of reinforcement

The last chapter of the history is being written by neuroscience. In 1953 James Olds and Peter Milner found that a rat would press a lever for hours to deliver a pulse of current to its own brain — the first demonstration that reinforcement had an identifiable neural substrate.[42] In 1997 Wolfram Schultz, Peter Dayan, and Read Montague showed that midbrain dopamine neurons report a prediction error — firing to unexpected reward, falling silent to predicted reward, dipping when a predicted reward fails — precisely the quantity in the Rescorla–Wagner model and in temporal-difference learning.[30][43] Kent Berridge and Terry Robinson then separated "wanting" from "liking": animals without dopamine still enjoy sugar but no longer work for it.[44] The operant laboratory's strangest findings — the grip of the variable-ratio schedule, the fading of a fully predictable reinforcer — suddenly had a mechanism. What happens in the brain, in depth ›

Operant conditioning in artificial intelligence

Reinforcement learning is the branch of machine learning in which an agent learns by acting in an environment and receiving a numerical reward. Its founders were explicit about the lineage: Richard Sutton and Andrew Barto's textbook opens with Thorndike's law of effect, and their early work with Charles Anderson on "neuronlike adaptive elements" was an attempt to build a learning rule that behaved like an animal under reinforcement.[31][45] Sutton's 1988 temporal-difference learning algorithm updates predictions from the difference between successive predictions — the computational cousin of Rescorla–Wagner — and it was TD error that Montague, Dayan, and Sejnowski proposed in 1996 as the thing dopamine neurons encode, a year before Schultz's data confirmed it.[43][46]

The vocabulary crossed over intact. AI researchers speak of reward, policy, exploration and exploitation, and reward shaping — the practice of adding intermediate rewards to guide an agent toward a hard-to-reach goal, named for Skinner's procedure.[47] Deep reinforcement learning, which pairs these algorithms with neural networks, learned Atari games from raw pixels in 2015 and beat the world's best Go players in 2016; reinforcement learning from human feedback is now a standard stage in training large language models, with human preferences standing in for the food hopper.[32][33]

The differences matter too. An RL agent's reward is written by an engineer, not discovered functionally; the agent can explore millions of episodes an animal never could; and the algorithms include explicit value estimates that Skinner would have called explanatory fictions. But the central claim — that a system with no built-in knowledge of a task can acquire complex, purposeful behavior purely from the consequences of its actions — is Thorndike's and Skinner's, vindicated at a scale neither imagined.

Key takeaways

Check yourself

Thorndike or Skinner: who discovered operant conditioning? Make the case for each.

Thorndike discovered the principle: his 1898 puzzle-box experiments showed, with learning curves rather than anecdotes, that consequences change the future probability of a response, and he named it the law of effect. Skinner named the process in 1937, defined the operant as a class of behavior controlled by its consequences, built the chamber and cumulative recorder, and discovered schedules, shaping, extinction, and stimulus control. Thorndike found the law; Skinner turned it into a field.

Behaviorism began with Watson's 1913 manifesto, so it seems safe to say early behaviorist learning theory was built on Thorndike's law of effect. Is that right?

No. Watson's early learning theory rested on Pavlov's reflex, not on Thorndike's law. It was Skinner, in the 1930s, who borrowed Pavlov's vocabulary of reinforcement, extinction, generalization, and discrimination and applied it to a different kind of learning, behavior controlled by its consequences rather than elicited by a prior stimulus.

A textbook states that Chomsky's 1959 review and the cognitive revolution killed behaviorism. What does the historical record show?

Mainstream psychology's attention did shift to mental processes, and findings such as latent learning, observational learning, and biological constraints showed that consequences are not the whole story. But those findings bounded the law of effect rather than overturning it: reinforcement, schedule effects, extinction bursts, shaping, and stimulus control still replicate, and behavior analysts kept publishing in their own journals, founding an applied journal in 1968 and a credentialing body in 1998. The "death of behaviorism" was mostly a change in who wrote the introductory textbook.

Sutton and Barto open their reinforcement learning textbook with Thorndike. What does a reinforcement learning agent share with a cat in a puzzle box, and where does the comparison break down?

Both acquire complex, purposeful behavior with no built-in knowledge of the task, purely from the consequences of their actions, and the vocabulary of reward, exploration, and reward shaping crossed over intact. The differences: an agent's reward is written by an engineer rather than discovered functionally, the agent can explore millions of episodes an animal never could, and its algorithms carry explicit value estimates that Skinner would have called explanatory fictions.

Explain it to a friend. Explain why the cognitive revolution did not overturn the law of effect, without using the words "behaviorism" or "cognitive."

Frequently asked questions

Who discovered operant conditioning?

Edward Thorndike discovered the underlying principle, the law of effect, in his 1898 puzzle-box experiments with cats. B. F. Skinner named the process "operant conditioning" in 1937, invented the experimental methods to study it, and founded the science with The Behavior of Organisms (1938). Both are correctly credited.

When was operant conditioning discovered?

The law of effect dates to Thorndike's 1898 monograph and was named in his 1911 book. The term "operant conditioning" dates to 1937, and the systematic science to 1938.

Who is the father of operant conditioning?

B. F. Skinner is usually called the father of operant conditioning because he named it, built the apparatus and methods, and developed the concepts — reinforcement schedules, shaping, stimulus control, the operant itself. Thorndike is the grandfather: he found the law of effect on which everything rests.

What is the difference between Thorndike and Skinner?

Thorndike studied discrete trials, measured time to respond, and explained learning as connections "stamped in" by satisfaction. Skinner let animals respond freely, measured rate of response, defined reinforcement purely by its effect, and rejected explanations in terms of satisfaction, drive, or connections.

What are the origins of behaviorism?

Behaviorism as a movement began with John B. Watson's 1913 paper "Psychology as the Behaviorist Views It," which argued that psychology should study observable behavior rather than consciousness. Its roots include Pavlov's conditioned reflexes, Thorndike's animal learning, and Morgan's canon of parsimony. Skinner's radical behaviorism, from 1945, is a later and different philosophy.

Did the cognitive revolution disprove operant conditioning?

No. It showed that reinforcement is not the only way organisms learn — observation, latent learning, and biological preparedness matter too — and it shifted mainstream psychology's attention to mental processes. The findings of operant conditioning still replicate, and the field continued as behavior analysis, with its own journals, profession, and applications.

How is operant conditioning related to artificial intelligence?

Reinforcement learning, a major branch of machine learning, is a mathematical formalization of learning from consequences. Its founders, Sutton and Barto, cite Thorndike's law of effect directly; temporal-difference learning parallels the Rescorla–Wagner model; and "reward shaping" takes its name from Skinner's procedure. The same prediction-error signal appears in the firing of dopamine neurons.

Is behaviorism still used today?

Yes, as applied behavior analysis (a licensed profession), reinforcement-based animal training, contingency management for addiction, positive behavior supports in schools, organizational behavior management, and the self-management methods behind habit-building apps. Its concepts also run through behavioral economics, neuroscience, and AI.

References

  1. Bain, A. (1855). The Senses and the Intellect. John W. Parker. See also Bain, A. (1859). The Emotions and the Will. John W. Parker.
  2. Romanes, G. J. (1882). Animal Intelligence. Kegan Paul, Trench. Full text in the library ›
  3. Morgan, C. L. (1894). An Introduction to Comparative Psychology. Walter Scott.
  4. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. Psychological Review Monograph Supplement, 2(4), 1–109. Read the 1898 monograph as Chapter II of the 1911 book in the library ›
  5. Thorndike, E. L. (1911). Animal Intelligence: Experimental Studies. Macmillan. Full text in the library ›
  6. Thorndike, E. L. (1932). The Fundamentals of Learning. Teachers College, Columbia University.
  7. Watson, J. B. (1913). Psychology as the behaviorist views it. Psychological Review, 20(2), 158–177.
  8. Pavlov, I. P. (1927). Conditioned Reflexes (G. V. Anrep, Trans.). Oxford University Press.
  9. Miller, S., & Konorski, J. (1928). Sur une forme particulière des réflexes conditionnels. Comptes Rendus des Séances de la Société de Biologie, 99, 1155–1157. English translation: Miller, S., & Konorski, J. (1969). On a particular form of conditioned reflex. Journal of the Experimental Analysis of Behavior, 12(1), 187–189.
  10. Konorski, J., & Miller, S. (1937). On two types of conditioned reflex. Journal of General Psychology, 16, 264–272; Skinner, B. F. (1937). Two types of conditioned reflex: A reply to Konorski and Miller. Journal of General Psychology, 16, 272–279.
  11. Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
  12. Hull, C. L. (1943). Principles of Behavior: An Introduction to Behavior Theory. Appleton-Century.
  13. Estes, W. K. (1944). An experimental study of punishment. Psychological Monographs, 57(3), i–40.
  14. Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208. See also Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. University of California Publications in Psychology, 4, 257–275.
  15. Skinner, B. F. (1948). 'Superstition' in the pigeon. Journal of Experimental Psychology, 38(2), 168–172.
  16. Keller, F. S., & Schoenfeld, W. N. (1950). Principles of Psychology: A Systematic Text in the Science of Behavior. Appleton-Century-Crofts.
  17. Skinner, B. F. (1950). Are theories of learning necessary? Psychological Review, 57(4), 193–216.
  18. Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
  19. Skinner, B. F. (1956). A case history in scientific method. American Psychologist, 11(5), 221–233.
  20. Premack, D. (1959). Toward empirical behavior laws: I. Positive reinforcement. Psychological Review, 66(4), 219–233.
  21. Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267–272.
  22. Breland, K., & Breland, M. (1961). The misbehavior of organisms. American Psychologist, 16(11), 681–684.
  23. Chomsky, N. (1959). A review of B. F. Skinner's Verbal Behavior. Language, 35(1), 26–58.
  24. Ayllon, T., & Azrin, N. H. (1968). The Token Economy: A Motivational System for Therapy and Rehabilitation. Appleton-Century-Crofts.
  25. Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. Journal of Applied Behavior Analysis, 1(1), 91–97.
  26. Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical Conditioning II: Current Research and Theory (pp. 64–99). Appleton-Century-Crofts. See also Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. Journal of Comparative and Physiological Psychology, 66(1), 1–5.
  27. Iwata, B. A., Dorsey, M. F., Slifer, K. J., Bauman, K. E., & Richman, G. S. (1994). Toward a functional analysis of self-injury. Journal of Applied Behavior Analysis, 27(2), 197–209. (Reprinted from Analysis and Intervention in Developmental Disabilities, 2, 3–20, 1982.)
  28. Nevin, J. A., Mandell, C., & Atak, J. R. (1983). The analysis of behavioral momentum. Journal of the Experimental Analysis of Behavior, 39(1), 49–59. See also Nevin, J. A., & Grace, R. C. (2000). Behavioral momentum and the law of effect. Behavioral and Brain Sciences, 23(1), 73–90.
  29. Lovaas, O. I. (1987). Behavioral treatment and normal educational and intellectual functioning in young autistic children. Journal of Consulting and Clinical Psychology, 55(1), 3–9.
  30. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593–1599.
  31. Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. (First edition 1998.)
  32. Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533.
  33. Silver, D., Huang, A., Maddison, C. J., et al. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484–489. See also Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30.
  34. Skinner, B. F. (1935). The generic nature of the concepts of stimulus and response. Journal of General Psychology, 12, 40–65.
  35. Sidman, M. (1960). Tactics of Scientific Research: Evaluating Experimental Data in Psychology. Basic Books.
  36. Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2), 81–97; Neisser, U. (1967). Cognitive Psychology. Appleton-Century-Crofts.
  37. Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575–582.
  38. Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. Psychonomic Science, 4(1), 123–124. See also Seligman, M. E. P. (1970). On the generality of the laws of learning. Psychological Review, 77(5), 406–418.
  39. Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.
  40. Leaf, J. B., Cihon, J. H., Leaf, R., McEachin, J., Liu, N., Russell, N., Unumb, L., Shapiro, S., & Khosrowshahi, D. (2022). Concerns about ABA-based intervention: An evaluation and recommendations. Journal of Autism and Developmental Disorders, 52(6), 2838–2853. On the early use of aversives, see Lovaas, O. I., Schaeffer, B., & Simmons, J. Q. (1965). Building social behavior in autistic children by use of electric shock. Journal of Experimental Research in Personality, 1, 99–109.
  41. Ainslie, G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. Psychological Bulletin, 82(4), 463–496.
  42. Olds, J., & Milner, P. (1954). Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology, 47(6), 419–427.
  43. Montague, P. R., Dayan, P., & Sejnowski, T. J. (1996). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. Journal of Neuroscience, 16(5), 1936–1947.
  44. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience? Brain Research Reviews, 28(3), 309–369.
  45. Barto, A. G., Sutton, R. S., & Anderson, C. W. (1983). Neuronlike adaptive elements that can solve difficult learning control problems. IEEE Transactions on Systems, Man, and Cybernetics, 13(5), 834–846.
  46. Sutton, R. S. (1988). Learning to predict by the methods of temporal differences. Machine Learning, 3(1), 9–44.
  47. Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. In Proceedings of the Sixteenth International Conference on Machine Learning (pp. 278–287). Morgan Kaufmann.
  48. Aristotle. (c. 350 BCE). Nicomachean Ethics, Book II (W. D. Ross, Trans.). See especially 1103a–1103b on habituation and 1104b on pleasure and pain as signs of character.