History · Foundations

The Law of Effect: Thorndike's Puzzle Boxes, the Truncated Law, and What Skinner Kept

A hungry kitten in a wooden box, food outside, and a stopwatch. From that arrangement Edward Thorndike drew the first experimental law of learning, then cut it in half on his own evidence. Everything on this site about reinforcement descends from it, and so does the algorithm that trains modern AI.

Updated 17 min read

Definition

The law of effect is Edward L. Thorndike's principle that, of the several responses an animal makes to a situation, those "accompanied or closely followed by satisfaction" become more firmly connected with that situation and are more likely to recur, while those "accompanied or closely followed by discomfort" have their connections weakened and are less likely to recur.[1] Thorndike found the effect in his puzzle-box experiments of 1898, gave it this formal statement in 1911, and dropped the second half in 1932 when his own data showed that discomfort does not weaken what satisfaction strengthens.[2][3]

It is the ancestor of reinforcement. Skinner kept the finding, discarded the theory of connections and satisfactions that came with it, and rebuilt it as operant conditioning.

In brief

  • Thorndike's hungry cats escaped from latched boxes by acts that began as accidents; escape times fell gradually across trials, and he concluded that the resulting pleasure stamped the successful act in and the useless ones out.
  • The 1911 law has two halves, satisfaction strengthens and discomfort weakens, and Thorndike himself dropped the second in 1932 after finding that "wrong" did not weaken human responses the way "right" strengthened them.
  • Skinner replaced "satisfaction" with a reinforcer defined by its effect on rate, Herrnstein turned the law into an equation, and reinforcement learning in AI is its direct descendant.

The puzzle-box experiments of 1898

Thorndike began the work as William James's graduate student at Harvard, testing chicks in improvised pens, and finished it at Columbia, where the cats and dogs were added and the whole was published in 1898 as his doctoral thesis.[4][2] The method was plain. A hungry animal was shut in a box with food outside in sight and could get out only by some simple act: pulling a loop of cord, pressing a lever, turning a wooden button, stepping on a platform. Thorndike timed the escape, put the animal back, and repeated the trial until the time was short and constant; an animal that failed was taken out but, as he underlined, not fed. The subjects were about a dozen kittens, three small dogs, and chicks a few days old, whose motive was not hunger but, in his phrase, dislike of loneliness.[2]

A kitten dropped into a box did not study the latch. It squeezed at every opening, clawed and bit the bars, and thrust its paws through the gaps at anything within reach. Somewhere in that scramble a paw caught the loop, the door fell open, and the cat ate. On later trials the useless movements dropped away and the successful one came sooner, until the cat clawed the loop the moment it was put in. Thorndike's description gave the law its first vocabulary: the non-successful impulses are "stamped out," and the impulse leading to the successful act is "stamped in by the resulting pleasure."[2]

What made this an experiment rather than an anecdote was the time-curve, each animal's escape time plotted against trial number. Cat 12 in box A took 160 seconds on the first trial, 30 on the second, 90 on the third, and 6 or 7 seconds by the twenty-fourth.[2] The times, he said, were "facts which may be obtained by any observer who can tell time." In the easiest boxes every cat succeeded; in the thumb latch and the boxes requiring three separate acts, some never did.[2]

The argument against reasoning

Thorndike's target was the anecdotal literature, George Romanes in particular, which took a cat's opening a latched door as proof of reasoning. His answer was that all his cats opened latched doors, by accident, and that the curve showed what kind of process was at work: a cat that grasped how the box worked should show "a sudden vertical descent," and none did. Thorndike read the gradual slope as "the wearing smooth of a path in the brain, not the decisions of a rational consciousness."[2] Cats that watched a trained cat escape learned nothing from it, and cats whose paws he pressed onto the mechanism did not learn by being put through the act.[2] It was neither imitation nor insight but selection among the cat's own movements by what followed them. The monograph is republished as Chapter II of the 1911 book: read the 1898 monograph.

The law of effect as Thorndike stated it in 1911

The 1898 monograph described the process; the 1911 book, Animal Intelligence, stated it as a law. Chapter VI, "Laws and Hypotheses for Behavior," gives two provisional laws of learning. The first is the law of effect.

The law of effect (Thorndike, 1911, p. 244)

"Of several responses made to the same situation, those which are accompanied or closely followed by satisfaction to the animal will, other things being equal, be more firmly connected with the situation, so that, when it recurs, they will be more likely to recur; those which are accompanied or closely followed by discomfort to the animal will, other things being equal, have their connections with that situation weakened, so that, when it recurs, they will be less likely to occur. The greater the satisfaction or discomfort, the greater the strengthening or weakening of the bond."[1]

The second is the law of exercise: a response becomes more strongly connected with a situation in proportion to how often, how vigorously, and how long it has been connected with it.[1] Repetition strengthens; consequences select. Thorndike held that effect was the more fundamental, since an animal that mostly makes one response and only occasionally another can end up making the rare one every time if it alone is followed by satisfaction: "the law of effect is primary, irreducible to the law of exercise."[1]

What "satisfaction" meant

The word invites a reading in terms of feelings, and Thorndike headed it off. A satisfying state of affairs is "one which the animal does nothing to avoid, often doing such things as attain and preserve it"; an annoying one is "one which the animal commonly avoids and abandons."[1] Satisfiers "cannot be determined with precision and surety save by observation," and what satisfies is not what is good for the animal; he listed overeating and intoxication among the most potent satisfiers of man.[1] This is a functional definition, nearly three decades before Skinner's, and it is why the law survived the change of vocabulary.

Three "other things" had to be equal: exercise; closeness in time, which Thorndike illustrated rather than measured (a button that opened the door after one, five, fifty, or five hundred seconds would be learned fastest in the first case and in the fourth "almost certainly never"); and attention, whether the successful movement was "an eminent, emphatic part" of what the animal was doing.[1] Beneath the law sat a neural hypothesis, learning as a change in what he called the intimacy of the synapse, and the image of a path worn smooth in the brain echoes his teacher William James's chapter on habit.[1] Read the statement in context: Chapter VI.

Before Thorndike: Bain, Morgan, and trial and error

The idea was not new; the experiment was. In The Senses and the Intellect (1855) Alexander Bain described how an organism's spontaneous movements are sorted by their consequences: a movement that happens to coincide with pleasure is kept up and repeated, and one that coincides with pain is dropped.[5] That is the law of effect without the box, the curve, or the name, and the phrase "trial and error" is already in Bain.[5] C. Lloyd Morgan supplied the rule of evidence: his canon of 1894 holds that no animal action should be explained by a higher mental faculty if a lower one will account for it.[6] Thorndike quoted Morgan at length in 1898 as "the least offender" among the theorists he was attacking.[2] What he added was everything that turns a plausible principle into a law: a repeatable situation, many animals, a controlled motive, a measure independent of the observer, and a test that could have come out the other way. Where this fits in the longer story is on the history page.

The 1932 revision: the truncated law of effect

In a 1927 paper titled simply "The law of effect" Thorndike restated the law for adult human subjects and argued that an after-effect strengthens a connection directly and automatically, whether or not the learner thinks about it.[7] The experiments that followed changed the law. In a typical one a subject was shown a word, chose a number to go with it, and was told "right" or "wrong." Right strengthened. Wrong did little or nothing.[3] In The Fundamentals of Learning (1932) he concluded that annoyers do not act on connections the way satisfiers do: a punished response is not weakened in proportion to the punishment; at most, the annoyer leads the learner to vary, and the variation may then be rewarded.[3] The second half of the 1911 law was gone. This is the truncated law of effect. The law of exercise went with it: subjects who drew lines of a set length hundreds of times without being told how they were doing did not improve.[3]

The asymmetry became one of the field's most durable findings. Skinner reported in 1938 that when a rat's lever presses were punished at the start of extinction by having the lever slap back against its paws, responding was suppressed only while the slap was in force, and the punished rats made about as many responses as unpunished ones before they quit.[8] Estes extended the result with shock in 1944: suppression, but temporary.[9] The modern statement, that punishment is not reinforcement with the sign reversed, descends from Thorndike's own retraction; whether stronger or better-timed punishers do more than a spoken "wrong" is taken up on that page.

The criticisms

Is the law circular?

The most persistent objection is that the law explains nothing, because a satisfier is identified by the very strengthening it is supposed to explain. Leo Postman's 1947 review of the law's first half-century laid out this argument along with the others then in play, among them whether an after-effect is needed for learning or only for performance.[10] Paul Meehl's answer, in 1950, is the one still taught. Separate the weak law, a definition (a reinforcer is whatever strengthens the response it follows), from the strong law, an empirical claim (all learning requires such an event). Then notice that the weak law is not empty either, because reinforcers are trans-situational: an event shown to strengthen one response in one situation will strengthen other responses in other situations. That is a prediction, and it could fail; if food had strengthened loop-pulling in box A but not lever-pressing in box I, the law would have been in trouble.[11] Thorndike's definition of a satisfier by approach and avoidance had already pointed the same way.[1]

Was it trial and error at all?

In 1946 Edwin Guthrie and George Horton published Cats in a Puzzle Box, in which cats escaped by touching a pole in the center of the box and were photographed at the moment of escape. Each cat settled on its own way of hitting the pole and repeated it almost exactly, and Guthrie took the stereotypy as evidence for his rival theory: whatever movement the animal happens to be making when the situation changes stays attached to it by contiguity alone.[12] In 1979 Bruce Moore and Susan Stuttard pointed out that rubbing a flank or cheek against an upright object is the domestic cat's species-typical greeting. In a replica of Guthrie's box, cats rubbed the pole when a person was visible whether or not rubbing opened the door, and rarely did so when no one was in view. The response was not learned in the box; it was elicited by the experimenters standing in front of it.[13] The paper's subtitle, "Tripping over the cat," is fair.

The complication cuts both ways. Thorndike's boxes required acts that are no part of a greeting, and his cats' times fell across dozens of trials, so his curves are not explained away. But what an animal brings to the box is never a blank repertoire, and Thorndike had glimpsed this himself: cats released from box Z whenever they licked themselves learned to lick as soon as they were put in, but the lick shrank with practice to "a mere quick turn of the head," and he confessed himself ignorant of why.[2] Consequences work on a repertoire shaped by evolution, a point the shaping page returns to.

How Skinner rebuilt it: reinforcement and selection by consequences

B. F. Skinner read Thorndike's law as a fact in search of a better description, and The Behavior of Organisms (1938) supplied one. The "response" became the operant, a class of acts defined by their common effect rather than their form. "Satisfaction" was dropped: a reinforcer was any event that, following a response, raised its future rate, and nothing was said about how it felt. And the measure changed: Skinner's rats lived with the lever and pressed whenever they liked, so rate of responding, drawn by the cumulative recorder, replaced time to escape.[8] He called this Type R conditioning, to separate it from Pavlov's Type S (see operant versus classical conditioning), and the apparatus is on the Skinner box page.[8]

In "Selection by Consequences" (1981) Skinner argued that the law of effect is one instance of a causal mode peculiar to living things, in which variation is followed by selection: natural selection shapes the species, operant conditioning shapes the individual's behavior within its lifetime, and cultural practices are selected by their consequences for the group.[14] Read this way, the law of effect is Darwinian, and Thorndike had used the language in 1898: from among the cat's movements, "one is selected by success."[2]

Richard Herrnstein made it quantitative. In "On the Law of Effect" (1970) he extended the matching law to the single response in a Skinner box by treating even that as a choice between the measured response and everything else the animal could be doing. The result is a hyperbola: response rate rises with reinforcement rate at diminishing returns, to a ceiling set by the reinforcement available for everything else.[15]

Thorndike's law of effect vs. Skinner's reinforcement

AspectThorndike's law of effect (1898–1932)Skinner's reinforcement (1938 onward)
Unit of analysisA connection between a situation and a specific movementThe operant: a class of responses defined by its common consequence
MechanismSatisfaction "stamps in" a connection, physically a change at the synapse; later, a direct and automatic after-effectNone claimed; reinforcement is a functional relation between consequence and rate, explained as selection by consequences
Punishment1911: the mirror image of satisfaction. 1932: dropped, since annoyers do not weaken connections the way satisfiers strengthen themNot the mirror image of reinforcement: suppresses responding, often temporarily and with side effects; a separate process
MeasurementTime to escape on each discrete trial; a learning curveRate of responding in the free operant; a cumulative record
How the consequence is definedA state of affairs the animal does nothing to avoid and often acts to attainWhatever increases the rate of the response it follows

Why the law of effect still matters

Strip away the neurons and the word "satisfaction," and the law of effect is the working rule behind every reinforcement procedure on this site.

SettingSituationResponseAfter-effectNext time
Dog trainingDog at the closed doorSitsDoor opensSitting at doors comes sooner
ClassroomTeacher asks a questionStudent raises a handIs called on and answersHand-raising increases
ParentingToddler in the grocery cartWhinesParent hands over a snackWhining in carts increases
Self-managementSitting down to writeOpens the file and types one sentenceThe dread liftsStarting gets easier

Three of Thorndike's lessons transfer unchanged. Satisfiers are found by observation, not assumption, the insight behind the Premack principle. Closeness in time matters: the button that opens the door after fifty seconds teaches almost nothing. And learning is a curve, not a step; the trainer who expects a perfect sit after one treat is expecting the sudden vertical descent Thorndike never saw. A fourth lesson is negative: since "wrong" does not undo what "right" does, the tools for weakening a response are extinction and reinforcement of something else.[3]

Reinforcement learning in AI

Reinforcement learning, the branch of machine learning in which an agent learns by acting in an environment and receiving a numerical reward, descends from Thorndike's law directly, and its founders say so. Sutton and Barto open the field's history with the law of effect and note that it contains the two things trial-and-error learning requires: it is selectional, since among the actions tried those with better outcomes are kept, and associative, since the kept actions are tied to the situations in which they were tried.[16] An agent that tries actions, keeps the ones that pay, and learns which states they pay in is doing what the cat did in box A, with the reward written down as a number. The differences are real: an engineer chooses the reward, the agent can run millions of trials, and the algorithms keep explicit estimates of value. But the core claim is Thorndike's: a system with no knowledge of a task can acquire competent behavior from nothing but the consequences of its own actions.

What the evidence does not show

Key takeaways

Check yourself

A trainer says her dog "figured out" that sitting opens the back door, because it now sits the instant it reaches the door. What would Thorndike want to see before agreeing?

The curve. If the dog grasped the rule, the latency to sit should have dropped from long to short in a single trial and stayed there, the sudden vertical descent Thorndike looked for and never found in his cats. A gradual fall over many door-approaches, with the sit emerging from a scramble of other behaviors, is the law of effect at work: a response selected by its consequence. Thorndike would also note that a single, simple, definite act can be stamped in by one experience without any inference, so even a fast curve is not proof of understanding.

A student writes that the law of effect is circular because reinforcers are defined by what they reinforce. How does Meehl's analysis answer this?

By separating two laws. The weak law is indeed a definition: a reinforcer is whatever strengthens the response it follows. The strong law, that all learning requires reinforcement, is an empirical claim that could be false. And the weak law still has content, because reinforcers are trans-situational: an event that strengthens one response in one situation is predicted to strengthen other responses elsewhere. That prediction could fail and does not, which is what makes the law a law rather than a tautology.

A teacher marks every wrong answer with a red X and expects the errors to disappear. What did Thorndike's own data suggest, and what should the teacher do instead?

In the experiments behind the 1932 revision, telling a learner "wrong" did little or nothing to weaken the response, while "right" reliably strengthened the correct one. The X is the second half of the 1911 law, the half Thorndike withdrew. The teacher should make sure correct answers are followed promptly by something the student works for, and treat the errors by withholding that consequence and reinforcing the alternative, rather than expecting the mark itself to subtract the error.

Want more? The 20-question quiz covers every page on the site with instant explanations.

Frequently asked questions

What is the law of effect in simple terms?

Behavior that is followed by a satisfying result becomes more likely in that situation; behavior followed by an unpleasant result becomes less likely. Thorndike found it with hungry cats learning to escape from latched boxes: the movement that happened to open the door was repeated sooner on every trial. The first half of the law is the basis of reinforcement; Thorndike himself later dropped the second half.

Who discovered the law of effect?

Edward L. Thorndike, in puzzle-box experiments with cats, dogs, and chicks published in 1898 as his Columbia doctoral thesis. He stated the law formally in Animal Intelligence (1911). Alexander Bain had described the same principle in 1855, and Lloyd Morgan had argued for trial-and-error explanations in 1894, but Thorndike was the first to test it with repeated trials, many animals, and learning curves.

What is an example of the law of effect?

A cat in Thorndike's box A clawed at everything, happened to pull a loop, and the door opened; after two dozen trials it pulled the loop within seconds of being put in. Everyday examples: a dog that sits at the door and is let out sits sooner next time; a child who whines in the store and gets a snack whines more; a kick that unjams a vending machine is repeated the next time it jams.

What is the difference between the law of effect and reinforcement?

Reinforcement is Skinner's rebuilding of the law of effect. Thorndike connected a specific response to a situation through "satisfaction" and measured time to escape on separate trials. Skinner defined the operant as a class of responses, defined a reinforcer purely by its effect on the rate of responding, and measured rate in a free-operant chamber. Skinner also treated punishment as a separate process rather than the mirror image of reinforcement, which Thorndike had concluded in 1932.

What is the truncated law of effect?

The revised law Thorndike proposed in The Fundamentals of Learning (1932). In experiments with human learners told "right" or "wrong" after each response, "right" strengthened responses but "wrong" did little to weaken them. He concluded that annoyers do not weaken connections the way satisfiers strengthen them, and dropped the punishment half of his 1911 statement. The reward half is the truncated law.

What is the difference between the law of effect and the law of exercise?

Thorndike stated both in 1911. The law of exercise says a response becomes more strongly connected to a situation the more often, vigorously, and lastingly it has been made in it; repetition strengthens. The law of effect says consequences decide which responses are strengthened. Thorndike argued that effect was primary, and by 1932 he had concluded that repetition without an after-effect strengthens nothing, abandoning the law of exercise as an independent law.

Is the law of effect circular?

Only in its weak form, which is a definition: a reinforcer is whatever strengthens the response it follows. Paul Meehl showed in 1950 that the law still makes a testable claim, because reinforcers are trans-situational: an event that strengthens one response in one situation will strengthen other responses in other situations. That prediction could fail and has not. Thorndike had also defined satisfiers independently, by what the animal approaches and does not avoid.

How is the law of effect related to reinforcement learning in AI?

Reinforcement learning is the branch of machine learning in which an agent learns by acting and receiving a numerical reward. Sutton and Barto's textbook opens the field's history with Thorndike's law and notes that it contains both things trial-and-error learning needs: selection of actions by their outcomes, and association of those actions with the situations in which they occurred. An RL agent is a cat in a puzzle box with the reward written down as a number.

References

  1. Thorndike, E. L. (1911). Animal Intelligence: Experimental Studies. Macmillan. Full text in the library: /library/thorndike-animal-intelligence/
  2. Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. Psychological Review Monograph Supplement, 2(4), 1–109. Reprinted as Chapter II of Thorndike (1911).
  3. Thorndike, E. L. (1932). The Fundamentals of Learning. Teachers College, Columbia University.
  4. Chance, P. (1999). Thorndike's puzzle boxes and the origins of the experimental analysis of behavior. Journal of the Experimental Analysis of Behavior, 72(3), 433–440.
  5. Bain, A. (1855). The Senses and the Intellect. John W. Parker.
  6. Morgan, C. L. (1894). An Introduction to Comparative Psychology. Walter Scott.
  7. Thorndike, E. L. (1927). The law of effect. American Journal of Psychology, 39(1/4), 212–222.
  8. Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
  9. Estes, W. K. (1944). An experimental study of punishment. Psychological Monographs, 57(3), i–40.
  10. Postman, L. (1947). The history and present status of the law of effect. Psychological Bulletin, 44(6), 489–563.
  11. Meehl, P. E. (1950). On the circularity of the law of effect. Psychological Bulletin, 47(1), 52–75.
  12. Guthrie, E. R., & Horton, G. P. (1946). Cats in a Puzzle Box. Rinehart.
  13. Moore, B. R., & Stuttard, S. (1979). Dr. Guthrie and Felis domesticus or: Tripping over the cat. Science, 205(4410), 1031–1033.
  14. Skinner, B. F. (1981). Selection by consequences. Science, 213(4507), 501–504.
  15. Herrnstein, R. J. (1970). On the law of effect. Journal of the Experimental Analysis of Behavior, 13(2), 243–266.
  16. Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.