Recency and reinforcement learning
People repeat what recently felt good and avoid what recently felt bad. In a domain where outcomes are mostly noise, that machinery learns the noise — and it works on whole generations, not just individuals.
Chapter 8 · Advanced
The simplest learning rule there is: repeat behaviours that coincided with pleasure, avoid those that coincided with pain. It is how nearly all skill is acquired, and in investing it misfires in a specific, documented way.
The evidence that investors do it
The survey assembles several independent findings, each a different corner of the same mechanism.
Repurchasing past winners. Investors are more likely to buy back a stock they previously sold for a profit than one previously sold at a loss. The company's current prospects are the same either way; what differs is how the last encounter felt.
Trading more after success. Individual investors trade more actively when their most recent trades were successful. Note the interaction with chapter 5: a run of luck raises activity, and activity is the thing that costs money.
Industry-level extrapolation. Investors — particularly less sophisticated ones — are more likely to buy a stock in an industry if their previous investments in that industry beat the market.
IPO subscription. Investors are more likely to subscribe to a new issue if their personal experience with IPOs has been profitable. Not if IPOs in general have performed well; if theirs did.
Savings rates. Investors whose retirement accounts have seen higher returns or lower variance increase their saving rates — the propensity to save responding to recent market history rather than to need.
The cohort effect
The most striking version operates across decades rather than months: age cohorts who experienced high stock market returns through their lives are less risk averse and more likely to invest in equities.
Think about what that means. Someone who began earning in a long bull market and someone who began in a long flat market will hold different portfolios for the rest of their lives, having drawn different conclusions from a sample of one. Neither is reasoning badly by the lights of their own experience. Both are generalising from the only history they lived through.
This matters in India specifically, where a very large share of participants began investing after 2020. Their entire sample is drawn from one regime, and reinforcement learning will have taught them its lessons as though they were permanent.
Working the problem
Why the algorithm fails here.
Learning to cook works because the link between decision and outcome is tight, fast and mostly deterministic. Too much salt tastes salty, immediately, nearly every time. The feedback is about your decision.
Investing breaks all three properties.
The link is loose. A good decision frequently produces a bad outcome and a bad decision a good one, because most short-run variation is noise. So the feedback is largely about the noise, not about you.
The feedback is slow. Years can pass before a judgement is settled, by which time the reasoning is unavailable for inspection and you remember only the result.
It is asymmetric in what it teaches. Reckless behaviour that happens to pay is reinforced more strongly than careful behaviour that pays modestly, because the reward was larger. The mechanism therefore actively selects for risk-taking during good periods — exactly when risk-taking is least punished and most dangerous.
The condition that would make it work: outcomes would have to be a reliable signal of decision quality — high signal-to-noise and prompt. Where that holds, reinforcement learning is excellent. It holds for cooking and for most crafts. The one place it approximately holds in finance is process errors with immediate consequences: paying a charge you did not need to pay, missing a deadline, failing to diversify and seeing a single holding wipe out a year. Those are learnable from outcomes, because the outcome follows from the decision directly rather than through the market.
What replaces it
Judge the decision, not the result. Write down why, before. Review the reasoning against what you knew at the time, not against what happened.
Lengthen the sample deliberately. Your own experience is too short to learn from. Market history long enough to contain regimes you have not lived through is the substitute, and it is why reading about 2008, 2000 and 1992 is practical rather than historical.
Treat a winning streak as a risk factor. It will be raising your activity and your position sizes through a mechanism you cannot feel operating.
The point
Repeat-what-worked is an excellent learning rule wherever outcomes reliably reflect decisions, and investing is not such a domain: the link is loose, slow and asymmetric, so the rule learns noise and preferentially reinforces whatever risk happened to pay. Investors demonstrably do it anyway — repurchasing stocks they sold at a profit, trading more after wins, subscribing to IPOs because their own past IPOs paid — and the effect runs across whole cohorts, whose lifetime risk appetite tracks the market history they happened to live through. The substitutes are judging the reasoning rather than the result, and borrowing a longer sample than your own.
Check yourself
4 questions. Every answer is explained afterwards, including the ones you get right — guessing correctly is not the same as knowing. Score 70% or more and the chapter is marked done.
Question 1 of 4
0 of 4 answered. You can submit with questions unanswered — they simply score zero.
Now do it with your own numbers
Reinforcement learning — repeat what worked, avoid what did not — is a good algorithm for most of life. Explain precisely why it performs badly in investing, and identify the one condition that would have to hold for it to work.
Compare learning to cook with learning to invest. What is different about the relationship between the quality of the decision and the quality of the outcome?