Multiple regression, and what breaks it
More explanatory variables, more ways to be wrong. Multicollinearity, overfitting and omitted variables are the three that matter, and all three produce confident, wrong answers.
Chapter 12 · Advanced
Each coefficient is now the effect of its variable holding the others constant — a phrase that carries most of the difficulty in this chapter, because in financial data the others rarely hold constant.
What it is used for
Factor models. Rather than explaining a return with the market alone, explain it with several systematic exposures:
The alpha that survives is return explained by none of the included exposures. Chapter 10 of the Portfolio theory subject takes this further; here the concern is what goes wrong mechanically.
Problem one: adding variables always helps
never falls when a variable is added. Not "usually does not" — mathematically cannot, because the fit can always set the new coefficient to zero and do no worse. Add random noise as a seventh variable and rises.
So cannot be used to compare models with different numbers of variables. Adjusted penalises each added parameter:
This can fall, and falling is the signal that a variable earned nothing.
The problem posed above: six variables on four years of monthly data is 48 observations and seven parameters. An of 0.94 on that is close to guaranteed and tells you almost nothing. The questions worth asking are the adjusted figure, and what it does out of sample.
Problem two: multicollinearity
When explanatory variables are correlated with each other, the fit cannot attribute credit between them. The coefficients become unstable — large standard errors, signs that flip when a variable is added or a few observations change — while the overall fit stays fine.
That combination is the diagnostic: a model that predicts well but whose individual coefficients are nonsense.
In finance it is everywhere, because candidate variables are mostly measuring related things. Size, liquidity and volatility move together; so do most valuation ratios. A model with "P/E, P/B, EV/EBITDA and price-to-sales" has four partial views of one underlying quantity.
What to do: drop the redundant variables, or combine them, and resist reading an individual coefficient as a clean effect.
Problem three: omitted variables
The serious one, because it biases rather than merely blurs.
If a variable affects and is correlated with an included , its effect is attributed to . The coefficient on is then wrong, not just imprecise, and more data makes it more precisely wrong.
The classic finance case: regress returns on some characteristic without controlling for risk, and the characteristic appears to generate return when it was exposure to a risk factor all along. Nearly every "anomaly" that failed to replicate out of sample has this shape.
No diagnostic finds an omitted variable. The model cannot know about data it does not have. Only reasoning about the mechanism can.
Problem four: overfitting
With enough parameters any dataset can be fitted exactly, and a model that fits the noise describes the past rather than the process. The symptom is a large gap between in-sample and out-of-sample performance.
Defences, in rough order of value:
- Hold out data. Build on one period, test on another never looked at. Looking at it once and adjusting destroys the test.
- Keep parameters few relative to observations. The ratio matters more than either number.
- Prefer a reason. A variable included because there is a mechanism generalises; one included because it improved the fit does not.
- Expect decay. A relationship that is real and public gets traded away. That is markets working, not the model breaking.
This is the quantitative content of SEBI's requirement that every performance disclosure says past performance may or may not be sustained. A fitted relationship is a description of a sample, and the sample is over.
Reading somebody else's model
| Ask | Why |
|---|---|
| How many variables, how many observations? | Overfitting is a ratio |
| Adjusted or plain ? | Plain cannot fall |
| Out-of-sample results? | The only honest test |
| Are the variables correlated with each other? | Coefficients may be meaningless |
| What is missing, and does it matter? | The bias no diagnostic catches |
| Is there a mechanism? | Decides whether it should generalise |
The point
Multiple regression estimates each effect holding the others constant, which financial data rarely permits. cannot fall when variables are added, so only the adjusted figure or out-of-sample performance means anything. Multicollinearity makes coefficients unstable while the fit looks fine; omitted variables make them biased, and no diagnostic can detect what is absent. Few parameters, a reason for each, and a period the model has never seen.
Check yourself
4 questions. Every answer is explained afterwards, including the ones you get right — guessing correctly is not the same as knowing. Score 70% or more and the chapter is marked done.
Question 1 of 4
0 of 4 answered. You can submit with questions unanswered — they simply score zero.
Now do it with your own numbers
A model explaining fund returns uses six variables and reports an R-squared of 0.94 on four years of monthly data. State why that figure is close to uninformative, and what you would ask for instead.
Count the observations against the parameters. Then ask what the adjusted R-squared is, and what the model does on data it has never seen.