Sampling and the central limit theorem
Everything you measure in finance is a sample, and a sample is not the truth. The central limit theorem says how far off it is likely to be, which is the only reason any of it is usable.
Chapter 8 · Intermediate
You never observe a distribution. You observe a handful of draws from it and try to work backwards. Every return figure, every volatility estimate, every correlation in this course is a sample statistic standing in for a population parameter nobody can see.
The distinction, and the notation
| Population (unknown) | Sample (measured) |
|---|---|
| — true mean return | — average of what you saw |
| — true volatility | — measured volatility |
| — true correlation | — measured correlation |
A sample statistic is itself a random variable: run history again and you would get a different number. The question is how different.
The central limit theorem
For a sample of independent draws from any distribution with finite variance, the distribution of the sample mean approaches a normal distribution as grows:
Two claims in there, and the second is the useful one.
It does not matter what the underlying distribution is. Skewed, fat-tailed, discrete — the sample mean still tends to normality. This is why the normal distribution keeps appearing even where the data is clearly not normal, and it is a stronger result than it first sounds.
The spread of the sample mean shrinks as . The standard error:
What costs you
The square root is the brutal part. To halve your uncertainty you need four times the data; to divide it by ten you need a hundred times.
Work the problem. A fund with over years:
The measured 14% carries a standard error of nearly 10 percentage points. A two-standard-error range — roughly 95% confidence, chapter 9 — runs from about −6% to +34%.
Five years of data cannot distinguish a very good fund from a losing one. That is not a complaint about the fund; it is arithmetic. To get the standard error down to 2% would need around 120 years of returns, by which time the fund, its manager and its mandate are all different things.
This is the real content behind SEBI requiring every performance disclosure to carry the line that past performance may or may not be sustained. The regulator is not being lawyerly. The sample is too small to support the inference most readers draw from it.
Three ways the theorem does not apply
The CLT assumes independent, identically distributed draws with finite variance. Financial data violates all three.
Not independent. Volatility clusters: turbulent days follow turbulent days. Chapter 13 measures this. Dependence means the effective sample size is smaller than the count of observations, so the true standard error is larger than — the formula flatters you.
Not identically distributed. The market that produced 2008's returns is not the market of 2024. Averaging across regimes estimates a parameter that no longer exists.
Tails are heavy. Convergence to normality is slow when the underlying has fat tails, and for some distributions the sample mean never settles at all. Chapter 6's point, arriving again.
The bias no sample size fixes
Survivorship bias is not sampling error, and this is the key distinction: more data does not help, because the data is systematically wrong rather than noisy.
Funds that performed badly are merged away or closed. A list of funds available today is a list of the ones that survived, so an average computed across it overstates what an investor would actually have experienced. The same mechanism inflates index backtests, strategy backtests and any study of "companies that did X".
Sampling error is random and shrinks with . Selection bias is systematic and does not. Collecting ten times more of a biased sample gives a more precise estimate of the wrong number.
What this changes in practice
- Treat any performance figure from under a decade as approximately uninformative about skill.
- Prefer long samples, and then distrust them for a different reason — regimes change.
- Ask what is missing from a dataset before asking what it says.
- When two funds differ by less than a standard error, they have not been distinguished.
The point
Every statistic in finance is a sample standing in for something unobservable. The central limit theorem says the sample mean is approximately normal around the truth with a standard error of — so five years of a 22%-volatility fund gives a 95% range roughly 40 percentage points wide. Dependence, changing regimes and fat tails all make that understate the real uncertainty, and survivorship bias is a separate problem that more data cannot fix.
Check yourself
4 questions. Every answer is explained afterwards, including the ones you get right — guessing correctly is not the same as knowing. Score 70% or more and the chapter is marked done.
Question 1 of 4
0 of 4 answered. You can submit with questions unanswered — they simply score zero.
Now do it with your own numbers
A fund has returned 14% a year over five years, with an annual standard deviation of 22%. Work out the standard error of that mean. Then say what range of true long-run returns is consistent with what you observed.
The standard error is σ divided by the square root of n, with n = 5. The answer is large enough that "this fund returns 14%" is not a claim the data supports.