Hypothesis testing
A formal way to ask whether a result could be chance. It is also the machinery behind most false findings in finance, because testing enough ideas guarantees some of them pass.
Chapter 10 · Advanced
A hypothesis test asks one question: could this result plausibly have arisen by chance if there were no real effect?
It is a useful discipline and a dangerous one, because its machinery is easy to run and easy to abuse, and in finance the abuse is the norm rather than the exception.
The structure
The null hypothesis is the boring claim: no effect, no difference, no skill. The alternative is what you suspect.
The null is always the position of "nothing is happening", and that asymmetry is deliberate. A test can reject the null or fail to reject it. It never proves it.
The statistic says how many standard errors the observation sits from what the null predicts. Large values are hard to explain by chance.
The p-value, stated carefully
The p-value is the probability of seeing a result at least this extreme, if the null were true.
What it is not, and each of these is a common misreading:
- Not the probability the null is true.
- Not the probability your finding is real.
- Not a measure of how large or important the effect is.
That first one is the base-rate error from chapter 4 wearing different clothes: is not . Converting between them needs a prior — how plausible the hypothesis was before the data — and a strategy picked from thousands has a very low prior.
The 0.05 threshold is a convention with no mathematical standing whatever. A p-value of 0.049 and one of 0.051 are the same evidence.
Two ways to be wrong
| true | false | |
|---|---|---|
| Reject | Type I error | correct |
| Fail to reject | correct | Type II error |
Type I is a false positive: finding a pattern that is not there. Its rate is the significance level you chose, so testing at 0.05 means accepting a 1-in-20 false positive rate per test.
Type II is a false negative: missing a real effect. Its rate depends on power, which rises with sample size and effect size.
The trade-off is fixed: tightening one loosens the other at a given sample size. The only way to improve both is more data — and chapter 8 showed how expensive that is at .
The multiple comparisons problem
This is the one that matters in finance.
At a 5% significance level, 1 test in 20 clears the bar with no effect present. Test 60 strategy variations and you expect three to look significant by chance. Report only the winner and you have a publishable finding built on nothing.
| Tests | Chance of at least one false positive |
|---|---|
| 1 | 5% |
| 10 | 40% |
| 20 | 64% |
| 60 | 95% |
| 100 | 99.4% |
At 60 tests, finding something significant is nearly certain. The p-value of the winner is not evidence of anything, because the selection happened after the testing.
The dishonest version of this has names — p-hacking, data dredging. The far more common version is innocent: a researcher tries ideas until one works, which is a reasonable way to explore and a terrible way to conclude.
How to tell whether it happened
- How many variants were tried? If the answer is unknown, the p-value is uninterpretable.
- Was the rule fixed before the data was seen? A rule chosen after is not a test.
- Does it hold out of sample, on a period not used to build it? The only real check.
- Is there a reason it should work? A mechanism beats a correlation, because mechanisms generalise.
Statistical and practical significance
A significant result can be trivially small. With enough data, a strategy beating its benchmark by 0.1% a year can reach — and lose to costs. Chapter 11 of the Derivatives subject makes the same point in rupees: the measured edge has to survive brokerage, impact and tax before it is an edge at all.
The reverse also holds. A large effect in a small sample may fail to reach significance while being the best estimate available. Failing to reject the null is not evidence that the null is true; it is often evidence that the sample was too small to tell.
Why this chapter sits under the technical analysis subject
SEBI's measurement that the great majority of individual derivatives traders lose money is not a hypothesis test on a strategy — it is a census of outcomes. That distinction is the chapter's practical lesson: an observed population result beats an inferred one. Where you can count, count.
The point
A hypothesis test asks whether chance alone could produce this result, and the p-value answers — not the probability the finding is real. Testing at 5% accepts one false positive in twenty per test, so sixty variations make a "significant" winner near-certain with nothing there. Ask how many things were tried, whether the rule was fixed in advance, whether it survives out of sample, and whether the effect is large enough to matter after costs.
Check yourself
4 questions. Every answer is explained afterwards, including the ones you get right — guessing correctly is not the same as knowing. Score 70% or more and the chapter is marked done.
Question 1 of 4
0 of 4 answered. You can submit with questions unanswered — they simply score zero.
Now do it with your own numbers
A backtest finds a strategy beating its benchmark with a p-value of 0.03. You later learn the researcher tested 60 variations and reported the best. Work out roughly how many of the 60 you would expect to reach p < 0.05 by chance alone, and say what the 0.03 is now worth.
Under the null, p-values are uniform — about 5% of 60 tests clear 0.05 with no effect present at all. Then ask what the researcher would have reported had a different variation won.