Skip to content
FreeFinance

Covariance, and where diversification comes from

Chapter 2's cross term, taken seriously. It explains why adding holdings reduces risk, why the reduction stops, and exactly where it stops — which turns out to depend on one number that nobody controls.

Chapter 3 · Beginner

The Risk subject said diversification removes the risk you are not paid for. This chapter is the arithmetic behind that sentence, and the arithmetic says something the sentence does not: exactly how much, and where it stops.

The quantity

Covariance measures how two returns move together:

σij=E[(Ri−μi)(Rj−μj)]\sigma_{ij} = E\big[(R_i - \mu_i)(R_j - \mu_j)\big]

Read the product inside. When both assets are above their means, or both below, the product is positive. When one is above and the other below, it is negative. Covariance is the average of that product, so it is positive when they tend to move together and negative when they tend to move oppositely.

Correlation is covariance made comparable, by dividing out the two assets' own sizes:

ρij=σijσiσj,−1≤ρij≤1\rho_{ij} = \frac{\sigma_{ij}}{\sigma_i \sigma_j}, \qquad -1 \le \rho_{ij} \le 1

The Quantitative methods subject derives both and the bounds. What is new here is what they do inside a portfolio.

The cross term, read carefully

From chapter 2:

σp2=w12σ12+w22σ22+2w1w2 ρ12 σ1σ2⏟the only term you did not expect\sigma_p^2 = w_1^2\sigma_1^2 + w_2^2\sigma_2^2 + \underbrace{2w_1w_2\,\rho_{12}\,\sigma_1\sigma_2}_{\text{the only term you did not expect}}

The first two terms are always positive. Squares of real numbers.

The third carries the sign of ρ\rho. When the correlation is negative, the third term is negative, and it subtracts from the total. That is the mechanism of diversification, in one line: a portfolio can be less risky than the sum of its parts because one term in the sum can be negative.

And even when ρ\rho is positive — as it almost always is for equities — the third term is smaller than it would be at ρ=1\rho = 1, which is enough. Chapter 2's table showed the whole range.

Here is the cleanest demonstration. Two assets, each with a 20% standard deviation, held equally:

Correlation Portfolio standard deviation
+1.0 20.00%
+0.5 17.32%
0.0 14.14%
−0.5 10.00%
−1.0 0.00%

Three rows deserve comment.

At ρ=+1\rho = +1 you get 20% — no benefit at all. Two assets that move identically are one asset, whatever their names.

At ρ=0\rho = 0 you get 14.14%, which is 20/220/\sqrt{2}. Combining two independent risks of equal size divides risk by the square root of two, not by two. That N\sqrt{N} is the shape of the whole result, and it is the same square root that appears in the Quantitative methods subject's standard error.

At ρ=−1\rho = -1 you get zero. Two perfectly opposed risky assets combine into a riskless one. It does not happen in practice, but it tells you the mechanism is capable of complete cancellation — the benefit is not a rounding effect.

Many holdings, and where it stops

Now the case that matters. Take NN holdings in equal weights, each with standard deviation σ\sigma, and an average pairwise correlation ρˉ\bar\rho. Splitting the double sum into its diagonal and off-diagonal parts:

σp2=σ2N⏟own variances+(1−1N)ρˉ σ2⏟covariances\sigma_p^2 = \underbrace{\frac{\sigma^2}{N}}_{\text{own variances}} + \underbrace{\left(1 - \frac{1}{N}\right)\bar\rho\,\sigma^2}_{\text{covariances}}

This formula is the subject's most useful single result, because of what each term does as NN grows.

The first term goes to zero. Divide by NN and it vanishes. This is the risk specific to each holding, and it can be removed completely, for free, by holding more things.

The second term does not. As N→∞N \to \infty, (1−1/N)→1(1 - 1/N) \to 1, so the term approaches ρˉ σ2\bar\rho\,\sigma^2. Taking the square root:

σp⟶σρˉ\sigma_p \longrightarrow \sigma\sqrt{\bar\rho}

That is the floor. No amount of diversification takes portfolio risk below σρˉ\sigma\sqrt{\bar\rho}, because the covariances never go away — they are what the holdings share.

What the floor looks like

Thirty-percent volatility stocks, equally weighted, at three different average correlations:

N ρˉ=0\bar\rho = 0 ρˉ=0.2\bar\rho = 0.2 ρˉ=0.4\bar\rho = 0.4
1 30.0% 30.0% 30.0%
2 21.2% 23.2% 25.1%
5 13.4% 18.0% 21.6%
10 9.5% 15.9% 20.3%
30 5.5% 14.3% 19.4%
100 3.0% 13.7% 19.1%
1,000 0.9% 13.4% 19.0%
limit 0% 13.4% 19.0%

Read across the bottom row. With uncorrelated holdings, risk goes to zero. At an average correlation of 0.2, it stops at 13.4% — and a thousand holdings get you essentially nothing that thirty did not.

Read down the ρˉ=0.2\bar\rho = 0.2 column. Going from 1 holding to 10 takes risk from 30% to 15.9%, roughly halving it. Going from 10 to 1,000 takes it from 15.9% to 13.4%. Almost the entire benefit arrives early, which is the arithmetic behind the Risk subject's answer to "how much is enough".

And the correlation matters more than the count. Thirty holdings at ρˉ=0.4\bar\rho = 0.4 carry 19.4% risk; ten holdings at ρˉ=0\bar\rho = 0 carry 9.5%. Three times fewer holdings, half the risk — because what you hold together matters more than how many things you hold.

The two kinds of risk, now defined rather than asserted

The Risk subject named two kinds of risk. The formula defines them.

Specific risk is the σ2/N\sigma^2/N term — the part unique to each holding, which averages away.

Systematic risk is the ρˉ σ2\bar\rho\,\sigma^2 term — the part holdings share, which does not.

This is where chapter 7's central claim will come from. If specific risk can be removed at no cost by anyone willing to hold more things, then nobody needs to be paid to bear it — and in equilibrium, nobody is. Only the shared part earns a return. That argument is CAPM in embryo, and it is already visible in the second term of this formula.

The correlations you actually get

The theory needs ρˉ\bar\rho as an input, and India's published index data shows what sort of numbers appear.

The Nifty200 Quality 30 — thirty companies selected on quality scores — has a five-year correlation with the Nifty 50 of 0.84. A correlation that high between a thirty-stock selection and a fifty-stock index means the two are largely the same bet. Squaring it, about 71% of the Quality index's variance is shared with the Nifty, leaving 29% that is not.

Two honest cautions about using any such number.

Measured correlations are estimates from a sample, and they move with the window. The Quality index's correlation with the Nifty is 0.84 over five years and 0.91 since inception — the same pair of indices, two different numbers.

And they rise when it matters. The Risk subject's chapter on correlation is specifically about this: the correlations in the table are measured in ordinary markets and tend towards 1 in the markets you diversified for. The formula is exactly right and the input is unreliable in precisely the state you most want it. Chapter 12 returns to this as one of the theory's real failures rather than a quibble.

Working the problem

Thirty stocks, each 30% volatility, equal weights.

At ρˉ=0.2\bar\rho = 0.2:

σp2=90030+(1−130)(0.2)(900)=30+174=204\sigma_p^2 = \frac{900}{30} + \left(1 - \frac{1}{30}\right)(0.2)(900) = 30 + 174 = 204

σp=204=14.3%\sigma_p = \sqrt{204} = 14.3\%

At ρˉ=0\bar\rho = 0:

σp2=90030+0=30,σp=5.5%\sigma_p^2 = \frac{900}{30} + 0 = 30, \qquad \sigma_p = 5.5\%

What the thirtieth stock removed. Going from 29 to 30 holdings, the specific-risk term falls from 900/29=31.03900/29 = 31.03 to 900/30=30.00900/30 = 30.00 — about one unit of variance. In standard-deviation terms the portfolio goes from 14.312% to 14.283%, an improvement of under three hundredths of a percentage point.

What it could not remove. The 174 units of covariance variance, which is 85% of the total. The thirtieth stock did nothing to that, and neither will the three-hundredth.

So the answer to "how many holdings" is really a question about correlation. At ρˉ=0.2\bar\rho = 0.2 a thirty-stock portfolio has already captured essentially all the available benefit: it sits at 14.28% against a floor of 13.42%, so the remaining 0.87 points is all that unlimited further diversification could ever buy. Adding holdings past this point adds cost, admission and monitoring for a benefit in the second decimal place.

And the comparison between the two columns is the real lesson. The same thirty stocks give 14.3% or 5.5% depending on nothing but their average correlation. An investor who holds thirty companies in one sector and an investor who holds thirty across unrelated sectors have done the same amount of work and bought very different amounts of protection — which is why the Risk subject insists that a long list of holdings is not by itself diversification, and why the Nifty 50's 37.45% weight in financial services is worth knowing.

The point

Covariance enters portfolio variance through a cross term that carries the sign of the correlation, so a portfolio can be less risky than the weighted average of its parts — at ρ=0\rho = 0 two equal 20% assets combine to 14.14%, and at ρ=−1\rho = -1 to zero. With NN equally weighted holdings the variance splits into σ2/N\sigma^2/N, which vanishes as holdings are added, and (1−1/N)ρˉσ2(1-1/N)\bar\rho\sigma^2, which does not — so risk falls towards a floor of σρˉ\sigma\sqrt{\bar\rho}. That is the difference between specific and systematic risk, defined rather than asserted, and it is already the seed of CAPM: the part that can be removed for free is the part nobody is paid to bear. Most of the benefit arrives in the first ten or so holdings, and average correlation matters far more than the count.

Check yourself

4 questions. Every answer is explained afterwards, including the ones you get right — guessing correctly is not the same as knowing. Score 70% or more and the chapter is marked done.

Question 1 of 4

RiskModerate
Two assets each have a standard deviation of 20% and are uncorrelated. Held in equal weights, what is the portfolio’s standard deviation, in per cent?

0 of 4 answered. You can submit with questions unanswered — they simply score zero.

Now do it with your own numbers

Thirty stocks, each with a standard deviation of 30%, held in equal weights. Work out the portfolio's standard deviation if the average pairwise correlation is 0.2, then if it is 0. Say which part of the risk the thirtieth stock removed and which part it could never remove.

Split the double sum into the diagonal terms and the off-diagonal ones, and see what each does as N grows.

Sources