Skip to content
VidiMaster it, module by module
Module 10/Covariance & correlation

Covariance & correlation

Covariance measures how two variables move together; correlation rescales it to [−1,1][-1,1]. Covariance is bilinear — just like an inner product — which is exactly the bridge back to linear algebra and PCA.

Before you start — give these a try

Attempting first primes your brain for the lesson — even if you miss. Nothing is graded or saved; it's just a warm-up.

E[XY]=10, E[X]=2, E[Y]=4E[XY]=10,\ E[X]=2,\ E[Y]=4. Compute Cov(X,Y)=E[XY]−E[X]E[Y]\mathrm{Cov}(X,Y)=E[XY]-E[X]E[Y].

Var(X)=4, Var(Y)=9, Cov(X,Y)=3\mathrm{Var}(X)=4,\ \mathrm{Var}(Y)=9,\ \mathrm{Cov}(X,Y)=3. Compute Var(X+Y)\mathrm{Var}(X+Y).

What you’ll be able to do

  • Compute covariance via Cov(X,Y)=E[XY]−E[X]E[Y]\mathrm{Cov}(X,Y)=E[XY]-E[X]E[Y].
  • Use Var(X+Y)=Var(X)+Var(Y)+2Cov(X,Y)\mathrm{Var}(X+Y)=\mathrm{Var}(X)+\mathrm{Var}(Y)+2\mathrm{Cov}(X,Y) and bilinearity.
  • Compute the correlation coefficient and know independent ⟹ uncorrelated (but not the converse).

In your course

· MATH2015 · Linear Algebra & Probability
§15.5 Covariance and Correlation
  • Definition 15.7Covariance Cov(X,Y)=E[(X−μX)(Y−μY)]\mathrm{Cov}(X,Y)=E[(X-\mu_X)(Y-\mu_Y)]
  • Proposition 15.13Cov(X,Y)=E[XY]−μXμY\mathrm{Cov}(X,Y)=E[XY]-\mu_X\mu_Y
  • Proposition 15.14Var(X+Y)=Var(X)+Var(Y)+2 Cov(X,Y)\mathrm{Var}(X+Y)=\mathrm{Var}(X)+\mathrm{Var}(Y)+2\,\mathrm{Cov}(X,Y)
  • Proposition 15.16Independent ⟹ uncorrelated (converse false)
  • Proposition 15.17Bilinearity of covariance (like an inner product)
  • Definition 15.9 / Theorem 15.1Correlation coefficient, ρ∈[−1,1]\rho\in[-1,1]
Bilinearity of covariance (Remark 15.7) mirrors the dot product; the covariance matrix is what PCA (§15.6, covered in Module 5) diagonalizes — the bridge from probability back to linear algebra.
1

Covariance

Cov(X,Y)=E[(X−μX)(Y−μY)]=E[XY]−μXμY\mathrm{Cov}(X,Y)=E[(X-\mu_X)(Y-\mu_Y)]=E[XY]-\mu_X\mu_Y (Definition 15.7, Proposition 15.13). It is positive when X,YX,Y tend to be large together, negative when one is high as the other is low, and Cov(X,X)=Var(X)\mathrm{Cov}(X,X)=\mathrm{Var}(X).

2

Variance of a sum, in general

Var(X+Y)=Var(X)+Var(Y)+2 Cov(X,Y)\mathrm{Var}(X+Y)=\mathrm{Var}(X)+\mathrm{Var}(Y)+2\,\mathrm{Cov}(X,Y)\n\n(Proposition 15.14). The covariance term is what makes variances fail to simply add when variables are correlated.

3

Independent ⟹ uncorrelated (not conversely)

If X,YX,Y are independent then Cov(X,Y)=0\mathrm{Cov}(X,Y)=0 (Proposition 15.16). The converse fails: e.g. XX uniform on {−1,0,1}\{-1,0,1\} and Y=X2Y=X^2 have Cov(X,Y)=0\mathrm{Cov}(X,Y)=0 but are clearly dependent.

4

Correlation coefficient & bilinearity

To remove scale, use the correlation ρ(X,Y)=Cov(X,Y)Var(X) Var(Y)∈[−1,1]\rho(X,Y)=\dfrac{\mathrm{Cov}(X,Y)}{\sqrt{\mathrm{Var}(X)\,\mathrm{Var}(Y)}}\in[-1,1] (Definition 15.9, Theorem 15.1); ±1\pm1 means a perfect linear relationship. Covariance is bilinear — Cov(aX,Y)=a Cov(X,Y)\mathrm{Cov}(aX,Y)=a\,\mathrm{Cov}(X,Y) and it distributes over sums (Prop 15.17) — mirroring the dot product, which is why the covariance matrix feeds PCA.

Proposition 15.13 — Covariance formula

Cov(X,Y)=E[XY]−E[X] E[Y]\mathrm{Cov}(X,Y)=E[XY]-E[X]\,E[Y].

Intuition. Expand (X−μX)(Y−μY)(X-\mu_X)(Y-\mu_Y) and use linearity.
Proposition 15.14 — Variance of a sum

Var(X+Y)=Var(X)+Var(Y)+2 Cov(X,Y)\mathrm{Var}(X+Y)=\mathrm{Var}(X)+\mathrm{Var}(Y)+2\,\mathrm{Cov}(X,Y); more generally add 2∑i<jCov(Xi,Xj)2\sum_{i<j}\mathrm{Cov}(X_i,X_j).

Intuition. Correlated parts reinforce (or cancel) the spread.
Theorem 15.1 — Correlation bounds

For positive finite variances, ρ(X,Y)=Cov(X,Y)Var(X)Var(Y)∈[−1,1]\rho(X,Y)=\dfrac{\mathrm{Cov}(X,Y)}{\sqrt{\mathrm{Var}(X)\mathrm{Var}(Y)}}\in[-1,1], with ±1\pm1 iff X,YX,Y are perfectly linearly related.

Intuition. Correlation is the cosine of an ‘‘angle’’ between the centered variables.

Worked examples

Example 1

E[XY]=10, E[X]=2, E[Y]=4E[XY]=10,\ E[X]=2,\ E[Y]=4. Find Cov(X,Y)\mathrm{Cov}(X,Y).

  1. 1

    Cov(X,Y)=E[XY]−E[X]E[Y]=10−2⋅4\mathrm{Cov}(X,Y)=E[XY]-E[X]E[Y]=10-2\cdot 4.

  2. 2

    =10−8=2=10-8=2.

Answer. Cov(X,Y)=2\mathrm{Cov}(X,Y)=2.
Example 2

XX is uniform on {−1,0,1}\{-1,0,1\} and Y=X2Y=X^2. Show they are uncorrelated.