Skip to content
VidiMaster it, module by module
Module 8/Expectation & Variance

Variance & Standard Deviation

The mean locates a random variable; the variance measures how far it typically strays from that centre. For XX with mean μ=E[X]\mu=\mathbb{E}[X], Definition 13.4 sets Var(X)=E[(X−μ)2]\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2] — the average squared deviation from the mean — written σ2\sigma^2, with its square root the standard deviation σ=SD(X)=Var(X)\sigma=\mathrm{SD}(X)=\sqrt{\mathrm{Var}(X)} restoring the original units. Squaring the deviation inside the expectation makes Proposition 13.2 a sum for a discrete variable, Var(X)=∑k(k−μ)2 P(X=k)\mathrm{Var}(X)=\sum_k(k-\mu)^2\,\mathbb{P}(X=k), but expanding that square yields the far handier computational formula of Proposition 13.3, Var(X)=E[X2]−μ2\mathrm{Var}(X)=\mathbb{E}[X^2]-\mu^2 — the mean of the square minus the square of the mean. From the definition flow the transformation rules of Proposition 13.4: a shift moves the centre but not the spread, while a scale factor comes out squared, Var(aX+b)=a2Var(X)\mathrm{Var}(aX+b)=a^2\mathrm{Var}(X) (so SD(aX+b)=∣a∣ SD(X)\mathrm{SD}(aX+b)=|a|\,\mathrm{SD}(X)). Finally Proposition 13.5 reads off the extreme case — Var(X)=0\mathrm{Var}(X)=0 precisely when XX is almost surely constant. (The specific variances of the Bernoulli, binomial, uniform and geometric families, Examples 13.9–13.12, are collected in a separate lesson on the moments of named distributions.)

Before you start — give these a try

Attempting first primes your brain for the lesson — even if you miss. Nothing is graded or saved; it's just a warm-up.

A random variable XX has p.m.f. P(X=1)=0.25\mathbb{P}(X=1)=0.25, P(X=3)=0.5\mathbb{P}(X=3)=0.5, P(X=5)=0.25\mathbb{P}(X=5)=0.25. Find the standard deviation SD(X)\mathrm{SD}(X). Give your answer as a decimal to 33 places.

Which of the following is the computational formula for the variance of a random variable XX (Proposition 13.3)?

What you’ll be able to do

  • State Definition 13.4: the variance Var(X)=E[(X−μ)2]\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2] with μ=E[X]\mu=\mathbb{E}[X], the notation σ2=Var(X)\sigma^2=\mathrm{Var}(X), and the standard deviation σ=SD(X)=Var(X)\sigma=\mathrm{SD}(X)=\sqrt{\mathrm{Var}(X)}.
  • Compute a variance directly from a p.m.f. using the discrete formula Var(X)=∑k(k−μ)2 P(X=k)\mathrm{Var}(X)=\sum_k(k-\mu)^2\,\mathbb{P}(X=k) (Proposition 13.2).
  • Apply the computational formula Var(X)=E[X2]−μ2=E[X2]−(E[X])2\mathrm{Var}(X)=\mathbb{E}[X^2]-\mu^2=\mathbb{E}[X^2]-(\mathbb{E}[X])^2 (Proposition 13.3) as the fast route to a variance from the first two moments.
  • Use the scaling law Var(aX+b)=a2 Var(X)\mathrm{Var}(aX+b)=a^2\,\mathrm{Var}(X) (Proposition 13.4) and read off that an additive shift leaves spread unchanged while SD(aX+b)=∣a∣ SD(X)\mathrm{SD}(aX+b)=|a|\,\mathrm{SD}(X).
  • Interpret variance and standard deviation as measures of spread about the mean: Var(X)≥0\mathrm{Var}(X)\ge 0 always, with Var(X)=0\mathrm{Var}(X)=0 exactly when XX is almost surely constant (Proposition 13.5).

In your course

· MATH2015 · Linear Algebra & Probability
§13.2 Variance
  • Definition 13.3Moments
    The nnth moment of XX is E[Xn]\mathbb{E}[X^n]; the second moment E[X2]\mathbb{E}[X^2] is called the mean square.
  • Definition 13.4Variance
    Var(X)=E[(X−μ)2]=σ2\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2]=\sigma^2, with standard deviation σ=SD(X)=Var(X)\sigma=\mathrm{SD}(X)=\sqrt{\mathrm{Var}(X)}.
  • Proposition 13.2Variance from the p.m.f. or density
    Discrete: Var(X)=∑k(k−μ)2P(X=k)\mathrm{Var}(X)=\sum_k(k-\mu)^2\mathbb{P}(X=k); density ff: Var(X)=∫−∞∞(x−μ)2f(x) dx\mathrm{Var}(X)=\int_{-\infty}^{\infty}(x-\mu)^2 f(x)\,dx.
  • Proposition 13.3Computational formula for variance
    Var(X)=E[X2]−(E[X])2=E[X2]−μ2\mathrm{Var}(X)=\mathbb{E}[X^2]-(\mathbb{E}[X])^2=\mathbb{E}[X^2]-\mu^2.
  • Proposition 13.4Mean and variance of aX+baX+b
    E(aX+b)=aE(X)+b\mathbb{E}(aX+b)=a\mathbb{E}(X)+b and Var(aX+b)=a2Var(X)\mathrm{Var}(aX+b)=a^2\mathrm{Var}(X).
  • Proposition 13.5Zero variance
    Var(X)≥0\mathrm{Var}(X)\ge 0, with Var(X)=0\mathrm{Var}(X)=0 if and only if XX is almost surely constant (P(X=μ)=1\mathbb{P}(X=\mu)=1).
Per-distribution variances — Bernoulli p(1−p)p(1-p), binomial np(1−p)np(1-p), uniform (b−a)212\tfrac{(b-a)^2}{12}, geometric 1−pp2\tfrac{1-p}{p^2} (Examples 13.9–13.12) — are kept light here and developed in the dedicated moments-of-distributions lesson.
1

Variance: the average squared deviation

The mean μ=E[X]\mu=\mathbb{E}[X] says where a random variable sits; it says nothing about how tightly its values cluster there. Variance fills that gap. Measuring a typical deviation X−μX-\mu directly is useless because departures above and below the mean cancel: E[X−μ]=μ−μ=0\mathbb{E}[X-\mu]=\mu-\mu=0 for every XX. The remedy is to square the deviation before averaging, so that every discrepancy counts positively. Definition 13.4 therefore sets Var(X)=E[(X−μ)2],\mathrm{Var}(X)=\mathbb{E}\big[(X-\mu)^2\big], the mean squared deviation from μ\mu, written σ2\sigma^2. For a discrete XX, Proposition 13.2 turns this expectation into a weighted sum over the possible values, Var(X)=∑k(k−μ)2 P(X=k)\mathrm{Var}(X)=\sum_k (k-\mu)^2\,\mathbb{P}(X=k) (with the integral ∫−∞∞(x−μ)2f(x) dx\int_{-\infty}^{\infty}(x-\mu)^2 f(x)\,dx when XX has density ff). Each term weighs how far a value lies from the mean, squared, by how likely that value is: a variable that can land far from μ\mu with appreciable probability has a large variance, and one glued near μ\mu a small one. Because the deviation is squared, variance carries the square of XX's units — one reason the standard deviation is often reported in its place.

2

The computational formula $\mathbb{E}[X^2]-\mu^2$

Summing (k−μ)2P(X=k)(k-\mu)^2\mathbb{P}(X=k) works, but it is clumsy: it needs μ\mu first and then a fresh pass squaring each centred value. Expanding the square gives a shortcut. Since (X−μ)2=X2−2μX+μ2(X-\mu)^2=X^2-2\mu X+\mu^2 and μ\mu is a constant, linearity of expectation yields Var(X)=E[X2]−2μ E[X]+μ2=E[X2]−2μ2+μ2,\mathrm{Var}(X)=\mathbb{E}[X^2]-2\mu\,\mathbb{E}[X]+\mu^2=\mathbb{E}[X^2]-2\mu^2+\mu^2, which collapses to the computational formula of Proposition 13.3: Var(X)=E[X2]−μ2=E[X2]−(E[X])2.\mathrm{Var}(X)=\mathbb{E}[X^2]-\mu^2=\mathbb{E}[X^2]-(\mathbb{E}[X])^2. In words, variance is the mean of the square minus the square of the mean. In practice you read two ordinary expectations off the p.m.f. — μ=E[X]=∑kk P(X=k)\mu=\mathbb{E}[X]=\sum_k k\,\mathbb{P}(X=k) and the second moment E[X2]=∑kk2 P(X=k)\mathbb{E}[X^2]=\sum_k k^2\,\mathbb{P}(X=k) — and subtract μ2\mu^2. This is almost always the quicker computation by hand, and it is the route used throughout this lesson. It even carries a free sanity check: because Var(X)≥0\mathrm{Var}(X)\ge 0, you must always find E[X2]≥(E[X])2\mathbb{E}[X^2]\ge(\mathbb{E}[X])^2, so a negative answer signals an arithmetic slip.

3

Standard deviation: spread in the original units

Variance is measured in the square of XX's units — square dollars, square seconds — which makes its size awkward to interpret next to the mean. Taking the square root cures this. The standard deviation is σ=SD(X)=Var(X),\sigma=\mathrm{SD}(X)=\sqrt{\mathrm{Var}(X)}, a nonnegative number back in the original units of XX, representing a typical distance of XX from its mean. A small σ\sigma means the distribution is concentrated near μ\mu; a large σ\sigma means it is widely spread. Variance and standard deviation carry exactly the same information — each determines the other — but σ\sigma is the one you can sensibly compare to μ\mu, or mark off on the same axis as the values of XX. In practice you finish most variance computations by taking a square root to report σ\sigma.

4

Shifting, scaling, and zero variance

Variance responds to the two basic transformations of a random variable in a memorable way. Proposition 13.4 states that for constants a,ba,b, Var(aX+b)=a2 Var(X).\mathrm{Var}(aX+b)=a^2\,\mathrm{Var}(X). Two things stand out. First, the additive constant bb disappears: adding bb slides the whole distribution along the line without changing how spread out it is, so Var(X+b)=Var(X)\mathrm{Var}(X+b)=\mathrm{Var}(X) — spread is about distances between values, and a rigid shift preserves those distances. Second, the multiplier aa comes out squared, not linearly: stretching values by aa stretches every deviation by aa, hence every squared deviation by a2a^2. At the level of the standard deviation this reads cleanly as SD(aX+b)=∣a∣ SD(X)\mathrm{SD}(aX+b)=|a|\,\mathrm{SD}(X) (the absolute value because a standard deviation is never negative). A useful consequence is that Var(2X)=4 Var(X)\mathrm{Var}(2X)=4\,\mathrm{Var}(X), not 2 Var(X)2\,\mathrm{Var}(X). At the opposite extreme, Proposition 13.5 pins down when there is no spread at all: since Var(X)=E[(X−μ)2]\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2] is the expectation of a nonnegative quantity it is always ≥0\ge 0, and it equals 00 exactly when XX is almost surely constant, i.e. P(X=μ)=1\mathbb{P}(X=\mu)=1. (The variances of the standard families — p(1−p)p(1-p) for a Bernoulli, np(1−p)np(1-p) for a binomial, (b−a)212\tfrac{(b-a)^2}{12} for a uniform, 1−pp2\tfrac{1-p}{p^2} for a geometric — are gathered in a separate lesson on the moments of named distributions.)

Definition 13.4 — Variance

Let XX be a random variable with mean μ=E[X]\mu=\mathbb{E}[X]. Its variance is Var(X)=E[(X−μ)2],\mathrm{Var}(X)=\mathbb{E}\big[(X-\mu)^2\big], also written σ2\sigma^2. The standard deviation is σ=SD(X)=Var(X)\sigma=\mathrm{SD}(X)=\sqrt{\mathrm{Var}(X)}. For a discrete XX, Var(X)=∑k(k−μ)2 P(X=k)\mathrm{Var}(X)=\sum_k(k-\mu)^2\,\mathbb{P}(X=k) (Proposition 13.2), and ∫−∞∞(x−μ)2f(x) dx\int_{-\infty}^{\infty}(x-\mu)^2 f(x)\,dx when XX has density ff.

Intuition. Variance is the average squared distance from the mean, so it quantifies spread. Squaring stops positive and negative deviations from cancelling (their unsquared average is always 00) and penalises large departures heavily. The standard deviation undoes the squaring to report that spread in the original units of XX.
Proposition 13.3 — Computational formula for variance

For any random variable with a well-defined mean μ=E[X]\mu=\mathbb{E}[X], Var(X)=E[X2]−(E[X])2=E[X2]−μ2.\mathrm{Var}(X)=\mathbb{E}[X^2]-(\mathbb{E}[X])^2=\mathbb{E}[X^2]-\mu^2.

Intuition. Expanding (X−μ)2=X2−2μX+μ2(X-\mu)^2=X^2-2\mu X+\mu^2 and using linearity collapses the definition to 'the mean of the square minus the square of the mean'. It is the fastest hand computation: read E[X]\mathbb{E}[X] and E[X2]\mathbb{E}[X^2] off the p.m.f., then subtract μ2\mu^2. Since variance is ≥0\ge 0, you always get E[X2]≥μ2\mathbb{E}[X^2]\ge\mu^2.
Proposition 13.4 — Mean and variance of $aX+b$

Let a,b∈Ra,b\in\mathbb{R}. Provided the mean and variance of XX exist, E(aX+b)=a E(X)+b,Var(aX+b)=a2 Var(X).\mathbb{E}(aX+b)=a\,\mathbb{E}(X)+b,\qquad \mathrm{Var}(aX+b)=a^2\,\mathrm{Var}(X). Equivalently, SD(aX+b)=∣a∣ SD(X)\mathrm{SD}(aX+b)=|a|\,\mathrm{SD}(X).

Intuition. A shift by bb moves the mean but not the spread, so bb drops out of the variance. A scale factor aa multiplies every deviation by aa, hence every squared deviation by a2a^2 — so the scale appears squared. In particular Var(2X)=4 Var(X)\mathrm{Var}(2X)=4\,\mathrm{Var}(X), never 2 Var(X)2\,\mathrm{Var}(X).
Proposition 13.5 — Zero variance

For any random variable, Var(X)≥0\mathrm{Var}(X)\ge 0, and Var(X)=0\mathrm{Var}(X)=0 if and only if XX is almost surely constant, i.e. P(X=μ)=1\mathbb{P}(X=\mu)=1 where μ=E[X]\mu=\mathbb{E}[X].

Intuition. Variance is the expectation of the nonnegative quantity (X−μ)2(X-\mu)^2, so it can never be negative. It hits 00 only when (X−μ)2(X-\mu)^2 is 00 with probability one — that is, when XX takes the single value μ\mu almost surely and there is no spread to measure.

Worked examples

Example 1

A random variable XX has p.m.f. P(X=0)=14\mathbb{P}(X=0)=\tfrac14, P(X=1)=12\mathbb{P}(X=1)=\tfrac12, P(X=2)=14\mathbb{P}(X=2)=\tfrac14. Find E[X]\mathbb{E}[X], compute Var(X)\mathrm{Var}(X) using the computational formula E[X2]−μ2\mathbb{E}[X^2]-\mu^2, and give the standard deviation.

  1. 1

    Mean. μ=E[X]=0⋅14+1⋅12+2⋅14=12+12=1.\mu=\mathbb{E}[X]=0\cdot\tfrac14+1\cdot\tfrac12+2\cdot\tfrac14=\tfrac12+\tfrac12=1.

  2. 2

    Second moment. E[X2]=02⋅14+12⋅12+22⋅14=12+1=32.\mathbb{E}[X^2]=0^2\cdot\tfrac14+1^2\cdot\tfrac12+2^2\cdot\tfrac14=\tfrac12+1=\tfrac32.

  3. 3

    Computational formula (Proposition 13.3). Var(X)=E[X2]−μ2=32−12=12=0.5.\mathrm{Var}(X)=\mathbb{E}[X^2]-\mu^2=\tfrac32-1^2=\tfrac12=0.5.

  4. 4

    Cross-check with the definition. Var(X)=∑k(k−1)2P(X=k)=(−1)214+0212+1214=14+14=12\mathrm{Var}(X)=\sum_k(k-1)^2\mathbb{P}(X=k)=(-1)^2\tfrac14+0^2\tfrac12+1^2\tfrac14=\tfrac14+\tfrac14=\tfrac12 — the same value.

  5. 5

    Standard deviation. σ=Var(X)=12≈0.707.\sigma=\sqrt{\mathrm{Var}(X)}=\sqrt{\tfrac12}\approx0.707.

Answer. μ=1\mu=1, Var(X)=12=0.5\mathrm{Var}(X)=\tfrac12=0.5, and SD(X)=0.5≈0.707\mathrm{SD}(X)=\sqrt{0.5}\approx0.707.
Example 2

Let XX have p.m.f. P(X=−2)=12\mathbb{P}(X=-2)=\tfrac12, P(X=2)=12\mathbb{P}(X=2)=\tfrac12. (a) Find Var(X)\mathrm{Var}(X) and SD(X)\mathrm{SD}(X). (b) Use Proposition 13.4 to find Var(3X−1)\mathrm{Var}(3X-1) and SD(3X−1)\mathrm{SD}(3X-1). (c) What is Var(X−1)\mathrm{Var}(X-1)?

Example 3

Let X∼Ber(0.3)X\sim\mathrm{Ber}(0.3), so P(X=1)=0.3\mathbb{P}(X=1)=0.3 and P(X=0)=0.7\mathbb{P}(X=0)=0.7. Compute Var(X)\mathrm{Var}(X) directly, confirm it matches the Bernoulli formula p(1−p)p(1-p), and give SD(X)\mathrm{SD}(X).