Skip to content
VidiMaster it, module by module
Module 8/Expectation & Variance

Mean and Variance of the Common Distributions

Two numbers summarise any distribution: the mean μ=E[X]\mu=\mathbb{E}[X] (Definition 13.1, the first moment), which says where the distribution is centred, and the variance σ2=Var(X)=E[(X−μ)2]\sigma^2=\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2] (Definition 13.4), which says how widely it spreads about that centre; its square root σ=SD(X)\sigma=\mathrm{SD}(X) is the standard deviation. For hand computation the variance is almost always found from the computational formula Var(X)=E[X2]−(E[X])2\mathrm{Var}(X)=\mathbb{E}[X^2]-(\mathbb{E}[X])^2 (Proposition 13.3). This lesson is a consolidated reference: it collects, once and for all, the mean and variance of the four distributions you meet most often, each derived in the course's own notes. For the discrete families these are the Bernoulli Ber(p)\mathrm{Ber}(p) with mean pp (Example 13.2) and variance p(1−p)p(1-p) (Example 13.9); the binomial Bin(n,p)\mathrm{Bin}(n,p) with mean npnp (Example 13.3) and variance np(1−p)np(1-p) (Example 13.10); and the geometric Geom(p)\mathrm{Geom}(p) (support k≥1k\ge 1, p.m.f. (1−p)k−1p(1-p)^{k-1}p) with mean 1p\frac{1}{p} (Example 13.4) and variance 1−pp2\frac{1-p}{p^2} (Example 13.12). For the continuous world we record the uniform Unif[a,b]\mathrm{Unif}[a,b] with mean a+b2\frac{a+b}{2} (Example 13.7) and variance (b−a)212\frac{(b-a)^2}{12} (Example 13.11). Once you can name the family and read off its parameters, every mean and variance is a one-line substitution.

Before you start — give these a try

Attempting first primes your brain for the lesson — even if you miss. Nothing is graded or saved; it's just a warm-up.

Let X∼Bin(12,0.25)X\sim\mathrm{Bin}(12,0.25). Find Var(X)\mathrm{Var}(X). Give your answer as a decimal to 33 places.

Let X∼Geom(0.4)X\sim\mathrm{Geom}(0.4) (trials until the first success, p.m.f. (1−p)k−1p(1-p)^{k-1}p for k≥1k\ge 1). Find Var(X)\mathrm{Var}(X). Give your answer as a decimal to 33 places.

What you’ll be able to do

  • Recall that the mean is the first moment μ=E[X]\mu=\mathbb{E}[X] (Definitions 13.1-13.2) and the variance is σ2=Var(X)=E[(X−μ)2]\sigma^2=\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2] (Definition 13.4), and compute variances through the computational formula Var(X)=E[X2]−(E[X])2\mathrm{Var}(X)=\mathbb{E}[X^2]-(\mathbb{E}[X])^2 (Proposition 13.3).
  • State and apply the Bernoulli moments: X∼Ber(p)X\sim\mathrm{Ber}(p) has mean E[X]=p\mathbb{E}[X]=p (Example 13.2) and variance Var(X)=p(1−p)\mathrm{Var}(X)=p(1-p) (Example 13.9).
  • State and apply the binomial moments: X∼Bin(n,p)X\sim\mathrm{Bin}(n,p) has mean E[X]=np\mathbb{E}[X]=np (Example 13.3) and variance Var(X)=np(1−p)\mathrm{Var}(X)=np(1-p) (Example 13.10).
  • State and apply the geometric moments: X∼Geom(p)X\sim\mathrm{Geom}(p) has mean E[X]=1p\mathbb{E}[X]=\frac{1}{p} (Example 13.4) and variance Var(X)=1−pp2\mathrm{Var}(X)=\frac{1-p}{p^2} (Example 13.12).
  • State and apply the continuous uniform moments: X∼Unif[a,b]X\sim\mathrm{Unif}[a,b] has mean E[X]=a+b2\mathbb{E}[X]=\frac{a+b}{2} (Example 13.7) and variance Var(X)=(b−a)212\mathrm{Var}(X)=\frac{(b-a)^2}{12} (Example 13.11), and obtain the standard deviation σ=Var(X)\sigma=\sqrt{\mathrm{Var}(X)}.

In your course

· MATH2015 · Linear Algebra & Probability
§13.1 Expectation§13.2 Variance
  • Definition 13.1Expectation (first moment)
    μ=E[X]=∑kk P(X=k)\mu=\mathbb{E}[X]=\sum_k k\,\mathbb{P}(X=k) for discrete XX (and ∫xf(x) dx\int x f(x)\,dx for continuous XX); the expectation is the first moment.
  • Definition 13.3Moments
    The nnth moment of XX is E[Xn]\mathbb{E}[X^n]; the second moment E[X2]\mathbb{E}[X^2] is the mean square.
  • Definition 13.4Variance and standard deviation
    Var(X)=E[(X−μ)2]\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2], with σ=SD(X)=Var(X)\sigma=\mathrm{SD}(X)=\sqrt{\mathrm{Var}(X)}.
  • Proposition 13.3Computational formula for the variance
    Var(X)=E[X2]−(E[X])2\mathrm{Var}(X)=\mathbb{E}[X^2]-(\mathbb{E}[X])^2.
  • Proposition 13.4Mean and variance under affine maps
    E(aX+b)=aE(X)+b\mathbb{E}(aX+b)=a\mathbb{E}(X)+b and Var(aX+b)=a2Var(X)\mathrm{Var}(aX+b)=a^2\mathrm{Var}(X).
  • Example 13.2Mean of a Bernoulli random variable
    X∼Ber(p)⇒E[X]=pX\sim\mathrm{Ber}(p)\Rightarrow\mathbb{E}[X]=p.
  • Example 13.3Mean of a binomial random variable
    X∼Bin(n,p)⇒E[X]=npX\sim\mathrm{Bin}(n,p)\Rightarrow\mathbb{E}[X]=np.
  • Example 13.4Mean of a geometric random variable
    X∼Geom(p)⇒E[X]=1pX\sim\mathrm{Geom}(p)\Rightarrow\mathbb{E}[X]=\frac{1}{p}.
  • Example 13.7Mean of a uniform random variable
    X∼Unif[a,b]⇒E[X]=a+b2X\sim\mathrm{Unif}[a,b]\Rightarrow\mathbb{E}[X]=\frac{a+b}{2}.
  • Example 13.9Variance of a Bernoulli random variable
    X∼Ber(p)⇒Var(X)=p(1−p)X\sim\mathrm{Ber}(p)\Rightarrow\mathrm{Var}(X)=p(1-p); for an indicator, Var(1A)=P(A)P(Ac)\mathrm{Var}(\mathbf{1}_A)=\mathbb{P}(A)\mathbb{P}(A^c).
  • Example 13.10Variance of a binomial random variable
    X∼Bin(n,p)⇒Var(X)=np(1−p)X\sim\mathrm{Bin}(n,p)\Rightarrow\mathrm{Var}(X)=np(1-p).
  • Example 13.11Variance of a uniform random variable
    X∼Unif[a,b]⇒Var(X)=(b−a)212X\sim\mathrm{Unif}[a,b]\Rightarrow\mathrm{Var}(X)=\frac{(b-a)^2}{12}.
  • Example 13.12Variance of a geometric random variable
    X∼Geom(p)⇒Var(X)=1−pp2X\sim\mathrm{Geom}(p)\Rightarrow\mathrm{Var}(X)=\frac{1-p}{p^2}.
Means are from §13.1 (Examples 13.2-13.7) and variances from §13.2 (Examples 13.9-13.12). The geometric here counts trials up to and including the first success (support k≥1k\ge 1, p.m.f. (1−p)k−1p(1-p)^{k-1}p), giving mean 1/p1/p; a convention that starts the count at 00 would shift the mean by 11. All numerical answers were recomputed and cross-checked.
1

The mean and the variance: a distribution's two moments

The mean (or expectation, or first moment) μ=E[X]\mu=\mathbb{E}[X] is the probability-weighted average of the values of XX: a sum ∑kk P(X=k)\sum_k k\,\mathbb{P}(X=k) in the discrete case (Definition 13.1) and an integral ∫−∞∞xf(x) dx\int_{-\infty}^{\infty} x f(x)\,dx in the continuous case (Definition 13.2). It locates the centre of the distribution. The variance σ2=Var(X)=E[(X−μ)2]\sigma^2=\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2] (Definition 13.4) is the mean squared distance from that centre, so it measures spread; being an average of squares it is never negative, and its square root σ=SD(X)=Var(X)\sigma=\mathrm{SD}(X)=\sqrt{\mathrm{Var}(X)} — the standard deviation — restores the original units. In practice you rarely expand E[(X−μ)2]\mathbb{E}[(X-\mu)^2] directly. Instead you use the computational formula (Proposition 13.3) Var(X)=E[X2]−(E[X])2,\mathrm{Var}(X)=\mathbb{E}[X^2]-(\mathbb{E}[X])^2, which trades the deviation for the second moment E[X2]\mathbb{E}[X^2] (the mean square, Definition 13.3) minus the square of the mean. Finally, both summaries behave predictably under an affine change of variable (Proposition 13.4): E(aX+b)=aE(X)+b\mathbb{E}(aX+b)=a\mathbb{E}(X)+b, while Var(aX+b)=a2Var(X)\mathrm{Var}(aX+b)=a^2\mathrm{Var}(X) — shifting by bb moves the centre but not the spread, and scaling by aa multiplies the variance by a2a^2.

2

The three discrete families: Bernoulli, binomial, geometric

Three named discrete distributions cover most discrete models, and their moments are derived in §13.1-§13.2. A Bernoulli variable X∼Ber(p)X\sim\mathrm{Ber}(p) is a single trial scoring 11 with probability pp and 00 otherwise; its mean is E[X]=1⋅p+0⋅(1−p)=p\mathbb{E}[X]=1\cdot p+0\cdot(1-p)=p (Example 13.2) and its variance is Var(X)=p(1−p)\mathrm{Var}(X)=p(1-p) (Example 13.9), largest at p=12p=\tfrac12 (maximum uncertainty) and zero at p=0p=0 or p=1p=1 (a sure outcome). A binomial variable X∼Bin(n,p)X\sim\mathrm{Bin}(n,p) counts the successes in nn such independent trials, so both moments simply scale by nn: E[X]=np\mathbb{E}[X]=np (Example 13.3) and Var(X)=np(1−p)\mathrm{Var}(X)=np(1-p) (Example 13.10). Notice Var(X)=E[X]⋅(1−p)\mathrm{Var}(X)=\mathbb{E}[X]\cdot(1-p), so for a binomial the variance is always the mean times 1−p1-p. A geometric variable X∼Geom(p)X\sim\mathrm{Geom}(p) counts the number of independent trials up to and including the first success, with p.m.f. P(X=k)=(1−p)k−1p\mathbb{P}(X=k)=(1-p)^{k-1}p for k≥1k\ge 1; summing the series gives E[X]=1p\mathbb{E}[X]=\frac{1}{p} (Example 13.4), and with E[X2]=1+qp2\mathbb{E}[X^2]=\frac{1+q}{p^2} (where q=1−pq=1-p) the computational formula yields Var(X)=qp2=1−pp2\mathrm{Var}(X)=\frac{q}{p^2}=\frac{1-p}{p^2} (Example 13.12). A rare success (small pp) means a long, highly variable wait, which is why both 1p\frac{1}{p} and 1−pp2\frac{1-p}{p^2} blow up as p→0p\to 0.

3

The continuous uniform on $[a,b]$

The continuous uniform distribution X∼Unif[a,b]X\sim\mathrm{Unif}[a,b] spreads probability evenly over the interval [a,b][a,b], with constant density f(x)=1b−af(x)=\frac{1}{b-a} there and 00 outside. Its mean is found by integration (Example 13.7): E[X]=∫abx⋅1b−a dx=b2−a22(b−a)=a+b2,\mathbb{E}[X]=\int_a^b x\cdot\frac{1}{b-a}\,dx=\frac{b^2-a^2}{2(b-a)}=\frac{a+b}{2}, exactly the midpoint of the interval, as symmetry demands. For the variance (Example 13.11) first compute the second moment E[X2]=∫abx2⋅1b−a dx=b3−a33(b−a)=a2+ab+b23,\mathbb{E}[X^2]=\int_a^b x^2\cdot\frac{1}{b-a}\,dx=\frac{b^3-a^3}{3(b-a)}=\frac{a^2+ab+b^2}{3}, and then the computational formula gives Var(X)=E[X2]−(E[X])2=a2+ab+b23−(a+b2)2=(b−a)212.\mathrm{Var}(X)=\mathbb{E}[X^2]-(\mathbb{E}[X])^2=\frac{a^2+ab+b^2}{3}-\left(\frac{a+b}{2}\right)^2=\frac{(b-a)^2}{12}. The variance depends only on the width b−ab-a, never on where the interval sits: sliding [a,b][a,b] along the line (replacing XX by X+cX+c) leaves the spread unchanged, in line with Var(aX+b)=a2Var(X)\mathrm{Var}(aX+b)=a^2\mathrm{Var}(X). The standard deviation is σ=b−a12\sigma=\frac{b-a}{\sqrt{12}}.

4

A reference table, and how to use it

The four results together form a compact lookup table. Here p∈(0,1)p\in(0,1) is a success probability, nn a number of trials, and a<ba<b the endpoints of an interval.

| Distribution | Mean E[X]\mathbb{E}[X] | Variance Var(X)\mathrm{Var}(X) | | --- | --- | --- | | Ber(p)\mathrm{Ber}(p) | pp | p(1−p)p(1-p) | | Bin(n,p)\mathrm{Bin}(n,p) | npnp | np(1−p)np(1-p) | | Geom(p)\mathrm{Geom}(p) | 1p\dfrac{1}{p} | 1−pp2\dfrac{1-p}{p^2} | | Unif[a,b]\mathrm{Unif}[a,b] | a+b2\dfrac{a+b}{2} | (b−a)212\dfrac{(b-a)^2}{12} |

Applying it is a three-step recipe. First, name the family and read off its parameters from the wording (a single yes/no trial is Bernoulli; a count of successes in a fixed number of trials is binomial; a count of trials until the first success is geometric; an evenly-spread continuous quantity on an interval is uniform). Second, substitute the parameters into the mean and variance formulas — no summation or integration needed, since the derivations are already done. Third, if a spread in the original units is wanted, take σ=Var(X)\sigma=\sqrt{\mathrm{Var}(X)}. For instance, Bin(10,0.3)\mathrm{Bin}(10,0.3) has mean 10(0.3)=310(0.3)=3 and variance 10(0.3)(0.7)=2.110(0.3)(0.7)=2.1; Geom(0.25)\mathrm{Geom}(0.25) has mean 10.25=4\frac{1}{0.25}=4 and variance 0.750.252=12\frac{0.75}{0.25^2}=12; and Unif[2,8]\mathrm{Unif}[2,8] has mean 2+82=5\frac{2+8}{2}=5 and variance (8−2)212=3\frac{(8-2)^2}{12}=3.

Examples 13.2 & 13.9 — Mean and variance of a Bernoulli

If X∼Ber(p)X\sim\mathrm{Ber}(p) with P(X=1)=p\mathbb{P}(X=1)=p and P(X=0)=1−p\mathbb{P}(X=0)=1-p, then E[X]=p,Var(X)=p(1−p).\mathbb{E}[X]=p,\qquad \mathrm{Var}(X)=p(1-p). In particular, for the indicator 1A\mathbf{1}_A of an event AA, E[1A]=P(A)\mathbb{E}[\mathbf{1}_A]=\mathbb{P}(A) and Var(1A)=P(A)P(Ac)\mathrm{Var}(\mathbf{1}_A)=\mathbb{P}(A)\mathbb{P}(A^c).

Intuition. The mean is just the success probability. The variance p(1−p)p(1-p) is a downward parabola in pp: it is 00 when the outcome is certain (p=0p=0 or p=1p=1) and peaks at p=12p=\tfrac12, where a single trial is as unpredictable as possible.
Examples 13.3 & 13.10 — Mean and variance of a binomial

If X∼Bin(n,p)X\sim\mathrm{Bin}(n,p), then E[X]=np,Var(X)=np(1−p).\mathbb{E}[X]=np,\qquad \mathrm{Var}(X)=np(1-p).

Intuition. A binomial is the sum of nn independent Ber(p)\mathrm{Ber}(p) trials, so its mean and variance are nn times the Bernoulli values pp and p(1−p)p(1-p). Equivalently Var(X)=E[X]⋅(1−p)\mathrm{Var}(X)=\mathbb{E}[X]\cdot(1-p): the variance is the mean discounted by the failure probability.
Examples 13.4 & 13.12 — Mean and variance of a geometric

If X∼Geom(p)X\sim\mathrm{Geom}(p) with p.m.f. P(X=k)=(1−p)k−1p\mathbb{P}(X=k)=(1-p)^{k-1}p for k≥1k\ge 1, then E[X]=1p,Var(X)=1−pp2.\mathbb{E}[X]=\frac{1}{p},\qquad \mathrm{Var}(X)=\frac{1-p}{p^2}.

Intuition. XX is the waiting time for the first success. On average it takes 1/p1/p trials — one in ten when p=0.1p=0.1. Because rare successes produce long and erratic waits, both the mean and the variance grow without bound as p→0p\to 0.
Examples 13.7 & 13.11 — Mean and variance of a continuous uniform

If X∼Unif[a,b]X\sim\mathrm{Unif}[a,b] with density f(x)=1b−af(x)=\frac{1}{b-a} on [a,b][a,b], then E[X]=a+b2,Var(X)=(b−a)212.\mathbb{E}[X]=\frac{a+b}{2},\qquad \mathrm{Var}(X)=\frac{(b-a)^2}{12}.

Intuition. By symmetry the mean sits at the midpoint of the interval. The variance depends only on the width b−ab-a — sliding the interval left or right changes the centre but not the spread — and grows with the square of the width, matching Var(aX+b)=a2Var(X)\mathrm{Var}(aX+b)=a^2\mathrm{Var}(X).

Worked examples

Example 1

Binomial. In a quality-control batch, each of 2020 independent items is defective with probability 0.40.4. Let XX be the number of defective items, so X∼Bin(20,0.4)X\sim\mathrm{Bin}(20,0.4). Find the mean, variance, and standard deviation of XX.

  1. 1

    Identify the family and parameters. A count of successes (defectives) in a fixed number n=20n=20 of independent trials, each with success probability p=0.4p=0.4, is binomial: X∼Bin(n,p)X\sim\mathrm{Bin}(n,p) with n=20n=20, p=0.4p=0.4.

  2. 2

    Mean (Example 13.3). E[X]=np=20×0.4=8.\mathbb{E}[X]=np=20\times 0.4=8.

  3. 3

    Variance (Example 13.10). Var(X)=np(1−p)=20×0.4×0.6=4.8.\mathrm{Var}(X)=np(1-p)=20\times 0.4\times 0.6=4.8. (Check: Var(X)=E[X]⋅(1−p)=8×0.6=4.8\mathrm{Var}(X)=\mathbb{E}[X]\cdot(1-p)=8\times 0.6=4.8.)

  4. 4

    Standard deviation. σ=Var(X)=4.8≈2.191.\sigma=\sqrt{\mathrm{Var}(X)}=\sqrt{4.8}\approx 2.191.

Answer. E[X]=8\mathbb{E}[X]=8, Var(X)=4.8\mathrm{Var}(X)=4.8, and σ=4.8≈2.191\sigma=\sqrt{4.8}\approx 2.191.
Example 2

Geometric. A game is won on each independent attempt with probability 0.20.2. Let XX be the number of attempts up to and including the first win, so X∼Geom(0.2)X\sim\mathrm{Geom}(0.2). Find the mean, variance, and standard deviation of XX.

Example 3

Continuous uniform. A bus arrives at a time spread evenly between minute 33 and minute 1111 after you reach the stop; let X∼Unif[3,11]X\sim\mathrm{Unif}[3,11] be that arrival time. Find E[X]\mathbb{E}[X] and Var(X)\mathrm{Var}(X), and cross-check the variance via the computational formula.