Skip to content
VidiMaster it, module by module
Module 8/More Distributions

The Normal Distribution N(μ,σ2)N(\mu,\sigma^2)

The standard normal Z∼N(0,1)Z\sim N(0,1) is the reference bell curve: it has mean 00 and variance 11 (Proposition 13.6), and its cumulative distribution function is written Φ(z)=P(Z≤z)\Phi(z)=\mathbb{P}(Z\le z). Every other normal is an affine stretch-and-shift of it. For μ∈R\mu\in\mathbb{R} and σ>0\sigma>0, setting X=σZ+μX=\sigma Z+\mu produces the general normal distribution X∼N(μ,σ2)X\sim N(\mu,\sigma^2) (Definition 13.5), with mean μ\mu, variance σ2\sigma^2, and the bell-shaped density f(x)=1σ2πexp⁡ ⁣(−(x−μ)22σ2),f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), centred at μ\mu with its spread set by σ\sigma. Affine maps keep you inside the family (Proposition 13.7): if X∼N(μ,σ2)X\sim N(\mu,\sigma^2) then aX+b∼N(aμ+b, a2σ2)aX+b\sim N(a\mu+b,\,a^2\sigma^2), and the special case Z=X−μσZ=\frac{X-\mu}{\sigma} — standardization — turns any normal back into the standard one. That is the computational workhorse: P(X≤x)=Φ ⁣(x−μσ)\mathbb{P}(X\le x)=\Phi\!\left(\frac{x-\mu}{\sigma}\right) collapses every normal probability to a single Φ\Phi-value read from a standard-normal table or software, as in Example 13.14 (X∼N(−3,4)X\sim N(-3,4) gives P(X≤−1.7)=Φ(0.65)≈0.742\mathbb{P}(X\le-1.7)=\Phi(0.65)\approx0.742). The same device yields the 68–95–99.7 rule: a normal variable falls within 11, 22, or 33 standard deviations of its mean with probability about 0.6830.683, 0.9540.954, and 0.9970.997 (Example 13.15).

Before you start — give these a try

Attempting first primes your brain for the lesson — even if you miss. Nothing is graded or saved; it's just a warm-up.

Let X∼N(10,4)X\sim N(10,4). Find P(X≤13)\mathbb{P}(X\le13). Use the standard normal; give your answer to 33 decimals.

For the standard normal Z∼N(0,1)Z\sim N(0,1), use the standard normal to find P(−1≤Z≤1)\mathbb{P}(-1\le Z\le1) (the '6868' of the 6868–9595–99.799.7 rule). Give your answer to 33 decimals.

What you’ll be able to do

  • State Definition 13.5: X∼N(μ,σ2)X\sim N(\mu,\sigma^2) (with μ∈R\mu\in\mathbb{R}, σ>0\sigma>0) has density f(x)=1σ2πexp⁡ ⁣(−(x−μ)22σ2)f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), and read off that the first parameter μ\mu is the mean and the second parameter σ2\sigma^2 is the variance (so σ\sigma is the standard deviation).
  • Recall Proposition 13.6: the standard normal Z∼N(0,1)Z\sim N(0,1) has E[Z]=0\mathbb{E}[Z]=0 and Var(Z)=E[Z2]=1\mathrm{Var}(Z)=\mathbb{E}[Z^2]=1, and its c.d.f. Φ(z)=P(Z≤z)\Phi(z)=\mathbb{P}(Z\le z) is the object all normal probabilities are expressed through.
  • Apply Proposition 13.7 (linear transformations): if X∼N(μ,σ2)X\sim N(\mu,\sigma^2) and a≠0a\neq0, then aX+b∼N(aμ+b, a2σ2)aX+b\sim N(a\mu+b,\,a^2\sigma^2) — an affine image of a normal is normal — with the special case X=σZ+μX=\sigma Z+\mu building N(μ,σ2)N(\mu,\sigma^2) from N(0,1)N(0,1).
  • Use standardization Z=X−μσ∼N(0,1)Z=\frac{X-\mu}{\sigma}\sim N(0,1) to compute probabilities as P(X≤x)=Φ ⁣(x−μσ)\mathbb{P}(X\le x)=\Phi\!\left(\frac{x-\mu}{\sigma}\right), reading Φ\Phi from a standard-normal table or software, and reproduce Example 13.14 (X∼N(−3,4)X\sim N(-3,4): P(X≤−1.7)=Φ(0.65)≈0.742\mathbb{P}(X\le-1.7)=\Phi(0.65)\approx0.742).
  • State and apply the 68–95–99.7 rule — P(∣X−μ∣≤σ)≈0.683\mathbb{P}(|X-\mu|\le\sigma)\approx0.683, P(∣X−μ∣≤2σ)≈0.954\mathbb{P}(|X-\mu|\le2\sigma)\approx0.954, P(∣X−μ∣≤3σ)≈0.997\mathbb{P}(|X-\mu|\le3\sigma)\approx0.997 — as the standardized values 2Φ(1)−12\Phi(1)-1, 2Φ(2)−12\Phi(2)-1, 2Φ(3)−12\Phi(3)-1 (Example 13.15).

In your course

· MATH2015 · Linear Algebra & Probability
§13.3 More on Common Probability Distributions
  • Proposition 13.6Standard normal: mean and variance
    If Z∼N(0,1)Z\sim N(0,1) then E[Z]=0\mathbb{E}[Z]=0 and Var(Z)=E[Z2]=1\mathrm{Var}(Z)=\mathbb{E}[Z^2]=1.
  • Definition 13.5Normal distribution
    For μ∈R\mu\in\mathbb{R}, σ>0\sigma>0, X∼N(μ,σ2)X\sim N(\mu,\sigma^2) has density f(x)=1σ2πexp⁡ ⁣(−(x−μ)22σ2)f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), with mean μ\mu and variance σ2\sigma^2.
  • Proposition 13.7Linear transformations of a normal
    If X∼N(μ,σ2)X\sim N(\mu,\sigma^2) and a≠0a\neq0, then aX+b∼N(aμ+b, a2σ2)aX+b\sim N(a\mu+b,\,a^2\sigma^2); in particular Z=X−μσ∼N(0,1)Z=\frac{X-\mu}{\sigma}\sim N(0,1).
  • Example 13.14A normal probability by standardization
    For X∼N(−3,4)X\sim N(-3,4) (σ=2\sigma=2), P(X≤−1.7)=Φ ⁣(−1.7+32)=Φ(0.65)≈0.742\mathbb{P}(X\le-1.7)=\Phi\!\big(\frac{-1.7+3}{2}\big)=\Phi(0.65)\approx0.742.
  • Example 13.15Two standard deviations from the mean
    For X∼N(μ,σ2)X\sim N(\mu,\sigma^2), P(∣X−μ∣>2σ)=2(1−Φ(2))≈0.046\mathbb{P}(|X-\mu|>2\sigma)=2(1-\Phi(2))\approx0.046, so P(∣X−μ∣≤2σ)≈0.954\mathbb{P}(|X-\mu|\le2\sigma)\approx0.954.
The standard-normal c.d.f. Φ\Phi has no elementary closed form, so Φ\Phi-values are obtained from a standard-normal table or software; Φ(−z)=1−Φ(z)\Phi(-z)=1-\Phi(z) by symmetry. The 6868–9595–99.799.7 rule is the mnemonic for 2Φ(k)−12\Phi(k)-1 at k=1,2,3k=1,2,3 (≈0.683, 0.954, 0.997\approx0.683,\,0.954,\,0.997). Numeric answers here were recomputed (not taken from OCR).
1

The standard normal $N(0,1)$ and its c.d.f. $\Phi$

The standard normal distribution Z∼N(0,1)Z\sim N(0,1) is the bell curve every other normal is measured against. Its density is φ(z)=12πe−z2/2\varphi(z)=\frac{1}{\sqrt{2\pi}}e^{-z^2/2} — symmetric about 00, peaked at the origin, and decaying rapidly in both tails. Proposition 13.6 records its two defining moments: E[Z]=0,Var(Z)=E[Z2]=1.\mathbb{E}[Z]=0,\qquad \mathrm{Var}(Z)=\mathbb{E}[Z^2]=1. The mean is 00 because zφ(z)z\varphi(z) is an odd integrand on a symmetric domain, so its integral vanishes; the variance is 11 because ∫−∞∞z2φ(z) dz=1\int_{-\infty}^{\infty}z^2\varphi(z)\,dz=1 (a standard integration by parts), and Var(Z)=E[Z2]−(E[Z])2=1−0=1\mathrm{Var}(Z)=\mathbb{E}[Z^2]-(\mathbb{E}[Z])^2=1-0=1. There is no elementary antiderivative for φ\varphi, so probabilities are not computed by hand but read from the standard-normal c.d.f. Φ(z)=P(Z≤z)=∫−∞z12πe−t2/2 dt,\Phi(z)=\mathbb{P}(Z\le z)=\int_{-\infty}^{z}\frac{1}{\sqrt{2\pi}}e^{-t^2/2}\,dt, tabulated or built into software. Symmetry gives the frequently used identity Φ(−z)=1−Φ(z)\Phi(-z)=1-\Phi(z), so a table of positive zz-values suffices for all of them.

2

The general normal $N(\mu,\sigma^2)$ as a stretch-and-shift (Definition 13.5)

The whole normal family is produced from ZZ by affine transformations. Fix a mean μ∈R\mu\in\mathbb{R} and a standard deviation σ>0\sigma>0 and set X=σZ+μ.X=\sigma Z+\mu. Scaling by σ\sigma widens or narrows the bell and shifting by μ\mu slides its centre, giving E[X]=σ E[Z]+μ=μ\mathbb{E}[X]=\sigma\,\mathbb{E}[Z]+\mu=\mu and Var(X)=σ2 Var(Z)=σ2\mathrm{Var}(X)=\sigma^2\,\mathrm{Var}(Z)=\sigma^2. Transforming the density accordingly, Definition 13.5 says XX has the normal distribution with mean μ\mu and variance σ2\sigma^2, written X∼N(μ,σ2)X\sim N(\mu,\sigma^2), if f(x)=1σ2πexp⁡ ⁣(−(x−μ)22σ2).f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right). Read the notation carefully: the first slot is the mean μ\mu and the second slot is the variance σ2\sigma^2 — not the standard deviation. So N(−3,4)N(-3,4) has mean −3-3, variance 44, and standard deviation σ=4=2\sigma=\sqrt{4}=2. The graph is a symmetric bell centred at μ\mu; a larger σ\sigma makes it shorter and wider, a smaller σ\sigma taller and narrower (compare N(0,1)N(0,1) with N(2,2.25)N(2,2.25), where σ=1.5\sigma=1.5).

3

Linear transformations and standardization (Proposition 13.7)

Normality is preserved by any affine map. Proposition 13.7: if X∼N(μ,σ2)X\sim N(\mu,\sigma^2) and a≠0a\neq0, b∈Rb\in\mathbb{R}, then aX+b∼N ⁣(aμ+b,  a2σ2).aX+b\sim N\!\big(a\mu+b,\;a^2\sigma^2\big). The new mean is aμ+ba\mu+b (apply E[aX+b]=aE[X]+b\mathbb{E}[aX+b]=a\mathbb{E}[X]+b) and the new variance is a2σ2a^2\sigma^2 (apply Var(aX+b)=a2Var(X)\mathrm{Var}(aX+b)=a^2\mathrm{Var}(X)); crucially, the shape stays a bell, so the result is again exactly normal. The most important special case runs the construction X=σZ+μX=\sigma Z+\mu in reverse. Taking a=1σa=\tfrac1\sigma and b=−μσb=-\tfrac{\mu}{\sigma} gives the standardized variable Z=X−μσ∼N(0,1),Z=\frac{X-\mu}{\sigma}\sim N(0,1), which re-expresses the value xx as a zz-score — how many standard deviations xx lies above (or below) the mean. Subtracting μ\mu recentres to 00; dividing by σ\sigma rescales the variance to 11. Standardization is the bridge from any normal to the single tabulated distribution N(0,1)N(0,1).

4

Computing probabilities, and the 68–95–99.7 rule (Examples 13.14–13.15)

Standardization turns every normal probability into a lookup in Φ\Phi. Since X≤xX\le x exactly when Z=X−μσ≤x−μσZ=\frac{X-\mu}{\sigma}\le\frac{x-\mu}{\sigma}, P(X≤x)=Φ ⁣(x−μσ),\mathbb{P}(X\le x)=\Phi\!\left(\frac{x-\mu}{\sigma}\right), and a two-sided probability is a difference, P(a≤X≤b)=Φ ⁣(b−μσ)−Φ ⁣(a−μσ)\mathbb{P}(a\le X\le b)=\Phi\!\big(\tfrac{b-\mu}{\sigma}\big)-\Phi\!\big(\tfrac{a-\mu}{\sigma}\big). For X∼N(−3,4)X\sim N(-3,4) (so σ=2\sigma=2), Example 13.14 gives P(X≤−1.7)=Φ ⁣(−1.7+32)=Φ(0.65)≈0.742\mathbb{P}(X\le-1.7)=\Phi\!\big(\tfrac{-1.7+3}{2}\big)=\Phi(0.65)\approx0.742. Applying this to the symmetric band ∣X−μ∣≤kσ|X-\mu|\le k\sigma gives the 68–95–99.7 (empirical) rule: because ∣X−μ∣≤kσ|X-\mu|\le k\sigma is the same event as ∣Z∣≤k|Z|\le k, P(∣X−μ∣≤kσ)=2Φ(k)−1,\mathbb{P}(|X-\mu|\le k\sigma)=2\Phi(k)-1, which equals about 0.6830.683 for k=1k=1, 0.9540.954 for k=2k=2, and 0.9970.997 for k=3k=3. Equivalently, the chance of straying more than 2σ2\sigma from the mean is 2(1−Φ(2))≈0.0462\big(1-\Phi(2)\big)\approx0.046 (Example 13.15) — so a normal variable sits within two standard deviations of its mean over 95%95\% of the time.

Proposition 13.6 — Standard normal: mean and variance

Let Z∼N(0,1)Z\sim N(0,1), the standard normal with density φ(z)=12πe−z2/2\varphi(z)=\frac{1}{\sqrt{2\pi}}e^{-z^2/2}. Then E[Z]=0andVar(Z)=E[Z2]=1.\mathbb{E}[Z]=0\qquad\text{and}\qquad \mathrm{Var}(Z)=\mathbb{E}[Z^2]=1.

Intuition. The mean is 00 by symmetry: zφ(z)z\varphi(z) is an odd function integrated over a symmetric domain, so the integral cancels to 00. The variance is 11 because E[Z2]=∫z2φ(z) dz=1\mathbb{E}[Z^2]=\int z^2\varphi(z)\,dz=1 (integration by parts) and Var(Z)=E[Z2]−(E[Z])2=1−0\mathrm{Var}(Z)=\mathbb{E}[Z^2]-(\mathbb{E}[Z])^2=1-0. These two values are what the name 'standard' refers to — a normal already centred at 00 and scaled to unit spread, the yardstick for all the others.
Definition 13.5 — Normal distribution

Let μ∈R\mu\in\mathbb{R} and σ>0\sigma>0. A random variable XX has the normal distribution with mean μ\mu and variance σ2\sigma^2 if it has density f(x)=1σ2πexp⁡ ⁣(−(x−μ)22σ2),x∈R.f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right),\qquad x\in\mathbb{R}. We write X∼N(μ,σ2)X\sim N(\mu,\sigma^2). Equivalently, X=σZ+μX=\sigma Z+\mu for a standard normal ZZ, which gives E[X]=μ\mathbb{E}[X]=\mu and Var(X)=σ2\mathrm{Var}(X)=\sigma^2.

Intuition. The density is a bell centred at μ\mu whose width is governed by σ\sigma: the factor e−(x−μ)2/(2σ2)e^{-(x-\mu)^2/(2\sigma^2)} falls off as xx moves away from μ\mu, and the constant 1σ2π\frac{1}{\sigma\sqrt{2\pi}} normalises the total area to 11. Mind the convention — the second argument is the variance σ2\sigma^2, not σ\sigma — so for N(−3,4)N(-3,4) the standard deviation is 4=2\sqrt{4}=2.
Proposition 13.7 — Linear transformations of a normal

Let X∼N(μ,σ2)X\sim N(\mu,\sigma^2) and let Y=aX+bY=aX+b with a≠0a\neq0 and b∈Rb\in\mathbb{R}. Then Y∼N ⁣(aμ+b,  a2σ2).Y\sim N\!\big(a\mu+b,\;a^2\sigma^2\big). In particular, the standardized variable Z=X−μσZ=\frac{X-\mu}{\sigma} has the standard normal distribution N(0,1)N(0,1).

Intuition. An affine map rescales and shifts a bell curve into another bell curve, so normality is preserved; only the parameters move, following the general rules E[aX+b]=aμ+b\mathbb{E}[aX+b]=a\mu+b and Var(aX+b)=a2σ2\mathrm{Var}(aX+b)=a^2\sigma^2. Choosing a=1σ, b=−μσa=\tfrac1\sigma,\ b=-\tfrac{\mu}{\sigma} recentres to mean 00 and rescales to variance 11, which is exactly why standardization lets any normal probability be read from the single table of Φ\Phi.

Worked examples

Example 1

Example 13.14. Suppose X∼N(−3,4)X\sim N(-3,4). Find P(X≤−1.7)\mathbb{P}(X\le-1.7), using the standard normal c.d.f. Φ\Phi (give your answer to 33 decimals).

  1. 1

    Read off the parameters. In N(−3,4)N(-3,4) the mean is μ=−3\mu=-3 and the variance is σ2=4\sigma^2=4, so the standard deviation is σ=4=2\sigma=\sqrt{4}=2.

  2. 2

    Standardize (Proposition 13.7). The variable Z=X−μσ=X+32Z=\dfrac{X-\mu}{\sigma}=\dfrac{X+3}{2} is standard normal, Z∼N(0,1)Z\sim N(0,1).

  3. 3

    Convert the event to a zz-score. X≤−1.7X\le-1.7 is the same event as Z≤−1.7−(−3)2=1.32=0.65Z\le\dfrac{-1.7-(-3)}{2}=\dfrac{1.3}{2}=0.65, so P(X≤−1.7)=P(Z≤0.65)=Φ(0.65)\mathbb{P}(X\le-1.7)=\mathbb{P}(Z\le0.65)=\Phi(0.65).

  4. 4

    Evaluate Φ\Phi. From a standard-normal table (or software), Φ(0.65)≈0.7422\Phi(0.65)\approx0.7422, i.e. 0.7420.742 to three decimals.

Answer. P(X≤−1.7)=Φ ⁣(−1.7+32)=Φ(0.65)≈0.742\mathbb{P}(X\le-1.7)=\Phi\!\big(\tfrac{-1.7+3}{2}\big)=\Phi(0.65)\approx0.742.
Example 2

Example 13.15 — Two standard deviations from the mean. Let X∼N(μ,σ2)X\sim N(\mu,\sigma^2). What is the probability that XX deviates from its mean by more than 2σ2\sigma, i.e. P(∣X−μ∣>2σ)\mathbb{P}(|X-\mu|>2\sigma)? What does this say about staying within 2σ2\sigma?

Example 3

Linear transformation then a probability (Proposition 13.7). Let X∼N(2,2.25)X\sim N(2,2.25) (so σ=1.5\sigma=1.5) and define Y=2X+1Y=2X+1. (a) Find the distribution of YY. (b) Compute P(Y≤6)\mathbb{P}(Y\le6) to 33 decimals.