Skip to content
VidiMaster it, module by module
Module 7/Random Variables & Distributions

The Cumulative Distribution Function: F(x)=P{X≤x}F(x)=\mathbb{P}\{X\le x\}

Unlike the probability mass function (discrete variables only) or the density (continuous variables only), the cumulative distribution function (c.d.f.) is defined for every random variable (Definition 12.6): F(x)=P{X≤x}F(x)=\mathbb{P}\{X\le x\} records the total probability accumulated up to the point xx — the 'sum-or-area-so-far' function. The '≤\le' is deliberate (Remark 12.2): it includes the endpoint and makes the c.d.f. deliver interval probabilities, P{a<X≤b}=F(b)−F(a)\mathbb{P}\{a<X\le b\}=F(b)-F(a). For a discrete variable the c.d.f. is a step function (Remark 12.3) that is flat between the possible values and jumps by P{X=x}\mathbb{P}\{X=x\} at each one, as for the change-in-wealth variable of Example 12.7; for a continuous variable it rises smoothly, obtained by integrating the density, as for the Unif[1,3]\mathrm{Unif}[1,3] variable of Example 12.8. Whatever its type, every c.d.f. shares the same three shape properties (Proposition 12.2): it is nondecreasing, right-continuous, and climbs from 00 at −∞-\infty to 11 at +∞+\infty.

Before you start — give these a try

Attempting first primes your brain for the lesson — even if you miss. Nothing is graded or saved; it's just a warm-up.

For the same variable XX (p(0)=0.1, p(1)=0.3, p(2)=0.4, p(3)=0.2p(0)=0.1,\ p(1)=0.3,\ p(2)=0.4,\ p(3)=0.2), use the interval formula P{a<X≤b}=F(b)−F(a)\mathbb{P}\{a<X\le b\}=F(b)-F(a) to find P{1<X≤3}\mathbb{P}\{1<X\le3\}. Give a decimal to 33 places.

Let Y∼Unif[0,4]Y\sim\mathrm{Unif}[0,4], whose c.d.f. is F(y)=y−04−0=y4F(y)=\dfrac{y-0}{4-0}=\dfrac{y}{4} for 0≤y≤40\le y\le4 (and 00 below 00, 11 above 44). Find F(3)=P{Y≤3}F(3)=\mathbb{P}\{Y\le3\}. Give a decimal to 33 places.

What you’ll be able to do

  • State Definition 12.6: the c.d.f. of a random variable XX is the function F(x)=P{X≤x}F(x)=\mathbb{P}\{X\le x\}, defined for all x∈Rx\in\mathbb{R}, and explain why it applies to any random variable — discrete, continuous, or mixed — unlike the p.m.f. or density.
  • Use the interval formula P{a<X≤b}=F(b)−F(a)\mathbb{P}\{a<X\le b\}=F(b)-F(a) (Remark 12.2) and explain why the '≤\le' in the definition makes the relevant interval the half-open (a,b](a,b].
  • Build the step-function c.d.f. of a discrete variable from its p.m.f. by cumulative summation F(x)=∑k≤xp(k)F(x)=\sum_{k\le x}p(k), reading off the upward jump of size P{X=x}\mathbb{P}\{X=x\} at each possible value (Remark 12.3, Example 12.7).
  • Build the c.d.f. of a continuous variable by integrating its density, F(x)=∫−∞xf(t) dtF(x)=\int_{-\infty}^{x} f(t)\,dt, and in particular derive the piecewise-linear c.d.f. F(x)=x−ab−aF(x)=\dfrac{x-a}{b-a} of Unif[a,b]\mathrm{Unif}[a,b] (Example 12.8).
  • State the defining properties of a c.d.f. (Proposition 12.2) — nondecreasing, right-continuous, with lim⁡x→−∞F(x)=0\lim_{x\to-\infty}F(x)=0 and lim⁡x→+∞F(x)=1\lim_{x\to+\infty}F(x)=1 — and use them to decide whether a given function can be a c.d.f.

In your course

· MATH2015 · Linear Algebra & Probability
§12.2 Cumulative Distribution Function
  • Definition 12.6Cumulative distribution function
    F(x)=P{X≤x}F(x)=\mathbb{P}\{X\le x\} for all x∈Rx\in\mathbb{R}; defined for every random variable.
  • Remark 12.2The '≤\le' inequality and interval probabilities
    P{a<X≤b}=F(b)−F(a)\mathbb{P}\{a<X\le b\}=F(b)-F(a), the probability of (a,b](a,b].
  • Remark 12.3The discrete c.d.f. is a step function
    F(x)=∑k≤xp(k)F(x)=\sum_{k\le x}p(k); jumps by P{X=x}\mathbb{P}\{X=x\} at each value; right-continuous.
  • Proposition 12.2Properties of the c.d.f.
    Nondecreasing with 0≤F≤10\le F\le1; lim⁡x→−∞F=0\lim_{x\to-\infty}F=0, lim⁡x→+∞F=1\lim_{x\to+\infty}F=1; right-continuous.
  • Example 12.7c.d.f. of the change in wealth WW (a step function)
  • Example 12.8c.d.f. of X∼Unif[1,3]X\sim\mathrm{Unif}[1,3] (a linear ramp)
The course notes write the c.d.f. argument as F(s)F(s); here we use F(x)F(x), and probabilities are written P\mathbb{P}. The continuous-variable formula F(x)=∫−∞xf(t) dtF(x)=\int_{-\infty}^{x}f(t)\,dt underlies Example 12.8.
1

The cumulative distribution function $F(x)=\mathbb{P}\{X\le x\}$

A probability mass function describes only discrete variables and a density only continuous ones, but the cumulative distribution function (c.d.f.) describes them all. For a random variable XX it is the function F(x)=P{X≤x},x∈R(Definition 12.6),F(x)=\mathbb{P}\{X\le x\},\qquad x\in\mathbb{R}\quad(\textbf{Definition 12.6}), the probability that XX lands at or below the level xx. Think of F(x)F(x) as the accumulated probability — the 'sum-or-area-so-far' function — sweeping a threshold from −∞-\infty rightward and recording how much probability has been collected by the time it reaches xx. Because F(x)F(x) is a probability it always lies in [0,1][0,1], and because sweeping right can only add probability, FF can never decrease. This single object works for a die (discrete), a waiting time (continuous), or a payout that is partly a lump and partly spread out (mixed) — which is exactly why the c.d.f., rather than the p.m.f. or density, is the universal description of a random variable's distribution.

2

Why the '$\le$' matters: interval probabilities $F(b)-F(a)$

The inequality in F(x)=P{X≤x}F(x)=\mathbb{P}\{X\le x\} is less-than-or-equal, so F(x)F(x) is the probability of the half-line (−∞,x](-\infty,x] including the endpoint xx (Remark 12.2). This is what lets the c.d.f. measure any interval: for a<ba<b, P{a<X≤b}=P{X≤b}−P{X≤a}=F(b)−F(a).\mathbb{P}\{a<X\le b\}=\mathbb{P}\{X\le b\}-\mathbb{P}\{X\le a\}=F(b)-F(a). Subtracting F(a)=P{X≤a}F(a)=\mathbb{P}\{X\le a\} strips off everything up to and including aa, leaving the half-open interval (a,b](a,b] — open on the left, closed on the right. For a continuous variable a single point has probability 00, so swapping << for ≤\le at an endpoint changes nothing. For a discrete variable it matters a great deal: an endpoint can carry positive mass, so P{X≤b}\mathbb{P}\{X\le b\} and P{X<b}\mathbb{P}\{X<b\} differ by exactly P{X=b}\mathbb{P}\{X=b\}. Keeping track of which endpoints are included is therefore the whole game when computing discrete interval probabilities from FF.

3

The c.d.f. of a discrete variable is a step function

If XX is discrete with p.m.f. p(k)=P{X=k}p(k)=\mathbb{P}\{X=k\}, then the definition becomes a sum, F(x)=P{X≤x}=∑k≤xp(k),F(x)=\mathbb{P}\{X\le x\}=\sum_{k\le x}p(k), taken over the possible values kk that are ≤x\le x (Remark 12.3). As the threshold xx moves right it picks up a new term only when it crosses a possible value, so FF is a step function: perfectly flat between consecutive values and jumping upward by exactly p(x)=P{X=x}p(x)=\mathbb{P}\{X=x\} at each possible value xx. The jumps add up to the total mass 11, so FF climbs in a staircase from 00 to 11. A discrete c.d.f. is therefore not continuous — it is only right-continuous, lim⁡s→x+F(s)=F(x)\lim_{s\to x^{+}}F(s)=F(x): at a value xx the function already sits at the post-jump (upper) level, because the '≤\le' includes xx itself. Reading the picture in reverse recovers the p.m.f.: the mass at a point is the height of the jump there, p(x)=F(x)−lim⁡s→x−F(s)p(x)=F(x)-\lim_{s\to x^{-}}F(s). Example 12.7 (the change in wealth WW) is the model case.

4

The continuous case, and the properties every c.d.f. shares (Proposition 12.2)

For a continuous variable with density ff, the sum becomes an integral, F(x)=P{X≤x}=∫−∞xf(t) dt,F(x)=\mathbb{P}\{X\le x\}=\int_{-\infty}^{x} f(t)\,dt, the area under the density up to xx; here FF rises smoothly instead of in jumps. For X∼Unif[a,b]X\sim\mathrm{Unif}[a,b] the density is the constant 1b−a\tfrac{1}{b-a} on [a,b][a,b], and integrating gives the straight ramp F(x)=x−ab−aF(x)=\dfrac{x-a}{b-a} for a≤x≤ba\le x\le b (with F=0F=0 below aa and F=1F=1 above bb), as in Example 12.8. Discrete or continuous, every c.d.f. obeys the same three laws (Proposition 12.2): it is nondecreasing with 0≤F(x)≤10\le F(x)\le1; it satisfies the limits lim⁡x→−∞F(x)=0\lim_{x\to-\infty}F(x)=0 (no probability has accumulated yet) and lim⁡x→+∞F(x)=1\lim_{x\to+\infty}F(x)=1 (all of it eventually does); and it is right-continuous. These properties are also a checklist: a function that decreases somewhere, escapes [0,1][0,1], or fails the end-limits cannot be a c.d.f.

Definition 12.6 — Cumulative distribution function

The cumulative distribution function of a random variable XX is the function F:R→[0,1]F:\mathbb{R}\to[0,1] given by F(x)=P{X≤x}for all x∈R.F(x)=\mathbb{P}\{X\le x\}\qquad\text{for all } x\in\mathbb{R}. It is defined for every random variable, whether discrete, continuous, or mixed.

Intuition. F(x)F(x) is the probability accumulated up to the threshold xx — the 'sum-or-area-so-far' function. Unlike the p.m.f. (discrete only) or the density (continuous only), this one object exists for any distribution, which is why it is the universal way to describe a random variable.
Remark 12.2 — The '$\le$' inequality and interval probabilities

Because the definition uses '≤\le', F(x)=P{X≤x}F(x)=\mathbb{P}\{X\le x\} is the probability of the closed half-line (−∞,x](-\infty,x] (the endpoint xx is included). Consequently, for any a<ba<b, P{a<X≤b}=P{X≤b}−P{X≤a}=F(b)−F(a),\mathbb{P}\{a<X\le b\}=\mathbb{P}\{X\le b\}-\mathbb{P}\{X\le a\}=F(b)-F(a), the probability of the half-open interval (a,b](a,b].

Intuition. Subtracting F(a)F(a) removes all probability up to and including aa, so the left end is excluded and the right end included — hence (a,b](a,b]. For continuous variables ≤\le and << agree (points have probability 00); for discrete variables they can differ by the mass P{X=b}\mathbb{P}\{X=b\} sitting at an endpoint.
Remark 12.3 — The c.d.f. of a discrete variable is a step function

If XX is discrete with p.m.f. p(k)=P{X=k}p(k)=\mathbb{P}\{X=k\}, then F(x)=∑k≤xp(k).F(x)=\sum_{k\le x}p(k). This FF is a step function: constant between consecutive possible values and increasing by p(x)=P{X=x}p(x)=\mathbb{P}\{X=x\} at each possible value xx. It is right-continuous, lim⁡s→x+F(s)=F(x)\lim_{s\to x^{+}}F(s)=F(x).

Intuition. Sliding the threshold right adds a new term only when it passes a possible value, producing a staircase whose riser heights are the masses. Right-continuity reflects the '≤\le': at a jump point xx, the value F(x)F(x) is already the upper (post-jump) level. Reversing this reads the mass off the graph: p(x)=F(x)−lim⁡s→x−F(s)p(x)=F(x)-\lim_{s\to x^{-}}F(s).
Proposition 12.2 — Properties of the c.d.f.

Every cumulative distribution function FF satisfies: (1) 0≤F(x)≤10\le F(x)\le1, and FF is nondecreasing (a≤b⇒F(a)≤F(b)a\le b\Rightarrow F(a)\le F(b)); (2) the limits lim⁡x→−∞F(x)=0\lim_{x\to-\infty}F(x)=0 and lim⁡x→+∞F(x)=1\lim_{x\to+\infty}F(x)=1; (3) FF is right-continuous, lim⁡s→x+F(s)=F(x)\lim_{s\to x^{+}}F(s)=F(x). Conversely, any function with these three properties is the c.d.f. of some random variable.

Intuition. FF is a probability so it lives in [0,1][0,1]; sweeping the threshold right only accumulates more, so it cannot decrease; by the time the threshold passes −∞-\infty nothing is collected and past +∞+\infty everything is, forcing the limits 00 and 11. These three properties are both the signature of a c.d.f. and a quick test: fail any one and the function cannot be a c.d.f.

Worked examples

Example 1

Example 12.7 (change in wealth). A random variable WW (a change in wealth) takes the values −1, 1, 3-1,\,1,\,3 with probability mass function pW(−1)=12, pW(1)=16, pW(3)=13p_W(-1)=\tfrac12,\ p_W(1)=\tfrac16,\ p_W(3)=\tfrac13. Find the cumulative distribution function F(x)=P{W≤x}F(x)=\mathbb{P}\{W\le x\} and describe its graph.

  1. 1

    Check it is a p.m.f. The masses are nonnegative and sum to 11: 12+16+13=36+16+26=66=1.\tfrac12+\tfrac16+\tfrac13=\tfrac36+\tfrac16+\tfrac26=\tfrac66=1.

  2. 2

    Below the smallest value. For x<−1x<-1 no possible value is ≤x\le x, so F(x)=∑k≤xpW(k)=0.F(x)=\sum_{k\le x}p_W(k)=0.

  3. 3

    Accumulate across each value. For −1≤x<1-1\le x<1 only −1-1 qualifies: F(x)=pW(−1)=12F(x)=p_W(-1)=\tfrac12. For 1≤x<31\le x<3 both −1-1 and 11 qualify: F(x)=12+16=23F(x)=\tfrac12+\tfrac16=\tfrac23. For x≥3x\ge3 all three qualify: F(x)=12+16+13=1.F(x)=\tfrac12+\tfrac16+\tfrac13=1.

  4. 4

    Read the graph. FF is a step function with upward jumps of 12\tfrac12 at x=−1x=-1, 16\tfrac16 at x=1x=1, and 13\tfrac13 at x=3x=3 — each jump equals P{W=x}\mathbb{P}\{W=x\}. It is flat in between and right-continuous (the value at a jump is the upper level).

Answer. F(x)=0F(x)=0 for x<−1x<-1;  F(x)=12\ F(x)=\tfrac12 for −1≤x<1-1\le x<1;  F(x)=23\ F(x)=\tfrac23 for 1≤x<31\le x<3;  F(x)=1\ F(x)=1 for x≥3x\ge3. A staircase whose jump at each possible value is the p.m.f. there.
Example 2

Example 12.8 (uniform). Let X∼Unif[1,3]X\sim\mathrm{Unif}[1,3], so its density is f(x)=12f(x)=\tfrac12 for 1≤x≤31\le x\le3 and 00 otherwise. Find the c.d.f. F(x)=P{X≤x}=∫−∞xf(t) dtF(x)=\mathbb{P}\{X\le x\}=\int_{-\infty}^{x} f(t)\,dt.

Example 3

Example (an interval probability from a step c.d.f.). A discrete random variable XX takes the values 0,1,2,30,1,2,3 with p.m.f. p(0)=0.1, p(1)=0.3, p(2)=0.4, p(3)=0.2p(0)=0.1,\ p(1)=0.3,\ p(2)=0.4,\ p(3)=0.2. Compute F(0)F(0) and F(2)F(2), then use the interval formula to find P{0<X≤2}\mathbb{P}\{0<X\le2\}.