Skip to content
VidiMaster it, module by module
Module 7/Common Distributions

The Bernoulli and Binomial Distributions

Repeated independent trials with just two outcomes — success (probability pp) and failure (probability 1−p1-p) — are the seed from which two fundamental discrete distributions grow. The Bernoulli distribution (Definition 12.7) records a single such trial: a random variable X∈{0,1}X\in\{0,1\} with P(X=1)=p\mathbb{P}(X=1)=p and P(X=0)=1−p\mathbb{P}(X=0)=1-p, written X∼Ber(p)X\sim\mathrm{Ber}(p). Run the trial nn independent times and you obtain nn independent Bernoulli variables X1,…,XnX_1,\dots,X_n; their sum Sn=X1+⋯+XnS_n=X_1+\cdots+X_n counts the successes, and its distribution is the Binomial (Definition 12.8): P(X=k)=(nk)pk(1−p)n−k\mathbb{P}(X=k)=\binom{n}{k}p^k(1-p)^{n-k} for k=0,1,…,nk=0,1,\dots,n, written X∼Bin(n,p)X\sim\mathrm{Bin}(n,p). The binomial coefficient (nk)\binom{n}{k} counts the arrangements of kk successes among the nn trials, each arrangement carrying probability pk(1−p)n−kp^k(1-p)^{n-k}, and the binomial theorem guarantees these probabilities sum to 11. We close with Example 12.9: the chance that five rolls of a fair die yield two or three sixes, with S5∼Bin(5,16)S_5\sim\mathrm{Bin}(5,\tfrac16), equal to 15007776≈0.193\tfrac{1500}{7776}\approx 0.193.

Before you start — give these a try

Attempting first primes your brain for the lesson — even if you miss. Nothing is graded or saved; it's just a warm-up.

Let X∼Bin(6,0.5)X\sim\mathrm{Bin}(6,0.5) (for example, the number of heads in six tosses of a fair coin). Find P(X=4)\mathbb{P}(X=4). Give your answer as a decimal to 33 places.

A spinner lands on red with probability 0.20.2 on each independent spin. In n=10n=10 spins, let X∼Bin(10,0.2)X\sim\mathrm{Bin}(10,0.2) be the number of reds. Find P(X≥1)\mathbb{P}(X\ge 1), the probability of at least one red. Give your answer as a decimal to 33 places.

What you’ll be able to do

  • State Definition 12.7: a random variable XX is Bernoulli with parameter pp (where 0≤p≤10\le p\le 1) if X∈{0,1}X\in\{0,1\} with P(X=1)=p\mathbb{P}(X=1)=p and P(X=0)=1−p\mathbb{P}(X=0)=1-p, written X∼Ber(p)X\sim\mathrm{Ber}(p), and recognise a single success/failure trial as the setting it models.
  • Explain how nn independent repeated trials give rise to nn independent Bernoulli variables X1,…,XnX_1,\dots,X_n, and compute the probability p j(1−p) n−jp^{\,j}(1-p)^{\,n-j} of any specific outcome string having exactly jj successes.
  • State Definition 12.8: the Binomial distribution X∼Bin(n,p)X\sim\mathrm{Bin}(n,p) has P(X=k)=(nk)pk(1−p)n−k\mathbb{P}(X=k)=\binom{n}{k}p^k(1-p)^{n-k} for k=0,1,…,nk=0,1,\dots,n, and derive it by counting the (nk)\binom{n}{k} arrangements of kk successes among nn trials.
  • Compute binomial probabilities P(X=k)\mathbb{P}(X=k), cumulative probabilities P(X≤k)\mathbb{P}(X\le k), and tail probabilities P(X≥k)\mathbb{P}(X\ge k) (often via the complement 1−P(X=0)1-\mathbb{P}(X=0)), and reproduce Example 12.9 (two or three sixes in five rolls of a fair die, ≈0.193\approx 0.193).
  • Decide whether a described experiment is a binomial setup — a fixed number nn of trials, independent of one another, each with the same success probability pp and only two outcomes — and identify the parameters nn and pp.

In your course

· MATH2015 · Linear Algebra & Probability
§12.3 Common Probability Distributions
  • Definition 12.7Bernoulli distribution
    For 0≤p≤10\le p\le 1, X∼Ber(p)X\sim\mathrm{Ber}(p) takes values in {0,1}\{0,1\} with P(X=1)=p\mathbb{P}(X=1)=p and P(X=0)=1−p\mathbb{P}(X=0)=1-p.
  • Definition 12.8Binomial distribution
    For a positive integer nn and 0≤p≤10\le p\le 1, X∼Bin(n,p)X\sim\mathrm{Bin}(n,p) has P(X=k)=(nk)pk(1−p)n−k\mathbb{P}(X=k)=\binom{n}{k}p^k(1-p)^{n-k} for k=0,1,…,nk=0,1,\dots,n.
  • Example 12.9Two or three sixes in five rolls of a fair die
    For S5∼Bin(5,16)S_5\sim\mathrm{Bin}(5,\tfrac16), P(S5∈{2,3})=P(S5=2)+P(S5=3)=15007776≈0.193\mathbb{P}(S_5\in\{2,3\})=\mathbb{P}(S_5=2)+\mathbb{P}(S_5=3)=\dfrac{1500}{7776}\approx 0.193.
Section 12.3 continues with the geometric, negative binomial, uniform and normal distributions (Definitions 12.9–12.12), which are treated in separate lessons.
1

Repeated trials with binary outcomes

The simplest experiments involving independence are repeated trials, each with just two outcomes: success (probability pp) and failure (probability 1−p1-p). We encode success as 11 and failure as 00. Repeating the trial nn times, an outcome is a binary string ω=(s1,…,sn)\omega=(s_1,\dots,s_n) with each si∈{0,1}s_i\in\{0,1\}, so the sample space is Ω={(s1,…,sn):si∈{0,1}}\Omega=\{(s_1,\dots,s_n):s_i\in\{0,1\}\}, of size 2n2^n. Crucially we assume the trials are independent, which makes the probability of a whole string multiply across the individual trials: an outcome with some 11's and the rest 00's has probability P{ω}=p (# of 1’s) (1−p) (# of 0’s).\mathbb{P}\{\omega\}=p^{\,(\#\,\text{of }1\text{'s})}\,(1-p)^{\,(\#\,\text{of }0\text{'s})}. For instance, the specific string (0,1,0,0,1,0)(0,1,0,0,1,0) over six trials — two successes and four failures — has probability p2(1−p)4p^2(1-p)^4. Notice this depends only on how many successes occurred, not on where they fell: every string with the same number of 11's is equally likely. That single fact is the engine behind everything that follows.

2

The Bernoulli distribution (Definition 12.7)

A single trial with two outcomes is modelled by the Bernoulli distribution. For a parameter 0≤p≤10\le p\le 1, a random variable XX is Bernoulli with success probability pp if it takes only the values X∈{0,1}X\in\{0,1\} with P(X=1)=p,P(X=0)=1−p,\mathbb{P}(X=1)=p,\qquad \mathbb{P}(X=0)=1-p, and we write X∼Ber(p)X\sim\mathrm{Ber}(p). Here XX is the indicator of success: X=1X=1 when the trial succeeds and X=0X=0 when it fails. A fair coin's heads/tails is Ber(12)\mathrm{Ber}(\tfrac12); rolling a six on a fair die is Ber(16)\mathrm{Ber}(\tfrac16). The Bernoulli distribution is the atom of the theory: a sequence of nn independent trials gives rise to nn independent Bernoulli random variables X1,…,XnX_1,\dots,X_n, all sharing the same parameter pp. By independence, a joint outcome such as P(X1=0,X2=1,X3=0,X4=0,X5=1,X6=0)=p2(1−p)4\mathbb{P}(X_1=0,X_2=1,X_3=0,X_4=0,X_5=1,X_6=0)=p^2(1-p)^4 — exactly the product structure of the previous concept.

3

Counting successes: from Bernoulli to Binomial

Given nn independent Bernoulli trials X1,…,XnX_1,\dots,X_n, the natural question is how many succeeded. Define Sn=X1+X2+⋯+Xn,S_n=X_1+X_2+\cdots+X_n, which counts the number of 11's (successes) in the sample. To find P(Sn=k)\mathbb{P}(S_n=k), note that each particular string with exactly kk successes has probability pk(1−p)n−kp^k(1-p)^{n-k} (the product rule from independence). How many such strings are there? Choosing which kk of the nn positions are the successes is a combination, so there are (nk)\binom{n}{k} of them. Since these strings are disjoint events of equal probability, we add them: P(Sn=k)=(nk)pk(1−p)n−k.\mathbb{P}(S_n=k)=\binom{n}{k}p^k(1-p)^{n-k}. The two ingredients are worth separating: pk(1−p)n−kp^k(1-p)^{n-k} is the probability of one way to get kk successes, while (nk)\binom{n}{k} counts how many ways there are. Their product is the binomial probability.

4

The Binomial distribution (Definition 12.8)

The distribution of SnS_n is the Binomial. For a positive integer nn and 0≤p≤10\le p\le 1, a random variable XX has the binomial distribution with parameters nn and pp if P(X=k)=(nk)pk(1−p)n−k,k=0,1,…,n,\mathbb{P}(X=k)=\binom{n}{k}p^k(1-p)^{n-k},\qquad k=0,1,\dots,n, and we write X∼Bin(n,p)X\sim\mathrm{Bin}(n,p). The two parameters carry plain meanings: nn is the number of trials and pp is the per-trial success probability. These values form a genuine distribution because they sum to 11 — directly by the binomial theorem, ∑k=0n(nk)pk(1−p)n−k=(p+(1−p))n=1n=1.\sum_{k=0}^{n}\binom{n}{k}p^k(1-p)^{n-k}=\bigl(p+(1-p)\bigr)^n=1^n=1. A Bernoulli variable is just the special case n=1n=1: Bin(1,p)=Ber(p)\mathrm{Bin}(1,p)=\mathrm{Ber}(p). To compute a cumulative probability P(X≤k)\mathbb{P}(X\le k) you sum the individual terms from 00 to kk; for the tail P(X≥1)\mathbb{P}(X\ge 1) the complement 1−P(X=0)=1−(1−p)n1-\mathbb{P}(X=0)=1-(1-p)^n is almost always quicker.

Definition 12.7 — Bernoulli distribution

Let 0≤p≤10\le p\le 1. A random variable XX has the Bernoulli distribution with success probability pp if X∈{0,1}X\in\{0,1\} and P(X=1)=p\mathbb{P}(X=1)=p, P(X=0)=1−p\mathbb{P}(X=0)=1-p. We write X∼Ber(p)X\sim\mathrm{Ber}(p).

Intuition. This is the mathematical model of one yes/no trial, with XX the indicator that records 11 for success and 00 for failure. Everything about it is fixed by the single number pp: since XX takes only two values and their probabilities must sum to 11, knowing P(X=1)=p\mathbb{P}(X=1)=p forces P(X=0)=1−p\mathbb{P}(X=0)=1-p. It is the building block — nn independent copies are what the binomial counts.
Definition 12.8 — Binomial distribution

Let nn be a positive integer and 0≤p≤10\le p\le 1. A random variable XX has the binomial distribution with parameters nn and pp if P(X=k)=(nk)pk(1−p)n−k,k=0,1,…,n.\mathbb{P}(X=k)=\binom{n}{k}p^k(1-p)^{n-k},\qquad k=0,1,\dots,n. We write X∼Bin(n,p)X\sim\mathrm{Bin}(n,p).

Intuition. This gives the probability of getting exactly kk successes in nn independent trials, each succeeding with probability pp. Read it as 'count ×\times weight': (nk)\binom{n}{k} is the number of ways to place the kk successes among the nn trials, and pk(1−p)n−kp^k(1-p)^{n-k} is the probability of any one such arrangement. The case n=1n=1 recovers the Bernoulli distribution.
Binomial as a sum of independent Bernoullis

If X1,…,XnX_1,\dots,X_n are independent Bernoulli variables with common parameter pp, then their sum Sn=X1+⋯+XnS_n=X_1+\cdots+X_n counts the successes and has Sn∼Bin(n,p)S_n\sim\mathrm{Bin}(n,p). There are (nk)\binom{n}{k} outcome strings with exactly kk successes, each of probability pk(1−p)n−kp^k(1-p)^{n-k}.

Intuition. A binomial variable is literally a tally of independent yes/no trials. The p.m.f. reads straight off this: each specific string of kk successes and n−kn-k failures has probability pk(1−p)n−kp^k(1-p)^{n-k} by independence, and there are (nk)\binom{n}{k} equally likely such strings, so adding them yields the binomial formula. This is precisely why the binomial coefficient appears.
Normalization via the binomial theorem

The binomial probabilities sum to one: ∑k=0n(nk)pk(1−p)n−k=(p+(1−p))n=1.\sum_{k=0}^{n}\binom{n}{k}p^k(1-p)^{n-k}=\bigl(p+(1-p)\bigr)^n=1.

Intuition. A list of numbers is a probability distribution only if it is nonnegative and totals 11. For the binomial, the total is the binomial-theorem expansion of (p+(1−p))n=1n=1(p+(1-p))^n=1^n=1 — the very identity after which the distribution is named. So whatever the values of nn and pp, the n+1n+1 probabilities P(X=0),…,P(X=n)\mathbb{P}(X=0),\dots,\mathbb{P}(X=n) account for all the probability exactly once.

Worked examples

Example 1

Bernoulli trials and a specific sequence. A biased coin lands heads (a success, X=1X=1) with probability p=0.3p=0.3 on each independent toss. (a) Write down P(X=1)\mathbb{P}(X=1) and P(X=0)\mathbb{P}(X=0) for a single toss. (b) The coin is tossed six times; find the probability of the exact sequence tails, heads, tails, tails, heads, tails — that is, (X1,…,X6)=(0,1,0,0,1,0)(X_1,\dots,X_6)=(0,1,0,0,1,0).

  1. 1

    (a) The single Bernoulli trial. By Definition 12.7, X∼Ber(0.3)X\sim\mathrm{Ber}(0.3), so P(X=1)=p=0.3\mathbb{P}(X=1)=p=0.3 and P(X=0)=1−p=0.7\mathbb{P}(X=0)=1-p=0.7.

  2. 2

    (b) Identify successes and failures. The sequence (0,1,0,0,1,0)(0,1,0,0,1,0) has two 11's (successes) and four 00's (failures).

  3. 3

    Use independence to multiply. The six tosses are independent Bernoulli trials, so the joint probability is the product of the per-toss probabilities: P(0,1,0,0,1,0)=(1−p) p (1−p) (1−p) p (1−p)=p2(1−p)4.\mathbb{P}(0,1,0,0,1,0)=(1-p)\,p\,(1-p)\,(1-p)\,p\,(1-p)=p^2(1-p)^4.

  4. 4

    Substitute p=0.3p=0.3. p2(1−p)4=(0.3)2(0.7)4=0.09×0.2401=0.021609.p^2(1-p)^4=(0.3)^2(0.7)^4=0.09\times 0.2401=0.021609.

Answer. P(X=1)=0.3\mathbb{P}(X=1)=0.3 and P(X=0)=0.7\mathbb{P}(X=0)=0.7; the specific sequence has probability p2(1−p)4=(0.3)2(0.7)4≈0.022p^2(1-p)^4=(0.3)^2(0.7)^4\approx 0.022. Only the number of successes matters, not their positions.
Example 2

Example 12.9. What is the probability that five rolls of a fair die yield two or three sixes?

Example 3

A tail probability by complement. Items coming off a production line are defective independently with probability p=0.1p=0.1. In a sample of n=8n=8 items, let XX be the number of defectives, so X∼Bin(8,0.1)X\sim\mathrm{Bin}(8,0.1). Find P(X≥1)\mathbb{P}(X\ge 1), the probability of at least one defective.