Skip to content
VidiMaster it, module by module

Modules

Showing Linear Algebra & Probability.

← All modules

Mathematics

Linear Algebra & Probability

MATH201510 modules

Module 1 · Vectors, Spaces & Linear Maps

Algebra

Vectors in ℝⁿ, vector spaces and subspaces, linear maps, and matrices.

0 of 19 lessons complete0%
2Vector Spaces & Subspaces
Subspaces

A subspace is a subset of a vector space that is itself a vector space under the very same addition and scalar multiplication. You never recheck all the axioms — you run one quick three-part test: it must contain the zero vector, and be closed under addition and under scalar multiplication.

Not started·0/6 solved
Linear Combinations & Span

A linear combination is what you get by scaling vectors and adding them; the span collects every possible linear combination of a set of vectors. Span is always a subspace through the origin, and its geometric shape — point, line, plane, or all of space — reflects how many independent directions the vectors supply.

Not started·0/6 solved
Is a Vector in the Span?

Deciding whether a vector bb lies in span{v1,…,vk}\text{span}\{v_1,\ldots,v_k\} comes down to one question: is the linear system [ v1 ⋯ vk∣b ][\,v_1\ \cdots\ v_k\mid b\,] consistent? This lesson shows how to set it up, read the answer from row reduction, and recover the coefficients when bb is in the span.

Not started·0/6 solved
Linear Independence

A set of vectors is linearly independent when none of them is "redundant" — no vector can be built from the others. This lesson defines independence precisely, shows how to test for it with the homogeneous system, rank, and determinant, and teaches you to extract an explicit dependency relation when one exists.

Not started·0/6 solved
Basis

A basis is the smallest set of vectors that still describes an entire space: linearly independent (no redundancy) and spanning (nothing left out). This lesson shows how to recognize a basis, why every basis of Rn\mathbb{R}^n has exactly nn vectors, and how to test a given set fast.

Not started·0/6 solved
Dimension

The dimension of a vector space is the number of vectors in any basis — a single number that counts its independent directions, equals the rank of any spanning set, and is the same no matter which basis you pick.

Not started·0/6 solved
Coordinates Relative to a Basis

A basis gives every vector a unique "address" — its coordinate vector. This lesson shows how to compute [v]B[v]_B by solving Bc=vBc = v and how to rebuild vv from its coordinates, all in R2\mathbb{R}^2 and R3\mathbb{R}^3.

Not started·0/6 solved
3Linear Maps
4Matrices
Matrices ↔ Linear Maps

Every linear map T:Rn→RmT:\mathbb{R}^n\to\mathbb{R}^m is "secretly" multiplication by a single matrix AA, built by seeing what TT does to the standard basis vectors. This lesson shows how to build that matrix, use it to apply the map, and why composing maps means multiplying matrices.

Not started·0/6 solved
Geometric Transformations in ℝ²

Every scaling, rotation, reflection, and shear of the plane is captured by a single 2×22\times 2 matrix, and applying the transformation is just matrix-vector multiplication. This lesson teaches you to read a matrix as a geometric motion and to build the matrix for the motion you want.

Not started·0/6 solved
Matrix Operations

Matrices can be added, scaled, and multiplied, but the rules are not the ones from ordinary arithmetic. This lesson covers entrywise addition and scalar multiplication, the dimension rule and row-by-column recipe for matrix multiplication, and why ABAB and BABA usually differ.

Not started·0/6 solved
Matrix Inverse

The inverse A−1A^{-1} is the matrix that undoes AA. This lesson shows you when a 2×22\times 2 matrix has an inverse, how to compute it with a one-line formula, and how inverses behave under multiplication.

Not started·0/6 solved
Transpose; Symmetric & Skew-Symmetric

The transpose flips a matrix across its main diagonal, swapping every row with the corresponding column. This lesson shows how to compute it, the clean algebra it obeys (including the order-reversing product rule), and how it defines symmetric matrices (A=ATA=A^T) and skew-symmetric matrices (A=−ATA=-A^T).

Not started·0/6 solved
Rank & Nullity (Matrices)

Rank counts the independent directions a matrix produces (its pivots), while nullity counts the directions it collapses. The Rank–Nullity Theorem ties them together: for an m×nm\times n matrix, rank⁡(A)+nullity⁡(A)=n\operatorname{rank}(A)+\operatorname{nullity}(A)=n.

Not started·0/6 solved

Module 2 · Gaussian Elimination & Determinants

Algebra

Elementary row operations, row reduction, linear systems, matrix inverses, and determinants.

0 of 10 lessons complete0%
1Systems & Gaussian Elimination
Elementary Row Operations & Elementary Matrices

Master the three elementary row operations, encode each one as an elementary matrix, and see Theorem 5.2 turn every row operation into a single left-multiplication A→EAA \to EA — the engine behind Gaussian elimination.

Not started·0/6 solved
Row-Echelon Form & the Gauss Algorithm

The staircase shape at the heart of Gaussian elimination: what row-echelon form means, how the recursive Gauss algorithm produces it step by step, and how counting pivots gives you the rank of any matrix.

Not started·0/6 solved
Reduced Row-Echelon Form & Gauss–Jordan

Row echelon form gives you a staircase of pivots, but reduced row echelon form (rref) goes further: every pivot is scaled to a leading 1 and stands alone in its column. This lesson covers Definition 5.4, the Gauss–Jordan sweep that produces the rref, and Remark 5.2 — the fact that the rref of a matrix is unique.

Not started·0/6 solved
Computing the Inverse by Row Reduction

Once you can row-reduce a matrix, you can invert one. This lesson turns Gauss–Jordan elimination into a machine for computing A−1A^{-1}: augment AA with the identity, reduce the left block to InI_n, and read the inverse off the right block. You will also learn to spot, mid-reduction, exactly when a matrix has no inverse at all.

Not started·0/6 solved
Linear Systems: Matrix Form & Solvability

Write any linear system as Ax=bAx=b, decide whether it is consistent using the rank criterion rank⁡(A)=rank⁡([A ∣ b])\operatorname{rank}(A)=\operatorname{rank}([A\,|\,b]), and describe its whole solution set as a particular solution plus the kernel.

Not started·0/6 solved
Solving Systems by Gaussian Elimination

How to find every solution of a linear system Ax=bAx=b. We justify why row-reducing the augmented matrix [A∣b][A\mid b] never changes the solution set (Lemma 5.1 and Theorem 5.5), read off consistency by comparing rank⁡(A)\operatorname{rank}(A) with rank⁡([A∣b])\operatorname{rank}([A\mid b]), and run the course's 4-step procedure to write the full solution as a particular solution plus a basis of ker⁡A\ker A. Along the way we meet the three possible outcomes: a unique solution, infinitely many, or none.

Not started·0/6 solved
2Determinants
The Determinant Function: 2×2, 3×3 & Axioms

The determinant condenses a square matrix into a single number that tells you whether the matrix is invertible and measures the area or volume its rows span. This lesson builds the determinant from the 2×22\times 2 formula ad−bcad-bc, its geometric meaning, the three defining axioms (Definition 6.1), and the Rule of Sarrus for 3×33\times 3 matrices.

Not started·0/6 solved
Computing Determinants: Row Reduction & Cofactors

Two reliable ways to compute the determinant of a square matrix: row-reduce to triangular form and multiply the pivots, or expand in cofactors along a convenient row or column. Includes the exact rules for how each elementary row operation changes the value, and why det⁡(A)≠0\det(A)\neq 0 means AA is invertible.

Not started·0/6 solved
Properties of Determinants

The determinant is more than a number you grind out by row reduction: it is a multiplicative invariant that detects invertibility and behaves predictably under products, powers, inverses, transposes, and change of basis. This lesson (§6.2) collects the key algebraic properties of det⁡\det and turns them into fast computational tools.

Not started·0/6 solved
Leibniz Definition, Block Matrices & Cramer's Rule

An enrichment tour beyond the three axioms (D1)-(D3): the Leibniz (permutation) formula for the determinant, why block-triangular matrices simply multiply their determinants, and how Cramer's rule and the Vandermonde determinant turn determinants into problem-solving tools.

Not started·0/6 solved

Module 3 · Eigenvalues & Diagonalization

Algebra

Eigenvalues and eigenvectors, the characteristic polynomial, diagonalization, and applications — dynamical systems, Markov chains, and PageRank.

0 of 6 lessons complete0%
1Eigenvalues & Diagonalization
Eigenvalues, Eigenvectors & the Characteristic Polynomial

Meet the equation Av=λvA\mathbf{v}=\lambda\mathbf{v}. This lesson explains what eigenvalues and eigenvectors are, why the eigenspace EλE_\lambda is exactly the kernel of A−λIA-\lambda I, and how the characteristic polynomial cA(λ)=det⁡(A−λI)c_A(\lambda)=\det(A-\lambda I) turns the search for eigenvalues into solving a single polynomial equation.

Not started·0/6 solved
Multiplicities, Trace, Determinant & Similarity

Two eigenvalues can be the same number yet behave differently. This lesson separates algebraic multiplicity (how often λ\lambda is a root of the characteristic polynomial) from geometric multiplicity (how many independent eigenvectors it has), uses the eigenvalues to read off the trace and determinant, pins down the bound 1≤gemu⁡(λ)≤almu⁡(λ)≤n1\le\operatorname{gemu}(\lambda)\le\operatorname{almu}(\lambda)\le n, and shows which spectral data survive transposing or passing to a similar matrix.

Not started·0/6 solved
Diagonalization

A square matrix is diagonalizable when it is similar to a diagonal matrix, A=SDS−1A=SDS^{-1}. This lesson shows how an eigenbasis supplies the columns of SS and the diagonal of DD, how to decide whether a matrix is diagonalizable by matching geometric and algebraic multiplicities, and why the factorization makes powers AkA^k cheap.

Not started·0/6 solved
Invariant Subspaces

An enrichment look at subspaces that a matrix maps into themselves. We define AA-invariant subspaces, see why one-dimensional invariant subspaces are exactly the eigenlines, describe the invariant subspaces of a diagonalizable matrix, and read off the axis and perpendicular plane of a 3D rotation.

Not started·0/6 solved
2Applications

Module 4 · Orthogonality

Algebra

Inner products, norms and angles; orthonormal bases; orthogonal projections and least squares; the Gram–Schmidt process and QR factorization; and orthogonal maps.

0 of 6 lessons complete0%
1Inner Products & Orthonormal Bases
2Projections & Least Squares
Orthogonal Projections & the Orthogonal Complement

Every vector x∈Rn\mathbf{x}\in\mathbb{R}^n splits uniquely as x=x∥+x⊥\mathbf{x}=\mathbf{x}_{\parallel}+\mathbf{x}_{\perp}: a part x∥=proj⁡Vx\mathbf{x}_{\parallel}=\operatorname{proj}_V\mathbf{x} lying in a subspace VV and a part x⊥\mathbf{x}_{\perp} orthogonal to it. This lesson builds that decomposition from an orthonormal basis, turns it into the symmetric projection matrix P=QQT=A(ATA)−1ATP=QQ^{T}=A(A^{T}A)^{-1}A^{T}, measures the distance from x\mathbf{x} to VV as ∥x−proj⁡Vx∥\|\mathbf{x}-\operatorname{proj}_V\mathbf{x}\|, and relates the orthogonal complement V⊥V^{\perp} to VV through dim⁡V+dim⁡V⊥=n\dim V+\dim V^{\perp}=n and im⁡(A)⊥=ker⁡(AT)\operatorname{im}(A)^{\perp}=\ker(A^{T}).

Not started·0/6 solved
Least-Squares Approximation & the Normal Equations

When Ax=bA\mathbf{x}=\mathbf{b} is inconsistent — b∉im⁡(A)\mathbf{b}\notin\operatorname{im}(A) — there is no exact solution, so we settle for the next best thing: the vector x^\hat{\mathbf{x}} that makes Ax^A\hat{\mathbf{x}} as close to b\mathbf{b} as possible. This lesson shows that the closest point in a subspace is the orthogonal projection (Theorem 8.12), that the minimiser x^\hat{\mathbf{x}} is found by solving the normal equation ATAx^=ATbA^{\mathsf{T}}A\hat{\mathbf{x}}=A^{\mathsf{T}}\mathbf{b} (Theorem 8.13), and how this turns messy, over-determined data into a single clean line of best fit.

Not started·0/6 solved
3Gram-Schmidt & Orthogonal Maps

Module 5 · Symmetric Matrices & SVD

Algebra

Symmetric matrices and the Spectral Theorem (orthogonal diagonalization), singular values and the Singular Value Decomposition, and principal component analysis.

0 of 5 lessons complete0%
1Symmetric Matrices & the Spectral Theorem
2Singular Value Decomposition

Module 6 · Probability Foundations

Algebra

Probability spaces and the axioms, the rules of probability and inclusion–exclusion, counting and sampling, conditional probability, Bayes’ formula, and independence.

0 of 6 lessons complete0%
1Probability Spaces & Rules
Probability Spaces: Sample Space, Events & Axioms

Every probability question begins with a probability space (Ω,F,P)(\Omega,\mathcal{F},\mathbb{P}) (Definition 10.1): a sample space Ω\Omega holding every possible outcome ω\omega, a collection of events F\mathcal{F} (the subsets of Ω\Omega we assign probabilities to), and a probability measure P\mathbb{P} obeying three axioms — every probability lies in [0,1][0,1], the certain event has P(Ω)=1\mathbb{P}(\Omega)=1 (and the impossible event P(∅)=0\mathbb{P}(\varnothing)=0), and the probabilities of disjoint events add. This lesson translates experiments into the language of sets, unpacks the axioms and their first consequences (Remark 10.1), and — when Ω\Omega is finite with equally likely outcomes — collapses probability into pure counting via P(A)=∣A∣∣Ω∣\mathbb{P}(A)=\tfrac{|A|}{|\Omega|} (Proposition 10.1). We put the machinery to work on a fair die (Example 10.1), a pair of distinguishable dice (Example 10.2), and an urn draw (Example 10.3).

Not started·0/6 solved
The Rules of Probability: Complement, Addition & Inclusion–Exclusion

Once the axioms are in place, almost every probability computation runs on a short list of rules, and this lesson assembles them. The complement rule P(Ac)=1−P(A)\mathbb{P}(A^c)=1-\mathbb{P}(A) turns a hard event into an easy one \u2014 the engine behind the complement trick P(at least one)=1−P(none)\mathbb{P}(\text{at least one})=1-\mathbb{P}(\text{none}), used in Example 10.4 to find the chance a value repeats when a fair die is rolled four times. For mutually exclusive (disjoint) events the Disjoint Sum Rule adds probabilities outright; when events overlap, the addition rule (the General Sum Rule) corrects for the double-counted intersection, P(A∪B)=P(A)+P(B)−P(A∩B)\mathbb{P}(A\cup B)=\mathbb{P}(A)+\mathbb{P}(B)-\mathbb{P}(A\cap B). Remark 10.3 rearranges this to recover P(A∩B)\mathbb{P}(A\cap B) and extends it to the inclusion\u2013exclusion principle for three events, which (for two) we apply in Example 10.5.

Not started·0/6 solved
Counting and the Four Sampling Methods

Probability on a finite, equally likely sample space collapses to counting: P(A)=#A/#ΩP(A)=\#A/\#\Omega. This lesson organizes that counting around two yes/no questions about how a sample of size kk is drawn from nn objects --- does order matter? and is each object replaced before the next draw? --- giving a 2×22\times2 grid of four sampling schemes. Three of them produce equally likely outcomes and power every probability here: ordered-with-replacement (nkn^k, Example 10.6), ordered-without-replacement (n!/(n−k)!n!/(n-k)!, Example 10.7), and unordered-without-replacement ((nk)\binom{n}{k}, Example 10.8). We count each scheme, see how dividing an ordered count by k!k! produces combinations, and turn the counts into probabilities --- including Jordan's books solved two ways (Example 10.9, Remark 10.4) and the three contrasting scenarios of Example 10.10.

Not started·0/6 solved
2Conditional Probability & Independence
Conditional Probability & the Law of Total Probability

New information reshapes a probability model: once we learn that an event BB has occurred, outcomes outside BB become impossible and the survivors must be renormalized so their probabilities again sum to 11. Definition 11.1 packages this as P(A∣B)=P(A∩B)/P(B)\mathbb{P}(A\mid B)=\mathbb{P}(A\cap B)/\mathbb{P}(B), valid whenever P(B)>0\mathbb{P}(B)>0. Rearranging gives the multiplication rule P(A∩B)=P(A∣B)P(B)\mathbb{P}(A\cap B)=\mathbb{P}(A\mid B)\mathbb{P}(B) — the engine behind sequential draws without replacement — and splitting Ω\Omega into a partition (Definition 11.2) yields the Law of Total Probability (Proposition 11.1), P(A)=∑iP(A∣Bi)P(Bi)\mathbb{P}(A)=\sum_i\mathbb{P}(A\mid B_i)\mathbb{P}(B_i), the backbone of the two-stage, two-urn reasoning in Example 11.4.

Not started·0/6 solved
Bayes' Formula: Reversing the Order of Conditioning

Conditional probability answers P(A∣B)P(A\mid B); Bayes' formula (Proposition 11.2) answers the reverse question P(B∣A)P(B\mid A) -- the probability of a cause given an observed effect. Assembled from the definition of conditional probability (Definition 11.1) and the law of total probability (Proposition 11.1), it updates a prior belief P(B)P(B) into a posterior P(B∣A)P(B\mid A) once evidence AA arrives (Remark 11.2). The centerpiece is the medical-test / base-rate problem (Example 11.6), where a highly accurate test for a rare disease returns a surprisingly low posterior -- the base-rate fallacy that routinely fools intuition.

Not started·0/6 solved
Independence

Two events are independent when knowing one tells you nothing about the other: P(A∩B) = P(A)·P(B).

Not started·0/6 solved

Module 7 · Random Variables & Distributions

Algebra

Random variables and probability mass functions, continuous variables and densities, the cumulative distribution function, and the common distributions: Bernoulli, binomial, geometric, negative binomial, uniform, and normal.

0 of 6 lessons complete0%
1Random Variables & Distributions
Random Variables & the Probability Mass Function

A random variable (Definition 12.1) is not a number but a function X:Ω→RX:\Omega\to\mathbb{R} that reads each outcome ω\omega of an experiment and reports a real value X(ω)X(\omega) — the sum of two dice, a gambler's change in wealth, a count. Writing {X∈B}={ω∈Ω:X(ω)∈B}\{X\in B\}=\{\omega\in\Omega:X(\omega)\in B\} turns a question about values into an event with a probability P{X∈B}\mathbb{P}\{X\in B\}, and the whole family of these probabilities is the probability distribution of XX (Definition 12.2). This lesson concentrates on discrete random variables (Definition 12.3), whose values form a finite or countably infinite list that carries all the probability. For them the distribution collapses to a single object, the probability mass function p(k)=P{X=k}p(k)=\mathbb{P}\{X=k\} (Definition 12.4): every probability becomes a sum of masses, P{X∈B}=∑k∈Bp(k)\mathbb{P}\{X\in B\}=\sum_{k\in B}p(k), and a function is a legitimate p.m.f. exactly when it is nonnegative and sums to 11. We build p.m.f.s and read probabilities off them for a pair of dice (Examples 12.1, 12.3), a wealth/payoff game (Examples 12.2, 12.4), and a biased spinner. (Continuous random variables and densities are left for a later lesson.)

Not started·0/6 solved
Continuous Random Variables & Probability Density Functions

A continuous random variable XX (Definition 12.5) draws its probabilities not from a list of point masses but from the area under a curve: there is a probability density function (pdf) f:R→Rf:\mathbb{R}\to\mathbb{R} with P(X≤b)=∫−∞bf(x) dxfor all b∈R,\mathbb{P}(X\le b)=\int_{-\infty}^{b}f(x)\,dx\qquad\text{for all }b\in\mathbb{R}, and more generally P(X∈B)=∫Bf(x) dx\mathbb{P}(X\in B)=\int_B f(x)\,dx. This lesson pins down what makes a function a legitimate density (Remark 12.1: f≥0f\ge 0 everywhere and total area ∫−∞∞f(x) dx=1\int_{-\infty}^{\infty}f(x)\,dx=1), reads probabilities off as areas P(a≤X≤b)=∫abf(x) dx\mathbb{P}(a\le X\le b)=\int_a^b f(x)\,dx, and draws the defining contrast with the discrete world — for a continuous variable every single point is negligible, P(X=c)=0\mathbb{P}(X=c)=0 (Proposition 12.1), so endpoints never matter: P(a≤X≤b)=P(a<X<b)\mathbb{P}(a\le X\le b)=\mathbb{P}(a<X<b). We anchor everything in the uniform distribution on [0,1][0,1] (Example 12.5), where f(x)=1f(x)=1 on [0,1][0,1] and a probability is just a length, and in a worked exponential density (Example 12.6).

Not started·0/6 solved
The Cumulative Distribution Function: F(x)=P{X≤x}F(x)=\mathbb{P}\{X\le x\}

Unlike the probability mass function (discrete variables only) or the density (continuous variables only), the cumulative distribution function (c.d.f.) is defined for every random variable (Definition 12.6): F(x)=P{X≤x}F(x)=\mathbb{P}\{X\le x\} records the total probability accumulated up to the point xx — the 'sum-or-area-so-far' function. The '≤\le' is deliberate (Remark 12.2): it includes the endpoint and makes the c.d.f. deliver interval probabilities, P{a<X≤b}=F(b)−F(a)\mathbb{P}\{a<X\le b\}=F(b)-F(a). For a discrete variable the c.d.f. is a step function (Remark 12.3) that is flat between the possible values and jumps by P{X=x}\mathbb{P}\{X=x\} at each one, as for the change-in-wealth variable of Example 12.7; for a continuous variable it rises smoothly, obtained by integrating the density, as for the Unif[1,3]\mathrm{Unif}[1,3] variable of Example 12.8. Whatever its type, every c.d.f. shares the same three shape properties (Proposition 12.2): it is nondecreasing, right-continuous, and climbs from 00 at −∞-\infty to 11 at +∞+\infty.

Not started·0/6 solved
2Common Distributions
The Bernoulli and Binomial Distributions

Repeated independent trials with just two outcomes — success (probability pp) and failure (probability 1−p1-p) — are the seed from which two fundamental discrete distributions grow. The Bernoulli distribution (Definition 12.7) records a single such trial: a random variable X∈{0,1}X\in\{0,1\} with P(X=1)=p\mathbb{P}(X=1)=p and P(X=0)=1−p\mathbb{P}(X=0)=1-p, written X∼Ber(p)X\sim\mathrm{Ber}(p). Run the trial nn independent times and you obtain nn independent Bernoulli variables X1,…,XnX_1,\dots,X_n; their sum Sn=X1+⋯+XnS_n=X_1+\cdots+X_n counts the successes, and its distribution is the Binomial (Definition 12.8): P(X=k)=(nk)pk(1−p)n−k\mathbb{P}(X=k)=\binom{n}{k}p^k(1-p)^{n-k} for k=0,1,…,nk=0,1,\dots,n, written X∼Bin(n,p)X\sim\mathrm{Bin}(n,p). The binomial coefficient (nk)\binom{n}{k} counts the arrangements of kk successes among the nn trials, each arrangement carrying probability pk(1−p)n−kp^k(1-p)^{n-k}, and the binomial theorem guarantees these probabilities sum to 11. We close with Example 12.9: the chance that five rolls of a fair die yield two or three sixes, with S5∼Bin(5,16)S_5\sim\mathrm{Bin}(5,\tfrac16), equal to 15007776≈0.193\tfrac{1500}{7776}\approx 0.193.

Not started·0/6 solved
The Geometric & Negative Binomial Distributions

Independent repeated Bernoulli trials — each a success with probability pp, a failure with probability 1−p1-p — raise a question the binomial does not answer: not how many successes in a fixed number of trials, but how long until a success. The geometric distribution (Definition 12.9) answers the first version: X∼Geom(p)X\sim\mathrm{Geom}(p) is the trial on which the first success occurs, with probability mass function P{X=k}=(1−p)k−1p\mathbb{P}\{X=k\}=(1-p)^{k-1}p for k=1,2,3,…k=1,2,3,\dots — the price of k−1k-1 failures followed by one success. Its masses form a geometric series summing to 11, and its tail has a clean closed form P{X>k}=(1−p)k\mathbb{P}\{X>k\}=(1-p)^k (all of the first kk trials fail), so P{X≤k}=1−(1−p)k\mathbb{P}\{X\le k\}=1-(1-p)^k; Example 12.10 uses this to find the chance of needing more than seven rolls of a fair die for the first six, (5/6)7≈0.279(5/6)^7\approx0.279. The negative binomial distribution (Definition 12.10) generalises the idea to the trial of the rr-th success: X∼NegBin(r,p)X\sim\mathrm{NegBin}(r,p) has P{X=k}=(k−1r−1)pr(1−p)k−r\mathbb{P}\{X=k\}=\binom{k-1}{r-1}p^r(1-p)^{k-r} for k=r,r+1,…k=r,r+1,\dots, the coefficient (k−1r−1)\binom{k-1}{r-1} counting the arrangements of the r−1r-1 earlier successes among the first k−1k-1 trials. Remark 12.4 ties the two together: the geometric is exactly NegBin(1,p)\mathrm{NegBin}(1,p).

Not started·0/6 solved
Continuous Distributions: The Uniform and the Standard Normal

These are the first two continuous distributions of the course, and both compute probabilities as areas under a density curve rather than by summing point masses. The continuous uniform distribution X∼Unif(a,b)X\sim\mathrm{Unif}(a,b) (Definition 12.11) spreads probability evenly over an interval [a,b][a,b]: its density is the constant 1b−a\frac{1}{b-a}, so the chance of landing in a subinterval is just its length ratio, P{c≤X≤d}=d−cb−aP\{c\le X\le d\}=\frac{d-c}{b-a}. The standard normal (Gaussian) distribution Z∼N(0,1)Z\sim N(0,1) (Definition 12.12) is the famous bell curve, with density φ(x)=12πe−x2/2\varphi(x)=\frac{1}{\sqrt{2\pi}}e^{-x^2/2} symmetric about 00. Its cumulative distribution function Φ\Phi has no closed form, so we read Φ\Phi from a standard-normal table or technology and assemble every probability from P{a≤Z≤b}=Φ(b)−Φ(a)P\{a\le Z\le b\}=\Phi(b)-\Phi(a) together with the symmetry rule Φ(−x)=1−Φ(x)\Phi(-x)=1-\Phi(x). We close with Example 12.11, computing P(−1≤Z≤1.5)≈0.7745P(-1\le Z\le 1.5)\approx 0.7745, and record the landmark values P(−1≤Z≤1)≈0.683P(-1\le Z\le1)\approx0.683 and P(−2≤Z≤2)≈0.954P(-2\le Z\le2)\approx0.954.

Not started·0/6 solved

Module 8 · Expectation & Variance

Algebra

Expectation and the law of the unconscious statistician, variance and standard deviation, the means and variances of the common distributions, and the normal, Poisson, and exponential distributions.

0 of 5 lessons complete0%
1Expectation & Variance
Expectation: The Mean of a Random Variable

The expectation (or mean) of a random variable is the single number that says where its probability is centred — a weighted average of the values, each weighted by its probability. For a discrete variable this is E[X]=∑kk P{X=k}\mathbb{E}[X]=\sum_k k\,\mathbb{P}\{X=k\} (Definition 13.1), also called the first moment and written μ=E[X]\mu=\mathbb{E}[X]; rolling a fair die gives E[X]=16(1+⋯+6)=3.5\mathbb{E}[X]=\tfrac16(1+\cdots+6)=3.5 (Example 13.1) — already a sign that the mean need not be a value XX can actually take. For a continuous variable the sum becomes an integral against the density, E[X]=∫−∞∞x f(x) dx\mathbb{E}[X]=\int_{-\infty}^{\infty} x\,f(x)\,dx (Definition 13.2), which for X∼Unif[a,b]X\sim\mathrm{Unif}[a,b] returns the midpoint a+b2\tfrac{a+b}{2} (Example 13.7). To average a function of XX you do not need the distribution of g(X)g(X): the law of the unconscious statistician (Proposition 13.1) gives E[g(X)]=∑kg(k) P{X=k}\mathbb{E}[g(X)]=\sum_k g(k)\,\mathbb{P}\{X=k\}, and the choice g(x)=xng(x)=x^n produces the moments E[Xn]\mathbb{E}[X^n] (Definition 13.3). Expectation is powerful but not automatic: it is well-defined only when its sum or integral settles on a value, finite or ±∞\pm\infty (Remark 13.1), and it can be infinite (Examples 13.5, 13.8) or even undefined (Example 13.6). (The means of the named distributions — Bernoulli, binomial, geometric — are developed in a later lesson; here we keep them light.)

Not started·0/6 solved
Variance & Standard Deviation

The mean locates a random variable; the variance measures how far it typically strays from that centre. For XX with mean μ=E[X]\mu=\mathbb{E}[X], Definition 13.4 sets Var(X)=E[(X−μ)2]\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2] — the average squared deviation from the mean — written σ2\sigma^2, with its square root the standard deviation σ=SD(X)=Var(X)\sigma=\mathrm{SD}(X)=\sqrt{\mathrm{Var}(X)} restoring the original units. Squaring the deviation inside the expectation makes Proposition 13.2 a sum for a discrete variable, Var(X)=∑k(k−μ)2 P(X=k)\mathrm{Var}(X)=\sum_k(k-\mu)^2\,\mathbb{P}(X=k), but expanding that square yields the far handier computational formula of Proposition 13.3, Var(X)=E[X2]−μ2\mathrm{Var}(X)=\mathbb{E}[X^2]-\mu^2 — the mean of the square minus the square of the mean. From the definition flow the transformation rules of Proposition 13.4: a shift moves the centre but not the spread, while a scale factor comes out squared, Var(aX+b)=a2Var(X)\mathrm{Var}(aX+b)=a^2\mathrm{Var}(X) (so SD(aX+b)=∣a∣ SD(X)\mathrm{SD}(aX+b)=|a|\,\mathrm{SD}(X)). Finally Proposition 13.5 reads off the extreme case — Var(X)=0\mathrm{Var}(X)=0 precisely when XX is almost surely constant. (The specific variances of the Bernoulli, binomial, uniform and geometric families, Examples 13.9–13.12, are collected in a separate lesson on the moments of named distributions.)

Not started·0/6 solved
Mean and Variance of the Common Distributions

Two numbers summarise any distribution: the mean μ=E[X]\mu=\mathbb{E}[X] (Definition 13.1, the first moment), which says where the distribution is centred, and the variance σ2=Var(X)=E[(X−μ)2]\sigma^2=\mathrm{Var}(X)=\mathbb{E}[(X-\mu)^2] (Definition 13.4), which says how widely it spreads about that centre; its square root σ=SD(X)\sigma=\mathrm{SD}(X) is the standard deviation. For hand computation the variance is almost always found from the computational formula Var(X)=E[X2]−(E[X])2\mathrm{Var}(X)=\mathbb{E}[X^2]-(\mathbb{E}[X])^2 (Proposition 13.3). This lesson is a consolidated reference: it collects, once and for all, the mean and variance of the four distributions you meet most often, each derived in the course's own notes. For the discrete families these are the Bernoulli Ber(p)\mathrm{Ber}(p) with mean pp (Example 13.2) and variance p(1−p)p(1-p) (Example 13.9); the binomial Bin(n,p)\mathrm{Bin}(n,p) with mean npnp (Example 13.3) and variance np(1−p)np(1-p) (Example 13.10); and the geometric Geom(p)\mathrm{Geom}(p) (support k≥1k\ge 1, p.m.f. (1−p)k−1p(1-p)^{k-1}p) with mean 1p\frac{1}{p} (Example 13.4) and variance 1−pp2\frac{1-p}{p^2} (Example 13.12). For the continuous world we record the uniform Unif[a,b]\mathrm{Unif}[a,b] with mean a+b2\frac{a+b}{2} (Example 13.7) and variance (b−a)212\frac{(b-a)^2}{12} (Example 13.11). Once you can name the family and read off its parameters, every mean and variance is a one-line substitution.

Not started·0/6 solved
2More Distributions
The Normal Distribution N(μ,σ2)N(\mu,\sigma^2)

The standard normal Z∼N(0,1)Z\sim N(0,1) is the reference bell curve: it has mean 00 and variance 11 (Proposition 13.6), and its cumulative distribution function is written Φ(z)=P(Z≤z)\Phi(z)=\mathbb{P}(Z\le z). Every other normal is an affine stretch-and-shift of it. For μ∈R\mu\in\mathbb{R} and σ>0\sigma>0, setting X=σZ+μX=\sigma Z+\mu produces the general normal distribution X∼N(μ,σ2)X\sim N(\mu,\sigma^2) (Definition 13.5), with mean μ\mu, variance σ2\sigma^2, and the bell-shaped density f(x)=1σ2πexp⁡ ⁣(−(x−μ)22σ2),f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), centred at μ\mu with its spread set by σ\sigma. Affine maps keep you inside the family (Proposition 13.7): if X∼N(μ,σ2)X\sim N(\mu,\sigma^2) then aX+b∼N(aμ+b, a2σ2)aX+b\sim N(a\mu+b,\,a^2\sigma^2), and the special case Z=X−μσZ=\frac{X-\mu}{\sigma} — standardization — turns any normal back into the standard one. That is the computational workhorse: P(X≤x)=Φ ⁣(x−μσ)\mathbb{P}(X\le x)=\Phi\!\left(\frac{x-\mu}{\sigma}\right) collapses every normal probability to a single Φ\Phi-value read from a standard-normal table or software, as in Example 13.14 (X∼N(−3,4)X\sim N(-3,4) gives P(X≤−1.7)=Φ(0.65)≈0.742\mathbb{P}(X\le-1.7)=\Phi(0.65)\approx0.742). The same device yields the 68–95–99.7 rule: a normal variable falls within 11, 22, or 33 standard deviations of its mean with probability about 0.6830.683, 0.9540.954, and 0.9970.997 (Example 13.15).

Not started·0/6 solved
The Poisson & Exponential Distributions

Two distributions govern rare, randomly timed events, and this lesson treats them as a pair. The Poisson distribution (Definition 13.6) counts how many such events occur in a fixed window — calls at a switchboard, decays of a sample, requests to a server — when they happen independently at a constant average rate λ>0\lambda>0. A Poisson(λ)(\lambda) variable takes values k=0,1,2,…k=0,1,2,\dots with mass P{X=k}=e−λλkk!\mathbb{P}\{X=k\}=\dfrac{e^{-\lambda}\lambda^{k}}{k!}; summing the series gives e−λeλ=1e^{-\lambda}e^{\lambda}=1, so it is a genuine p.m.f., and a short computation (Proposition 13.8) shows the single parameter λ\lambda is both the mean and the variance, E[X]=Var(X)=λ\mathbb{E}[X]=\mathrm{Var}(X)=\lambda (Example 13.16: 55 requests per second give P{X=3}≈0.140\mathbb{P}\{X=3\}\approx0.140). The exponential distribution (Definition 13.7) is the continuous companion that measures the waiting time until the next such event. An Exp(λ)(\lambda) variable has density f(x)=λe−λxf(x)=\lambda e^{-\lambda x} for x≥0x\ge 0, tail P{X>t}=e−λt\mathbb{P}\{X>t\}=e^{-\lambda t}, c.d.f. 1−e−λt1-e^{-\lambda t}, mean 1/λ1/\lambda and variance 1/λ21/\lambda^{2} (Examples 13.17–13.19). Its signature feature is the memoryless property (Proposition 13.9): P{X>t+s∣X>t}=P{X>s}\mathbb{P}\{X>t+s\mid X>t\}=\mathbb{P}\{X>s\} — having already waited tells you nothing about how much longer you must wait (Example 13.20, the turtle and the highway).

Not started·0/6 solved

Module 9 · Approximations & the CLT

Algebra

The Central Limit Theorem and the normal approximation to the binomial (with continuity correction), the Poisson approximation and the law of rare events, choosing between them, and applications: confidence intervals and the law of large numbers.

0 of 4 lessons complete0%
1Normal approximation & the CLT

Module 10 · Joint Distributions & Covariance

Algebra

Joint and marginal distributions of discrete random variables, independence and exchangeability, linearity of expectation and the variance of sums, and covariance and correlation — the bridge from probability back to linear algebra.

0 of 5 lessons complete0%