Modules
Browse by subject and course — or jump straight to any lesson.
Mathematics
Module 1 · Vectors, Spaces & Linear Maps
AlgebraVectors in ℝⁿ, vector spaces and subspaces, linear maps, and matrices.
1Vectors
Vector Operations in ℝⁿ
Vectors in are lists of numbers you can add and scale component by component. This lesson builds the algebra and geometry of vector addition, scalar multiplication, and linear combinations, then pins down the properties that make these operations so predictable.
Lines and Planes in ℝ³
Learn to describe lines and planes in three-dimensional space using points and direction vectors, move fluidly between parametric, vector, and normal forms, and test whether a given point lies on a line or plane.
2Vector Spaces & Subspaces
Subspaces
A subspace is a subset of a vector space that is itself a vector space under the very same addition and scalar multiplication. You never recheck all the axioms — you run one quick three-part test: it must contain the zero vector, and be closed under addition and under scalar multiplication.
Linear Combinations & Span
A linear combination is what you get by scaling vectors and adding them; the span collects every possible linear combination of a set of vectors. Span is always a subspace through the origin, and its geometric shape — point, line, plane, or all of space — reflects how many independent directions the vectors supply.
Is a Vector in the Span?
Deciding whether a vector lies in comes down to one question: is the linear system consistent? This lesson shows how to set it up, read the answer from row reduction, and recover the coefficients when is in the span.
Linear Independence
A set of vectors is linearly independent when none of them is "redundant" — no vector can be built from the others. This lesson defines independence precisely, shows how to test for it with the homogeneous system, rank, and determinant, and teaches you to extract an explicit dependency relation when one exists.
Basis
A basis is the smallest set of vectors that still describes an entire space: linearly independent (no redundancy) and spanning (nothing left out). This lesson shows how to recognize a basis, why every basis of has exactly vectors, and how to test a given set fast.
Dimension
The dimension of a vector space is the number of vectors in any basis — a single number that counts its independent directions, equals the rank of any spanning set, and is the same no matter which basis you pick.
Coordinates Relative to a Basis
A basis gives every vector a unique "address" — its coordinate vector. This lesson shows how to compute by solving and how to rebuild from its coordinates, all in and .
3Linear Maps
Linear Maps: Definition
A linear map is a function between vector spaces that preserves addition and scalar multiplication. This lesson defines it precisely, shows how to test a formula for linearity, and separates genuine linear maps from affine and nonlinear impostors.
Kernel & Injectivity
The kernel of a linear map collects every input that gets sent to zero, and its size tells you instantly whether the map is injective. Learn to compute it by solving and to read off injectivity from its dimension.
Image & Surjectivity
The image of a linear map is the set of all outputs it can actually produce — exactly the span of its matrix's columns. Measuring that image with the rank tells you instantly whether the map is surjective and whether a particular vector can be reached.
Rank–Nullity (Linear Maps)
The Rank–Nullity Theorem says that for a linear map out of a finite-dimensional space, the dimensions of its kernel and image always add up to the dimension of the domain. This one equation lets you compute any missing count and instantly rule out impossible injective or surjective maps.
4Matrices
Matrices ↔ Linear Maps
Every linear map is "secretly" multiplication by a single matrix , built by seeing what does to the standard basis vectors. This lesson shows how to build that matrix, use it to apply the map, and why composing maps means multiplying matrices.
Geometric Transformations in ℝ²
Every scaling, rotation, reflection, and shear of the plane is captured by a single matrix, and applying the transformation is just matrix-vector multiplication. This lesson teaches you to read a matrix as a geometric motion and to build the matrix for the motion you want.
Matrix Operations
Matrices can be added, scaled, and multiplied, but the rules are not the ones from ordinary arithmetic. This lesson covers entrywise addition and scalar multiplication, the dimension rule and row-by-column recipe for matrix multiplication, and why and usually differ.
Matrix Inverse
The inverse is the matrix that undoes . This lesson shows you when a matrix has an inverse, how to compute it with a one-line formula, and how inverses behave under multiplication.
Transpose; Symmetric & Skew-Symmetric
The transpose flips a matrix across its main diagonal, swapping every row with the corresponding column. This lesson shows how to compute it, the clean algebra it obeys (including the order-reversing product rule), and how it defines symmetric matrices () and skew-symmetric matrices ().
Rank & Nullity (Matrices)
Rank counts the independent directions a matrix produces (its pivots), while nullity counts the directions it collapses. The Rank–Nullity Theorem ties them together: for an matrix, .
Module 2 · Gaussian Elimination & Determinants
AlgebraElementary row operations, row reduction, linear systems, matrix inverses, and determinants.
1Systems & Gaussian Elimination
Elementary Row Operations & Elementary Matrices
Master the three elementary row operations, encode each one as an elementary matrix, and see Theorem 5.2 turn every row operation into a single left-multiplication — the engine behind Gaussian elimination.
Row-Echelon Form & the Gauss Algorithm
The staircase shape at the heart of Gaussian elimination: what row-echelon form means, how the recursive Gauss algorithm produces it step by step, and how counting pivots gives you the rank of any matrix.
Reduced Row-Echelon Form & Gauss–Jordan
Row echelon form gives you a staircase of pivots, but reduced row echelon form (rref) goes further: every pivot is scaled to a leading 1 and stands alone in its column. This lesson covers Definition 5.4, the Gauss–Jordan sweep that produces the rref, and Remark 5.2 — the fact that the rref of a matrix is unique.
Computing the Inverse by Row Reduction
Once you can row-reduce a matrix, you can invert one. This lesson turns Gauss–Jordan elimination into a machine for computing : augment with the identity, reduce the left block to , and read the inverse off the right block. You will also learn to spot, mid-reduction, exactly when a matrix has no inverse at all.
Linear Systems: Matrix Form & Solvability
Write any linear system as , decide whether it is consistent using the rank criterion , and describe its whole solution set as a particular solution plus the kernel.
Solving Systems by Gaussian Elimination
How to find every solution of a linear system . We justify why row-reducing the augmented matrix never changes the solution set (Lemma 5.1 and Theorem 5.5), read off consistency by comparing with , and run the course's 4-step procedure to write the full solution as a particular solution plus a basis of . Along the way we meet the three possible outcomes: a unique solution, infinitely many, or none.
2Determinants
The Determinant Function: 2×2, 3×3 & Axioms
The determinant condenses a square matrix into a single number that tells you whether the matrix is invertible and measures the area or volume its rows span. This lesson builds the determinant from the formula , its geometric meaning, the three defining axioms (Definition 6.1), and the Rule of Sarrus for matrices.
Computing Determinants: Row Reduction & Cofactors
Two reliable ways to compute the determinant of a square matrix: row-reduce to triangular form and multiply the pivots, or expand in cofactors along a convenient row or column. Includes the exact rules for how each elementary row operation changes the value, and why means is invertible.
Properties of Determinants
The determinant is more than a number you grind out by row reduction: it is a multiplicative invariant that detects invertibility and behaves predictably under products, powers, inverses, transposes, and change of basis. This lesson (§6.2) collects the key algebraic properties of and turns them into fast computational tools.
Leibniz Definition, Block Matrices & Cramer's Rule
An enrichment tour beyond the three axioms (D1)-(D3): the Leibniz (permutation) formula for the determinant, why block-triangular matrices simply multiply their determinants, and how Cramer's rule and the Vandermonde determinant turn determinants into problem-solving tools.
Module 3 · Eigenvalues & Diagonalization
AlgebraEigenvalues and eigenvectors, the characteristic polynomial, diagonalization, and applications — dynamical systems, Markov chains, and PageRank.
1Eigenvalues & Diagonalization
Eigenvalues, Eigenvectors & the Characteristic Polynomial
Meet the equation . This lesson explains what eigenvalues and eigenvectors are, why the eigenspace is exactly the kernel of , and how the characteristic polynomial turns the search for eigenvalues into solving a single polynomial equation.
Multiplicities, Trace, Determinant & Similarity
Two eigenvalues can be the same number yet behave differently. This lesson separates algebraic multiplicity (how often is a root of the characteristic polynomial) from geometric multiplicity (how many independent eigenvectors it has), uses the eigenvalues to read off the trace and determinant, pins down the bound , and shows which spectral data survive transposing or passing to a similar matrix.
Diagonalization
A square matrix is diagonalizable when it is similar to a diagonal matrix, . This lesson shows how an eigenbasis supplies the columns of and the diagonal of , how to decide whether a matrix is diagonalizable by matching geometric and algebraic multiplicities, and why the factorization makes powers cheap.
Invariant Subspaces
An enrichment look at subspaces that a matrix maps into themselves. We define -invariant subspaces, see why one-dimensional invariant subspaces are exactly the eigenlines, describe the invariant subspaces of a diagonalizable matrix, and read off the axis and perpendicular plane of a 3D rotation.
2Applications
Dynamical Systems: Discrete & Continuous
Use eigenvalues and diagonalization to solve discrete systems and continuous systems in closed form, and to read off their stability and long-run behavior directly from the eigenvalues.
Markov Chains, PageRank & Path Counting
A Markov chain moves between states by fixed probabilities, encoded in a column-stochastic transition matrix that updates the state via . When is regular, the chain always settles to a unique equilibrium distribution (the normalized eigenvector for ) -- the same idea behind Google's PageRank. The matrix-power trick also counts paths in a directed graph.
Module 4 · Orthogonality
AlgebraInner products, norms and angles; orthonormal bases; orthogonal projections and least squares; the Gram–Schmidt process and QR factorization; and orthogonal maps.
1Inner Products & Orthonormal Bases
Inner Products, Norms & Angles
An inner product turns two vectors into a single scalar, and from it we build the norm — unlocking length, distance, the Cauchy–Schwarz inequality, orthogonality, the Pythagorean theorem, and the angle between vectors.
Orthonormal Bases & Coordinates
Finding a vector's coordinates in a basis usually means solving a linear system — tedious and error-prone in high dimensions. This lesson shows why an orthonormal basis makes that work vanish: each coordinate is a single dot product, (Theorem 8.5). Along the way we define orthogonal and orthonormal sets (Definition 8.7), see why orthonormal vectors are automatically linearly independent (Proposition 8.1) and form bases (Theorem 8.4), and learn to normalize any orthogonal basis (Remark 8.4).
2Projections & Least Squares
Orthogonal Projections & the Orthogonal Complement
Every vector splits uniquely as : a part lying in a subspace and a part orthogonal to it. This lesson builds that decomposition from an orthonormal basis, turns it into the symmetric projection matrix , measures the distance from to as , and relates the orthogonal complement to through and .
Least-Squares Approximation & the Normal Equations
When is inconsistent — — there is no exact solution, so we settle for the next best thing: the vector that makes as close to as possible. This lesson shows that the closest point in a subspace is the orthogonal projection (Theorem 8.12), that the minimiser is found by solving the normal equation (Theorem 8.13), and how this turns messy, over-determined data into a single clean line of best fit.
3Gram-Schmidt & Orthogonal Maps
The Gram–Schmidt Process & QR Factorization
Orthonormal bases are wonderful to compute with — but where do they come from? The Gram–Schmidt process manufactures one from any basis, normalizing and then, for each later vector, subtracting its projection onto the span of the vectors already processed before normalizing what remains (Theorem 8.15). The result depends on the order of the inputs (Remark 8.6), and when the vectors are the columns of a matrix the whole computation repackages into the QR factorization , with orthonormal and upper triangular with positive diagonal (Theorem 8.16).
Orthogonal Maps & Orthogonal Matrices
Orthogonal maps are the linear transformations of that preserve the inner product: . This lesson shows why that single condition is the same as preserving length (Theorem 8.17) and angle (Corollary 8.2), why it is equivalent to sending the standard basis to an orthonormal basis (Theorem 8.18), and how it becomes the matrix criterion , equivalently with orthonormal columns (Theorems 8.19, 8.21). Rotations and reflections are the picture, and every orthogonal matrix has (Corollary 8.3).
Module 5 · Symmetric Matrices & SVD
AlgebraSymmetric matrices and the Spectral Theorem (orthogonal diagonalization), singular values and the Singular Value Decomposition, and principal component analysis.
1Symmetric Matrices & the Spectral Theorem
Symmetric Matrices: Real Eigenvalues & Orthogonal Eigenvectors
Symmetric matrices () are the best-behaved operators in all of linear algebra, and this lesson uncovers why. Two structural miracles set the stage for the Spectral Theorem: every eigenvalue of a symmetric matrix is real (Proposition 9.2), and eigenvectors belonging to distinct eigenvalues are automatically orthogonal (Lemma 9.1). Along the way we see that orthogonal diagonalizability forces symmetry (Remark 9.1), and that the orthogonal complement of an eigenvector is -invariant (Lemma 9.2) — the geometric engine that will power the full theorem next time.
The Spectral Theorem
A real square matrix can be diagonalized by an orthogonal change of basis exactly when it is symmetric. The Spectral Theorem (Theorem 9.1) packages a symmetric as , where holds the (always real) eigenvalues and the columns of the orthogonal matrix form an orthonormal eigenbasis. This lesson shows why symmetry is the exact dividing line, and how to build and in practice.
2Singular Value Decomposition
Singular Values of a Matrix
Diagonalization is reserved for square, symmetric matrices — yet every matrix has singular values. The trick is to pass to the symmetric matrix , whose eigenvalues are always real and nonnegative (Lemma 9.3). Their square roots are the singular values (Definition 9.2); an orthonormal eigenbasis of sends the unit sphere to an ellipse with semi-axes , because (Theorem 9.2); and the number of nonzero is exactly the rank of (Proposition 9.3).
The Singular Value Decomposition
The Spectral Theorem factors a symmetric matrix as , but it is confined to square — indeed symmetric — matrices. The singular value decomposition lifts that restriction: every matrix, square or rectangular, factors as with and orthogonal and a diagonal matrix of singular values . This lesson builds the SVD from the eigenvectors of , rewrites it as a sum of rank-one pieces , and reads off its geometry: sends the unit sphere to an ellipsoid whose semi-axes are the singular values.
Module 6 · Probability Foundations
AlgebraProbability spaces and the axioms, the rules of probability and inclusion–exclusion, counting and sampling, conditional probability, Bayes’ formula, and independence.
1Probability Spaces & Rules
Probability Spaces: Sample Space, Events & Axioms
Every probability question begins with a probability space (Definition 10.1): a sample space holding every possible outcome , a collection of events (the subsets of we assign probabilities to), and a probability measure obeying three axioms — every probability lies in , the certain event has (and the impossible event ), and the probabilities of disjoint events add. This lesson translates experiments into the language of sets, unpacks the axioms and their first consequences (Remark 10.1), and — when is finite with equally likely outcomes — collapses probability into pure counting via (Proposition 10.1). We put the machinery to work on a fair die (Example 10.1), a pair of distinguishable dice (Example 10.2), and an urn draw (Example 10.3).
The Rules of Probability: Complement, Addition & Inclusion–Exclusion
Once the axioms are in place, almost every probability computation runs on a short list of rules, and this lesson assembles them. The complement rule turns a hard event into an easy one \u2014 the engine behind the complement trick , used in Example 10.4 to find the chance a value repeats when a fair die is rolled four times. For mutually exclusive (disjoint) events the Disjoint Sum Rule adds probabilities outright; when events overlap, the addition rule (the General Sum Rule) corrects for the double-counted intersection, . Remark 10.3 rearranges this to recover and extends it to the inclusion\u2013exclusion principle for three events, which (for two) we apply in Example 10.5.
Counting and the Four Sampling Methods
Probability on a finite, equally likely sample space collapses to counting: . This lesson organizes that counting around two yes/no questions about how a sample of size is drawn from objects --- does order matter? and is each object replaced before the next draw? --- giving a grid of four sampling schemes. Three of them produce equally likely outcomes and power every probability here: ordered-with-replacement (, Example 10.6), ordered-without-replacement (, Example 10.7), and unordered-without-replacement (, Example 10.8). We count each scheme, see how dividing an ordered count by produces combinations, and turn the counts into probabilities --- including Jordan's books solved two ways (Example 10.9, Remark 10.4) and the three contrasting scenarios of Example 10.10.
2Conditional Probability & Independence
Conditional Probability & the Law of Total Probability
New information reshapes a probability model: once we learn that an event has occurred, outcomes outside become impossible and the survivors must be renormalized so their probabilities again sum to . Definition 11.1 packages this as , valid whenever . Rearranging gives the multiplication rule — the engine behind sequential draws without replacement — and splitting into a partition (Definition 11.2) yields the Law of Total Probability (Proposition 11.1), , the backbone of the two-stage, two-urn reasoning in Example 11.4.
Bayes' Formula: Reversing the Order of Conditioning
Conditional probability answers ; Bayes' formula (Proposition 11.2) answers the reverse question -- the probability of a cause given an observed effect. Assembled from the definition of conditional probability (Definition 11.1) and the law of total probability (Proposition 11.1), it updates a prior belief into a posterior once evidence arrives (Remark 11.2). The centerpiece is the medical-test / base-rate problem (Example 11.6), where a highly accurate test for a rare disease returns a surprisingly low posterior -- the base-rate fallacy that routinely fools intuition.
Independence
Two events are independent when knowing one tells you nothing about the other: P(A∩B) = P(A)·P(B).
Module 7 · Random Variables & Distributions
AlgebraRandom variables and probability mass functions, continuous variables and densities, the cumulative distribution function, and the common distributions: Bernoulli, binomial, geometric, negative binomial, uniform, and normal.
1Random Variables & Distributions
Random Variables & the Probability Mass Function
A random variable (Definition 12.1) is not a number but a function that reads each outcome of an experiment and reports a real value — the sum of two dice, a gambler's change in wealth, a count. Writing turns a question about values into an event with a probability , and the whole family of these probabilities is the probability distribution of (Definition 12.2). This lesson concentrates on discrete random variables (Definition 12.3), whose values form a finite or countably infinite list that carries all the probability. For them the distribution collapses to a single object, the probability mass function (Definition 12.4): every probability becomes a sum of masses, , and a function is a legitimate p.m.f. exactly when it is nonnegative and sums to . We build p.m.f.s and read probabilities off them for a pair of dice (Examples 12.1, 12.3), a wealth/payoff game (Examples 12.2, 12.4), and a biased spinner. (Continuous random variables and densities are left for a later lesson.)
Continuous Random Variables & Probability Density Functions
A continuous random variable (Definition 12.5) draws its probabilities not from a list of point masses but from the area under a curve: there is a probability density function (pdf) with and more generally . This lesson pins down what makes a function a legitimate density (Remark 12.1: everywhere and total area ), reads probabilities off as areas , and draws the defining contrast with the discrete world — for a continuous variable every single point is negligible, (Proposition 12.1), so endpoints never matter: . We anchor everything in the uniform distribution on (Example 12.5), where on and a probability is just a length, and in a worked exponential density (Example 12.6).
The Cumulative Distribution Function:
Unlike the probability mass function (discrete variables only) or the density (continuous variables only), the cumulative distribution function (c.d.f.) is defined for every random variable (Definition 12.6): records the total probability accumulated up to the point — the 'sum-or-area-so-far' function. The '' is deliberate (Remark 12.2): it includes the endpoint and makes the c.d.f. deliver interval probabilities, . For a discrete variable the c.d.f. is a step function (Remark 12.3) that is flat between the possible values and jumps by at each one, as for the change-in-wealth variable of Example 12.7; for a continuous variable it rises smoothly, obtained by integrating the density, as for the variable of Example 12.8. Whatever its type, every c.d.f. shares the same three shape properties (Proposition 12.2): it is nondecreasing, right-continuous, and climbs from at to at .
2Common Distributions
The Bernoulli and Binomial Distributions
Repeated independent trials with just two outcomes — success (probability ) and failure (probability ) — are the seed from which two fundamental discrete distributions grow. The Bernoulli distribution (Definition 12.7) records a single such trial: a random variable with and , written . Run the trial independent times and you obtain independent Bernoulli variables ; their sum counts the successes, and its distribution is the Binomial (Definition 12.8): for , written . The binomial coefficient counts the arrangements of successes among the trials, each arrangement carrying probability , and the binomial theorem guarantees these probabilities sum to . We close with Example 12.9: the chance that five rolls of a fair die yield two or three sixes, with , equal to .
The Geometric & Negative Binomial Distributions
Independent repeated Bernoulli trials — each a success with probability , a failure with probability — raise a question the binomial does not answer: not how many successes in a fixed number of trials, but how long until a success. The geometric distribution (Definition 12.9) answers the first version: is the trial on which the first success occurs, with probability mass function for — the price of failures followed by one success. Its masses form a geometric series summing to , and its tail has a clean closed form (all of the first trials fail), so ; Example 12.10 uses this to find the chance of needing more than seven rolls of a fair die for the first six, . The negative binomial distribution (Definition 12.10) generalises the idea to the trial of the -th success: has for , the coefficient counting the arrangements of the earlier successes among the first trials. Remark 12.4 ties the two together: the geometric is exactly .
Continuous Distributions: The Uniform and the Standard Normal
These are the first two continuous distributions of the course, and both compute probabilities as areas under a density curve rather than by summing point masses. The continuous uniform distribution (Definition 12.11) spreads probability evenly over an interval : its density is the constant , so the chance of landing in a subinterval is just its length ratio, . The standard normal (Gaussian) distribution (Definition 12.12) is the famous bell curve, with density symmetric about . Its cumulative distribution function has no closed form, so we read from a standard-normal table or technology and assemble every probability from together with the symmetry rule . We close with Example 12.11, computing , and record the landmark values and .
Module 8 · Expectation & Variance
AlgebraExpectation and the law of the unconscious statistician, variance and standard deviation, the means and variances of the common distributions, and the normal, Poisson, and exponential distributions.
1Expectation & Variance
Expectation: The Mean of a Random Variable
The expectation (or mean) of a random variable is the single number that says where its probability is centred — a weighted average of the values, each weighted by its probability. For a discrete variable this is (Definition 13.1), also called the first moment and written ; rolling a fair die gives (Example 13.1) — already a sign that the mean need not be a value can actually take. For a continuous variable the sum becomes an integral against the density, (Definition 13.2), which for returns the midpoint (Example 13.7). To average a function of you do not need the distribution of : the law of the unconscious statistician (Proposition 13.1) gives , and the choice produces the moments (Definition 13.3). Expectation is powerful but not automatic: it is well-defined only when its sum or integral settles on a value, finite or (Remark 13.1), and it can be infinite (Examples 13.5, 13.8) or even undefined (Example 13.6). (The means of the named distributions — Bernoulli, binomial, geometric — are developed in a later lesson; here we keep them light.)
Variance & Standard Deviation
The mean locates a random variable; the variance measures how far it typically strays from that centre. For with mean , Definition 13.4 sets — the average squared deviation from the mean — written , with its square root the standard deviation restoring the original units. Squaring the deviation inside the expectation makes Proposition 13.2 a sum for a discrete variable, , but expanding that square yields the far handier computational formula of Proposition 13.3, — the mean of the square minus the square of the mean. From the definition flow the transformation rules of Proposition 13.4: a shift moves the centre but not the spread, while a scale factor comes out squared, (so ). Finally Proposition 13.5 reads off the extreme case — precisely when is almost surely constant. (The specific variances of the Bernoulli, binomial, uniform and geometric families, Examples 13.9–13.12, are collected in a separate lesson on the moments of named distributions.)
Mean and Variance of the Common Distributions
Two numbers summarise any distribution: the mean (Definition 13.1, the first moment), which says where the distribution is centred, and the variance (Definition 13.4), which says how widely it spreads about that centre; its square root is the standard deviation. For hand computation the variance is almost always found from the computational formula (Proposition 13.3). This lesson is a consolidated reference: it collects, once and for all, the mean and variance of the four distributions you meet most often, each derived in the course's own notes. For the discrete families these are the Bernoulli with mean (Example 13.2) and variance (Example 13.9); the binomial with mean (Example 13.3) and variance (Example 13.10); and the geometric (support , p.m.f. ) with mean (Example 13.4) and variance (Example 13.12). For the continuous world we record the uniform with mean (Example 13.7) and variance (Example 13.11). Once you can name the family and read off its parameters, every mean and variance is a one-line substitution.
2More Distributions
The Normal Distribution
The standard normal is the reference bell curve: it has mean and variance (Proposition 13.6), and its cumulative distribution function is written . Every other normal is an affine stretch-and-shift of it. For and , setting produces the general normal distribution (Definition 13.5), with mean , variance , and the bell-shaped density centred at with its spread set by . Affine maps keep you inside the family (Proposition 13.7): if then , and the special case — standardization — turns any normal back into the standard one. That is the computational workhorse: collapses every normal probability to a single -value read from a standard-normal table or software, as in Example 13.14 ( gives ). The same device yields the 68–95–99.7 rule: a normal variable falls within , , or standard deviations of its mean with probability about , , and (Example 13.15).
The Poisson & Exponential Distributions
Two distributions govern rare, randomly timed events, and this lesson treats them as a pair. The Poisson distribution (Definition 13.6) counts how many such events occur in a fixed window — calls at a switchboard, decays of a sample, requests to a server — when they happen independently at a constant average rate . A Poisson variable takes values with mass ; summing the series gives , so it is a genuine p.m.f., and a short computation (Proposition 13.8) shows the single parameter is both the mean and the variance, (Example 13.16: requests per second give ). The exponential distribution (Definition 13.7) is the continuous companion that measures the waiting time until the next such event. An Exp variable has density for , tail , c.d.f. , mean and variance (Examples 13.17–13.19). Its signature feature is the memoryless property (Proposition 13.9): — having already waited tells you nothing about how much longer you must wait (Example 13.20, the turtle and the highway).
Module 9 · Approximations & the CLT
AlgebraThe Central Limit Theorem and the normal approximation to the binomial (with continuity correction), the Poisson approximation and the law of rare events, choosing between them, and applications: confidence intervals and the law of large numbers.
1Normal approximation & the CLT
The CLT & the normal approximation
For large , a count behaves like a normal with the same mean and variance. That is the Central Limit Theorem at work — and it lets us estimate binomial probabilities with the standard normal.
Poisson approximation & the law of rare events
When successes are rare — large , tiny — the binomial count is approximately Poisson with mean . This is the law of rare events, with a clean error bound.
Normal vs. Poisson: choosing the approximation
Two approximations, two regimes. Moderate with large → normal. Tiny with small → Poisson. Picking the wrong one can be badly off.
Module 10 · Joint Distributions & Covariance
AlgebraJoint and marginal distributions of discrete random variables, independence and exchangeability, linearity of expectation and the variance of sums, and covariance and correlation — the bridge from probability back to linear algebra.
1Joint distributions
Joint distributions & marginals
The joint p.m.f. gives the probability of every combination of values of several random variables at once; summing it over one variable recovers the marginal distribution of the other.
Independence of random variables
Random variables are independent exactly when their joint p.m.f. factors into the product of the marginals — for every combination of values, not just one.
Exchangeable random variables
A sequence is exchangeable when reordering it doesn't change its distribution. That symmetry lets you compute positional probabilities — like ‘‘the 23rd card is a spade’’ — as if they were the first.
2Multivariate expectation & variance
Physics
Physics 1 · Classical Mechanics
PhysicsKinematics, Newton's laws, energy & momentum