Principal Component Analysis
PCA is this module's capstone: it takes a cloud of high-dimensional data and finds the few directions that matter. Center the features into a matrix , form the covariance matrix , and — because is symmetric — use the Spectral Theorem (or the SVD) to diagonalize it, . The orthonormal eigenvectors are the principal components; their eigenvalues are the variances captured along each one. Keeping the components with the largest eigenvalues compresses the data with minimal loss — turning a point cloud in into a picture in .
Before you start — give these a try
Attempting first primes your brain for the lesson — even if you miss. Nothing is graded or saved; it's just a warm-up.
Two features are recorded for subjects; with subjects as columns the raw data is . Center each feature and compute the covariance matrix (divide by ). Enter as a matrix.
For the covariance matrix , list its eigenvalues — the variances carried by the two principal components — in decreasing order.
What you’ll be able to do
- Build the data matrix of a PCA — objects in the columns, features in the rows — then center each feature (and, when scales differ, divide by its standard deviation) to obtain the normalized matrix (Remark 15.9, steps 1–2).
- Form the covariance matrix and read its entries as sample variances (diagonal) and covariances (off-diagonal), knowing it is symmetric and positive semidefinite (Definition 15.7, §15.5).
- Apply the Spectral Theorem (Theorem 9.1) to diagonalize , identify the principal components as the orthonormal eigenvectors (columns of ), and explain why they are determined only up to sign.
- Interpret each eigenvalue as the variance along its principal component, compute the total variance and the proportion of variance explained , and see that the change of basis makes the new features uncorrelated ().
- Perform dimensionality reduction by keeping the components with the largest eigenvalues and discarding near-zero ones (redundant features), and recognize the equivalent SVD route with (Theorem 9.3).
In your course
· MATH2015 · Linear Algebra & Probability- Definition 15.7Covariance, with ; assembling all pairwise sample covariances of the centered features gives the covariance matrix .
- Proposition 15.17Properties of covariance (bilinearity)Covariance is symmetric and bilinear: and — the reason is symmetric (Remark 15.7 notes the analogy with inner products).
- Remark 15.8The key step of PCAThe central mathematical step is a change of basis that makes the features uncorrelated; one then focuses on the principal components to understand the data's structure.
- Remark 15.9The PCA procedureThe seven-step recipe: build the feature matrix, normalize to , form , eigen-decompose (sorted eigenvalues, unit eigenvectors as columns of an orthogonal ), transform , keep the leading components, and interpret.
- Theorem 9.1Spectral TheoremA real symmetric matrix has an orthonormal basis of eigenvectors: with orthogonal and diagonal. This is what makes the covariance matrix orthogonally diagonalizable in step 4, and §15.6 cites it explicitly.
- Theorem 9.3Singular value decomposition; for centered the left singular vectors are the principal components and , giving PCA without forming .
Centering (and scaling) the data: the matrix $A$
PCA starts from raw measurements and strips away the parts that carry no shape information. Arrange the data as a matrix whose columns are the objects (subjects, samples) and whose rows are the features (Remark 15.9, step 1). The first normalization step is centering: from each feature (row) subtract its sample mean, so every row of the resulting matrix has mean . If the features live on very different scales (heights in cm versus dimensionless ratios, say), also divide each centered row by its sample standard deviation , giving rows of unit variance — the standardized version, which makes a correlation matrix. The normalized data is the matrix . Centering is not optional: covariance is measured about the mean, and skipping it folds the mean vector into the first principal component and corrupts every eigenvalue.
The covariance matrix $C=\tfrac{1}{n-1}AA^{\mathsf T}$
To see how features move together we build the covariance matrix for features (Remark 15.9, step 3). Its entry is exactly the sample covariance of features and (Definition 15.7, §15.5), and the diagonal entries are the feature variances. Dividing by rather than is the unbiased sample convention used in the course. Two structural facts drive everything that follows. First, is symmetric, , because — covariance is symmetric and bilinear, behaving like an inner product on centered features (Proposition 15.17, Remark 15.7). Second, is positive semidefinite, so all its eigenvalues are — reassuring, since they will turn out to be variances. A large signals redundancy between two features; a near-zero one means they are linearly unrelated.
Diagonalizing $C$: the principal components
The whole point of PCA is to change coordinates so the features become uncorrelated, i.e. so the covariance matrix becomes diagonal (Remark 15.8). Because is real and symmetric, the Spectral Theorem (Theorem 9.1) guarantees an orthonormal basis of eigenvectors and an orthogonal matrix — its columns the eigenvectors — with Rewrite the data in the new basis via (Remark 15.9, step 5; here since is orthogonal). The covariance matrix of the new features is which is diagonal: every covariance is now . The principal components are these orthonormal eigenvectors of (the columns of ), sorted so that . They are orthogonal by construction but not unique: an eigenvector is defined only up to sign (and, within a repeated-eigenvalue subspace, up to rotation), so and are equally valid principal components.
Eigenvalues as variance; dimensionality reduction (and the SVD route)
In the diagonalized covariance matrix , the eigenvalue is the variance of the data along the -th principal component. This gives a clean budget for information: the total variance is and the proportion of variance explained by component is . Dimensionality reduction keeps the components with the largest eigenvalues and discards the smallest: if are tiny, the matching new features barely vary across subjects and carry almost no information, so projecting onto the top components (the first rows of ) loses very little (Remark 15.9, steps 6–7). A near-zero eigenvalue flags a redundant feature — in the course's example the covariance matrix of height, weight and BMI has an eigenvalue because BMI is determined by the other two. The same principal directions come directly from the singular value decomposition (Theorem 9.3): if , the left singular vectors (columns of ) are the principal components and . Working from the SVD of avoids forming explicitly and is the numerically preferred route.
For random variables , , and . For centered sample data stored in (features objects), the covariance matrix is , whose entry is the sample covariance of features and .
Every real symmetric matrix has an orthonormal basis of eigenvectors with real eigenvalues; equivalently there is an orthogonal matrix () and a diagonal with .
(1) Build the feature matrix (objects in columns, features in rows). (2) Normalize: subtract each row's mean, and if the row variances differ greatly divide by the row's standard deviation, giving . (3) Compute . (4) Find the eigenvalues and eigenvectors of ; sort the eigenvalues decreasingly and normalize the eigenvectors to length , placing them in the columns of the orthogonal matrix . (5) Transform: . (6) From the eigenvalue sizes decide how many leading components to keep — the principal components are the first rows of . (7) Interpret them.
Every matrix factors as with orthogonal and diagonal carrying the nonnegative singular values .
Worked examples
A tiny study records two features for subjects. With subjects as columns, the raw data matrix is (row 1 feature , row 2 feature ). Center each feature and compute the covariance matrix (sample convention, divide by ).
- 1
Row means (Remark 15.9, step 2): and .
- 2
Subtract each row mean to center: ; each row now has mean .
- 3
Form : the entry is , the entry is , and the off-diagonal is , so .
- 4
Divide by : . The diagonal gives the sample variances ( each) and the off-diagonal the covariance (Definition 15.7).
Continuing with , use the Spectral Theorem to find the variances carried by the principal components, the total variance, and the proportion of variance explained by the first principal component. What are the principal components themselves?
Three features are measured on subjects; the already-centered data matrix is , where row 3 is exactly row 1 row 2. Compute , find its eigenvalues, and decide how many principal components to keep to capture at least of the variance.