PCA & Dimensionality Reduction
The curse of dimensionality
As you add more features to a dataset, something counterintuitive happens: the data doesn't become richer, it becomes sparser. In high dimensions, data points spread out so much that almost every pair of points ends up roughly equidistant from each other — the notion of "nearest neighbor" starts to lose meaning, and distance-based algorithms like K-Means or KNN degrade badly.
This is the curse of dimensionality. Each additional feature exponentially increases the volume of the space that needs to be covered by data to remain densely sampled — a dataset that's dense in 2 dimensions becomes vanishingly sparse in 100 dimensions unless it grows exponentially in size to match. Beyond hurting algorithm performance, high dimensionality also means more noise, more redundant or correlated features, and more compute and storage cost. Dimensionality reduction techniques address this by compressing many features into fewer, more informative ones.
PCA: finding the axes of maximum variance
Principal Component Analysis (PCA) is the classical dimensionality reduction technique. Its goal: find a new set of axes — linear combinations of the original features — ordered by how much variance in the data they capture, then keep only the first few.
The first principal component is the direction in feature space along which the data varies the most. The second principal component is the direction of next-highest variance, constrained to be orthogonal (perpendicular) to the first. This continues, with each subsequent component capturing the most remaining variance while staying orthogonal to all previous ones.
Because the components are ordered by how much information (variance) they capture, you can often represent a dataset's essential structure with a small number of components — sometimes just 2 or 3 — even when it originally had hundreds of features.
The linear algebra underneath
PCA is computed from the data's covariance matrix, which captures how every pair of features varies together. The principal components turn out to be the eigenvectors of this covariance matrix, and each eigenvector's corresponding eigenvalue tells you how much variance lies along that direction.
Intuitively: an eigenvector of a matrix is a direction that the matrix doesn't rotate, only stretches or shrinks — and the eigenvalue is the amount of that stretching. Applied to a covariance matrix, the eigenvectors point along the natural axes of spread in the data, and the eigenvalues measure how much the data spreads along each one. Sorting eigenvectors by eigenvalue, largest first, gives you the principal components in order of importance. This is why PCA is sometimes introduced purely as an eigendecomposition problem — that's literally the computation happening under the hood.
Note
Note
What PCA is used for
Visualization: reducing data to 2 or 3 dimensions so it can be plotted and inspected visually, revealing clusters or outliers that aren't obvious in the raw high-dimensional form.
Preprocessing: feeding a lower-dimensional, decorrelated representation into a downstream model, which can speed up training and sometimes reduce overfitting by removing redundant, correlated features.
Compression and noise reduction: discarding low-variance components often removes noise along with dimensionality, since noise tends to be spread thinly across many directions while true signal concentrates in a few.
Anomaly detection: points that reconstruct poorly from their low-dimensional PCA representation (high reconstruction error) are often anomalies — the compression that works well for typical data fails for unusual data.
When PCA isn't enough: t-SNE and UMAP
PCA is a linear technique — it can only find straight-line axes of variance. Many real datasets have important structure that's fundamentally nonlinear: think of data that lies on a curved manifold, like a Swiss roll shape. PCA will flatten this poorly, mixing points that were actually far apart on the manifold.
t-SNE (t-distributed Stochastic Neighbor Embedding) and UMAP (Uniform Manifold Approximation and Projection) are nonlinear dimensionality reduction techniques built specifically for visualization. Rather than preserving global variance like PCA, they focus on preserving local neighborhood structure — points that are close together in high-dimensional space stay close in the 2D or 3D projection, even if the global geometry gets distorted in the process.
UMAP has largely superseded t-SNE for most practical use because it's significantly faster, scales better to large datasets, and better preserves some global structure alongside local structure. Neither is typically used as a preprocessing step for downstream models the way PCA is — they're primarily visualization and exploratory analysis tools, and the axes they produce don't have the direct interpretability of PCA's variance-ranked components.
PCA vs. t-SNE vs. UMAP
| Technique | Linear/Nonlinear | Preserves | Speed | Typical use |
|---|---|---|---|---|
| PCA | Linear | Global variance structure | Fast, scales to huge data | Preprocessing, compression, quick visualization |
| t-SNE | Nonlinear | Local neighborhood structure | Slow, struggles beyond ~10k points | High-quality 2D visualization of clusters |
| UMAP | Nonlinear | Local + some global structure | Much faster than t-SNE | Visualization at scale, exploratory analysis |
See clustering structure that dimensionality reduction reveals
Dimensionality reduction and clustering are often used together — reduce to 2-3 dimensions with PCA or UMAP, then cluster the result. This K-Means visualization shows how cluster structure appears once data is projected into a low-dimensional space where groups become visually separable.
📊 K-Means Clustering Simulator
InteractiveWhat's next
PCA reduces dimensionality by finding variance-maximizing linear combinations of raw features. Modern deep learning takes a related but more powerful approach: learning dense embeddings directly from data via a neural network, optimized for a specific downstream task rather than variance alone. See the Embeddings lesson for how that works.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.