AI & MLDeep
Advanced

PCA & Dimensionality Reduction

12 min read

Learn
Deep Reading
Estimated 12 mins
Prereq
Advanced
Requires advanced math/coding
Interactive
Live Simulator
Contains interactive viz

The curse of dimensionality

As you add more features to a dataset, something counterintuitive happens: the data doesn't become richer, it becomes sparser. In high dimensions, data points spread out so much that almost every pair of points ends up roughly equidistant from each other — the notion of "nearest neighbor" starts to lose meaning, and distance-based algorithms like K-Means or KNN degrade badly.

This is the curse of dimensionality. Each additional feature exponentially increases the volume of the space that needs to be covered by data to remain densely sampled — a dataset that's dense in 2 dimensions becomes vanishingly sparse in 100 dimensions unless it grows exponentially in size to match. Beyond hurting algorithm performance, high dimensionality also means more noise, more redundant or correlated features, and more compute and storage cost. Dimensionality reduction techniques address this by compressing many features into fewer, more informative ones.

PCA: finding the axes of maximum variance

Principal Component Analysis (PCA) is the classical dimensionality reduction technique. Its goal: find a new set of axes — linear combinations of the original features — ordered by how much variance in the data they capture, then keep only the first few.

The first principal component is the direction in feature space along which the data varies the most. The second principal component is the direction of next-highest variance, constrained to be orthogonal (perpendicular) to the first. This continues, with each subsequent component capturing the most remaining variance while staying orthogonal to all previous ones.

Because the components are ordered by how much information (variance) they capture, you can often represent a dataset's essential structure with a small number of components — sometimes just 2 or 3 — even when it originally had hundreds of features.

The linear algebra underneath

PCA is computed from the data's covariance matrix, which captures how every pair of features varies together. The principal components turn out to be the eigenvectors of this covariance matrix, and each eigenvector's corresponding eigenvalue tells you how much variance lies along that direction.

Intuitively: an eigenvector of a matrix is a direction that the matrix doesn't rotate, only stretches or shrinks — and the eigenvalue is the amount of that stretching. Applied to a covariance matrix, the eigenvectors point along the natural axes of spread in the data, and the eigenvalues measure how much the data spreads along each one. Sorting eigenvectors by eigenvalue, largest first, gives you the principal components in order of importance. This is why PCA is sometimes introduced purely as an eigendecomposition problem — that's literally the computation happening under the hood.

python

Note

A common workflow: plot cumulative variance explained against number of components, and keep however many are needed to retain 90-95% of total variance. This is PCA's version of the elbow method — it lets you compress the data as much as possible while quantifying exactly how much information you're discarding.

Note

PCA finds directions of maximum variance, and variance is scale-dependent. A feature measured in thousands (income) will dominate a feature measured in single digits (years of education) purely due to units, not actual importance. Standardizing all features to zero mean and unit variance before running PCA is essential — skipping this step is one of the most common PCA mistakes.

What PCA is used for

Visualization: reducing data to 2 or 3 dimensions so it can be plotted and inspected visually, revealing clusters or outliers that aren't obvious in the raw high-dimensional form.

Preprocessing: feeding a lower-dimensional, decorrelated representation into a downstream model, which can speed up training and sometimes reduce overfitting by removing redundant, correlated features.

Compression and noise reduction: discarding low-variance components often removes noise along with dimensionality, since noise tends to be spread thinly across many directions while true signal concentrates in a few.

Anomaly detection: points that reconstruct poorly from their low-dimensional PCA representation (high reconstruction error) are often anomalies — the compression that works well for typical data fails for unusual data.

When PCA isn't enough: t-SNE and UMAP

PCA is a linear technique — it can only find straight-line axes of variance. Many real datasets have important structure that's fundamentally nonlinear: think of data that lies on a curved manifold, like a Swiss roll shape. PCA will flatten this poorly, mixing points that were actually far apart on the manifold.

t-SNE (t-distributed Stochastic Neighbor Embedding) and UMAP (Uniform Manifold Approximation and Projection) are nonlinear dimensionality reduction techniques built specifically for visualization. Rather than preserving global variance like PCA, they focus on preserving local neighborhood structure — points that are close together in high-dimensional space stay close in the 2D or 3D projection, even if the global geometry gets distorted in the process.

UMAP has largely superseded t-SNE for most practical use because it's significantly faster, scales better to large datasets, and better preserves some global structure alongside local structure. Neither is typically used as a preprocessing step for downstream models the way PCA is — they're primarily visualization and exploratory analysis tools, and the axes they produce don't have the direct interpretability of PCA's variance-ranked components.

PCA vs. t-SNE vs. UMAP

TechniqueLinear/NonlinearPreservesSpeedTypical use
PCALinearGlobal variance structureFast, scales to huge dataPreprocessing, compression, quick visualization
t-SNENonlinearLocal neighborhood structureSlow, struggles beyond ~10k pointsHigh-quality 2D visualization of clusters
UMAPNonlinearLocal + some global structureMuch faster than t-SNEVisualization at scale, exploratory analysis

See clustering structure that dimensionality reduction reveals

Dimensionality reduction and clustering are often used together — reduce to 2-3 dimensions with PCA or UMAP, then cluster the result. This K-Means visualization shows how cluster structure appears once data is projected into a low-dimensional space where groups become visually separable.

📊 K-Means Clustering Simulator

Interactive
Points: 0Iteration: 0Inertia: N/A

What's next

PCA reduces dimensionality by finding variance-maximizing linear combinations of raw features. Modern deep learning takes a related but more powerful approach: learning dense embeddings directly from data via a neural network, optimized for a specific downstream task rather than variance alone. See the Embeddings lesson for how that works.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.