Support Vector Machines (SVM)
Finding the best boundary, not just a boundary
Logistic regression finds a decision boundary that separates classes, but it doesn't ask whether that boundary is the *best possible* one — any line that separates the training data reasonably well will do. Support Vector Machines (SVMs) ask a sharper question: of all the boundaries that separate the classes, which one leaves the largest possible gap on either side?
That gap is called the margin, and the boundary that maximizes it is the max-margin classifier. The intuition is robustness: a boundary that barely squeezes between the two classes is fragile — a new point close to either class could easily land on the wrong side. A boundary with maximum breathing room on both sides is far more likely to classify new, unseen points correctly.
Support vectors: the points that matter
Once an SVM finds the max-margin boundary, most training points turn out to be irrelevant to it — you could delete them and the boundary wouldn't move. What determines the boundary are only the points closest to it on each side, the ones that "touch" the margin. These are the support vectors, and they give the algorithm its name.
This is a structurally different way of learning than most algorithms: instead of fitting to every data point, an SVM's final model is defined entirely by a small subset of the most informative, borderline examples. This is part of why SVMs generalize well even with limited data — the model complexity is tied to the number of support vectors, not the total dataset size.
Soft margins: allowing some mistakes
Real data is rarely perfectly separable — a few points usually fall on the wrong side of any reasonable boundary, whether from noise or genuine overlap between classes. A hard margin SVM, which requires zero violations, would either fail to converge or find a tiny, overfit margin squeezed around outliers.
The practical solution is the soft margin: allow some points to violate the margin (or even be misclassified), but penalize each violation. The hyperparameter C controls the tradeoff. A large C penalizes violations heavily, producing a narrower margin that fits the training data closely — more prone to overfitting. A small C tolerates more violations in exchange for a wider, more generalizable margin. Tuning C is usually the single most important step in fitting an SVM.
The kernel trick: solving nonlinear problems
So far this describes a *linear* classifier — it can only separate classes with a straight line (or flat hyperplane). Many real datasets aren't linearly separable at all: imagine one class forming a ring around another. No straight line can separate them.
The classical fix would be to manually transform the data into a higher-dimensional space where it becomes linearly separable — for the ring example, adding a feature like $x_1^2 + x_2^2$ would let a flat plane separate the ring from the center. The problem is that useful transformations can require enormous, even infinite-dimensional feature spaces, which is computationally infeasible to construct directly.
The kernel trick sidesteps this. SVM training and prediction only ever need the *dot product* between pairs of transformed points — never the transformed points themselves. A kernel function computes that dot product directly in the original space, without ever explicitly constructing the high-dimensional representation. This makes it possible to work in extremely high (or infinite) dimensional feature spaces at essentially the same computational cost as the original one.
Common SVM kernels
| Kernel | Formula intuition | Good for | Key hyperparameter |
|---|---|---|---|
| Linear | Standard dot product, no transformation | High-dimensional, already-linear data (e.g. text with TF-IDF) | C |
| Polynomial | Dot product raised to a power | Data with interaction effects between features | degree, C |
| RBF (Gaussian) | Similarity decays smoothly with distance | General-purpose nonlinear boundaries, most common default | gamma, C |
| Sigmoid | Behaves like a neural network activation | Rarely used; niche cases resembling shallow neural nets | gamma, C |
Note
gamma determines how much a single training example influences the decision boundary near it. A small gamma means influence reaches far, producing smoother, simpler boundaries. A large gamma restricts influence to a tight neighborhood, producing highly flexible boundaries that can overfit. C and gamma are typically tuned together via grid or random search.When SVMs beat neural networks
Deep learning dominates headlines, but SVMs remain a genuinely strong choice in specific conditions. Small datasets (hundreds to low thousands of examples) are where SVMs shine — neural networks need substantial data to learn useful representations and regularize away overfitting, while an SVM's margin-maximizing objective is inherently resistant to overfitting even with limited examples.
High-dimensional, low-sample-count data — think gene expression data with 20,000 features and 200 patients — is a classic SVM stronghold, particularly with a linear kernel. Clean, well-structured tabular data without the spatial or sequential structure that convolutional or transformer architectures are built to exploit also tends to favor SVMs and other classical methods (along with gradient-boosted trees) over deep learning.
The tradeoffs: SVMs scale poorly to very large datasets (training complexity is roughly quadratic to cubic in the number of samples), they don't naturally extend to multi-class problems (requiring one-vs-rest or one-vs-one schemes), and they lack the automatic feature learning that makes deep learning so effective on raw images, audio, and text.
Note
What's next
SVMs are discriminative — they learn a boundary directly, without modeling how the data was generated. Naive Bayes takes the opposite approach: it models the probability of each class generating the observed features, and remains remarkably effective for text classification despite a famously unrealistic assumption. That's the next lesson.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.