Naive Bayes Classifier
A classifier built on Bayes' theorem
Naive Bayes takes a fundamentally different approach to classification than logistic regression or SVMs. Instead of learning a decision boundary directly, it models how likely each class is to have *generated* the observed features, then picks the class with the highest posterior probability.
The foundation is Bayes' theorem:
$$P(y \mid x) = \frac{P(x \mid y) \, P(y)}{P(x)}$$
In words: the probability of class $y$ given features $x$ (what we want) equals the probability of seeing those features if $y$ were true, times how common $y$ is overall, divided by how common those features are in general. Since $P(x)$ is the same for every class we're comparing, classification just requires finding the $y$ that maximizes $P(x \mid y) \, P(y)$.
The "naive" assumption
Computing $P(x \mid y)$ directly is intractable for anything beyond a handful of features — it would require knowing the joint probability of every feature combination for every class, which needs exponentially more data than you'll ever have.
Naive Bayes makes a bold simplification: it assumes every feature is conditionally independent of every other feature, given the class. That is, once you know the class, knowing the value of one feature tells you nothing about the value of another. This is almost never literally true — in a spam email, the word "free" and the word "discount" are correlated, not independent, even after you know it's spam. But the assumption makes the math trivial: instead of one enormous joint probability, you multiply together simple, individually estimable per-feature probabilities:
$$P(x_1, x_2, \dots, x_n \mid y) \approx \prod_{i=1}^{n} P(x_i \mid y)$$
Remarkably, this simplification barely hurts classification accuracy in practice, even when the independence assumption is clearly violated — the model just needs to rank classes correctly, not estimate calibrated probabilities.
Note
Spam filtering: the canonical example
Naive Bayes became famous through early spam filters, and the mechanics show why it fits the problem so well. Each email is represented as a bag of words. For each word, the model estimates $P(\text{word} \mid \text{spam})$ and $P(\text{word} \mid \text{not spam})$ from training data — simply the frequency of that word in spam versus legitimate emails.
To classify a new email, multiply together the per-word probabilities for every word it contains, under both the spam and not-spam hypotheses, then compare. An email full of words disproportionately common in spam ("viagra," "winner," "free," "urgent") accumulates a high spam score; ordinary vocabulary keeps the not-spam score competitive.
This approach is fast to train (just counting), fast to predict (just multiplying), and works well even with a modest amount of labeled data — properties that made it the default spam-filtering technique for years before more sophisticated methods took over.
Note
alpha parameter in scikit-learn's implementation, and it should almost always be left on.Variants for different data types
"Naive Bayes" is really a family of models, differing in what distribution they assume each feature follows given the class. Multinomial Naive Bayes assumes features are counts (word frequencies) — the standard choice for text classification. Bernoulli Naive Bayes assumes binary features (word present or absent, ignoring frequency) — useful for shorter texts where presence matters more than count. Gaussian Naive Bayes assumes continuous features follow a normal distribution within each class — used for numeric, tabular data rather than text.
Naive Bayes vs. Logistic Regression
| Aspect | Naive Bayes | Logistic Regression |
|---|---|---|
| Model type | Generative (models how data is produced) | Discriminative (models the boundary directly) |
| Training speed | Very fast — closed-form counting | Slower — iterative optimization |
| Data efficiency | Strong with very little data | Needs more data to match NB's small-data performance |
| Handles feature correlation | Poorly — assumes independence | Well — learns feature interactions implicitly |
| Probability calibration | Often poorly calibrated | Well calibrated by design |
| Typical use case | Text classification, spam filtering, quick baselines | Tabular classification, interpretable coefficients |
Why it's still worth using in 2026
Despite decades of more sophisticated alternatives — logistic regression with TF-IDF, gradient-boosted trees, and now transformer-based text classifiers — Naive Bayes remains a legitimate production choice for high-volume, latency-sensitive text classification, and an excellent first baseline for any new classification problem. It trains in milliseconds on datasets that would take a neural network hours, requires no GPU, and its errors are easy to inspect and explain, since each class's decision is just a sum of interpretable per-word contributions.
What's next
Naive Bayes works directly on raw word counts, but modern NLP systems usually first convert text into dense embeddings that capture meaning rather than mere word presence. If you haven't already, the Embeddings lesson covers how that representation works and why it improves on bag-of-words approaches for most tasks beyond simple filtering.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.