AI & MLDeep
Intermediate

Naive Bayes Classifier

10 min read

Learn
Deep Reading
Estimated 10 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Static Playbook
Static guide & reference tables

A classifier built on Bayes' theorem

Naive Bayes takes a fundamentally different approach to classification than logistic regression or SVMs. Instead of learning a decision boundary directly, it models how likely each class is to have *generated* the observed features, then picks the class with the highest posterior probability.

The foundation is Bayes' theorem:

$$P(y \mid x) = \frac{P(x \mid y) \, P(y)}{P(x)}$$

In words: the probability of class $y$ given features $x$ (what we want) equals the probability of seeing those features if $y$ were true, times how common $y$ is overall, divided by how common those features are in general. Since $P(x)$ is the same for every class we're comparing, classification just requires finding the $y$ that maximizes $P(x \mid y) \, P(y)$.

The "naive" assumption

Computing $P(x \mid y)$ directly is intractable for anything beyond a handful of features — it would require knowing the joint probability of every feature combination for every class, which needs exponentially more data than you'll ever have.

Naive Bayes makes a bold simplification: it assumes every feature is conditionally independent of every other feature, given the class. That is, once you know the class, knowing the value of one feature tells you nothing about the value of another. This is almost never literally true — in a spam email, the word "free" and the word "discount" are correlated, not independent, even after you know it's spam. But the assumption makes the math trivial: instead of one enormous joint probability, you multiply together simple, individually estimable per-feature probabilities:

$$P(x_1, x_2, \dots, x_n \mid y) \approx \prod_{i=1}^{n} P(x_i \mid y)$$

Remarkably, this simplification barely hurts classification accuracy in practice, even when the independence assumption is clearly violated — the model just needs to rank classes correctly, not estimate calibrated probabilities.

Note

Naive Bayes is a textbook example of a model being useful despite a false premise. Because classification only requires comparing $P(x \mid y) P(y)$ across classes rather than computing exact probabilities, systematic errors from the independence assumption often affect every class similarly and cancel out at decision time. The ranking survives even when the magnitudes are wrong.

Spam filtering: the canonical example

Naive Bayes became famous through early spam filters, and the mechanics show why it fits the problem so well. Each email is represented as a bag of words. For each word, the model estimates $P(\text{word} \mid \text{spam})$ and $P(\text{word} \mid \text{not spam})$ from training data — simply the frequency of that word in spam versus legitimate emails.

To classify a new email, multiply together the per-word probabilities for every word it contains, under both the spam and not-spam hypotheses, then compare. An email full of words disproportionately common in spam ("viagra," "winner," "free," "urgent") accumulates a high spam score; ordinary vocabulary keeps the not-spam score competitive.

This approach is fast to train (just counting), fast to predict (just multiplying), and works well even with a modest amount of labeled data — properties that made it the default spam-filtering technique for years before more sophisticated methods took over.

python

Note

If a word never appeared in the spam training examples, its raw estimate for $P(\text{word} \mid \text{spam})$ would be exactly zero — and since probabilities multiply, a single unseen word would zero out the entire spam score regardless of every other word. Laplace (additive) smoothing adds a small constant count to every word-class pair before computing probabilities, ensuring no probability is ever exactly zero. This is the alpha parameter in scikit-learn's implementation, and it should almost always be left on.

Variants for different data types

"Naive Bayes" is really a family of models, differing in what distribution they assume each feature follows given the class. Multinomial Naive Bayes assumes features are counts (word frequencies) — the standard choice for text classification. Bernoulli Naive Bayes assumes binary features (word present or absent, ignoring frequency) — useful for shorter texts where presence matters more than count. Gaussian Naive Bayes assumes continuous features follow a normal distribution within each class — used for numeric, tabular data rather than text.

Naive Bayes vs. Logistic Regression

AspectNaive BayesLogistic Regression
Model typeGenerative (models how data is produced)Discriminative (models the boundary directly)
Training speedVery fast — closed-form countingSlower — iterative optimization
Data efficiencyStrong with very little dataNeeds more data to match NB's small-data performance
Handles feature correlationPoorly — assumes independenceWell — learns feature interactions implicitly
Probability calibrationOften poorly calibratedWell calibrated by design
Typical use caseText classification, spam filtering, quick baselinesTabular classification, interpretable coefficients

Why it's still worth using in 2026

Despite decades of more sophisticated alternatives — logistic regression with TF-IDF, gradient-boosted trees, and now transformer-based text classifiers — Naive Bayes remains a legitimate production choice for high-volume, latency-sensitive text classification, and an excellent first baseline for any new classification problem. It trains in milliseconds on datasets that would take a neural network hours, requires no GPU, and its errors are easy to inspect and explain, since each class's decision is just a sum of interpretable per-word contributions.

What's next

Naive Bayes works directly on raw word counts, but modern NLP systems usually first convert text into dense embeddings that capture meaning rather than mere word presence. If you haven't already, the Embeddings lesson covers how that representation works and why it improves on bag-of-words approaches for most tasks beyond simple filtering.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.