Hyperparameter Tuning
Parameters vs. hyperparameters
A model's parameters are learned automatically during training — the weights in a neural network, the coefficients in a linear regression. Its hyperparameters are the settings you choose before training even starts: learning rate, regularization strength, number of layers, tree depth, batch size, dropout rate.
Hyperparameters control how learning happens, not what is learned. Get them wrong and even a well-designed model architecture will underperform, take far longer to train, or fail to converge at all. Hyperparameter tuning is the systematic search for the combination of settings that produces the best validation performance.
Note
Grid search: exhaustive and expensive
Grid search defines a discrete set of values for each hyperparameter and evaluates every combination. If you're tuning learning rate over 5 values and regularization strength over 4 values, grid search trains and evaluates 20 models.
The appeal is completeness — you know you've checked every combination in your grid. The problem is that cost grows exponentially with the number of hyperparameters. Tuning 5 hyperparameters with 5 values each means 5^5 = 3,125 model trainings. For anything beyond 2-3 hyperparameters with small grids, this becomes computationally impractical.
Random search: better coverage for less cost
Random search samples hyperparameter combinations randomly from specified distributions, for a fixed number of iterations, rather than exhaustively covering a grid. Counterintuitively, this usually outperforms grid search for the same compute budget.
The reason is dimensionality: in most problems, only a couple of hyperparameters meaningfully affect performance, while others barely matter. Grid search wastes evaluations varying the unimportant ones across every value of the important ones. Random search, by sampling independently, is far more likely to land on a good value for the hyperparameters that actually matter, because every trial explores a fresh combination rather than repeating fixed grid lines.
Bayesian optimization: search informed by past results
Both grid and random search treat each trial independently — they never learn from previous evaluations. Bayesian optimization does. It builds a probabilistic model (commonly a Gaussian Process, or a tree-based surrogate as in TPE) of how hyperparameters map to validation performance, and uses that model to decide which combination to try next.
The process balances exploration (trying regions of the hyperparameter space it's uncertain about) with exploitation (refining regions that already look promising), using an acquisition function to score candidates. This makes it dramatically more sample-efficient than random search — it can often find a near-optimal configuration in a fraction of the trials, which matters enormously when each trial means training a large model for hours.
Popular libraries implementing this include Optuna, Hyperopt, and scikit-optimize. Bayesian optimization is the default approach for tuning expensive deep learning models where every trial has a real dollar cost.
Search strategies compared
| Strategy | Sample efficiency | Handles many hyperparameters | Parallelizable | Best for |
|---|---|---|---|---|
| Grid search | Low | Poorly — exponential blowup | Fully | 1-2 cheap hyperparameters, need exhaustive coverage |
| Random search | Medium | Well | Fully | Cheap-to-moderate models, several hyperparameters |
| Bayesian optimization | High | Well | Partially — sequential by nature | Expensive models, deep learning, limited trial budget |
Learning rate schedules
For neural networks, the learning rate is often the single most consequential hyperparameter, and the best practice isn't to pick one fixed value — it's to schedule it. A learning rate schedule changes the learning rate over the course of training.
Common patterns: step decay drops the rate by a factor every N epochs. Cosine annealing smoothly decreases the rate following a cosine curve, often used with periodic "warm restarts." Warmup starts with a very small learning rate for the first few hundred steps before ramping up — this stabilizes training early on, when weights are randomly initialized and large updates can be destructive, and is standard practice when training transformers.
The intuition behind decaying the rate over time: large steps early in training make fast progress across the broad loss landscape, while small steps late in training allow fine-grained convergence into a sharp minimum without overshooting it.
See why step size matters
This gradient descent visualization lets you observe convergence behavior directly. A learning rate that's too large causes the parameter to oscillate or diverge; too small and convergence crawls. Learning rate schedules exist precisely to get the benefits of a large rate early and a small rate late, within a single training run.
📉 Gradient Descent Optimization
InteractiveNote
What's next
Hyperparameter tuning optimizes settings for a model architecture you've already chosen. For some problems — especially small, structured datasets — the right architecture choice matters more than any amount of tuning. The next lesson, Support Vector Machines, covers a classical algorithm that still outperforms neural networks on small, high-dimensional data.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.