AI & MLDeep
Intermediate

Hyperparameter Tuning

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Live Simulator
Contains interactive viz

Parameters vs. hyperparameters

A model's parameters are learned automatically during training — the weights in a neural network, the coefficients in a linear regression. Its hyperparameters are the settings you choose before training even starts: learning rate, regularization strength, number of layers, tree depth, batch size, dropout rate.

Hyperparameters control how learning happens, not what is learned. Get them wrong and even a well-designed model architecture will underperform, take far longer to train, or fail to converge at all. Hyperparameter tuning is the systematic search for the combination of settings that produces the best validation performance.

Note

Every hyperparameter search evaluates candidates on a validation set (or via cross-validation folds), not the held-out test set. If you pick hyperparameters based on test performance, the test set stops measuring generalization — it just becomes another thing you optimized against.

Grid search: exhaustive and expensive

Grid search defines a discrete set of values for each hyperparameter and evaluates every combination. If you're tuning learning rate over 5 values and regularization strength over 4 values, grid search trains and evaluates 20 models.

The appeal is completeness — you know you've checked every combination in your grid. The problem is that cost grows exponentially with the number of hyperparameters. Tuning 5 hyperparameters with 5 values each means 5^5 = 3,125 model trainings. For anything beyond 2-3 hyperparameters with small grids, this becomes computationally impractical.

python

Random search: better coverage for less cost

Random search samples hyperparameter combinations randomly from specified distributions, for a fixed number of iterations, rather than exhaustively covering a grid. Counterintuitively, this usually outperforms grid search for the same compute budget.

The reason is dimensionality: in most problems, only a couple of hyperparameters meaningfully affect performance, while others barely matter. Grid search wastes evaluations varying the unimportant ones across every value of the important ones. Random search, by sampling independently, is far more likely to land on a good value for the hyperparameters that actually matter, because every trial explores a fresh combination rather than repeating fixed grid lines.

python

Bayesian optimization: search informed by past results

Both grid and random search treat each trial independently — they never learn from previous evaluations. Bayesian optimization does. It builds a probabilistic model (commonly a Gaussian Process, or a tree-based surrogate as in TPE) of how hyperparameters map to validation performance, and uses that model to decide which combination to try next.

The process balances exploration (trying regions of the hyperparameter space it's uncertain about) with exploitation (refining regions that already look promising), using an acquisition function to score candidates. This makes it dramatically more sample-efficient than random search — it can often find a near-optimal configuration in a fraction of the trials, which matters enormously when each trial means training a large model for hours.

Popular libraries implementing this include Optuna, Hyperopt, and scikit-optimize. Bayesian optimization is the default approach for tuning expensive deep learning models where every trial has a real dollar cost.

python

Search strategies compared

StrategySample efficiencyHandles many hyperparametersParallelizableBest for
Grid searchLowPoorly — exponential blowupFully1-2 cheap hyperparameters, need exhaustive coverage
Random searchMediumWellFullyCheap-to-moderate models, several hyperparameters
Bayesian optimizationHighWellPartially — sequential by natureExpensive models, deep learning, limited trial budget

Learning rate schedules

For neural networks, the learning rate is often the single most consequential hyperparameter, and the best practice isn't to pick one fixed value — it's to schedule it. A learning rate schedule changes the learning rate over the course of training.

Common patterns: step decay drops the rate by a factor every N epochs. Cosine annealing smoothly decreases the rate following a cosine curve, often used with periodic "warm restarts." Warmup starts with a very small learning rate for the first few hundred steps before ramping up — this stabilizes training early on, when weights are randomly initialized and large updates can be destructive, and is standard practice when training transformers.

The intuition behind decaying the rate over time: large steps early in training make fast progress across the broad loss landscape, while small steps late in training allow fine-grained convergence into a sharp minimum without overshooting it.

See why step size matters

This gradient descent visualization lets you observe convergence behavior directly. A learning rate that's too large causes the parameter to oscillate or diverge; too small and convergence crawls. Learning rate schedules exist precisely to get the benefits of a large rate early and a small rate late, within a single training run.

📉 Gradient Descent Optimization

Interactive
Parameter (w)0.03.0 (Min)6.0Loss
Learning Rate (η)0.10
Steps taken:0
Current Weight (w):0.6000
Loss Value:6.2600
Gradient (dL/dw):-4.8000

Note

Every hyperparameter trial for a large model can cost real GPU-hours. Before running a wide search, use a small subset of data or a smaller model variant to narrow the promising region, then run the full search budget on the best candidates. This staged approach, sometimes called successive halving or Hyperband, cuts costs substantially without sacrificing much search quality.

What's next

Hyperparameter tuning optimizes settings for a model architecture you've already chosen. For some problems — especially small, structured datasets — the right architecture choice matters more than any amount of tuning. The next lesson, Support Vector Machines, covers a classical algorithm that still outperforms neural networks on small, high-dimensional data.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.