AI & MLDeep
Advanced

Fine-Tuning LLMs

12 min read

Learn
Deep Reading
Estimated 12 mins
Prereq
Advanced
Requires advanced math/coding
Interactive
Live Simulator
Contains interactive viz

What fine-tuning actually is

A pre-trained LLM has already learned grammar, facts, reasoning patterns, and world knowledge from a massive, general-purpose corpus. Fine-tuning takes that pre-trained model and continues training it on a smaller, task-specific dataset — adjusting its weights so it specializes toward a narrower behavior: a particular tone, a particular domain (legal, medical, your company's support tickets), or a particular output format.

The key distinction from pre-training is scale and purpose. Pre-training builds general capability from trillions of tokens over weeks on huge clusters. Fine-tuning refines an already-capable model using thousands to millions of examples, often in hours on a handful of GPUs.

Full fine-tuning: updating every weight

The most straightforward approach is full fine-tuning: run backpropagation through the entire model and update every parameter. This gives maximum flexibility — the model can genuinely shift its behavior in deep ways — but it's expensive. A 70-billion-parameter model requires storing gradients and optimizer states for all 70 billion weights, which typically means several times the model's memory footprint in GPU RAM (Adam optimizer state alone roughly doubles or triples memory use).

Full fine-tuning also carries a real risk: catastrophic forgetting, where the model overwrites general capabilities it had during pre-training in order to fit the narrow fine-tuning dataset. A model fine-tuned heavily on legal contracts might get worse at casual conversation or general reasoning.

LoRA: fine-tune a fraction of the parameters

LoRA (Low-Rank Adaptation) is the dominant parameter-efficient fine-tuning (PEFT) technique. Instead of updating the full weight matrix $W$ of a layer, LoRA freezes $W$ entirely and injects two small trainable matrices, $A$ and $B$, such that the effective update is $\Delta W = BA$, where $A$ and $B$ have a much smaller inner dimension $r$ (the "rank") than $W$ itself.

During inference, the output becomes $Wx + BAx$. Because $r$ is typically 4–64 while $W$ might be a 4096×4096 matrix, the number of trainable parameters drops by 100–1000x. This makes fine-tuning feasible on a single consumer GPU, produces small checkpoint files (megabytes instead of gigabytes), and — critically — since the original weights $W$ are untouched, catastrophic forgetting is much less severe.

python

QLoRA: quantization plus LoRA

QLoRA pushes efficiency further by quantizing the frozen base model to 4-bit precision (using a technique called NF4, tuned for the statistical distribution of neural network weights) before applying LoRA adapters on top in full precision. This lets you fine-tune genuinely large models — 65B+ parameters — on a single GPU with 24–48GB of memory, something that would require dozens of GPUs with full fine-tuning.

The trade-off is a small amount of numerical precision lost to quantization, though QLoRA's authors showed this rarely translates into a meaningful quality gap versus full-precision LoRA or full fine-tuning on most tasks.

Note

LoRA and full fine-tuning describe *how* you update weights. Instruction tuning describes *what* you train on: a dataset of (instruction, response) pairs designed to teach the model to follow directions rather than just continue text statistically. It's what turns a raw next-token-predictor into something that reliably answers "summarize this" with a summary instead of more unrelated text. You can instruction-tune using either full fine-tuning or LoRA — they're orthogonal choices.

Fine-tuning approaches compared

MethodTrainable paramsGPU memoryForgetting riskBest for
Full fine-tuning100%Very high (multiples of model size)HighMaximum capability shift, large budgets
LoRA0.1–1%ModerateLowDomain adaptation, style, small teams
QLoRA0.1–1% (base quantized to 4-bit)LowLowFine-tuning large models on limited hardware
Instruction tuningDepends on underlying methodDependsDependsTeaching task-following behavior

Visualizing the fine-tuning objective

Fine-tuning is still gradient descent on a loss surface — it just starts from a pre-trained set of weights instead of random initialization, and typically uses a much lower learning rate to avoid disrupting existing knowledge. Use the demo below to build intuition for how optimizers navigate a loss landscape toward a minimum; fine-tuning applies the same mechanics on a smaller, more targeted dataset.

📉 Gradient Descent Optimization

Interactive
Parameter (w)0.03.0 (Min)6.0Loss
Learning Rate (η)0.10
Steps taken:0
Current Weight (w):5.4000
Loss Value:6.2600
Gradient (dL/dw):+4.8000

When to fine-tune vs. prompt engineering vs. RAG

Fine-tuning is often reached for too early. Before fine-tuning, ask what problem you're actually solving:

Use prompt engineering when the model already has the knowledge and capability but isn't being guided well — better instructions, examples (few-shot prompting), or output formatting requirements usually fix this cheaply and instantly, with no training required.

Use RAG when the problem is missing or outdated *knowledge* — the model doesn't know your company's internal docs, current pricing, or yesterday's data. Fine-tuning is a poor tool for injecting facts; it's slow to update and models tend to generalize or hallucinate around memorized facts rather than reciting them reliably. Retrieval solves this by handing relevant facts to the model at request time.

Use fine-tuning when the problem is *behavior*, not knowledge: you need a consistent output format across thousands of requests, a specific tone or persona, domain-specific reasoning patterns (e.g. following a particular legal analysis structure), or you need to bake in behavior that would otherwise require a long, expensive system prompt on every single call.

Note

A common mistake: fine-tuning a model on a company's product documentation, expecting it to reliably answer factual questions about that product. Fine-tuning nudges *how* the model responds, not what it reliably *knows*. For fact-heavy use cases, RAG — retrieving the relevant document at query time — is almost always more accurate, cheaper to keep current, and easier to audit than fine-tuning.

What's next

Once you've fine-tuned or chosen an approach, you need a way to know whether it actually improved anything. The next lesson, Evaluating LLMs, covers benchmarks, LLM-as-judge, and how to measure quality rigorously rather than by vibes.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.