Fine-Tuning LLMs
What fine-tuning actually is
A pre-trained LLM has already learned grammar, facts, reasoning patterns, and world knowledge from a massive, general-purpose corpus. Fine-tuning takes that pre-trained model and continues training it on a smaller, task-specific dataset — adjusting its weights so it specializes toward a narrower behavior: a particular tone, a particular domain (legal, medical, your company's support tickets), or a particular output format.
The key distinction from pre-training is scale and purpose. Pre-training builds general capability from trillions of tokens over weeks on huge clusters. Fine-tuning refines an already-capable model using thousands to millions of examples, often in hours on a handful of GPUs.
Full fine-tuning: updating every weight
The most straightforward approach is full fine-tuning: run backpropagation through the entire model and update every parameter. This gives maximum flexibility — the model can genuinely shift its behavior in deep ways — but it's expensive. A 70-billion-parameter model requires storing gradients and optimizer states for all 70 billion weights, which typically means several times the model's memory footprint in GPU RAM (Adam optimizer state alone roughly doubles or triples memory use).
Full fine-tuning also carries a real risk: catastrophic forgetting, where the model overwrites general capabilities it had during pre-training in order to fit the narrow fine-tuning dataset. A model fine-tuned heavily on legal contracts might get worse at casual conversation or general reasoning.
LoRA: fine-tune a fraction of the parameters
LoRA (Low-Rank Adaptation) is the dominant parameter-efficient fine-tuning (PEFT) technique. Instead of updating the full weight matrix $W$ of a layer, LoRA freezes $W$ entirely and injects two small trainable matrices, $A$ and $B$, such that the effective update is $\Delta W = BA$, where $A$ and $B$ have a much smaller inner dimension $r$ (the "rank") than $W$ itself.
During inference, the output becomes $Wx + BAx$. Because $r$ is typically 4–64 while $W$ might be a 4096×4096 matrix, the number of trainable parameters drops by 100–1000x. This makes fine-tuning feasible on a single consumer GPU, produces small checkpoint files (megabytes instead of gigabytes), and — critically — since the original weights $W$ are untouched, catastrophic forgetting is much less severe.
QLoRA: quantization plus LoRA
QLoRA pushes efficiency further by quantizing the frozen base model to 4-bit precision (using a technique called NF4, tuned for the statistical distribution of neural network weights) before applying LoRA adapters on top in full precision. This lets you fine-tune genuinely large models — 65B+ parameters — on a single GPU with 24–48GB of memory, something that would require dozens of GPUs with full fine-tuning.
The trade-off is a small amount of numerical precision lost to quantization, though QLoRA's authors showed this rarely translates into a meaningful quality gap versus full-precision LoRA or full fine-tuning on most tasks.
Note
Fine-tuning approaches compared
| Method | Trainable params | GPU memory | Forgetting risk | Best for |
|---|---|---|---|---|
| Full fine-tuning | 100% | Very high (multiples of model size) | High | Maximum capability shift, large budgets |
| LoRA | 0.1–1% | Moderate | Low | Domain adaptation, style, small teams |
| QLoRA | 0.1–1% (base quantized to 4-bit) | Low | Low | Fine-tuning large models on limited hardware |
| Instruction tuning | Depends on underlying method | Depends | Depends | Teaching task-following behavior |
Visualizing the fine-tuning objective
Fine-tuning is still gradient descent on a loss surface — it just starts from a pre-trained set of weights instead of random initialization, and typically uses a much lower learning rate to avoid disrupting existing knowledge. Use the demo below to build intuition for how optimizers navigate a loss landscape toward a minimum; fine-tuning applies the same mechanics on a smaller, more targeted dataset.
📉 Gradient Descent Optimization
InteractiveWhen to fine-tune vs. prompt engineering vs. RAG
Fine-tuning is often reached for too early. Before fine-tuning, ask what problem you're actually solving:
Use prompt engineering when the model already has the knowledge and capability but isn't being guided well — better instructions, examples (few-shot prompting), or output formatting requirements usually fix this cheaply and instantly, with no training required.
Use RAG when the problem is missing or outdated *knowledge* — the model doesn't know your company's internal docs, current pricing, or yesterday's data. Fine-tuning is a poor tool for injecting facts; it's slow to update and models tend to generalize or hallucinate around memorized facts rather than reciting them reliably. Retrieval solves this by handing relevant facts to the model at request time.
Use fine-tuning when the problem is *behavior*, not knowledge: you need a consistent output format across thousands of requests, a specific tone or persona, domain-specific reasoning patterns (e.g. following a particular legal analysis structure), or you need to bake in behavior that would otherwise require a long, expensive system prompt on every single call.
Note
What's next
Once you've fine-tuned or chosen an approach, you need a way to know whether it actually improved anything. The next lesson, Evaluating LLMs, covers benchmarks, LLM-as-judge, and how to measure quality rigorously rather than by vibes.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.