Fine-tuning & Training LLMs
1 · The lesson
readA reflex among engineers new to LLMs: "the model doesn't quite do what I want — I should fine-tune it." This is almost always the wrong instinct. In 2026, the order of operations is prompt engineering → RAG → fine-tuning, and most teams never reach the third step. Fine-tuning is more expensive, less flexible, and harder to debug than the alternatives. It is also genuinely the right tool for a narrow set of problems.
This lesson is about knowing the difference. When to fine-tune, when not to, how the modern parameter-efficient methods (LoRA, QLoRA) work, what a real training pipeline looks like, and why "I want the model to learn these new facts" is a sign you should be using RAG instead.
Run locally on a GPU machine with
pip install transformers peft datasets bitsandbytes accelerate. Training requires a CUDA GPU (24 GB+ VRAM for the QLoRA example with a 7B model). Inference of fine-tuned adapters works on smaller hardware. Code samples are illustrative — full runs take 30 minutes to several hours.
1. The 2026 Decision Tree
Need the model to behave differently? │ ├─ Does it lack information? → RAG (retrieve facts, don't train them in) ├─ Does it use the wrong format? → Prompt engineering, few-shot examples ├─ Is the style/voice off? → Fine-tune (style is learnable) ├─ Same task at lower latency/cost? → Fine-tune a smaller model ├─ Proprietary patterns it's never → Fine-tune │ seen in training? └─ Adversarial robustness? → Often fine-tune + careful data curation
The two things fine-tuning is good for in 2026:
1. Consistent style/format/voice — your brand tone, a specific output schema, refusal patterns. Style is what fine-tuning teaches best.
2. Cheap narrow models — fine-tuning a 7B open-weights model on your task may match a frontier API at 1/50th the cost. Worth it once volume is high enough.
The thing fine-tuning is bad for: teaching new facts. Models do not reliably absorb factual knowledge during fine-tuning the way they do during pretraining. You will spend days training, and the model will still sometimes get the facts wrong, and you have lost the ability to update those facts without retraining. RAG is strictly better for knowledge.
2. The Spectrum of "Training"
"Fine-tuning" is one point on a wider spectrum:
| Approach | What updates | Cost | When |
|---|---|---|---|
| Prompt engineering | Nothing | $ | Default — try this first |
| RAG | Nothing (you add a retrieval layer) | $$ | Knowledge / data access |
| Prompt tuning / prefix tuning | A few hundred learnable "virtual token" embeddings | $$ | Very small adaptation |
| LoRA / QLoRA | Small adapter matrices alongside frozen base | $$$ | The 2026 default for serious fine-tuning |
| Full fine-tuning | Every weight of the model | $$$$ | Rare; usually unnecessary |
| Continued pretraining | All weights, on more raw text | $$$$$ | Adding a new domain (medical, legal); expensive |
| From-scratch pretraining | All weights, from random init | $$$$$$+ | Only Anthropic, OpenAI, Google, Meta, etc. |
In practice, you spend 99% of your time on the first two and 1% on LoRA. The bottom of the table exists for context.
3. LoRA — The Key Insight
LoRA (Low-Rank Adaptation, Hu et al. 2021) is the trick that made fine-tuning practical for individual developers. The idea, in one sentence: instead of updating a giant weight matrix W, learn a small low-rank update BA and add it.
Original layer: y = W x (W is e.g. 4096 × 4096) LoRA-adapted layer: y = W x + (B A) x (A is r×4096, B is 4096×r, r is small) Where: W is FROZEN — the original pretrained weights, unchanged. A and B are TRAINABLE — but tiny. r is the "rank" — typically 8, 16, or 64.
Why it works: empirically, the change a fine-tune wants to make is itself low-rank. You don't need every weight in W to move; you need a small subspace of them to shift. BA is exactly that low-rank subspace, learned during training.
The numbers are dramatic. A 7B-parameter model has ~7B weights. A LoRA adapter on it (rank 16, applied to attention projections) has ~10M trainable parameters — roughly 700× fewer. You fit training in vastly less GPU memory, you save tiny adapter files (~40 MB) instead of duplicating the whole model (~14 GB), and you can serve dozens of fine-tunes on top of one shared base model.
QLoRA (Dettmers et al. 2023) goes further: load the base model in 4-bit quantisation (saving 4× memory), then train LoRA adapters in higher precision on top. The 4-bit quantisation is good enough that quality barely suffers, and you can now fine-tune a 70B model on a single 48 GB GPU instead of needing a small cluster.
4. Hugging Face PEFT — The Library
The PEFT (Parameter-Efficient Fine-Tuning) library implements LoRA, QLoRA, and friends on top of transformers. The shape of a real LoRA training script:
# Illustrative — requires a CUDA GPU and the libraries below. # pip install transformers peft datasets bitsandbytes accelerate trl from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training from trl import SFTTrainer, SFTConfig from datasets import load_dataset BASE = "mistralai/Mistral-7B-Instruct-v0.3" # 1. Load the base model in 4-bit (QLoRA) bnb = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype="bfloat16", ) model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto") model = prepare_model_for_kbit_training(model) tokenizer = AutoTokenizer.from_pretrained(BASE) # 2. Wrap with LoRA adapters lora = LoraConfig( r=16, lora_alpha=32, target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM", ) model = get_peft_model(model, lora) model.print_trainable_parameters() # trainable params: 13,631,488 || all params: 7,254,790,144 || trainable%: 0.188 # 3. Load your dataset (JSONL with chat-format messages) ds = load_dataset("json", data_files="my_training_data.jsonl") # 4. Train trainer = SFTTrainer( model=model, train_dataset=ds["train"], args=SFTConfig( output_dir="./lora-out", num_train_epochs=3, per_device_train_batch_size=4, gradient_accumulation_steps=4, learning_rate=2e-4, bf16=True, logging_steps=10, save_steps=200, ), ) trainer.train() # 5. Save just the adapter (tiny file) model.save_pretrained("./my-adapter")
The whole thing is maybe 30 lines of glue code. The libraries handle the heavy lifting. Most of the real work is upstream — getting the dataset right.
5. The Dataset — Format and Quality
Modern fine-tuning datasets are almost always in chat format, one example per line of a JSONL file:
{"messages": [
{"role": "system", "content": "You convert SQL to plain English."},
{"role": "user", "content": "SELECT name FROM users WHERE active = 1;"},
{"role": "assistant", "content": "Names of all active users."}
]}
{"messages": [
{"role": "system", "content": "You convert SQL to plain English."},
{"role": "user", "content": "SELECT COUNT(*) FROM orders WHERE created_at > NOW() - INTERVAL '7 days';"},
{"role": "assistant", "content": "Number of orders placed in the last seven days."}
]}The trainer reads the messages, applies the model's chat template to render them as a single string, and trains the model to produce the assistant message given the system + user messages.
Quality > quantity, by a wide margin. A famous Meta paper ("LIMA", 2023) showed that 1,000 carefully-curated examples beat 50,000 noisy ones for instruction tuning. The same lesson keeps being re-learned: spend the time hand-checking 500 examples rather than scraping 50K mediocre ones.
Practical dataset rules:
- Diversity — cover the range of inputs you actually expect. Don't have all examples about one sub-topic.
- Format consistency — every assistant response in the dataset must follow the format you want at inference time. Mixed formats teach the model nothing useful.
- No garbage examples — one example with a wrong answer can take 10 correct examples to undo. Hand-review before training.
- Match production distribution — if production inputs are short and direct, your training examples shouldn't all be long and chatty.
For RLHF-style preference data (used in DPO/IPO, the modern alternative to PPO-based RLHF), the format adds a "chosen" and "rejected" pair per prompt. Out of scope for this lesson; see the trl library's docs.
6. Cost — Fine-tuned 7B vs Frontier API
A concrete trade-off. You want a classifier for support tickets, 1M requests/month, ~500 input tokens, ~50 output tokens each.
Option A — Frontier API (Sonnet-tier):
1_000_000 × (500 × 2 + 50 × 10) / 1_000_000 = $1,500/month- No infrastructure. No training. Quality is excellent. Latency is whatever the API gives you.
Option B — Fine-tuned Mistral-7B on your own GPU:
- One-off training cost: ~$50-200 (a few hours on a rented H100 or A100)
- Inference: a single A10G or L40S handles ~5 req/sec → 13M reqs/month at one GPU, ~$300-500/month
- You own the deployment, the security model, the latency
- Quality may match or fall short — needs evaluation
The numbers cross over somewhere around 100K-500K requests/month depending on model tier and task complexity. Below that volume, fine-tuning is rarely worth the effort. Above 10M requests/month, it almost always is.
The other axis: privacy. If you cannot send the data to a third-party API at all (HIPAA, financial, defence), open-weights + fine-tuning is the only path. Cost analysis becomes secondary.
7. API-Based Fine-tuning — The Easier Path
You don't always need to manage the training yourself. Several providers offer fine-tuning as a service:
- OpenAI fine-tuning API — upload a JSONL, get a fine-tuned model deployed automatically. Mostly LoRA under the hood. Supports
gpt-4o-miniand similar; frontier models gated. - Anthropic — fine-tuning available for select models, via Bedrock/Vertex AI partners (often behind enterprise access).
- Google Vertex AI — fine-tune Gemini variants through their console.
- Together / Fireworks / Anyscale / Replicate — hosted fine-tuning + inference for open-weights models. The "easy mode" for Llama/Mistral fine-tuning.
- Hugging Face AutoTrain — point-and-click fine-tuning, including QLoRA, with deployment.
The tradeoff: lower control over hyperparameters, the dataset stays with the vendor, but vastly less infrastructure work. For a first fine-tune, the API path is usually right. Once you know the data is good and the approach works, the open-weights path opens up.
8. Evaluation — Did Fine-tuning Actually Help?
The mandatory first step before deploying any fine-tune: does it beat the base model + your best prompt? Surprisingly often, it does not — either the dataset wasn't good enough, or the base model's behaviour was already adequate with prompting.
Build an eval set (50-500 held-out test cases) and run both the base model with your best prompt and the fine-tuned model against it. Compare:
| Task type | Metric |
|---|---|
| Classification / extraction | Accuracy, F1, exact match |
| Translation / structured rewrite | BLEU, ROUGE, or exact match if there's a canonical answer |
| Code generation | Test-case pass rate (run the generated code) |
| Free-form generation | LLM-as-judge with a clear rubric, plus human spot-checks |
Track:
- Quality on each metric
- Latency (a 7B fine-tune is often faster than a frontier API)
- Cost per request
- Robustness on adversarial/out-of-distribution inputs
A fine-tune that wins on quality but loses on adversarial robustness is a footgun. Test broadly.
9. Catastrophic Forgetting
Fine-tune a model heavily on one narrow task and it can lose general capability. The model that aced "SQL → English" can no longer hold a normal conversation, write code, or follow a basic instruction outside its training distribution. This is catastrophic forgetting.
Three mitigations:
1. Don't over-train. LoRA on 1-3 epochs usually keeps the base capabilities intact. Watch validation loss; stop when it plateaus.
2. Mix in general data. Add a fraction (10-30%) of general instruction-following examples to your dataset so the model practices being a good assistant alongside learning your specific task.
3. Lower learning rate, lower rank. Both reduce how much the adapter can drift from base behaviour.
Keep adapters small and weight changes mild. If you find yourself needing to train aggressively to get the task right, the task is probably wrong for fine-tuning — go back to prompting or RAG.
10. RLHF — One Paragraph
RLHF (Reinforcement Learning from Human Feedback) is what turns an instruction-tuned model into the helpful, harmless, honest assistant you actually want to talk to. The original RLHF pipeline: train a reward model on human preference pairs (response_A vs response_B, which did the human prefer?), then optimise the LLM with PPO to maximise that reward. It is expensive, finicky, and not something individual developers do. DPO (Direct Preference Optimisation, Rafailov et al. 2023) is a much simpler alternative that skips the reward model and trains directly on preference pairs — it has largely replaced PPO-based RLHF for open-source fine-tuning. The trl library implements both. Worth knowing exists; rarely worth doing.
Common Mistakes
1. Fine-tuning to add facts
The single most common mistake. "I want the model to know our product catalogue" → fine-tune it. Three weeks later, the model occasionally invents products that don't exist, and the catalogue has changed and you'd need to retrain. RAG would have taken a day and updates for free.
2. Training on tiny noisy datasets
50 examples scraped together over a weekend produces an over-fit model that memorises your training set and fails on anything slightly different. Either curate a small high-quality set (200-1000 examples) deliberately, or use a larger, deliberately-filtered dataset (10K+).
3. Not having a held-out test set
Training loss going down means the model is learning something. It does not mean it generalises. Always hold out 10-20% of your data (or a separately collected eval set) and never train on it.
4. Deploying without comparing to the base model
The mandatory baseline is: same base model, your best zero/few-shot prompt. If the fine-tune doesn't beat that decisively on your eval set, you have just spent days and money for nothing.
5. Skipping the chat template
Each model has its own chat template (special tokens like [INST], <|im_start|>, <|user|>). The trainer applies it automatically when you use SFTTrainer with chat-format messages — but if you assemble prompts by hand, mismatched templates produce a model that won't respond properly at inference time. Trust the library.
6. Saving the full model instead of the adapter
A LoRA adapter is ~40 MB. Saving the full merged model is ~14 GB. Save the adapter and load it at inference time on top of the base; you save disk, bandwidth, and the option to swap base models later.
7. Ignoring catastrophic forgetting
A fine-tuned model that no longer handles general conversation is unusable in production unless your only call to it is the narrow trained task. Test the fine-tune on out-of-domain inputs as part of your eval suite.
🎯 Your Turn — Build a SQL-to-English Training Set
Write a small JSONL training file (10 examples) for a "convert SQL query to plain English" fine-tune, and describe the LoRA training command you would run.
Requirements:
- 10 examples, one per line of
sql_to_english.jsonl - Each example is a
{"messages": [...]}object with a system, user, and assistant message - System prompt: clear role + instruction
- User content: a SQL query
- Assistant content: a plain-English paraphrase, 1-2 sentences
- Cover variety: SELECT with WHERE, JOIN, COUNT/aggregate, GROUP BY, INSERT, UPDATE — don't make all 10 look the same
Then write a train.sh (or describe in comments) that would launch QLoRA fine-tuning on Mistral-7B using the trl library.
# Skeleton — write to sql_to_english.jsonl and document the training command. import json SYSTEM = "You convert SQL queries into a one-sentence plain-English description. Be precise. Do not add explanations." # TODO 1: list 10 (sql, english) tuples covering varied query types # TODO 2: write each to a JSONL line as {"messages": [system, user, assistant]} # TODO 3: write the conceptual training command EXAMPLES = [ # ("SELECT name FROM users WHERE active = 1;", "Names of all active users."), # ... fill in ]
Hint 1 — Covering variety
A good 10-example set covers: simple SELECT, SELECT with WHERE, JOIN across two tables, aggregate (COUNT/SUM/AVG), GROUP BY, ORDER BY + LIMIT, INSERT, UPDATE, DELETE, and a slightly tricky one (subquery or HAVING). If all 10 are SELECT-WHERE, the model overfits to that shape.Hint 2 — JSONL format
JSONL means one valid JSON object per line, no commas between lines, no enclosing array. Open the file in"w" mode, loop, f.write(json.dumps(obj) + "\n"). Verify by re-reading and parsing each line.
Show full solution
# Build a tiny SQL-to-English fine-tuning dataset and document the LoRA command. import json SYSTEM = ( "You convert SQL queries into a one-sentence plain-English description. " "Be precise. Do not add explanations." ) EXAMPLES = [ # Simple SELECT ("SELECT name FROM users;", "Names of all users."), # SELECT with WHERE ("SELECT name FROM users WHERE active = 1;", "Names of all active users."), # SELECT with multiple conditions ("SELECT email FROM users WHERE country = 'UK' AND signup_date > '2025-01-01';", "Emails of UK users who signed up after 1 January 2025."), # JOIN ("SELECT u.name, o.total FROM users u JOIN orders o ON o.user_id = u.id;", "Each user's name alongside their order totals."), # Aggregate ("SELECT COUNT(*) FROM orders WHERE status = 'pending';", "The number of pending orders."), # GROUP BY ("SELECT country, COUNT(*) FROM users GROUP BY country;", "The number of users in each country."), # ORDER BY + LIMIT ("SELECT product, sales FROM stats ORDER BY sales DESC LIMIT 5;", "The top five products by sales."), # HAVING (slightly trickier) ("SELECT customer_id, SUM(total) FROM orders GROUP BY customer_id HAVING SUM(total) > 1000;", "Customer IDs whose total order value exceeds 1000."), # INSERT ("INSERT INTO users (name, email) VALUES ('Bob', 's@example.com');", "Add a new user named Bob with email s@example.com."), # UPDATE ("UPDATE products SET price = price * 1.1 WHERE category = 'electronics';", "Increase the price of all electronics products by 10%."), ] with open("sql_to_english.jsonl", "w", encoding="utf-8") as f: for sql, english in EXAMPLES: record = { "messages": [ {"role": "system", "content": SYSTEM}, {"role": "user", "content": sql}, {"role": "assistant", "content": english}, ] } f.write(json.dumps(record) + "\n") print(f"wrote {len(EXAMPLES)} examples to sql_to_english.jsonl") # Sanity-check by reading it back with open("sql_to_english.jsonl", encoding="utf-8") as f: for line in f: obj = json.loads(line) assert len(obj["messages"]) == 3 print("validation: all examples parse and have 3 messages each") # ───────── Training command (conceptual) ───────── # Requires a CUDA GPU (24 GB+ VRAM for QLoRA on Mistral-7B). Run from the shell: # # pip install transformers peft datasets bitsandbytes accelerate trl # # python -c " # from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig # from peft import LoraConfig, prepare_model_for_kbit_training # from trl import SFTTrainer, SFTConfig # from datasets import load_dataset # # BASE = 'mistralai/Mistral-7B-Instruct-v0.3' # bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type='nf4', # bnb_4bit_compute_dtype='bfloat16') # model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map='auto') # model = prepare_model_for_kbit_training(model) # tokenizer = AutoTokenizer.from_pretrained(BASE) # # lora = LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05, bias='none', # task_type='CAUSAL_LM', # target_modules=['q_proj','k_proj','v_proj','o_proj']) # # ds = load_dataset('json', data_files='sql_to_english.jsonl', split='train') # # trainer = SFTTrainer( # model=model, train_dataset=ds, peft_config=lora, # args=SFTConfig(output_dir='./sql-en-adapter', num_train_epochs=3, # per_device_train_batch_size=2, gradient_accumulation_steps=4, # learning_rate=2e-4, bf16=True, logging_steps=1, save_steps=10), # ) # trainer.train() # trainer.save_model('./sql-en-adapter') # " # # Result: a ~40 MB LoRA adapter in ./sql-en-adapter that you load on top of # the base Mistral-7B-Instruct model at inference time.
What this dataset gets right:
- Variety across SQL shapes — SELECT, JOIN, aggregate, GROUP BY, ORDER BY+LIMIT, HAVING, INSERT, UPDATE. No single pattern dominates.
- Consistent assistant style — every output is a single short sentence. The model learns the format as well as the content.
- System prompt repeated per example — important for the model to associate the system instruction with this task at inference time.
- Plain English, not SQL jargon — "the top five products by sales", not "the result of ordering by sales descending and limiting to 5". The model learns to paraphrase, not echo.
Honest caveats:
- 10 examples is far too few to actually fine-tune well. A real fine-tune of this task wants 200-500+ varied examples. This exercise is about the format, not the volume.
- No DELETE example, no subqueries with EXISTS, no window functions. A real production dataset would cover the full surface area of SQL you expect to see.
- No evaluation set. You would set aside another 20+ held-out examples that never enter training, and measure exact-match or LLM-as-judge accuracy on them.
- Prompt-only baseline first. Before any of this, you would try the task with just a system prompt and 3-5 in-prompt few-shot examples on a frontier model. If that already works, do not fine-tune.
The shape — JSONL chat format, QLoRA via trl.SFTTrainer, save the adapter — is the working pattern for almost every modern open-weights fine-tune.
What You Learned
- Default to prompts and RAG. Fine-tune only when style consistency, latency/cost on a narrow task, or proprietary patterns demand it.
- Don't fine-tune to add facts — use RAG. Models absorb facts unreliably during fine-tuning, and you lose the ability to update.
- LoRA adds tiny low-rank adapter matrices (
BA) alongside frozen base weights — 100-1000× fewer trainable parameters than full fine-tuning. - QLoRA quantises the frozen base to 4-bit and trains LoRA on top — fits 70B-parameter fine-tuning on a single 48 GB GPU.
- Dataset format: JSONL of
{"messages": [system, user, assistant]}. Quality and diversity beat quantity. - Hugging Face PEFT + TRL is the standard open-source stack:
LoraConfig,SFTTrainer, save the adapter. - API-based fine-tuning (OpenAI, Vertex AI, Together) is the easy entry point — less control, far less infrastructure.
- Always evaluate against the base model + best prompt before deploying. Often, that baseline wins.
- Catastrophic forgetting is real — keep training light, mix in general data, eval on out-of-domain inputs.
- RLHF / DPO are the next layer for shaping preferences;
trlimplements both, but most teams never need them.
Next: Building Production AI Applications — turning a working prototype into something you can run at scale: latency, caching, evaluation pipelines, observability, and the security baseline.