Fine-Tuning & Adaptation
Marcus Feld leads machine learning at Veril, a legal-technology company whose product reviews commercial contracts and flags risky clauses for corporate legal teams. His general-purpose model handles most clauses well, but on Veril's specialty, complex indemnification and limitation-of-liability language, it misses subtle patterns that Veril's lawyers catch immediately. Marcus has 40,000 human-reviewed clause annotations built up over four years. His question is the one this lesson answers: should he fine-tune a model on that data, or is there a cheaper path to the same accuracy?
What Fine-Tuning Does
Fine-tuning is the process of adjusting a language model's parameters using task-specific training data. A general-purpose model was trained on vast quantities of internet text to be competent across thousands of tasks; a fine-tuned model is specialized to excel at particular tasks within your domain. Fine-tuning changes how the model behaves, embedding domain-specific knowledge, terminology, reasoning patterns and output expectations directly into its parameters rather than supplying them at request time.
The key contrast is with retrieval-augmented generation, or RAG, which provides relevant context to the model at inference time without changing the model itself. Fine-tuning bakes knowledge and behavior into the weights, which carries both advantages and risks. A fine-tuned model can develop consistent domain-specific reasoning and a reliable output format that prompting alone struggles to enforce, and that consistency is difficult to replicate with retrieval. It can also lose general capabilities it had before, and it will faithfully reproduce any bias present in the training data. For Marcus, the appeal is precisely the consistency: he wants every indemnification clause scored the same way every time, which is exactly the kind of behavioral regularity fine-tuning is good at producing.
Fine-Tuning Approaches
Not all fine-tuning is the same, and choosing the wrong approach is a common and expensive mistake. Four families cover most of what organizations actually do, and they solve different problems.
| Approach | What it does | Best for | Relative cost |
|---|---|---|---|
| Full fine-tuning | Updates all of the model's parameters on your data. | Deep domain shifts where behavior must change substantially. | Highest compute and data needs. |
| Parameter-efficient (for example LoRA) | Freezes the base model and trains a small set of added weights. | Most practical specialization tasks; cheap to run and to iterate. | A fraction of full fine-tuning. |
| Instruction fine-tuning | Teaches the model to follow a specific style or format of instruction. | Enforcing consistent output structure and task-following. | Moderate, depends on data volume. |
| Preference learning | Trains the model to prefer some outputs over others from ranked pairs. | Aligning tone, safety and quality judgments to your standard. | Higher, requires ranked preference data. |
Instruction fine-tuning works from pairs of instruction and response drawn from your own domain. A financial services firm might train on examples pairing an instruction such as analyze this quarterly filing for risk factors with the domain-specific risk analysis its analysts would actually write. A legal services firm might pair a request to draft a contract clause for a given scenario with properly formatted legal language. What makes this effective is that the model learns approaches rather than only facts: how to structure an analysis, which information to prioritize, what tone is appropriate. For domain-specific instruction following, that can produce dramatically better results than a general model.
Preference learning, often implemented as reinforcement learning from human feedback, works differently. Instead of supplying correct answers, you supply pairs of outputs and indicate which is better, and the model learns to prefer that direction. This is more flexible than instruction fine-tuning precisely because it does not demand a perfect output, only a relative judgment. A healthcare provider might show pairs of diagnosis explanations and mark which is more patient-friendly; a customer service organization might show pairs of response options and mark which tone better reflects its values. It is computationally more complex than instruction fine-tuning and can produce more nuanced alignment with domain expectations.
For Marcus, parameter-efficient fine-tuning with LoRA is the natural starting point. It lets him specialize the model on indemnification clauses while leaving the base model untouched, which keeps the compute bill low and lets him run several experiments quickly. He would reach for full fine-tuning only if the parameter-efficient version plateaued below the accuracy his legal team requires. Instruction fine-tuning is relevant to him too, since he wants a structured risk score and a rationale returned in the same format every time, but that is a formatting behavior he can layer on rather than his central problem.
How Fine-Tuning Fails
Fine-tuning fails in predictable ways, and a leader who knows the failure modes can avoid most of them. Data scarcity is the first and most common. Fine-tuning wants thousands of high-quality examples, and datasets of dozens or a few hundred usually fail to shift behavior in any useful direction. This creates a chicken-and-egg problem, because generating enough domain-specific data is itself expensive and slow, and many organizations underestimate the requirement and attempt a fine-tune with far too little to work with. The result looks like a failure of the technique when it is really a failure of preparation.
Overfitting follows directly from thin data or too many training passes: the model memorizes its training examples rather than learning the underlying pattern. It looks excellent on data it has seen and falls apart on anything new, so an overfitted financial model might handle its own company's transaction patterns beautifully and fail on patterns it has never encountered. Guarding against it means holding out a test set the model never trains on, using careful validation methodology, considering data augmentation where more real examples are unavailable, and sometimes choosing a smaller model, which is less prone to memorizing.
Catastrophic forgetting is the mirror image. Specializing hard on one task degrades the model's general abilities, so a model fine-tuned aggressively on indemnification clauses may get worse at ordinary summarization or at clause types it used to handle, and a model fine-tuned on highly specialized medical terminology may become less capable at general language understanding. This is a genuine trade-off rather than a bug: specialization is bought with generality, and the discipline is to tune just enough to gain the domain expertise without losing capabilities you still rely on.
Data quality and bias matter more than volume. Fine-tuning faithfully reproduces whatever is in the training data, including inconsistent human labels and historical bias. If two of Veril's reviewers scored similar clauses differently over the years, the model will learn the contradiction rather than resolve it, so cleaning and reconciling labels often does more good than adding examples. Staleness is the last failure mode: a fine-tuned model is a snapshot, and when a new regulation changes how a clause should be read, the model does not update itself. RAG can incorporate new information immediately; a fine-tuned model must be retrained. Taken together, these failures explain why fine-tuning is a program rather than an event. The realistic pattern is an iterative refinement cycle: start with a small parameter-efficient run, evaluate rigorously against a held-out set, find the gaps, collect or clean data to address them, and retrain. Each pass is cheap with LoRA and expensive with full fine-tuning, which is another reason to start parameter-efficient.
When It Is Worth the Investment
Fine-tuning is worth it when several conditions hold together: you have substantial high-quality domain data, on the order of thousands of clean examples; the task needs consistent domain-specific behavior or reasoning that prompting and RAG cannot reliably enforce; your domain is stable enough that the model will not be obsolete within months; you have budget for the compute; and the expected accuracy or consistency gain justifies building and maintaining it. It is the wrong choice when your data is thin, when the domain changes fast enough that you would be retraining constantly, when RAG or careful prompt engineering already reaches acceptable quality, or when the compute and maintenance cost outweighs the improvement. Many organizations discover, after honestly testing the alternatives, that they never needed fine-tuning at all.
The compute itself is not trivial. Instruction fine-tuning on a modern large model might cost hundreds or thousands of dollars in GPU time, and preference learning costs more again. But compute is rarely the deciding number. Consider Marcus's economics, with all figures hypothetical. A parameter-efficient experiment on his 40,000 annotations might cost roughly 6,000 dollars in compute for the training runs plus about three weeks of an engineer's time to prepare data and evaluate, call it 25,000 dollars all in for the first working version. If that lifts indemnification-clause accuracy from 82 percent to 93 percent, and each missed high-risk clause costs a client roughly 4,000 dollars in downstream legal exposure while eroding Veril's renewal rate, the improvement pays for itself quickly across a client base reviewing tens of thousands of clauses a month. If the same effort only moves accuracy from 82 percent to 84 percent, the honest conclusion is that RAG or better prompting was the right tool. The number that decides it is not the cost of fine-tuning; it is the value of the accuracy gain relative to the simpler alternatives.
The checklist Marcus applies before committing is short and deliberately obstructive:
- Have I first tried RAG and structured prompting, and measured where they fall short? If not, stop and do that.
- Do I have at least a few thousand clean, consistently labeled examples? If not, fix the data before training anything.
- Is the target behavior a consistency or format problem that prompting cannot enforce? If yes, fine-tuning is a strong candidate.
- Is my domain stable over the next year, or will retraining be constant? Constant retraining favours RAG.
- Does the modelled value of the accuracy gain clearly exceed build plus maintenance cost? If it is close, prefer the simpler approach.
Evaluating a Fine-Tuned Model
Evaluation is where fine-tuning projects succeed or quietly fail, because you have to measure two things at once: whether specialization improved the target task, and whether it damaged anything else. Measuring only domain accuracy is not enough, since a model can look better on your narrow metric while it has quietly lost general capability, and many organizations discover exactly that pattern after deployment, when a fine-tuned model beats the baseline on the headline number but fails on edge cases or on tasks nobody thought to retest.
A sound evaluation covers four things: domain-specific accuracy comparing the fine-tuned model against the untuned baseline on a held-out set; general-capability testing on standard tasks to confirm catastrophic forgetting did not occur; deliberate edge-case testing on the hard and unusual examples in your domain; and real user testing to confirm the improvement matters in practice rather than only on paper. Marcus runs all four, and his most important guardrail is the baseline comparison. If the fine-tuned model does not clearly beat the untuned model with good prompting on the same held-out clauses, the fine-tune has not earned its place in production. Many teams skip that comparison entirely and end up maintaining an expensive specialized model that a well-prompted baseline could have matched.
Anti-Patterns
- Fine-tuning before exhausting RAG and prompt engineering. Without a measured account of where the simpler methods fall short, you have no way to tell whether the fine-tune helped, and no defence when someone asks why you are maintaining a custom model.
- Training on the data you have rather than the data you need. Dozens or a few hundred examples will not move behavior usefully, and attempting it anyway produces a poor model plus the false conclusion that fine-tuning does not work for your domain.
- Adding volume to fix inconsistent labels. Contradictory human judgments in the training set teach the model the contradiction, so reconciling the labels you have usually beats collecting more of the same.
- Reaching for full fine-tuning first. It is the most expensive option to run and the most expensive to iterate on, which is precisely wrong for a first attempt whose main purpose is learning where the gaps are.
- Evaluating only on the target task. A model that gains on your narrow metric while losing general capability will pass your evaluation and fail in production, on inputs nobody thought to retest.
- Skipping the baseline comparison. Without measuring the fine-tuned model against a well-prompted untuned one on the same held-out set, you cannot know whether the specialization did anything the base model could not.
- Fine-tuning a fast-moving domain. Where regulations or conventions shift regularly, a trained snapshot goes stale between retraining cycles and the retraining bill never stops.
- Treating the fine-tune as finished. A deployed model with no held-out monitoring, no scheduled re-evaluation and no owner will drift out of usefulness quietly rather than visibly.
Practice Prompts
- Take a task you are considering fine-tuning for and write down exactly where RAG and structured prompting fall short, with a measurement rather than an impression.
- Count your genuinely clean, consistently labeled examples for that task, not your total records. The gap between the two numbers is your real data position.
- Sample a set of your existing labels and have two reviewers score them independently. Where they disagree, you have found what a fine-tune would learn.
- Write the evaluation plan before the training plan: which held-out set, which general-capability tasks, which edge cases, and which users will try it.
- Estimate the value of a plausible accuracy gain for one task, then compare it against a rough build-plus-maintenance cost. Note whether the answer is clear or close.
- Ask what would have to change in your domain for a fine-tuned model to go stale, and how quickly you would notice.
Reflection
Marcus's 40,000 annotations look like an obvious asset, and that is exactly what makes his decision difficult. Having the data creates pressure to use it, and fine-tuning is the visible, technically interesting way to use it, while the alternatives feel like settling. The discipline he applies is to treat the fine-tune as the thing that must prove itself against a well-prompted baseline, not the thing the baseline must justify replacing. Think about a domain dataset your own organization has accumulated. If someone proposed fine-tuning on it next quarter, what would you need to see measured first, and would anyone in the room be willing to conclude that the simpler approach was already good enough?
Glossary
- Fine-tuning: Adjusting a model's parameters using task-specific training data so that domain knowledge and behavior are embedded in the weights rather than supplied at request time.
- Retrieval-augmented generation (RAG): Supplying relevant context to a model at inference time without changing the model, which allows new information to be incorporated immediately.
- Full fine-tuning: Updating all of a model's parameters, appropriate for deep domain shifts and the most expensive option in both compute and data.
- Parameter-efficient fine-tuning: Freezing the base model and training a small set of added weights, of which LoRA is the common example, making experiments cheap enough to iterate on.
- Instruction fine-tuning: Training on instruction and response pairs from your domain so the model learns how work is structured, prioritized and phrased, not only what is true.
- Preference learning: Training on ranked pairs of outputs, often through reinforcement learning from human feedback, to align tone, safety and quality with your standard.
- Overfitting: Memorizing training examples instead of learning the pattern, producing excellent results on seen data and poor results on anything new.
- Catastrophic forgetting: The loss of general capability that follows aggressive specialization, and the reason evaluation must cover more than the target task.
- Held-out set: Examples deliberately excluded from training and reserved for evaluation, without which neither overfitting nor genuine improvement can be detected.
Related Lessons
- Retrieval-Augmented Generation (RAG) covers the alternative this lesson tells you to exhaust first, and the better answer for fast-moving domains.
- Prompt Engineering & In-Context Learning is the other simpler path that must be measured before a fine-tune can be justified.
- Comprehensive Evaluation Frameworks goes deeper into the two-sided measurement problem described here.
- Data Quality & Management addresses the label consistency that determines what a fine-tune will actually learn.
- Data Augmentation and Synthetic Data Techniques covers what to do when real domain examples are too few.
- Model Performance Risk Management deals with the drift and staleness that follow a model into production.
Closing
Fine-tuning is a powerful path to specialization, and it carries real costs and real failure modes. Success requires substantial clean data, careful two-sided evaluation, and honest expectations about the trade-off between specialization and general capability. Before committing, exhaust the simpler alternatives of RAG and prompt engineering and measure precisely where they fall short, because many organizations reach their goals without fine-tuning at all. For those where it genuinely fits, the discipline is the one Marcus follows: start small and parameter-efficient, evaluate against a well-prompted baseline rigorously, and iterate deliberately. Fine-tuning is a strategic investment to be justified by evidence, not a shortcut to reach for by default.
Key Takeaways
- Fine-tuning changes the weights; RAG changes the context. One produces consistent domain behavior baked into the model, the other incorporates new information immediately, and the choice between them is the first decision to make.
- Pick the approach to fit the problem. Parameter-efficient methods suit most specialization work, full fine-tuning suits deep domain shifts, instruction fine-tuning enforces structure, and preference learning aligns tone and judgment from ranked pairs.
- Data volume and label consistency decide the outcome. Thousands of clean examples are the working requirement, dozens or hundreds will not shift behavior, and contradictory labels get learned rather than resolved.
- Specialization is bought with generality. Catastrophic forgetting is a trade-off rather than a defect, which is why evaluation has to cover general capability alongside the target task.
- A fine-tuned model is a snapshot. It does not update itself when the domain moves, so a fast-changing domain favours retrieval over retraining.
- Compute cost is rarely the deciding number. What decides it is the value of the accuracy gain relative to what RAG and prompting already deliver.
- Evaluate on four fronts and always against a baseline. Domain accuracy, general capability, edge cases and real user testing, with the untuned well-prompted model as the comparison the fine-tune has to beat.
Frequently Asked Questions
How much training data do I need for fine-tuning? Typically at least 1,000 high-quality examples, and preferably 5,000 or more. With smaller datasets, fine-tuning risks overfitting rather than learning. With only dozens of examples, RAG or prompt engineering usually works better and costs far less, and the effort is better spent building the dataset than training on an inadequate one.
Can I fine-tune proprietary models I access only through an API? It depends on the provider. Some hosted models expose a fine-tuning interface; others do not allow it and expect you to use the base model with RAG or prompt engineering instead. Check your specific provider's current capabilities before you build a plan that assumes fine-tuning is available.
Is fine-tuning permanent? The trained model is a fixed artifact: you cannot easily reach in and undo what it learned. If you need different behavior, you retrain from the base model with revised data, which is why rigorous evaluation before you rely on a fine-tuned model in production matters so much.
Should I use instruction fine-tuning or preference learning? Instruction fine-tuning where you can write the correct output and want consistent structure and task-following. Preference learning where correctness is not the issue but tone, safety or quality judgment is, and where you can more easily say which of two outputs is better than produce a perfect one. Preference learning is computationally heavier and needs ranked data, so it is rarely the first thing to try.
Skill.re