Fine-Tuning and Custom Model Development
Helena Vasquez leads AI implementation at a mid-size commercial real estate firm. Her team spent four months perfecting prompts to get a general-purpose language model to write property valuation summaries in their firm's format: a specific section order, terminology borrowed from their 20-year methodology guide, and risk flags phrased the way their senior analysts phrase them. The prompts worked about 70% of the time. The other 30%, the model drifted into generic real estate language that their partners would immediately flag as "not our voice." Helena's CTO suggested fine-tuning. She had heard the term but was not sure what it actually meant, or whether it was worth it. This lesson answers both questions.
What Fine-Tuning Actually Is
A foundation model such as GPT-4, Claude, or Llama is trained on enormous amounts of general text: books, web pages, code, and articles. From that breadth it learns language patterns, facts, and reasoning strategies. But "general" is the operative word. The model knows how to write a property summary the way the average internet author would write one, not the way Helena's firm writes one, and no amount of instruction fully closes that gap when the target is a house style built up over two decades of practice.
Fine-tuning takes a pre-trained model and continues training it on a smaller, curated dataset specific to your task or domain. The useful analogy is finishing school for a broadly educated graduate: the foundational knowledge is already there and you are teaching professional specialization on top of it. Mechanically, what changes are the model's weights, the numerical parameters that determine its behavior. Training adjusts them so that patterns from your examples become more prominent in the model's outputs. After fine-tuning on 500 of Helena's firm's valuation summaries, the model does not merely follow a prompt asking for "our format." It has internalized what "our format" looks like, which is why the results stop depending on how carefully the instruction was worded.
What fine-tuning does not do is equally important, because misunderstanding it is the most common way the technique gets misapplied. Fine-tuning does not reliably inject new factual knowledge into a model. If you want the model to know about a policy that does not appear in its training data, fine-tuning is the wrong tool, and retrieval-augmented generation, which connects the model to a searchable document store, is usually better. Fine-tuning is about behavior, style, and format. It is not about facts.
| Approach | What it changes | Best suited to |
|---|---|---|
| Prompt engineering | The instruction sent with each request | Most problems, and always the first thing to try |
| Retrieval-augmented generation | The information available to the model at request time | Facts the model does not have, including policies and current documents |
| Fine-tuning | The model's weights, and therefore its default behavior | Style, format, and domain patterns that must hold consistently |
When Fine-Tuning Makes Sense
Fine-tuning is not always the right answer, and treating it as the standard response to disappointing output leads teams into a costly project when a cheaper one would have worked. Prompt engineering, meaning carefully written instructions, solves many problems faster and at lower cost. The signals below indicate that you have moved past what prompting can do for you.
- Prompt engineering has hit its ceiling. You have invested serious effort in prompting and still get inconsistent results. Helena's 70% success rate after four months of iteration is a classic signal, because the remaining failures are not the kind that a better sentence in the instruction will fix.
- You have consistent style or format requirements. Legal filings, clinical notes, technical specifications, and brand-voice content are all cases where the output must reliably match a pattern rather than simply be good in general. Reliability, not quality, is what fine-tuning buys.
- You need faster, cheaper inference. A fine-tuned smaller model often outperforms a larger general model on a specific task at a fraction of the API cost. A fine-tuned 7B-parameter model, meaning a mid-size model with seven billion parameters, may equal a 70B model on your particular task while costing considerably less to run.
- System prompts are getting unwieldy. If your prompt carries 2,000 words of style instructions that you send with every single request, fine-tuning can compress all of that into model behavior, reducing your token costs significantly and removing a long instruction that has to be maintained by hand.
The negative signals are just as clear. Fine-tuning is probably not worth it when your task is general enough that a well-prompted frontier model already handles it, when your use case changes frequently enough that the trained behavior would need constant revision, or when you have fewer than a few hundred high-quality examples to train on. That last constraint is not a soft preference. A fine-tuning project without enough good examples does not produce a slightly worse model; it produces a model whose behavior you cannot explain.
The Data Preparation Phase
This is where most fine-tuning projects succeed or fail. Model training is forgiving of many things, but it is not forgiving of bad data, and teams consistently underestimate how much of the total effort belongs here rather than in the training run itself. The standard format for supervised fine-tuning is input and output pairs: here is the input, meaning the prompt, and here is the ideal output, meaning the completion. For Helena's use case, each example is a property brief as input with the corresponding valuation summary as output. She needs a minimum of 50 to 100 high-quality examples to see meaningful improvement, and 500 or more is where results become reliably strong.
Quality matters more than quantity, and the most common mistake is dumping whatever examples happen to be available into the training data. If your historical examples include outputs your team considered mediocre, you will fine-tune mediocrity into the model, and it will reproduce that standard confidently and consistently. Before training, manually review a sample and remove or rewrite anything that falls below the quality bar you are aiming for. This step is tedious and it is not optional; the model has no way to distinguish an example you were proud of from one you tolerated.
Diversity in your training examples matters alongside quality. If all 200 of your examples are office building summaries, the model will be poorly calibrated for retail or industrial properties, and the failure will appear in production rather than in training, where everything looked fine. Sample across the full distribution of cases you expect to encounter once the model is live, including the awkward ones your team handles less often. On format, most major fine-tuning platforms, including OpenAI, Hugging Face, and Together AI, accept JSON Lines: one training example per line, structured as a JSON object. The exact schema varies by platform but the concept is consistent across them.
Training, Evaluation, and Iteration
Once the data is prepared, the actual training through an API-based service is relatively straightforward. You upload your training file, select a base model, configure a few hyperparameters, and run the job. Hyperparameters are the settings that control how the training proceeds, and the two most important are the learning rate and the number of training epochs. For a few hundred examples, training might take 30 minutes to a few hours and cost $50 to $500 depending on the model and the service. The modest cost is worth noting, because it means the expensive part of a fine-tuning project is the human work of curating data, not the compute.
Evaluation is where the real work happens. Before declaring a fine-tuned model ready, you need a held-out test set: examples that were not used in training, reserved specifically to measure how well the model generalizes to cases it has not seen. If you have 300 examples, a typical split is 250 for training and 50 for testing. Evaluating on your training data tells you only that the model memorized what you showed it, which is never the question you are trying to answer.
Evaluate on multiple dimensions rather than a single overall impression. Format compliance asks whether the output follows the required structure. Factual grounding asks whether the output stays within what the input provided or hallucinates additional information, which matters especially in domains where a confident invented detail is worse than an omission. Quality rating asks a domain expert to rate a sample of outputs on a 1 to 5 scale, blind, meaning without knowing whether a given output came from the fine-tuned model or the baseline. The blinding is what makes the exercise informative; an expert who knows which model produced which output will tend to confirm the expected result.
If results fall short, the most common culprits are insufficient training examples, inconsistent quality in the training data, or a base model that is not well suited to the task. In that order. Iterate on the data before assuming the model architecture is wrong, because data problems are both more likely and far cheaper to fix than a change of approach.
Deployment and Maintenance
A fine-tuned model is not a fire-and-forget solution, and it needs ongoing care for two distinct reasons. The first is distribution shift: the real-world inputs you encounter after deployment gradually diverge from the data you trained on. A model fine-tuned on 2022 to 2024 property briefs may struggle when market conditions change and the language professionals use changes with them. The degradation is gradual and quiet, which is exactly why it needs a scheduled check rather than an alert. Plan for quarterly or semi-annual re-evaluation against a current sample.
The second reason is base model updates. If you fine-tune on top of a specific model version and that version is later deprecated, you will need to re-run fine-tuning on the new version. Track which base model version you trained on, keep the training data in a state where it can be reused, and build re-training into your maintenance schedule as soon as a provider announces a deprecation rather than when the cutoff arrives. Fine-tuning is not a one-time investment; it is a recurring cost you commit to when you decide the quality improvement is worth it. Deciding that consciously at the start is what separates a fine-tuning program from a fine-tuning experiment that quietly becomes production infrastructure nobody owns.
Anti-Patterns
- Reaching for fine-tuning first. Treating it as the standard response to disappointing output, rather than the response after prompt engineering has genuinely hit its ceiling.
- Fine-tuning to teach facts. Using training to inject knowledge the model lacks, when retrieval-augmented generation is the tool designed for that job.
- Training on whatever examples exist. Historical outputs your team considered mediocre become the standard the model reproduces, confidently and consistently.
- A narrow training set. Examples drawn from one slice of your caseload produce a model that is poorly calibrated everywhere else, and the failure shows up in production rather than in training.
- Evaluating on training data. Without a held-out test set you learn only that the model memorized your examples, which was never in doubt.
- Unblinded expert review. A reviewer who knows which output came from the fine-tuned model will tend to confirm the result the project needs.
- Treating deployment as the finish line. Distribution shift and base model deprecation both arrive on their own schedule, and neither announces itself through an error.
Practice Prompts
- Diagnose the ceiling. Take a task where prompting underperforms and write down the recent failures. Decide for each whether a better instruction would have fixed it, and see whether fine-tuning is actually indicated.
- Classify the problem. For one underperforming use case, state whether the gap is style, format, or missing facts, then name which of prompting, retrieval, or fine-tuning matches it.
- Audit a sample. Pull a set of historical outputs you would use as training data and grade each against the standard you want the model to reach. Count how many you would have to rewrite.
- Check your coverage. List the categories of case your system will meet in production and mark which are represented in your candidate training data. Note the ones with no examples at all.
- Design the split. Decide in advance which examples are held out for testing and write down the criteria for that choice before you look at any results.
- Write the evaluation rubric. Define what format compliance and factual grounding mean for your specific outputs, precisely enough that two reviewers would score the same output the same way.
- Cost the maintenance. Estimate the re-evaluation and re-training work over the model's expected life and add it to the business case before the project is approved.
Reflection
Consider a task where your team is dissatisfied with model output and ask what kind of gap it actually is. Teams reach for fine-tuning when the model sounds wrong, but sounding wrong has several causes with different remedies. If the model lacks information, no amount of training on examples will supply it. If the instruction is ambiguous, the model is doing exactly what it was asked. Only when the target is a pattern that must hold consistently, and prompting has failed to make it hold, is fine-tuning the tool that matches the problem.
Then ask the maintenance question honestly. If the base model you fine-tuned on were deprecated next quarter, would you still have the training data, the split, and the evaluation rubric needed to reproduce the work? If the answer is no, the fine-tuned model in production is a one-off artifact rather than a maintained asset, and its quality will decay on a schedule nobody is watching.
Glossary
- Foundation model. A model trained on broad general text, learning language patterns, facts, and reasoning strategies before any task-specific adaptation.
- Fine-tuning. Continuing to train a pre-trained model on a smaller curated dataset so that patterns from your examples become the model's default behavior.
- Weights. The numerical parameters that determine a model's behavior, and the thing fine-tuning actually adjusts.
- Retrieval-augmented generation. Connecting a model to a searchable document store so it can draw on information it was never trained on.
- Supervised fine-tuning. Training on input and output pairs, where each example gives the prompt and the ideal completion.
- JSON Lines. The common training file format, with one training example per line structured as a JSON object.
- Hyperparameters. Settings that control how training proceeds, of which learning rate and number of epochs are the two most important.
- Held-out test set. Examples deliberately excluded from training and reserved to measure how well the model generalizes.
- Distribution shift. The gradual divergence of real-world inputs from the data a model was trained on.
- Base model deprecation. Retirement of the specific model version you fine-tuned on, which forces re-training on a successor.
Related Lessons
Fine-tuning sits inside the wider question of how to adapt a general model to specific work. Fine-Tuning & Adaptation introduces the technique at a foundational level, and Prompt Engineering & In-Context Learning covers the approach you should exhaust first. Retrieval-Augmented Generation (RAG) is the right tool for the knowledge problems fine-tuning cannot solve, and System Prompts and Persona Engineering addresses the long style instructions that fine-tuning can eventually replace. On the data side, Data Augmentation and Synthetic Data Techniques is relevant when you cannot assemble enough real examples. For the evaluation half of the work, see Comprehensive Evaluation Frameworks and Building Quality Rubrics for AI Outputs, and for the maintenance half, Monitoring & Optimization and Model Performance Risk Management.
Closing
Helena's four months of prompt engineering were not wasted. They established what the target actually was, produced a working definition of "our voice," and generated exactly the evidence needed to decide that the remaining 30% was not a prompting problem. That is the sequence worth copying. Fine-tuning is a reasonable technique with a narrow purpose: it makes a pattern hold that instructions could not make hold. It will not tell the model anything it does not know, it will not rescue mediocre examples, and it will not stay good on its own. If you go in with curated data, a held-out test set, a blind review, and a maintenance plan, it is a manageable piece of engineering. If you go in expecting the training run to be the hard part, you will find out otherwise.
Key Takeaways
- Fine-tuning adjusts behavior, not factual knowledge. It teaches a model style, format, and domain-specific patterns rather than new facts. For knowledge, retrieval-augmented generation is usually the right tool.
- Start with prompt engineering. Fine-tuning is not the first response to poor model output; it is the response when prompting has genuinely hit its ceiling after serious effort.
- Data quality is the single most important variable. Training on mediocre examples produces a mediocre model, so curate carefully and remove anything below your quality bar.
- You need enough diverse examples. Fifty to one hundred high-quality pairs show early improvement, and five hundred or more across a representative range of cases produces reliable results.
- Always hold out a test set. Reserve 15 to 20% of your examples for evaluation before training; it is the only way to measure objectively whether fine-tuning improved anything.
- Evaluate on several dimensions and blind the review. Format compliance, factual grounding, and expert quality rating measure different failures, and an unblinded reviewer confirms expectations.
- Plan for ongoing maintenance. Distribution shift and base model deprecation mean scheduled re-evaluation and periodic re-training, and those costs belong in the estimate upfront.
Frequently Asked Questions
How do I know whether my problem needs fine-tuning or retrieval? Ask what is missing from the output. If the model produces well-formed work that contains wrong or absent information, the gap is knowledge and retrieval is the answer. If the model has everything it needs and still writes in the wrong shape, order, or register, the gap is behavior and fine-tuning addresses it. The two are frequently combined in production, with retrieval supplying the facts and fine-tuning enforcing the format.
Is a fine-tuned small model really competitive with a large one? On a specific narrow task, often yes. A fine-tuned 7B-parameter model may equal a 70B model on the task it was trained for, at a fraction of the API cost. The qualifier matters: the advantage holds for the task you trained on and disappears outside it, so this is a trade of generality for efficiency rather than a free improvement.
What if we only have a few dozen examples? Then the honest answer is not yet. Meaningful improvement starts around 50 to 100 high-quality examples, and reliability arrives closer to 500 or more. Below that range, invest the effort in creating and curating examples, or in prompt engineering, rather than in a training run whose output you will not be able to interpret.
How often should we re-evaluate a fine-tuned model in production? Quarterly or semi-annual re-evaluation is a reasonable default for distribution shift, using a current sample rather than the original test set. Base model deprecation is event-driven rather than scheduled: when a provider announces that the version you trained on is being retired, re-training moves to the top of the queue regardless of where you are in the calendar.
Skill.re