AI, decoded

How do you prepare data for fine-tuning an LLM?

Build reviewed examples that match the inputs, conversations and outputs your model will face in use. Cover both routine cases and difficult boundaries, keep evaluation examples separate, and improve the dataset by reading the errors your trained model makes.

· Chain of Thought

Level 4: How AI systems are built · 4.5 Customizing models

Model TrainingAI Engineering

Start with the behavior you want

In episode 43, Maxime Labonne makes data preparation the center of post-training. His starting point is practical: examples should resemble how people will actually use the model. A collection of polished answers is less useful if the deployed assistant must handle follow-up questions, missing details or tool results that those examples never contain.

Write a short description of the behavior you want to change before collecting data. For a support assistant, that might be asking for an order number when it is missing, or producing the required escalation format. If you are still choosing an approach, start with prompting, RAG and fine-tuning.

Build complete, reviewable examples

For each example, save the input the model should see and a reviewed target response. Include the relevant instructions and earlier conversation when they affect the answer. Remove information the deployed model would not have. Otherwise, you are asking it to learn from clues that disappear in production.

Hugging Face’s SFT guide documents both prompt-completion pairs and conversations with speaker roles. For tool use, it includes tool definitions, calls and returned results. Those formats make the training example resemble the interaction you expect the model to handle.

Illustrative example: a returns assistant sees a customer ask to send something back. One example supplies the order number; another omits it. The first target response proceeds to the next permitted step. The second asks for the number. Add a follow-up where the customer corrects a mistaken number. Review whether the target response uses the corrected value.

Check coverage before adding volume

Labonne separates the accuracy of an individual sample from the diversity of the whole dataset. Turn that distinction into two reviews. First, check that each answer is correct for its input. Second, ask which kinds of situations are absent.

For the returns example, make a coverage list: ordinary requests, incomplete requests, contradictory information, unsupported requests and conversations with corrections. Inspect each group rather than accepting a large row count as evidence of coverage. Keep related conversations and paraphrases together when splitting data so a near-copy does not appear in both training and evaluation.

Reserve a set of reviewed cases before training. Compare the original model and the tuned model on those same cases. When you use an evaluation failure to create training examples, preserve a separate untouched check for the next comparison.

Where it falls short

A carefully formatted dataset can still teach the wrong behavior. Conflicting labels need an explicit policy decision; generating more examples will not settle it. Synthetic variations also need review, because different wording can hide the same underlying situation.

Labonne’s advice is to read actual model responses and training samples, then revise the data. Keep that loop narrow: identify a failure, inspect the examples that might explain it, make a targeted change and evaluate again. The useful result is a model that handles unseen cases better, not merely one that reproduces the training set.

Hear it from the guest

“Supervised fine tuning is when you give the models the questions and answers that you expect, and this should closely mimic how users actually use your models in real life.”
“What people don't do enough, in my experience, is reading the responses from the models and reading the samples of training data.”

Quotes lightly edited to remove filler words.

Go deeper

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Fine-TuningLarge Language ModelAccuracyPost-trainingRetrieval-Augmented Generation (RAG)Supervised Fine-TuningTool Use (Function Calling)