Most business tasks that companies hand to AI are narrow. Code an invoice, route an email to the right department, pull the fields out of a delivery note. With a frontier model you pay for broad knowledge you never touch.
AI model distillation is having a large model act as a teacher that demonstrates one well-defined task, then using those demonstrations to train a small model that does only that task, on your own hardware. The result is a model of a few hundred million to eight billion parameters that comes close to the teacher on that one task, at a fraction of the cost per item, without your data leaving the building.
How does distillation work in practice?
Distillation is less exotic than it sounds. The process has six steps, and the first two decide whether the rest is worth doing.
- Pin down the task. One input, one type of output, a fixed set of possible outcomes. "Classify incoming email into twelve categories" is a task. "Help customer service" is not.
- Build a gold set. Have two people who know the work judge 300 real cases. Where they disagree, the vagueness is in the task, not the model. Fix that first.
- Let the teacher label. A large model processes thousands of raw examples with a careful instruction. This becomes your training set.
- Train the student. A small model learns from the labelled set, with LoRA or QLoRA for a generative model or ordinary fine-tuning for a classifier.
- Measure against the gold set. Not against the teacher. The teacher makes mistakes too, and you want to know how the student compares to your people.
- Roll out with a safety net. Cases the student is unsure about go to a person or to a larger model.
The gold set is where projects fail. Skip it and you end up with a model and no idea whether it is any good. Our explainer on training your own AI model on company data covers that quality measurement and how fine-tuning differs from RAG.
Why can a small distilled model match a large one?
Because it doesn't need to know what the large model knows. A frontier model is trained to be reasonable at almost everything. A distilled model only has to do one thing well, and all of its capacity goes there.
The research is unusually consistent on this. In Google Research's Distilling Step-by-Step (2023), a fine-tuned 770-million-parameter T5 model beat the 540-billion-parameter PaLM while using only 80 percent of the available training data. That is a model seven hundred times smaller.
Predibase's LoRA Land (2024) tested 310 fine-tuned models across 31 tasks. Models trained with 4-bit LoRA scored 34 points above their base models on average and 10 points above GPT-4. On average, across narrow tasks. Broad reasoning is a different story, and we come back to it below.
One caveat tends to get lost when this research is summarised: these results hold for tasks with a clear right answer. Once the task opens up, the student's advantage disappears quickly.
Which student model fits which task?
Choosing the student is the main technical decision, and it depends on what has to come out of the model. A label, or text.
| Classifier (encoder) | Small generative model | |
|---|---|---|
| Example | RobBERT-2023 (355M) | a 7-8B model with LoRA |
| Output | a label or score | text or JSON |
| Good for | routing, labelling, sentiment, priority | field extraction, short summaries, rewriting |
| Training hardware | one consumer GPU | 6 GB VRAM (QLoRA) to 22 GB (LoRA) |
| Inference hardware | runs on a CPU | one GPU or an Apple Silicon machine |
| Biggest risk | the task turns out to need text after all | invents fields that are not in the source |
For Dutch text, RobBERT-2023 is the obvious starting point: a Dutch language model from KU Leuven, UGent and TU Berlin, MIT-licensed, scoring 18.6 points above BERTje on the Dutch DUMB benchmark. For English, any well-established encoder plays the same role. At 355 million parameters it is small enough to label thousands of emails a minute on ordinary server hardware.
If the output has to be text, move up to a generative model of seven to eight billion parameters. According to Unsloth's hardware requirements, QLoRA training of an 8B model fits in at least 6 GB of video memory, and 16-bit LoRA in 22 GB. That's one workstation, not a data centre. Our guide to AI hardware for business covers which machine fits.
Can you use ChatGPT, Claude or Gemini output to train a model?
This is the question nearly every distillation project asks too late. The teacher is a vendor's product, and that vendor has terms about what you may do with its output.
| Teacher | What the terms say | What it means for distillation |
|---|---|---|
| OpenAI API | No using output for competing models, with an exception for internal classifiers | An internal classifier is allowed; a generative model falls outside the exception |
| Anthropic (Claude) | No training competing AI models, no exception | Not without written approval |
| Google Gemini API | No developing models that compete with the service | Not without approval |
| DeepSeek-R1 | MIT, distillation explicitly allowed | Free to use |
| OpenAI gpt-oss, Qwen3, Gemma 4 | Apache 2.0 | Free to use |
| Llama 3.1 and later | Allowed | If distributed, the model name must start with "Llama" |
| Gemma 3 and earlier | A distilled model counts as a derivative | Gemma terms carry over |
The details matter. OpenAI's Services Agreement prohibits using output to develop models that compete with OpenAI, except under a "Permitted Exception": models primarily intended to categorise, classify or organise data, as long as you don't make them available to third parties. An internal email router fits. An internal model that writes answers does not. Anthropic's Commercial Terms have no such exception and prohibit training "competing AI models" without express approval. The Gemini API terms prohibit models that compete with the service.
None of the three defines what "competing" means. Which is exactly why you don't want to build on it.
The clean route is an open-weight teacher. DeepSeek-R1 is MIT-licensed and names "distillation for training other LLMs" as permitted use in so many words. OpenAI's own gpt-oss models and Qwen3 are Apache 2.0, and gpt-oss-120b runs on a single 80 GB GPU, which makes it a practical teacher. That brings a second benefit that outweighs the licence: you run an open teacher yourself, so the raw training data never leaves your environment either. It's the same logic we set out in private AI under the GDPR and the EU AI Act.
Does fine-tuning a model make you a provider under the AI Act?
Usually not. In its guidelines for providers of general-purpose AI models of 18 July 2025, the European Commission set out when someone who modifies an existing model becomes a provider in their own right. The indicative criterion: the compute used for the modification exceeds one third of the compute used to train the original model. The guidelines name fine-tuning explicitly as a form of modification.
A LoRA run on a few thousand examples sits many orders of magnitude below that. A distilled classifier usually doesn't fall under the definition of a general-purpose model at all.
That doesn't mean the AI Act leaves your application alone. Obligations for the system you use the model in, for example in decisions about staff or credit, apply regardless of who built the model.
What does a distilled model cost?
Below is a worked example for a classification task: routing incoming email to twelve departments. The assumptions are stated and the figures are indicative.
| Item | Assumption | Indication |
|---|---|---|
| Gold set | 300 cases, two reviewers, 1 minute per case | 10 hours of work |
| Labelling by the teacher | 5,000 emails, open-weight teacher on own hardware | electricity and an afternoon of compute |
| Labelling by hand (for comparison) | 5,000 emails, 30 seconds each | 42 hours of work |
| Training the classifier | RobBERT-2023 on one GPU | a few hours of compute |
| Build, measurement and integration | connection to mail server and ticketing | most of the budget |
| Running it | 200,000 emails a year on a CPU | no cost per item |
What stands out: compute is the smallest line. The money goes into defining the task, the gold set and the integration with your existing systems. That is also where the difference with a frontier API lies. The API charges per item, so every extension to a new mailbox or department raises the bill. Your own model costs the same at 200,000 messages a year as at 2 million.
[ TIME SAVED ]
Save 8 hours per week on manually forwarding and labelling incoming email
The broader arithmetic of running your own models versus paying API costs is in small language models for business, including how model routing combines a small and a large model.
When is distillation the wrong choice?
In three situations, and the last one gets missed most often.
Low volume. If you process a few hundred cases a month, building your own model won't pay back. An API with a good instruction is cheaper and goes live sooner.
A task that keeps changing. A distilled model learns a snapshot. If new categories or rules arrive every quarter, you retrain every quarter, and you update the gold set too.
A task that is actually broad. "Answer customer questions" looks like one task but is a hundred. A student that does well on the common questions fails on exactly the exceptions where a customer notices the difference. That belongs with a larger model, possibly with a distilled model as a filter in front of it. How to make that call across your organisation is covered in our guide to local AI for business.