METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅈTechnical words in the news

Supervised Fine-Tuning

A follow-up training stage where an already-trained model is shown paired question-and-answer examples to reteach it how to respond the way you want.

In plain words

Supervised Fine-Tuning is a training stage where a model that has already finished its initial training is shown pairs of questions and model answers, so it can be retaught to respond in a desired way.

Think of it like a fresh new hire. A new employee already has broad knowledge and language skills, but still needs to learn the company's preferred report format or how to talk to customers. Just as a company hands over a manual full of model answers for the employee to follow, supervised fine-tuning repeatedly shows the model a set of questions with correct answers attached, training it into the habit of responding in a specific tone, format, or with domain-specific knowledge.

Recently, the barrier to doing this has dropped sharply. It used to require expensive server-grade hardware to retrain a large model, but now lightweight methods that retrain only a small portion of the model have spread, making it possible to train a model in the tens-of-billions-of-parameters range using a single gaming graphics card. As a result, more individuals and small teams are building their own customized models tailored to specific languages or domains.

How it shows up in the news

In articles, this often appears abbreviated as SFT or 'supervised learning (SFT).' The training guide for Qwen3.8-27B explains that 'to do supervised learning (SFT) with text-only data, you just need to prepare a dataset with a text column rendered using the Qwen chat template,' while an article on the French-specialized model Luth-2 stated that it newly built a 3-billion-token supervised learning (SFT) dataset covering math, knowledge, code, tool calls, and multi-turn conversations. A common point of confusion is that supervised fine-tuning is not the process of building a model from scratch, but a follow-up stage that refines a model that already has baseline capabilities.

Try it yourself

  1. Find a public dataset with paired questions and answers.
  2. Open a free training notebook environment and choose an open-source model you want to retrain.
  3. Format the dataset to match the conversation format the model requires.
  4. Run supervised fine-tuning so the model learns to answer following the new examples.
  5. Save the trained model and ask it questions yourself to see how its way of answering has changed.

See also

Stories using this term

Browse every entry