Unsloth is an optimised fine-tuning library that makes LoRA training 2–5× faster with 70% less VRAM. This guide walks through fine-tuning a Llama or Mistral model on your own data using Unsloth on a free Google Colab GPU.
Fine-tuning is the right choice when: you have 500+ high-quality examples of the task; the task requires a consistent style or format that prompting cannot reliably produce; you need the lowest possible inference cost for a specialised task; or you want to inject domain knowledge not present in the base model.
Fine-tuning is overkill when: a well-crafted system prompt and few-shot examples achieve 90%+ of the quality you need; you have fewer than 200 training examples; or you are still exploring the problem space. Always try prompt engineering first.
Low-Rank Adaptation (LoRA) fine-tunes a model by adding small trainable adapter matrices to the attention layers, rather than updating all weights. The base model is frozen; only the adapters (0.1–2% of total parameters) are trained. This drastically reduces VRAM requirements and training time.
QLoRA combines LoRA with 4-bit quantisation of the base model, enabling fine-tuning of a 7B model on a single 16 GB GPU, or a 13B model on a 24 GB GPU. Unsloth implements an optimised version of QLoRA that is significantly faster than the reference implementation.
Open a new Colab notebook with a T4 (free) or A100 (Pro) GPU. Install Unsloth: `!pip install 'unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git'` followed by `!pip install --no-deps trl peft accelerate bitsandbytes`.
Load a model: `from unsloth import FastLanguageModel; model, tokenizer = FastLanguageModel.from_pretrained(model_name='unsloth/Phi-4-mini-instruct', max_seq_length=2048, dtype=None, load_in_4bit=True)`. Unsloth supports Llama 3, Mistral, Phi, Gemma, and Qwen families.
Format your data as a list of instruction-response pairs. Each example should be a dict with an `instruction` (the input) and `output` (the expected response). For chat models, use the model's chat template to format examples correctly.
Use `datasets` to load and process your data: `from datasets import Dataset; data = Dataset.from_list(your_list)`. Add a formatting function that applies the chat template and stores the result in a `text` field. Aim for at least 500 examples; 2000+ is better for reliable results.
Add LoRA adapters: `model = FastLanguageModel.get_peft_model(model, r=16, target_modules=['q_proj', 'v_proj'], lora_alpha=16, lora_dropout=0)`. `r=16` is a good starting value; higher r means more parameters but better fit.
Set up the Trainer with SFTTrainer from the `trl` library: `trainer = SFTTrainer(model=model, train_dataset=dataset, dataset_text_field='text', max_seq_length=2048, args=TrainingArguments(per_device_train_batch_size=2, num_train_epochs=3, ...))`. Start training with `trainer.train()`. On a T4 GPU, 1000 examples with a 7B model takes about 20–30 minutes.
Save as GGUF for local deployment: `model.save_pretrained_gguf('my-model', tokenizer, quantization_method='q4_k_m')`. This exports a GGUF file you can run directly with Ollama or LM Studio.
Push to Hugging Face Hub: `model.push_to_hub_gguf('your-username/my-model', tokenizer, quantization_method='q4_k_m', token='hf_...')`. Once uploaded, you can pull it with Ollama from any machine.
The quality of your training data is the biggest factor in fine-tuning success. This guide covers data collection strategies, formatting standards, quality filtering, and the minimum viable dataset size for different tasks.
Read guideKnowing when your fine-tuned model is actually better than the base model requires systematic evaluation. This guide covers benchmark datasets, LLM-as-judge evaluation, and metrics for task-specific assessment.
Read guideDPO trains models to prefer good responses over bad ones using human preference data — without the complexity of reinforcement learning. This guide covers collecting preference data, training with TRL's DPO trainer, and evaluating results.
Read guide