Add mlops/llm-fine-tuning
This commit is contained in:
@@ -0,0 +1,173 @@
|
||||
---
|
||||
name: llm-fine-tuning
|
||||
description: "LLM fine-tuning: framework comparison (Axolotl, TRL, Unsloth) and method selection (SFT, DPO, PPO, GRPO)."
|
||||
version: 1.0.0
|
||||
author: Hermes Agent
|
||||
license: MIT
|
||||
metadata:
|
||||
hermes:
|
||||
tags: [fine-tuning, llm, axolotl, trl, unsloth, lora, qlora, dpo, ppo, grpo, rlhf, sft, training]
|
||||
related_skills: [huggingface-hub, llm-inference]
|
||||
---
|
||||
|
||||
# LLM Fine-Tuning
|
||||
|
||||
Three frameworks for fine-tuning LLMs, each optimized for different use cases. Choose based on your training objective, then load the detailed reference for that framework.
|
||||
|
||||
## Choosing a Framework
|
||||
|
||||
| Need | Framework | Why |
|
||||
|------|-----------|-----|
|
||||
| Production-grade multi-method training | **Axolotl** | Most mature, YAML-driven, 100+ model presets, LoRA/QLoRA/DPO/KTO/ORPO/GRPO |
|
||||
| Reinforcement learning from human feedback | **TRL** | Native SFT/DPO/PPO/GRPO, HuggingFace-native, research-friendly |
|
||||
| Fastest training on limited hardware | **Unsloth** | 2-5x faster, 50-80% less memory, great for LoRA/QLoRA on single GPU |
|
||||
|
||||
## Choosing a Method
|
||||
|
||||
| Method | Use When | Training Data |
|
||||
|--------|----------|---------------|
|
||||
| **SFT** (Supervised Fine-Tuning) | Instruction tuning, domain adaptation | Prompt-completion pairs |
|
||||
| **DPO** (Direct Preference Optimization) | Align with preferences, no reward model needed | Chosen/rejected pairs |
|
||||
| **PPO** (Proximal Policy Optimization) | RLHF with online reward model | Prompt + reward model |
|
||||
| **GRPO** (Group Relative Policy Optimization) | Rule-based rewards, no value model needed | Prompt + rule-based reward |
|
||||
| **KTO/ORPO** | Simpler alignment, no pairwise preferences | Binary good/bad signals |
|
||||
| **LoRA/QLoRA** | Parameter-efficient, limited GPU VRAM | (applies to any method above) |
|
||||
|
||||
## Section 1: Axolotl — YAML-Driven Multi-Method Training
|
||||
|
||||
Production-grade fine-tuning framework. YAML configs, 100+ model presets, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, multimodal support.
|
||||
|
||||
**When to use:** Production training jobs, multi-method experiments, need reproducibility, working with large models.
|
||||
|
||||
**Key features:**
|
||||
- YAML-driven configuration (no Python scripts needed)
|
||||
- 100+ model presets (Llama, Mistral, Qwen, Gemma, etc.)
|
||||
- DeepSpeed FSDP integration for multi-GPU
|
||||
- LoRA/QLoRA/DPO/KTO/ORPO/GRPO support
|
||||
- Multimodal fine-tuning (vision-language models)
|
||||
- Cloud backends (Modal, RunPod)
|
||||
|
||||
**Quick start:**
|
||||
```bash
|
||||
pip install axolotl
|
||||
# Create a YAML config, then:
|
||||
accelerate launch -m axolotl.cli.train config.yaml
|
||||
```
|
||||
|
||||
**Key config pattern (FSDP v2):**
|
||||
```yaml
|
||||
fsdp_version: 2
|
||||
fsdp_config:
|
||||
offload_params: true
|
||||
state_dict_type: FULL_STATE_DICT
|
||||
auto_wrap_policy: TRANSFORMER_BASED_WRAP
|
||||
```
|
||||
|
||||
**Pitfalls:**
|
||||
- `context_parallel_size` must divide total GPUs
|
||||
- `save_compressed: true` reduces disk 40%, compatible with vLLM
|
||||
- Integration code need not go in `integrations/` folder
|
||||
- Handle both single-example and batched data shapes in custom collators
|
||||
|
||||
**See:** `references/axolotl.md` for full config reference, API docs, and dataset formats.
|
||||
|
||||
## Section 2: TRL — Transformer Reinforcement Learning
|
||||
|
||||
HuggingFace-native post-training library. SFT, DPO, PPO, GRPO, and reward model training.
|
||||
|
||||
**When to use:** RLHF workflows, preference alignment, research experiments, need tight HuggingFace integration.
|
||||
|
||||
**Key features:**
|
||||
- `SFTTrainer` — supervised instruction tuning
|
||||
- `DPOTrainer` — direct preference optimization
|
||||
- `PPOTrainer` — proximal policy optimization (classic RLHF)
|
||||
- `GRPOTrainer` — group relative policy optimization (no value model)
|
||||
- Reward model training and online RL methods
|
||||
- Works with any HuggingFace model + PEFT for LoRA
|
||||
|
||||
**Quick start (SFT):**
|
||||
```python
|
||||
from trl import SFTTrainer
|
||||
trainer = SFTTrainer(
|
||||
model="Qwen/Qwen2.5-0.5B",
|
||||
train_dataset=dataset,
|
||||
)
|
||||
trainer.train()
|
||||
```
|
||||
|
||||
**Quick start (DPO):**
|
||||
```python
|
||||
from trl import DPOTrainer, DPOConfig
|
||||
config = DPOConfig(output_dir="model-dpo", beta=0.1)
|
||||
trainer = DPOTrainer(
|
||||
model=model, args=config,
|
||||
train_dataset=preference_dataset,
|
||||
processing_class=tokenizer
|
||||
)
|
||||
trainer.train()
|
||||
```
|
||||
|
||||
**Pitfalls:**
|
||||
- DPO `beta` controls alignment strength — too high reduces diversity
|
||||
- PPO requires a reward model + reference model (memory-intensive)
|
||||
- GRPO eliminates the value model but needs careful reward normalization
|
||||
- Always use `processing_class=` instead of deprecated `tokenizer=`
|
||||
|
||||
**See:** `references/trl-fine-tuning.md` for full method references and `templates/` for starter scripts.
|
||||
|
||||
## Section 3: Unsloth — Fast & Memory-Efficient Fine-Tuning
|
||||
|
||||
2-5x faster training with 50-80% less memory. Optimized for LoRA/QLoRA on consumer GPUs.
|
||||
|
||||
**When to use:** Single-GPU training, limited VRAM (4-24GB), rapid iteration, learning/experimentation.
|
||||
|
||||
**Key features:**
|
||||
- 2-5x faster than standard HuggingFace training
|
||||
- 50-80% less VRAM usage
|
||||
- Auto-selects optimal LoRA/QLoRA parameters
|
||||
- Supports Llama, Mistral, Gemma, Qwen architectures
|
||||
- Drop-in replacement: same HuggingFace `Trainer` API
|
||||
|
||||
**Quick start:**
|
||||
```python
|
||||
from unsloth import FastLanguageModel
|
||||
model, tokenizer = FastLanguageModel.from_pretrained(
|
||||
model_name="unsloth/llama-3-8b-bnb-4bit",
|
||||
max_seq_length=2048,
|
||||
load_in_4bit=True,
|
||||
)
|
||||
model = FastLanguageModel.get_peft_model(
|
||||
model, r=16,
|
||||
lora_alpha=16, lora_dropout=0,
|
||||
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
|
||||
)
|
||||
```
|
||||
|
||||
**Pitfalls:**
|
||||
- Only supports specific architectures (Llama, Mistral, Gemma, Qwen)
|
||||
- May lag behind upstream Transformers version
|
||||
- Best for LoRA; full fine-tuning not the primary target
|
||||
|
||||
**See:** `references/unsloth.md` for full documentation.
|
||||
|
||||
## Comparison Table
|
||||
|
||||
| Feature | Axolotl | TRL | Unsloth |
|
||||
|---------|---------|-----|---------|
|
||||
| SFT | ✅ | ✅ | ✅ |
|
||||
| DPO | ✅ | ✅ | ❌ |
|
||||
| PPO/GRPO | ✅ | ✅ | ❌ |
|
||||
| LoRA/QLoRA | ✅ | ✅ (via PEFT) | ✅ (native) |
|
||||
| Full fine-tune | ✅ | ✅ | ⚠️ Limited |
|
||||
| Multi-GPU | ✅ DeepSpeed | ✅ Accelerate | ⚠️ Limited |
|
||||
| YAML config | ✅ | ❌ Python | ❌ Python |
|
||||
| Speed | Baseline | Baseline | 2-5x faster |
|
||||
| Memory | Baseline | Baseline | 50-80% less |
|
||||
| Cloud backends | Modal, RunPod | Any | Any |
|
||||
|
||||
## Workflow: Which Path?
|
||||
|
||||
1. **Need the fastest iteration on limited hardware?** → Start with Unsloth
|
||||
2. **Need RLHF / preference alignment?** → Use TRL (or Axolotl's TRL backend)
|
||||
3. **Need production reproducibility + multi-method?** → Use Axolotl
|
||||
4. **Training complete?** → Export to GGUF (see `llm-inference` skill) or serve with vLLM
|
||||
Reference in New Issue
Block a user