0Pricing
AI Engineering Academy · 课时

使用 Hugging Face PEFT 进行 LoRA 微调

配置 LoRA 秩、alpha 和目标模块,使用 TRL SFTTrainer 运行有监督微调,监控训练损失,并保存合并检查点和仅包含适配器的检查点。

使用 Hugging Face PEFT 进行 LoRA 微调 是 CoddyKit 上的免费 AI Engineering Academy 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 AI Engineering Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 AI Engineering Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why LoRA Instead of Full Fine-Tuning?

Full fine-tuning updates every parameter in a model. For a 7-billion parameter model at float32 precision, that requires ~28GB of GPU memory just for the weights, plus optimizer states, gradients, and activations — easily 80-120GB total. LoRA (Low-Rank Adaptation) instead adds a tiny number of trainable parameters (typically 0.1-1% of total) as low-rank matrix pairs that are applied to selected layers, reducing the GPU requirement by 10-50x while achieving comparable results.

# LoRA math intuition:
# Full fine-tuning: update W (large matrix, e.g., 4096 x 4096 = 16.7M parameters)
# LoRA: instead train W = W_0 + A @ B where:
#   A has shape (4096, r) - only r*4096 params
#   B has shape (r, 4096) - only r*4096 params
#   r (rank) is typically 4, 8, or 16 - much smaller than 4096

# Memory comparison for 7B model:
# Full fine-tuning: ~80GB GPU RAM
# LoRA (rank=8): ~12GB GPU RAM - fits on a single A100 or 3090
print('LoRA makes fine-tuning accessible without massive GPU clusters')

Installing the Required Libraries

LoRA fine-tuning with Hugging Face requires three libraries: transformers (model loading and tokenization), peft (Parameter-Efficient Fine-Tuning, which implements LoRA), and trl (Transformer Reinforcement Learning, which provides the SFTTrainer for supervised fine-tuning). Together these provide a high-level, production-ready fine-tuning workflow.

# pip install transformers peft trl accelerate bitsandbytes datasets

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, TaskType
from trl import SFTTrainer, SFTConfig
from datasets import Dataset
import torch

print('Libraries imported successfully')
print(f'GPU available: {torch.cuda.is_available()}')
if torch.cuda.is_available():
    print(f'GPU: {torch.cuda.get_device_name(0)}')
    print(f'GPU memory: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB')

Loading the Base Model with Quantization

For most LoRA fine-tuning, load the base model in 4-bit quantization using bitsandbytes. This reduces the base model's memory footprint by 75% (a 7B model goes from ~14GB to ~4GB) while retaining most of the model's capability. Only the LoRA adapter layers are trained in full precision. The combination of 4-bit quantization + LoRA is called QLoRA and makes fine-tuning 7B models accessible on consumer GPUs.

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch

MODEL_NAME = 'mistralai/Mistral-7B-Instruct-v0.2'  # or 'meta-llama/Llama-3.2-3B-Instruct'

# 4-bit quantization config (QLoRA)
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,                         # quantize base model to 4-bit
    bnb_4bit_quant_type='nf4',                 # NF4 quantization type
    bnb_4bit_compute_dtype=torch.float16,      # compute in float16
    bnb_4bit_use_double_quant=True             # double quantization for extra savings
)

# Load quantized model
model = AutoModelForCausalLM.from_pretrained(
    MODEL_NAME,
    quantization_config=bnb_config,
    device_map='auto',                         # automatically distribute across GPUs
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.pad_token = tokenizer.eos_token    # required for training batches
print(f'Model loaded. Parameters: {model.num_parameters():,}')

Configuring LoRA Hyperparameters

The three most important LoRA hyperparameters are: r (rank) — the dimension of the low-rank matrices; higher rank means more trainable parameters and more expressive adaptation but more memory and risk of overfitting. lora_alpha — a scaling factor usually set to 2x the rank. target_modules — which weight matrices to apply LoRA to; attention layers (q_proj, v_proj) are the most common choice. Start with r=8 and tune from there.

from peft import LoraConfig, TaskType

lora_config = LoraConfig(
    task_type=TaskType.CAUSAL_LM,         # causal language modeling
    r=8,                                   # rank: 4, 8, 16, 32 — higher = more params
    lora_alpha=16,                         # scaling: usually 2*r
    lora_dropout=0.1,                      # dropout on LoRA layers for regularization
    target_modules=['q_proj', 'v_proj',    # which weight matrices to adapt
                    'k_proj', 'o_proj',    # common to target all attention projections
                    'gate_proj', 'up_proj', 'down_proj'],  # and FFN layers
    bias='none',                           # do not adapt bias parameters
)

# Wrap the model with LoRA adapters
from peft import get_peft_model, prepare_model_for_kbit_training

model = prepare_model_for_kbit_training(model)  # prepares quantized model for training
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# Output: trainable params: 4,194,304 || all params: 3,752,071,168 || trainable%: 0.11%

Preparing the Dataset

The SFTTrainer expects datasets in a specific format. The simplest format is a dataset with a single text column containing the fully formatted prompt + response string. Use the tokenizer's chat template to format conversation examples consistently. The SFTTrainer handles tokenization, batching, and loss masking (computing loss only on the assistant's responses, not on the input prompt).

from datasets import Dataset
import json

def load_and_format_dataset(jsonl_path: str, tokenizer) -> Dataset:
    examples = []
    with open(jsonl_path) as f:
        for line in f:
            ex = json.loads(line.strip())
            # Apply chat template to format as expected by the model
            formatted = tokenizer.apply_chat_template(
                ex['messages'],
                tokenize=False,
                add_generation_prompt=False
            )
            examples.append({'text': formatted})
    
    return Dataset.from_list(examples)

# Load training and validation datasets
train_dataset = load_and_format_dataset('train.jsonl', tokenizer)
val_dataset = load_and_format_dataset('validation.jsonl', tokenizer)

print(f'Train examples: {len(train_dataset)}')
print(f'Validation examples: {len(val_dataset)}')
print('Sample formatted text:')
print(train_dataset[0]['text'][:300])

Training with SFTTrainer

The SFTTrainer from TRL wraps the Hugging Face Trainer with supervised fine-tuning defaults. Configure it with training hyperparameters: number of epochs (1-3 is usually sufficient for instruction fine-tuning), batch size, gradient accumulation steps (to simulate larger batches with limited memory), learning rate (2e-4 is a common starting point), and the output directory for checkpoints.

from trl import SFTTrainer, SFTConfig

training_args = SFTConfig(
    output_dir='./fine-tuned-model',
    num_train_epochs=2,                    # 2-3 epochs for instruction fine-tuning
    per_device_train_batch_size=4,         # increase if GPU memory allows
    gradient_accumulation_steps=4,         # effective batch size = 4*4 = 16
    learning_rate=2e-4,                    # typical LoRA learning rate
    warmup_ratio=0.05,                     # warmup for 5% of steps
    lr_scheduler_type='cosine',            # cosine decay learning rate schedule
    logging_steps=10,
    eval_strategy='steps',
    eval_steps=50,                         # evaluate on validation set every 50 steps
    save_steps=100,
    max_seq_length=2048,                   # max token length per example
    fp16=True,                             # mixed precision training
    report_to='none'                       # or 'wandb' for experiment tracking
)

trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset,
    eval_dataset=val_dataset,
)

print('Starting LoRA fine-tuning...')
trainer.train()

Monitoring Training Progress

During training, watch two key metrics: training loss should decrease steadily. validation loss should decrease at first, then plateau or slightly increase (overfitting). If validation loss increases significantly while training loss continues decreasing, stop training early — the model is memorizing training examples rather than generalizing. The optimal stopping point is just before validation loss starts rising.

# Reading training logs
# Training step logs look like:
# {'loss': 1.4523, 'grad_norm': 0.85, 'learning_rate': 0.0002, 'epoch': 0.2, 'step': 20}
# {'loss': 1.2341, 'grad_norm': 0.72, 'learning_rate': 0.00018, 'epoch': 0.4, 'step': 40}

# Validation results look like:
# {'eval_loss': 1.1823, 'eval_runtime': 12.3, 'eval_samples_per_second': 8.1, 'step': 50}
# {'eval_loss': 1.0923, 'eval_runtime': 12.1, 'eval_samples_per_second': 8.3, 'step': 100}
# {'eval_loss': 1.1234, 'eval_runtime': 12.4, 'eval_samples_per_second': 8.2, 'step': 150}
# ^ validation loss went UP at step 150 - overfitting starting

# Load the best checkpoint (lowest validation loss)
from transformers import TrainerCallback

print('Best model checkpoint is saved automatically by SFTTrainer (load_best_model_at_end=True)')

Saving LoRA Adapters and Merging

After training, save the LoRA adapter weights separately from the base model. The adapter is tiny (a few MB to a few hundred MB) and can be applied to the base model at inference time. For production deployment, you can also merge the LoRA weights into the base model, creating a single model file that does not require the PEFT library at inference time — this reduces inference overhead.

# Save only the LoRA adapter (tiny - typically 10-100MB)
model.save_pretrained('./lora-adapter-only')
tokenizer.save_pretrained('./lora-adapter-only')

# Load the adapter for inference (requires base model + peft)
from peft import PeftModel

base_model = AutoModelForCausalLM.from_pretrained(MODEL_NAME, torch_dtype=torch.float16, device_map='auto')
model_with_adapter = PeftModel.from_pretrained(base_model, './lora-adapter-only')

# Alternative: Merge adapter into base model (no PEFT needed at inference)
print('Merging LoRA weights into base model...')
merged_model = model_with_adapter.merge_and_unload()  # creates a regular model
merged_model.save_pretrained('./merged-model')
print('Merged model saved - can be used with standard transformers, no PEFT needed')

Running Inference on the Fine-Tuned Model

Test your fine-tuned model on a set of held-out prompts before declaring success. Compare outputs side-by-side with the base model on the same prompts to verify that the fine-tuning achieved its goals. Check both the cases it should handle better (the target task) and cases it should still handle well (general tasks that you do not want to have degraded).

def generate(model, tokenizer, prompt: str, max_new_tokens=512) -> str:
    inputs = tokenizer(prompt, return_tensors='pt').to(model.device)
    with torch.no_grad():
        outputs = model.generate(
            **inputs,
            max_new_tokens=max_new_tokens,
            temperature=0.1,
            do_sample=True,
            pad_token_id=tokenizer.eos_token_id
        )
    # Decode only the new tokens (not the input prompt)
    new_tokens = outputs[0][inputs['input_ids'].shape[1]:]
    return tokenizer.decode(new_tokens, skip_special_tokens=True)

# Compare base vs fine-tuned
test_prompt = 'Extract JSON from: "Alice Johnson, 28, software engineer in NYC since 2021"'
print('=== BASE MODEL ===')
print(generate(base_model, tokenizer, test_prompt))
print('\n=== FINE-TUNED MODEL ===')
print(generate(merged_model, tokenizer, test_prompt))

LoRA Key Hyperparameters Summary

To tune LoRA for your use case, start with these defaults and adjust one at a time. r=8 is a safe starting point — increase to 16 or 32 if the model needs more expressive capacity. lora_alpha = 2*r is a stable choice. Learning rate 2e-4 works for most instruction fine-tuning; decrease to 1e-4 if you see unstable training loss. Epochs 1-3: stop when validation loss stops decreasing.

When to Use OpenAI Fine-Tuning Instead

Hugging Face PEFT LoRA requires a GPU. If you do not have GPU infrastructure, OpenAI's fine-tuning API is a managed alternative that handles the training infrastructure for you. Upload your JSONL file, call the API to start a training run, and receive a fine-tuned model ID you can use in API calls. OpenAI fine-tuning supports GPT-4o-mini and GPT-3.5-turbo. It is more expensive per token but eliminates infrastructure management entirely.

from openai import OpenAI

client = OpenAI()

# Upload training file
train_file = client.files.create(
    file=open('train.jsonl', 'rb'),
    purpose='fine-tune'
)
val_file = client.files.create(
    file=open('validation.jsonl', 'rb'),
    purpose='fine-tune'
)

# Create fine-tuning job
job = client.fine_tuning.jobs.create(
    training_file=train_file.id,
    validation_file=val_file.id,
    model='gpt-4o-mini',  # base model to fine-tune
    hyperparameters={
        'n_epochs': 3,
        'batch_size': 'auto',
        'learning_rate_multiplier': 'auto'
    }
)
print(f'Fine-tuning job created: {job.id}')
# Monitor: client.fine_tuning.jobs.retrieve(job.id)
# Use: client.chat.completions.create(model=job.fine_tuned_model, ...)

Quick Check

Test your understanding of LoRA fine-tuning from this lesson.

Lesson Recap

In this lesson you learned: LoRA reduces fine-tuning GPU requirements by 10-50x by training only small low-rank adapter matrices instead of all model parameters, QLoRA (4-bit quantization + LoRA) makes 7B model fine-tuning accessible on consumer GPUs with ~12GB VRAM, and the SFTTrainer from TRL provides a high-level API that handles tokenization, batching, loss masking, and checkpoint saving. Next up we evaluate and deploy the fine-tuned model.

常见问题解答

「使用 Hugging Face PEFT 进行 LoRA 微调」课时是免费的吗?

是的 — 「使用 Hugging Face PEFT 进行 LoRA 微调」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 AI Engineering Academy 课程的其余内容,请升级到 CoddyKit PRO。 AI Engineering Academy 课程共包含 4 节课。

「使用 Hugging Face PEFT 进行 LoRA 微调」这节课中我会学到什么?

配置 LoRA 秩、alpha 和目标模块,使用 TRL SFTTrainer 运行有监督微调,监控训练损失,并保存合并检查点和仅包含适配器的检查点。 你通过在浏览器中直接运行的动手代码来练习 AI Engineering Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 AI Engineering Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 AI Engineering Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「使用 Hugging Face PEFT 进行 LoRA 微调」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 AI Engineering Academy 课中编写并运行代码吗?

能。每节 AI Engineering Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 微调何时胜过提示工程
  2. 准备高质量训练数据集
  3. 使用 Hugging Face PEFT 进行 LoRA 微调
  4. 评估并部署您的微调模型
← 返回 AI Engineering Academy