Error while fine-tuning distilbert model

AbdulrahmanAhmed · November 17, 2024, 12:24am

from transformers import AutoModelForSequenceClassification
from peft import LoraConfig, get_peft_model, TaskType
import torch

quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype= torch.float16
)

model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased" ,
    quantization_config = quantization_config
  )

lora_config = LoraConfig(
    r=8,
    lora_alpha=8,
    target_modules=["q_lin", "k_lin", "v_lin", "out_lin"],
    lora_dropout=0.1,
    task_type=TaskType.SEQ_CLS,
    bias="none"
)

model = get_peft_model(model, lora_config)

i tried to quantize and fine-tune distilbert-base-uncased using BnB and peft but I got this error message.

low_cpu_mem_usage` was None, now default to True since model is quantized.
Some weights of DistilBertForSequenceClassification were not initialized from the model checkpoint at distilbert-base-uncased and are newly initialized: ['classifier.bias', 'classifier.weight', 'pre_classifier.bias', 'pre_classifier.weight']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.
---------------------------------------------------------------------------
RuntimeError                              Traceback (most recent call last)
<ipython-input-30-df44ef690a88> in <cell line: 26>()
     24 )
     25 
---> 26 model = get_peft_model(model, lora_config)

5 frames
/usr/local/lib/python3.10/dist-packages/peft/mapping.py in get_peft_model(model, peft_config, adapter_name, mixed, autocast_adapter_dtype, revision)
    191     if peft_config.is_prompt_learning:
    192         peft_config = _prepare_prompt_learning_config(peft_config, model_config)
--> 193     return MODEL_TYPE_TO_PEFT_MODEL_MAPPING[peft_config.task_type](
    194         model, peft_config, adapter_name=adapter_name, autocast_adapter_dtype=autocast_adapter_dtype
    195     )

/usr/local/lib/python3.10/dist-packages/peft/peft_model.py in __init__(self, model, peft_config, adapter_name, **kwargs)
   1396 
   1397         # to make sure classifier layer is trainable; this may add a new ModulesToSaveWrapper
-> 1398         _set_trainable(self, adapter_name)
   1399 
   1400     def add_adapter(self, adapter_name: str, peft_config: PeftConfig) -> None:

/usr/local/lib/python3.10/dist-packages/peft/utils/other.py in _set_trainable(model, adapter_name)
    401                 target.set_adapter(target.active_adapter)
    402             else:
--> 403                 new_module = ModulesToSaveWrapper(target, adapter_name)
    404                 new_module.set_adapter(adapter_name)
    405                 setattr(parent, target_name, new_module)

/usr/local/lib/python3.10/dist-packages/peft/utils/other.py in __init__(self, module_to_save, adapter_name)
    197         self._active_adapter = adapter_name
    198         self._disable_adapters = False
--> 199         self.update(adapter_name)
    200         self.check_module()
    201 

/usr/local/lib/python3.10/dist-packages/peft/utils/other.py in update(self, adapter_name)
    261         self.original_module.requires_grad_(False)
    262         if adapter_name == self.active_adapter:
--> 263             self.modules_to_save[adapter_name].requires_grad_(True)
    264 
    265     def _create_new_hook(self, old_hook):

/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py in requires_grad_(self, requires_grad)
   2885         """
   2886         for p in self.parameters():
-> 2887             p.requires_grad_(requires_grad)
   2888         return self
   2889 

RuntimeError: only Tensors of floating point dtype can require gradients

it appears only when choosing task_type = TaskType.SEQ_CLS or TOKEN_CLS inside LoraConfig.
if I choose any other type like QUESTION_ANS or CAUSAL_LM no error occurs.

John6666 · November 17, 2024, 2:17am

Similar to this case.

github.com/bitsandbytes-foundation/bitsandbytes

"Only Tensors of floating point and complex dtype can require gradients", on FSDP, Accelerate, quatization

opened 04:09PM - 30 May 24 UTC

closed 12:25PM - 25 Jun 24 UTC

artkpv

### System Info Python 3.11.5 torch 2.3.0 tran…sformers 4.41.1 accelerate 0.30.1 ``` +---------------------------------------------------------------------------------------+ | NVIDIA-SMI 545.23.06 Driver Version: 545.23.06 CUDA Version: 12.3 | |-----------------------------------------+----------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+======================+======================| | 0 NVIDIA A100-SXM4-80GB Off | 00000000:D6:00.0 Off | 0 | | N/A 34C P0 61W / 400W | 4MiB / 81920MiB | 0% Default | | | | Disabled | +-----------------------------------------+----------------------+----------------------+ | 1 NVIDIA A100-SXM4-80GB Off | 00000000:DA:00.0 Off | 0 | | N/A 36C P0 60W / 400W | 4MiB / 81920MiB | 0% Default | | | | Disabled | +-----------------------------------------+----------------------+----------------------+ +---------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=======================================================================================| | No running processes found | +---------------------------------------------------------------------------------------+ ``` ### Reproduction I want to fine tune Llama 3 70B with HF TRL. I am trying Accelerate, bitsandbytes quantization, mixed precision, FSDP on two GPUs with 80 GBs each. Running this code ``` if rank == 0: hf_model = LlamaForCausalLM.from_pretrained( "meta-llama/Meta-Llama-3-70B-Instruct", load_in_8bit=True, device_map="auto", ) ``` via this command: ``` > accelerate launch --config_file ./accelerate_fsdp_config.yaml train.py ``` At Slurm managed cluster inside sbatch script with: ``` #SBATCH --nodes=1 #SBATCH --gpus-per-node=2 ``` with this fsdp config: ``` compute_environment: LOCAL_MACHINE debug: false distributed_type: FSDP downcast_bf16: 'no' enable_cpu_affinity: false fsdp_config: fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP fsdp_backward_prefetch: BACKWARD_PRE fsdp_cpu_ram_efficient_loading: true fsdp_forward_prefetch: false fsdp_offload_params: true fsdp_sharding_strategy: FULL_SHARD fsdp_state_dict_type: SHARDED_STATE_DICT fsdp_sync_module_states: true fsdp_use_orig_params: true machine_rank: 0 main_training_function: main mixed_precision: bf16 num_machines: 1 num_processes: 2 rdzv_backend: static same_network: true tpu_env: [] tpu_use_cluster: false tpu_use_sudo: false use_cpu: false ``` Results in: > [rank1]: RuntimeError: Only Tensors of floating point and complex dtype can require gradients the same for rank 0. The same when I do `load_in_8bit=True`. Error: ``` [rank0]: Traceback (most recent call last): [rank0]: File "/data/artyom_karpov/rl4steg/train.py", line 345, in <module> [rank0]: train(context) [rank0]: File "/data/artyom_karpov/rl4steg/train.py", line 83, in train [rank0]: hf_model = LlamaForCausalLM.from_pretrained( [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank0]: File "/data/artyom_karpov/rl4steg/.venv/lib/python3.11/site-packages/transformers/modeling_utils.py", line 3754, in from_pretrained [rank0]: ) = cls._load_pretrained_model( [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank0]: File "/data/artyom_karpov/rl4steg/.venv/lib/python3.11/site-packages/transformers/modeling_utils.py", line 4214, in _load_pretrained_model [rank0]: new_error_msgs, offload_index, state_dict_index = _load_state_dict_into_meta_model( [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank0]: File "/data/artyom_karpov/rl4steg/.venv/lib/python3.11/site-packages/transformers/modeling_utils.py", line 896, in _load_state_dict_into_meta_model [rank0]: value = type(value)(value.data.to("cpu"), **value.__dict__) [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank0]: File "/data/artyom_karpov/rl4steg/.venv/lib/python3.11/site-packages/bitsandbytes/nn/modules.py", line 297, in __new__ [rank0]: return torch.Tensor._make_subclass(cls, data, requires_grad) [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank0]: RuntimeError: Only Tensors of floating point and complex dtype can require gradients ``` ### Expected behavior I expect the model to be loaded.

Topic		Replies	Views
Peft Model For SequenceClassification failing _is_peft_model Beginners	0	872	February 11, 2024
Finetuning llama-2 for classification 🤗Transformers	2	1925	January 29, 2024
FineTuning 7B model on 3080 laptop (16GO VRAM) issues Beginners	1	30	May 16, 2025
How to do classification fine-tuning of quantized models? 🤗Transformers	0	476	February 2, 2024
RuntimeError when training: Expected floating point type for target with class probabilities, got Long Beginners	0	695	December 17, 2023

Error while fine-tuning distilbert model

Related topics