Inconsistent Tokenizer format when training Llama-2 with alignment handbook

elichen3051 · June 7, 2024, 12:42am

I’ve tried to use alignment handbook to train an instruction-tuned Llama-2-7b-hf by using SFT, and assigned the tokenizer_name_or_path=meta-llama/Llama-2-7b-chat-hf; However, after finishing training, I got a different tokenizer_config.json;

For example, the following is the tokenizer_config.json got from alignment handbook SFT training

{
  "add_bos_token": true,
  "add_eos_token": false,
  "added_tokens_decoder": {
    "0": {
      "content": "<unk>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "1": {
      "content": "<s>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "2": {
      "content": "</s>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    }
  },
  "bos_token": "<s>",
  "chat_template": "{% if messages[0]['role'] == 'system' %}{% set loop_messages = messages[1:] %}{% set system_message = messages[0]['content'] %}{% else %}{% set loop_messages = messages %}{% set system_message = false %}{% endif %}{% for message in loop_messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if loop.index0 == 0 and system_message != false %}{% set content = '<<SYS>>\\n' + system_message + '\\n<</SYS>>\\n\\n' + message['content'] %}{% else %}{% set content = message['content'] %}{% endif %}{% if message['role'] == 'user' %}{{ bos_token + '[INST] ' + content.strip() + ' [/INST]' }}{% elif message['role'] == 'assistant' %}{{ ' '  + content.strip() + ' ' + eos_token }}{% endif %}{% endfor %}",
  "clean_up_tokenization_spaces": false,
  "eos_token": "</s>",
  "legacy": false,
  "model_max_length": 2048,
  "pad_token": "</s>",
  "padding_side": "right",
  "sp_model_kwargs": {},
  "tokenizer_class": "LlamaTokenizer",
  "unk_token": "<unk>",
  "use_default_system_prompt": false
}

and the original one from meta-llama/Llama-2-7b-chat-hf is:

{
  "add_bos_token": true,
  "add_eos_token": false,
  "bos_token": {
    "__type": "AddedToken",
    "content": "<s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "chat_template": "{% if messages[0]['role'] == 'system' %}{% set loop_messages = messages[1:] %}{% set system_message = messages[0]['content'] %}{% else %}{% set loop_messages = messages %}{% set system_message = false %}{% endif %}{% for message in loop_messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if loop.index0 == 0 and system_message != false %}{% set content = '<<SYS>>\\n' + system_message + '\\n<</SYS>>\\n\\n' + message['content'] %}{% else %}{% set content = message['content'] %}{% endif %}{% if message['role'] == 'user' %}{{ bos_token + '[INST] ' + content.strip() + ' [/INST]' }}{% elif message['role'] == 'assistant' %}{{ ' '  + content.strip() + ' ' + eos_token }}{% endif %}{% endfor %}",
  "clean_up_tokenization_spaces": false,
  "eos_token": {
    "__type": "AddedToken",
    "content": "</s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "legacy": false,
  "model_max_length": 1000000000000000019884624838656,
  "pad_token": null,
  "padding_side": "right",
  "sp_model_kwargs": {},
  "tokenizer_class": "LlamaTokenizer",
  "unk_token": {
    "__type": "AddedToken",
    "content": "<unk>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  }
}

And here is the training recipe for llama2 added by myself.

# Model arguments
model_name_or_path: meta-llama/Llama-2-7b-hf
model_revision: main
torch_dtype: bfloat16
use_flash_attention_2: false

# Data training arguments
tokenizer_name_or_path: meta-llama/Llama-2-7b-chat-hf
dataset_mixer:
  HuggingFaceH4/ultrachat_200k: 1.0
  HuggingFaceH4/ultrafeedback_binarized: 1.0
dataset_splits:
  - train_sft
  - test_sft
preprocessing_num_workers: 16

# SFT trainer config
bf16: true
do_eval: true
evaluation_strategy: epoch
gradient_accumulation_steps: 8
gradient_checkpointing: true
gradient_checkpointing_kwargs:
  use_reentrant: False
hub_model_id: <repo>/llama2-7b-sft-full-chat-tokenizer-epoch-1
hub_strategy: every_save
hub_private_repo: true
learning_rate: 2.0e-05
log_level: info
logging_steps: 1
logging_strategy: steps
lr_scheduler_type: cosine
max_seq_length: 2048
max_steps: -1
num_train_epochs: 1
output_dir: model/llama2-7b-sft-full-chat-tokenizer-epoch-1
overwrite_output_dir: true
overwrite_output_dir: true
per_device_eval_batch_size: 4
per_device_train_batch_size: 4
push_to_hub: true
remove_unused_columns: true
report_to:
- tensorboard
- wandb
save_strategy: "epoch"
save_total_limit: 3
seed: 42
warmup_ratio: 0.1

I don’t know how to fix this problem, please help
thank you
Eli Chen

NitzanBar · June 25, 2024, 8:22am

Perhaps you should try remove this line:
tokenizer_name_or_path: meta-llama/Llama-2-7b-chat-hf

Topic		Replies	Views
How to train a LlamaTokenizer? 🤗Tokenizers	22	4082	August 20, 2024
Cannot load tokenizer for llama2 🤗Tokenizers	6	7211	September 13, 2024
Llama 2 fine tuning general questions (tokenizer, compute_metrics, labels)) Beginners	0	1514	October 28, 2023
Unable to load tokenizer 🤗Transformers	3	67	February 14, 2025
Can't set pad_token by adding special token to Llama's tokenizer 🤗Transformers	4	5958	August 12, 2024

Inconsistent Tokenizer format when training Llama-2 with alignment handbook

Related topics