Does anyone have an idea how we can run llama2 with multiple GPUs?

hzhuang0000 · September 27, 2023, 6:55am

The code is from here: Llama 2 is here - get it on Hugging Face

from transformers import AutoTokenizer
import transformers
import torch

model = "meta-llama/Llama-2-7b-chat-hf"

tokenizer = AutoTokenizer.from_pretrained(model)
pipeline = transformers.pipeline(
    "text-generation",
    model=model,
    torch_dtype=torch.float16,
    device_map="auto",
)

sequences = pipeline(
    'I liked "Breaking Bad" and "Band of Brothers". Do you have any recommendations of other shows I might like?\n',
    do_sample=True,
    top_k=10,
    num_return_sequences=1,
    eos_token_id=tokenizer.eos_token_id,
    max_length=200,
)
for seq in sequences:
    print(f"Result: {seq['generated_text']}")

Thanks

vbachi · October 26, 2023, 3:24pm

I am searching for the same too. Maybe accelerate library can help here but I do not know now how to use it. Have you found any solutions?

Topic		Replies	Views
Why transformers doesn't use Multiple GPUs (to increase tokens per second)? Beginners	7	583	September 22, 2024
Fine tunning llama2 with multiple GPUs and Hugging face trainer 🤗Transformers	1	3475	November 3, 2023
Multi-GPU inference with LLM produces gibberish 🤗Transformers	14	6550	September 28, 2024
Fine-tunning llama2 with multiple GPU hugging face trainer 🤗Transformers	8	3357	March 7, 2024
If I use llama 70b and 7b for speculative decoding, how should I put them on my multiple gpus in the code 🤗Transformers	0	46	October 11, 2024

Does anyone have an idea how we can run llama2 with multiple GPUs?

Related topics