CalmeRys-78B-Orpo-v0.1开源大语言模型 - 排行榜第一的高效对话选择

首页

Calmerys 78B Orpo V0.1

由 dfurman 开发

基于MaziyarPanahi/calme-2.4-rys-78b在mlabonne/orpo-dpo-mix-40k数据集上微调的大语言模型，在Open LLM Leaderboard上排名第一

大型语言模型

Transformers

英语开源协议:MIT #78B大模型 #ORPO微调 #多轮对话

下载量 353

发布时间 : 9/24/2024

模型简介

通用语言模型，适用于多种文本生成场景，包括代理能力、角色扮演、推理、多轮对话和长上下文连贯性

模型特点

高性能

在Open LLM Leaderboard上排名第一

多功能

支持多种文本生成场景，包括推理、对话和长上下文处理

微调优化

在精选数据集上进行ORPO微调，提升模型性能

模型能力

文本生成

多轮对话

逻辑推理

长上下文处理

角色扮演

使用案例

问答系统

数学问题解答

解决数学比较和计算问题

准确比较数字大小并展示计算过程

内容创作

食谱生成

生成详细的鸡尾酒配方

提供完整材料清单和分步制作指南

商业应用

销售数据分析

处理销售数据并计算剩余库存

以表格形式清晰展示计算过程和结果

🚀 CalmeRys-78B-Orpo-v0.1

CalmeRys-78B-Orpo-v0.1 是基于 MaziyarPanahi/calme-2.4-rys-78b 模型，在 mlabonne/orpo-dpo-mix-40k 数据集的 1.5k 行数据上微调得到的。它是一个通用语言模型，可用于多种文本生成场景，包括支持智能体能力、角色扮演、推理、多轮对话、长上下文连贯等。截至 2024 年 10 月，该模型在 Open LLM Leaderboard 上排名第一 🏆。

感谢 mlabonne、MaziyarPanahi 等人提供的源数据集和基础模型。

📦 安装指南

环境设置

!pip install -qU transformers accelerate bitsandbytes
!huggingface-cli download dfurman/CalmeRys-78B-Orpo-v0.1

from transformers import AutoTokenizer, BitsAndBytesConfig
import transformers
import torch


if torch.cuda.get_device_capability()[0] >= 8:
    !pip install -qqq flash-attn
    attn_implementation = "flash_attention_2"
    torch_dtype = torch.bfloat16
else:
    attn_implementation = "eager"
    torch_dtype = torch.float16

# # 必要时进行量化
# bnb_config = BitsAndBytesConfig(
#    load_in_4bit=True,
#    bnb_4bit_quant_type="nf4",
#    bnb_4bit_compute_dtype=torch_dtype,
#    bnb_4bit_use_double_quant=True,
# )

model = "dfurman/CalmeRys-78B-Orpo-v0.1"

tokenizer = AutoTokenizer.from_pretrained(model)
pipeline = transformers.pipeline(
    "text-generation",
    model=model,
    model_kwargs={
        "torch_dtype": torch_dtype,
        # "quantization_config": bnb_config,
        "device_map": "auto",
        "attn_implementation": attn_implementation,
    }
)

💻 使用示例

基础用法

question = "Is the number 9.11 larger than 9.9?"

messages = [
    {"role": "system", "content": "You are a helpful assistant that thinks step by step."},
    {"role": "user", "content": question},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
# print("***Prompt:\n", prompt)

outputs = pipeline(
    prompt, max_new_tokens=1000, do_sample=True, temperature=0.7, top_k=50, top_p=0.95
)
print("***Generation:")
print(outputs[0]["generated_text"][len(prompt) :])

***Generation:
To compare these two numbers, it's important to look at their decimal places after the whole number part, which is 9 in both cases. Comparing the tenths place, 9.11 has a '1' and 9.9 has a '9'. Since '9' is greater than '1', 9.9 is larger than 9.11.

高级用法

question = """The bakers at the Beverly Hills Bakery baked 200 loaves of bread on Monday morning. 
They sold 93 loaves in the morning and 39 loaves in the afternoon. 
A grocery store then returned 6 unsold loaves back to the bakery. 
How many loaves of bread did the bakery have left?
Respond as succinctly as possible. Format the response as a completion of this table:
|step|subquestion|procedure|result|
|:---|:----------|:--------|:-----:|"""


messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": question},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
# print("***Prompt:\n", prompt)

outputs = pipeline(prompt, max_new_tokens=1000, do_sample=True, temperature=0.7, top_k=50, top_p=0.95)
print("***Generation:")
print(outputs[0]["generated_text"][len(prompt):])

***Generation:
|1|Calculate total sold|Add morning and afternoon sales|132|
|2|Subtract sold from total|200 - 132|68|
|3|Adjust for returns|Add returned loaves to remaining|74|

复杂场景用法

question = "What's a good recipe for a spicy margarita?"

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": question},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
# print("***Prompt:\n", prompt)

outputs = pipeline(prompt, max_new_tokens=1000, do_sample=True, temperature=0.7, top_k=50, top_p=0.95)
print("***Generation:")
print(outputs[0]["generated_text"][len(prompt):])

***Generation:
To make a Spicy Margarita, you'll need to incorporate a chili or pepper element into your classic margarita recipe. Here’s a simple way to do it:

### Ingredients:
- 2 oz tequila (blanco or reposado)
- 1 oz fresh lime juice
- 1/2 oz triple sec (Cointreau or Grand Marnier)
- 1/2 oz agave syrup or simple syrup
- 1-2 slices of jalapeño (or more depending on how spicy you like it)
- Salt and/or chili powder for rimming the glass
- Ice
- Lime wheel for garnish

### Instructions:
1. **Muddle Jalapeño**: In a shaker, muddle the jalapeño slices slightly. This will release the oils and heat from the peppers.
2. **Add Remaining Ingredients**: Add the tequila, lime juice, triple sec, and agave syrup or simple syrup. 
3. **Shake and Strain**: Fill the shaker with ice and shake vigorously until cold. Strain into a salt and/or chili powder rimmed glass filled with ice.
4. **Garnish and Serve**: Garnish with a lime wheel and enjoy.

If you prefer a smoother spiciness that doesn't overpower the drink, you could also consider making a jalapeño-infused tequila by leaving the jalapeño slices in the bottle of tequila for several hours to a couple of days, adjusting the time based on desired level of spiciness. Then use this infused tequila instead of regular tequila in the recipe above. 

Another variation is to use a spicy syrup. To make this, combine equal parts water and sugar with a few sliced jalapeños in a saucepan. Bring to a boil, stirring occasionally to dissolve the sugar. Reduce heat and simmer for about 5 minutes. Let cool, strain out the jalapeños, then store in a sealed container in the refrigerator until ready to use. Use this spicy syrup instead of regular syrup in the recipe. 

As always, adjust the quantity of jalapeño or the type of chili used to suit your taste. Enjoy responsibly!

📚 详细文档

Open LLM Leaderboard评估结果

详细结果可查看此处

指标	数值
平均值	50.78
IFEval (0-Shot)	81.63
BBH (3-Shot)	61.92
MATH Lvl 5 (4-Shot)	37.92
GPQA (0-shot)	20.02
MuSR (0-shot)	36.37
MMLU-PRO (5-shot)	66.80

🦾 训练情况

你可以通过此链接在 W&B 上查看实验情况。以下是一些可视化结果：

image/png

📄 许可证

本模型采用 MIT 许可证。

属性	详情
模型类型	文本生成模型
基础模型	MaziyarPanahi/calme-2.4-rys-78b
训练数据	mlabonne/orpo-dpo-mix-40k
模型创建者	dfurman
量化者	dfurman
任务类型	文本生成
推理功能	否