EVA-Qwen2.5-72B-v0.2开源大语言模型 - 高效文本生成与指令跟随

首页

EVA Qwen2.5 72B V0.2

由 EVA-UNIT-01 开发

基于Qwen2.5-72B微调的大语言模型，专注于文本生成和指令跟随任务

大型语言模型

Transformers

开源协议:其他 #多任务指令微调 #72B大参数量 #创意写作优化

下载量 392

发布时间 : 11/21/2024

模型简介

该模型是基于Qwen2.5-72B架构进行微调的变体，主要用于文本生成、对话系统和指令跟随任务。通过多个高质量数据集的训练，增强了其理解和生成能力。

模型特点

大规模参数

拥有720亿参数，具备强大的语言理解和生成能力

多数据集微调

使用多个高质量数据集进行微调，包括指令跟随、写作和角色扮演等场景

指令优化

特别优化了对复杂指令的理解和执行能力

模型能力

文本生成

对话系统

指令跟随

创意写作

角色扮演

使用案例

内容创作

创意写作

生成小说、诗歌等创意文本

能够生成连贯且富有创意的文学作品

写作辅助

帮助用户完成各类写作任务

提供结构建议和内容扩展

对话系统

智能助手

构建能够理解复杂指令的对话系统

能够进行多轮有意义的对话

角色扮演

模拟特定角色进行互动

能够保持角色一致性并生成符合角色的回应

🚀 EVA Qwen2.5-72B v0.2

这是一个角色扮演/故事写作的专业模型，基于Qwen2.5-72B，在合成数据和自然数据的混合数据集上进行全参数微调。它使用了Celeste 70B 0.1数据混合，并进行了大幅扩展，以提升模型的通用性、创造力和独特风格。

此模型献给Nev。

✨ 主要特性

专业领域优化：专注于角色扮演和故事写作，在特定领域表现出色。
数据混合扩展：使用Celeste 70B 0.1数据混合并扩展，增强了模型的通用性、创造力和独特风格。
版本优化：0.2版本优化了训练超参数，增加了序列长度，在长上下文下指令遵循能力更好，减少了重复内容。

⚠️ 重要提示

⚠️ 重要提示

LLM - Compressor量化似乎无法正常工作，质量比正常情况差很多，之前的版本没有这种情况。GGUF和GPTQ似乎不受影响。

📚 详细文档

版本说明

0.2版本优化了训练超参数，增加了序列长度，在长上下文下指令遵循能力更好，减少了重复内容。

提示格式

提示格式为ChatML。

训练数据

Celeste 70B 0.1数据混合（减去Opus Instruct子集），详情见该模型的卡片。
Kalomaze的Opus_Instruct_25k数据集（过滤了拒绝回复的数据）。
Gryphe的ChatGPT - 4o - WritingPrompts的一个子集（1k行）。
Gryphe的Sonnet3.5 - Charcards - Roleplay的一个子集（2k行）。
Epiculous的Synthstruct和SynthRP数据集。
Dolphin - 2.9.3的一个子集，包括过滤后的not_samantha和一小部分systemchat。

训练时间和硬件

在8个H100 SXM上训练了17小时。

未来模型许可声明

对于所有未来的EVA - Unit - 01模型，许可证中将规定Infermatic及其任何员工或付费关联方不得使用、分发、下载或以其他方式使用EVA模型。虽然这不能追溯应用到我们现有的许可证，但我们正式要求Infermatic立即停止使用我们的模型获取不当利益，尽管我们知道此时这可能不会被遵守。EVA模型未来仍将在Featherless、ArliAI（未来）和其他愿意托管的平台上提供，也可用于本地和云端使用。

🔧 技术细节

查看Axolotl配置

Axolotl版本：0.4.1

base_model: Qwen/Qwen2.5-72B

load_in_8bit: false
load_in_4bit: false
strict: false

plugins:
  - axolotl.integrations.liger.LigerPlugin
liger_rope: true
liger_rms_norm: true
liger_swiglu: true
liger_fused_linear_cross_entropy: true

# plugins:
#   - axolotl.integrations.spectrum.SpectrumPlugin

# spectrum_top_fraction: 0.5
# # Optional if using a pre-scanned model as your base_model. Useful if using a model mirror
# spectrum_model_name: Qwen/Qwen2.5-32B

datasets:
  - path: datasets/Celeste_Filtered_utf8fix.jsonl
    type: sharegpt
  - path: datasets/deduped_not_samantha_norefusals.jsonl
    type: sharegpt
  - path: datasets/deduped_SynthRP-Gens_processed_ShareGPT_converted_cleaned.jsonl
    type: sharegpt
  - path: datasets/deduped_Synthstruct-Gens_processed_sharegpt_converted_cleaned.jsonl
    type: sharegpt
  - path: datasets/Gryphe-4o-WP-filtered-sharegpt_utf8fix.jsonl
    type: sharegpt
  - path: datasets/opus-instruct-22k-no_refusals-filtered_utf8fix.jsonl
    type: sharegpt
  - path: datasets/Sonnet3-5-charcard-names-filtered-sharegpt_utf8fix.jsonl
    type: sharegpt
  - path: datasets/SystemChat_subset_filtered_sharegpt_utf8fix.jsonl
    type: sharegpt

chat_template: chatml
shuffle_merged_datasets: true
val_set_size: 0.001
output_dir: EVA-Qwen2.5-72B-SFFT-v0.2

sequence_len: 10240
sample_packing: true
eval_sample_packing: false
pad_to_sequence_len: false

# adapter: qlora
# lora_model_dir:
# lora_r: 64
# lora_alpha: 128
# lora_dropout: 0.05
# lora_target_linear: true
# peft_use_dora: true

unfrozen_parameters:
- ^lm_head.weight$
- ^model.embed_tokens.weight$
# mlp.down_proj layers
- model.layers.62.mlp.down_proj
- model.layers.64.mlp.down_proj
- model.layers.63.mlp.down_proj
- model.layers.66.mlp.down_proj
- model.layers.65.mlp.down_proj
- model.layers.67.mlp.down_proj
- model.layers.68.mlp.down_proj
- model.layers.31.mlp.down_proj
- model.layers.60.mlp.down_proj
- model.layers.69.mlp.down_proj
- model.layers.61.mlp.down_proj
- model.layers.59.mlp.down_proj
- model.layers.30.mlp.down_proj
- model.layers.70.mlp.down_proj
- model.layers.32.mlp.down_proj
- model.layers.34.mlp.down_proj
- model.layers.33.mlp.down_proj
- model.layers.76.mlp.down_proj
- model.layers.72.mlp.down_proj
- model.layers.71.mlp.down_proj
- model.layers.58.mlp.down_proj
- model.layers.75.mlp.down_proj
- model.layers.29.mlp.down_proj
- model.layers.56.mlp.down_proj
- model.layers.26.mlp.down_proj
- model.layers.35.mlp.down_proj
- model.layers.28.mlp.down_proj
- model.layers.57.mlp.down_proj
- model.layers.77.mlp.down_proj
- model.layers.36.mlp.down_proj
- model.layers.27.mlp.down_proj
- model.layers.25.mlp.down_proj
- model.layers.78.mlp.down_proj
- model.layers.37.mlp.down_proj
- model.layers.73.mlp.down_proj
- model.layers.55.mlp.down_proj
- model.layers.54.mlp.down_proj
- model.layers.74.mlp.down_proj
- model.layers.24.mlp.down_proj
- model.layers.53.mlp.down_proj
# mlp.gate_proj layers
- model.layers.78.mlp.gate_proj
- model.layers.77.mlp.gate_proj
- model.layers.76.mlp.gate_proj
- model.layers.79.mlp.gate_proj
- model.layers.75.mlp.gate_proj
- model.layers.74.mlp.gate_proj
- model.layers.73.mlp.gate_proj
- model.layers.72.mlp.gate_proj
- model.layers.71.mlp.gate_proj
- model.layers.70.mlp.gate_proj
- model.layers.69.mlp.gate_proj
- model.layers.57.mlp.gate_proj
- model.layers.54.mlp.gate_proj
- model.layers.55.mlp.gate_proj
- model.layers.68.mlp.gate_proj
- model.layers.63.mlp.gate_proj
- model.layers.53.mlp.gate_proj
- model.layers.44.mlp.gate_proj
- model.layers.45.mlp.gate_proj
- model.layers.49.mlp.gate_proj
- model.layers.58.mlp.gate_proj
- model.layers.46.mlp.gate_proj
- model.layers.56.mlp.gate_proj
- model.layers.67.mlp.gate_proj
- model.layers.62.mlp.gate_proj
- model.layers.50.mlp.gate_proj
- model.layers.64.mlp.gate_proj
- model.layers.52.mlp.gate_proj
- model.layers.40.mlp.gate_proj
- model.layers.43.mlp.gate_proj
- model.layers.48.mlp.gate_proj
- model.layers.66.mlp.gate_proj
- model.layers.47.mlp.gate_proj
- model.layers.59.mlp.gate_proj
- model.layers.65.mlp.gate_proj
- model.layers.61.mlp.gate_proj
- model.layers.60.mlp.gate_proj
- model.layers.42.mlp.gate_proj
- model.layers.51.mlp.gate_proj
- model.layers.41.mlp.gate_proj
# mlp.up_proj layers
- model.layers.70.mlp.up_proj
- model.layers.69.mlp.up_proj
- model.layers.71.mlp.up_proj
- model.layers.68.mlp.up_proj
- model.layers.72.mlp.up_proj
- model.layers.67.mlp.up_proj
- model.layers.66.mlp.up_proj
- model.layers.73.mlp.up_proj
- model.layers.46.mlp.up_proj
- model.layers.63.mlp.up_proj
- model.layers.75.mlp.up_proj
- model.layers.76.mlp.up_proj
- model.layers.74.mlp.up_proj
- model.layers.45.mlp.up_proj
- model.layers.62.mlp.up_proj
- model.layers.64.mlp.up_proj
- model.layers.65.mlp.up_proj
- model.layers.44.mlp.up_proj
- model.layers.53.mlp.up_proj
- model.layers.47.mlp.up_proj
- model.layers.49.mlp.up_proj
- model.layers.48.mlp.up_proj
- model.layers.57.mlp.up_proj
- model.layers.43.mlp.up_proj
- model.layers.42.mlp.up_proj
- model.layers.56.mlp.up_proj
- model.layers.61.mlp.up_proj
- model.layers.54.mlp.up_proj
- model.layers.40.mlp.up_proj
- model.layers.55.mlp.up_proj
- model.layers.77.mlp.up_proj
- model.layers.60.mlp.up_proj
- model.layers.41.mlp.up_proj
- model.layers.35.mlp.up_proj
- model.layers.37.mlp.up_proj
- model.layers.58.mlp.up_proj
- model.layers.34.mlp.up_proj
- model.layers.38.mlp.up_proj
- model.layers.33.mlp.up_proj
- model.layers.39.mlp.up_proj
# self_attn.k_proj layers
- model.layers.36.self_attn.k_proj
- model.layers.79.self_attn.k_proj
- model.layers.35.self_attn.k_proj
- model.layers.34.self_attn.k_proj
- model.layers.37.self_attn.k_proj
- model.layers.33.self_attn.k_proj
- model.layers.38.self_attn.k_proj
- model.layers.39.self_attn.k_proj
- model.layers.74.self_attn.k_proj
- model.layers.77.self_attn.k_proj
- model.layers.41.self_attn.k_proj
- model.layers.69.self_attn.k_proj
- model.layers.32.self_attn.k_proj
- model.layers.78.self_attn.k_proj
- model.layers.30.self_attn.k_proj
- model.layers.70.self_attn.k_proj
- model.layers.25.self_attn.k_proj
- model.layers.42.self_attn.k_proj
- model.layers.29.self_attn.k_proj
- model.layers.31.self_attn.k_proj
- model.layers.68.self_attn.k_proj
- model.layers.66.self_attn.k_proj
- model.layers.22.self_attn.k_proj
- model.layers.65.self_attn.k_proj
- model.layers.44.self_attn.k_proj
- model.layers.40.self_attn.k_proj
- model.layers.63.self_attn.k_proj
- model.layers.23.self_attn.k_proj
- model.layers.28.self_attn.k_proj
- model.layers.24.self_attn.k_proj
- model.layers.26.self_attn.k_proj
- model.layers.67.self_attn.k_proj
- model.layers.75.self_attn.k_proj
- model.layers.27.self_attn.k_proj
- model.layers.57.self_attn.k_proj
- model.layers.64.self_attn.k_proj
- model.layers.71.self_attn.k_proj
- model.layers.61.self_attn.k_proj
- model.layers.72.self_attn.k_proj
- model.layers.73.self_attn.k_proj
# self_attn.o_proj layers
- model.layers.69.self_attn.o_proj
- model.layers.39.self_attn.o_proj
- model.layers.16.self_attn.o_proj
- model.layers.14.self_attn.o_proj
- model.layers.19.self_attn.o_proj
- model.layers.42.self_attn.o_proj
- model.layers.12.self_attn.o_proj
- model.layers.15.self_attn.o_proj
- model.layers.17.self_attn.o_proj
- model.layers.38.self_attn.o_proj
- model.layers.23.self_attn.o_proj
- model.layers.22.self_attn.o_proj
- model.layers.13.self_attn.o_proj
- model.layers.29.self_attn.o_proj
- model.layers.41.self_attn.o_proj
- model.layers.44.self_attn.o_proj
- model.layers.46.self_attn.o_proj
- model.layers.45.self_attn.o_proj
- model.layers.43.self_attn.o_proj
- model.layers.49.self_attn.o_proj
- model.layers.30.self_attn.o_proj
- model.layers.26.self_attn.o_proj
- model.layers.25.self_attn.o_proj
- model.layers.37.self_attn.o_proj
- model.layers.47.self_attn.o_proj
- model.layers.11.self_attn.o_proj
- model.layers.18.self_attn.o_proj
- model.layers.28.self_attn.o_proj
- model.layers.20.self_attn.o_proj
- model.layers.27.self_attn.o_proj
- model.layers.53.self_attn.o_proj
- model.layers.52.self_attn.o_proj
- model.layers.35.self_attn.o_proj
- model.layers.71.self_attn.o_proj
- model.layers.10.self_attn.o_proj
- model.layers.3.self_attn.o_proj
- model.layers.21.self_attn.o_proj
- model.layers.24.self_attn.o_proj
- model.layers.68.self_attn.o_proj
- model.layers.48.self_attn.o_proj
# self_attn.q_proj layers
- model.layers.1.self_attn.q_proj
- model.layers.2.self_attn.q_proj
- model.layers.3.self_attn.q_proj
- model.layers.0.self_attn.q_proj
- model.layers.5.self_attn.q_proj
- model.layers.4.self_attn.q_proj
- model.layers.6.self_attn.q_proj
- model.layers.8.self_attn.q_proj
- model.layers.7.self_attn.q_proj
- model.layers.9.self_attn.q_proj
- model.layers.10.self_attn.q_proj
- model.layers.68.self_attn.q_proj
- model.layers.25.self_attn.q_proj
- model.layers.12.self_attn.q_proj
- model.layers.54.self_attn.q_proj
- model.layers.55.self_attn.q_proj
- model.layers.61.self_attn.q_proj
- model.layers.18.self_attn.q_proj
- model.layers.49.self_attn.q_proj
- model.layers.66.self_attn.q_proj
- model.layers.72.self_attn.q_proj
- model.layers.11.self_attn.q_proj
- model.layers.52.self_attn.q_proj
- model.layers.64.self_attn.q_proj
- model.layers.15.self_attn.q_proj
- model.layers.60.self_attn.q_proj
- model.layers.50.self_attn.q_proj
- model.layers.59.self_attn.q_proj
- model.layers.53.self_attn.q_proj
- model.layers.48.self_attn.q_proj
- model.layers.57.self_attn.q_proj
- model.layers.70.self_attn.q_proj
- model.layers.17.self_attn.q_proj
- model.layers.67.self_attn.q_proj
- model.layers.71.self_attn.q_proj
- model.layers.62.self_attn.q_proj
- model.layers.51.self_attn.q_proj
- model.layers.19.self_attn.q_proj
- model.layers.58.self_attn.q_proj
- model.layers.13.self_attn.q_proj
# self_attn.v_proj layers
- model.layers.23.self_attn.v_proj
- model.layers.25.self_attn.v_proj
- model.layers.26.self_attn.v_proj
- model.layers.27.self_attn.v_proj
- model.layers.28.self_attn.v_proj
- model.layers.29.self_attn.v_proj
- model.layers.30.self_attn.v_proj
- model.layers.31.self_attn.v_proj
- model.layers.34.self_attn.v_proj
- model.layers.35.self_attn.v_proj
- model.layers.36.self_attn.v_proj
- model.layers.37.self_attn.v_proj
- model.layers.38.self_attn.v_proj
- model.layers.42.self_attn.v_proj
- model.layers.48.self_attn.v_proj
- model.layers.57.self_attn.v_proj
- model.layers.58.self_attn.v_proj
- model.layers.61.self_attn.v_proj
- model.layers.63.self_attn.v_proj
- model.layers.64.self_attn.v_proj
- model.layers.65.self_attn.v_proj
- model.layers.66.self_attn.v_proj
- model.layers.69.self_attn.v_proj
- model.layers.70.self_attn.v_proj
- model.layers.74.self_attn.v_proj
- model.layers.75.self_attn.v_proj
- model.layers.72.self_attn.v_proj
- model.layers.39.self_attn.v_proj
- model.layers.41.self_attn.v_proj
- model.layers.40.self_attn.v_proj
- model.layers.33.self_attn.v_proj
- model.layers.59.self_attn.v_proj
- model.layers.16.self_attn.v_proj
- model.layers.15.self_attn.v_proj
- model.layers.76.self_attn.v_proj
- model.layers.24.self_attn.v_proj
- model.layers.68.self_attn.v_proj
- model.layers.67.self_attn.v_proj
- model.layers.55.self_attn.v_proj
- model.layers.44.self_attn.v_proj

wandb_project: EVA-Qwen2.5-72B-SFFT-v0.2
wandb_entity:
wandb_watch:
wandb_name: Unit-02
wandb_log_model:

gradient_accumulation_steps: 8
micro_batch_size: 1
num_epochs: 3
optimizer: paged_ademamix_8bit
lr_scheduler: cosine
learning_rate: 0.00003
max_grad_norm: 1.5

train_on_inputs: false
group_by_length: false
bf16: auto
fp16:
tf32: false

gradient_checkpointing: "unsloth"
# gradient_checkpointing_kwargs:
#   use_reentrant: true
early_stopping_patience:
resume_from_checkpoint: EVA-Qwen2.5-72B-SFFT-v0.2/checkpoint-128
local_rank:
logging_steps: 1
xformers_attention:
flash_attention: true

warmup_steps: 20
evals_per_epoch: 4
saves_per_epoch: 4
save_safetensors: true
save_total_limit: 1
hub_model_id: 
hub_strategy: 
debug:
deepspeed: deepspeed_configs/zero3_bf16_cpuoffload_params.json
weight_decay: 0.12
# fsdp:
#   - full_shard
#   - auto_wrap
# fsdp_config:
#   fsdp_limit_all_gathers: true
#   fsdp_sync_module_states: false
#   fsdp_offload_params: true
#   fsdp_cpu_ram_efficient_loading: true
#   fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
#   fsdp_transformer_layer_cls_to_wrap: Qwen2DecoderLayer
#   fsdp_activation_checkpointing: true
#   fsdp_state_dict_type: SHARDED_STATE_DICT  # Changed from FULL_STATE_DICT
#   fsdp_sharding_strategy: FULL_SHARD
#   fsdp_forward_prefetch: false  # Added
#   fsdp_backward_prefetch: "BACKWARD_PRE"  # Added
#   fsdp_backward_prefetch_limit: 1  # Added
#   fsdp_mixed_precision: BF16  # Added