kamo - naoyuki开源英文语音识别模型 - 免费部署精准识别英文语音

首页

Kamo Naoyuki Mini An4 Asr Train Raw Bpe Valid.acc.best

由 espnet 开发

这是一个基于ESPnet2框架训练的自动语音识别(ASR)预训练模型，使用mini-an4数据集训练，支持英文语音识别。

语音识别英语#端到端语音识别 #BPE分词 #轻量级模型

下载量 425

发布时间 : 3/2/2022

模型简介

该模型是一个端到端的自动语音识别模型，能够将输入的语音信号转换为对应的文本内容。

模型特点

端到端语音识别

采用端到端架构，直接从语音信号转换为文本

基于ESPnet框架

使用ESPnet这一成熟的端到端语音处理工具包训练

BPE分词

使用字节对编码(BPE)进行文本处理

模型能力

英语语音识别

端到端语音转文本

使用案例

语音转录

会议记录转录

将英语会议录音自动转换为文字记录

语音指令识别

识别英语语音指令并转换为可执行命令

🚀 ESPnet2自动语音识别预训练模型

本模型是ESPnet2的自动语音识别（ASR）预训练模型，可用于音频的自动语音识别任务，基于espnet框架训练，具有较高的实用性和可扩展性。

🚀 快速开始

演示：如何在ESPnet2中使用

# coming soon

✨ 主要特性

预训练模型：该模型kamo-naoyuki/mini_an4_asr_train_raw_bpe_valid.acc.best是预训练好的，可直接用于相关任务。
来源清晰：♻️ 从 https://zenodo.org/record/3957940#.YN7zwJozZH4 导入，由kan - bayashi使用jsut/tts1配方在 espnet 中训练得到。

📚 详细文档

引用ESPnet

如果你在研究中使用了ESPnet，可以使用以下BibTex格式进行引用：

@inproceedings{watanabe2018espnet,
  author={Shinji Watanabe and Takaaki Hori and Shigeki Karita and Tomoki Hayashi and Jiro Nishitoba and Yuya Unno and Nelson {Enrique Yalta Soplin} and Jahn Heymann and Matthew Wiesner and Nanxin Chen and Adithya Renduchintala and Tsubasa Ochiai},
  title={{ESPnet}: End-to-End Speech Processing Toolkit},
  year={2018},
  booktitle={Proceedings of Interspeech},
  pages={2207--2211},
  doi={10.21437/Interspeech.2018-1456},
  url={http://dx.doi.org/10.21437/Interspeech.2018-1456}
}
@inproceedings{hayashi2020espnet,
  title={{Espnet-TTS}: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit},
  author={Hayashi, Tomoki and Yamamoto, Ryuichi and Inoue, Katsuki and Yoshimura, Takenori and Watanabe, Shinji and Toda, Tomoki and Takeda, Kazuya and Zhang, Yu and Tan, Xu},
  booktitle={Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  pages={7654--7658},
  year={2020},
  organization={IEEE}
}

或者使用arXiv格式：

@misc{watanabe2018espnet,
      title={ESPnet: End-to-End Speech Processing Toolkit}, 
      author={Shinji Watanabe and Takaaki Hori and Shigeki Karita and Tomoki Hayashi and Jiro Nishitoba and Yuya Unno and Nelson Enrique Yalta Soplin and Jahn Heymann and Matthew Wiesner and Nanxin Chen and Adithya Renduchintala and Tsubasa Ochiai},
      year={2018},
      eprint={1804.00015},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}

训练配置

完整的配置文件请参考 config.yaml

config: null
print_config: false
log_level: INFO
dry_run: false
iterator_type: sequence
output_dir: exp/asr_train_raw_bpe
ngpu: 1
seed: 0
num_workers: 1
num_att_plot: 3
dist_backend: nccl
dist_init_method: env://
dist_world_size: null
dist_rank: null
local_rank: 0
dist_master_addr: null
dist_master_port: null
dist_launcher: null
multiprocessing_distributed: false
cudnn_enabled: true
cudnn_benchmark: false
cudnn_deterministic: true