与变压器和源源过滤器进行转换的 " 快速转让学习 " 用于 " 强力低资源儿童演讲 " ASR " 的转让学习 (Transfer Learning for Robust Low-Resource Children's Speech ASR with Transformers and Source-Filter Warping) - 专知论文

会员服务 ·

0

Learning · 稳健性 · 语音识别 · MoDELS · 迁移学习 ·

2022 年 6 月 19 日

Transfer Learning for Robust Low-Resource Children's Speech ASR with Transformers and Source-Filter Warping

翻译：与变压器和源源过滤器进行转换的 " 快速转让学习 " 用于 " 强力低资源儿童演讲 " ASR " 的转让学习

Jenthe Thienpondt,Kris Demuynck

from arxiv, proceedings of INTERSPEECH 2022

Automatic Speech Recognition (ASR) systems are known to exhibit difficulties when transcribing children's speech. This can mainly be attributed to the absence of large children's speech corpora to train robust ASR models and the resulting domain mismatch when decoding children's speech with systems trained on adult data. In this paper, we propose multiple enhancements to alleviate these issues. First, we propose a data augmentation technique based on the source-filter model of speech to close the domain gap between adult and children's speech. This enables us to leverage the data availability of adult speech corpora by making these samples perceptually similar to children's speech. Second, using this augmentation strategy, we apply transfer learning on a Transformer model pre-trained on adult data. This model follows the recently introduced XLS-R architecture, a wav2vec 2.0 model pre-trained on several cross-lingual adult speech corpora to learn general and robust acoustic frame-level representations. Adopting this model for the ASR task using adult data augmented with the proposed source-filter warping strategy and a limited amount of in-domain children's speech significantly outperforms previous state-of-the-art results on the PF-STAR British English Children's Speech corpus with a 4.86% WER on the official test set.

翻译：据了解,自动语音识别系统在翻译儿童讲话时会遇到困难,这主要是因为没有大型儿童语言协会来培训强大的ASR模型,因此在用成人数据培训的系统解码儿童讲话时,没有培养强大的ASR模型,因此造成域错配。在本文件中,我们提出多项改进,以缓解这些问题。首先,我们提议基于发源过滤器语音模型的数据增强技术,以缩小成人与儿童讲话之间的域间差距。这使我们能够利用成人语言协会的数据提供量,使这些样本与儿童讲话有相似感。第二,利用这一增强战略,我们将学习应用在成人数据培训前的变换器模型上。这一模型遵循了最近推出的 XLS-R 结构,即 wav2vec 2.0 模型,对几个跨语言成人演讲协会进行了预先培训,以学习一般和稳健的声学框架级表达方式。我们采用这一模型来完成ASR任务,利用拟议的源过滤策略所强化的成人数据,以及有限的数量在英国官方演讲中,186 儿童演讲的成绩在英国官方演讲中明显超过英国的4.PFS-PRO 。

0

相关内容

Learning

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

125+阅读 · 2022年4月21日

【Manning新书】迁移学习自然语言处理，266页pdf，Transfer Learning for NLP

【Manning新书】迁移学习自然语言处理，266页pdf，Transfer Learning for NLP

专知会员服务

137+阅读 · 2021年11月6日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Industry Talk2

【ICIG2021】Latest News & Announcements of the Industry Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年7月29日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

孤独症儿童早期干预的同步TMS-EEG研究

国家自然科学基金

0+阅读 · 2017年12月31日

SATB2-Nanog-mTOR轴调控BMSCs衰老及其在颌骨增龄性骨量丢失中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

核苷肽类抗生素缩合机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

功能性遗传变异调控BARD1/BRCA1泛素化通路的机制及与儿童神经母细胞瘤的关联研究

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Legendre 级数多极边界元法理论研究

国家自然科学基金

0+阅读 · 2013年12月31日

PGC-1α调节骨骼肌脂肪酸代谢和胰岛素抵抗的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

精神分裂症超高危人群经颅磁刺激干预的磁共振研究

国家自然科学基金

0+阅读 · 2012年12月31日

暴露于电子垃圾污染物中的育龄人群基因组DNA甲基化状态的研究

国家自然科学基金

0+阅读 · 2011年12月31日

认知无线电系统中的非凸优化资源分配算法

国家自然科学基金

1+阅读 · 2009年12月31日

Continual Machine Reading Comprehension via Uncertainty-aware Fixed Memory and Adversarial Domain Adaptation

Arxiv

0+阅读 · 2022年8月10日

Few-shot Learning with Retrieval Augmented Language Models

Arxiv

0+阅读 · 2022年8月8日

MVDG: A Unified Multi-view Framework for Domain Generalization

Arxiv

0+阅读 · 2022年8月8日

DyTox: Transformers for Continual Learning with DYnamic TOken eXpansion

Arxiv

0+阅读 · 2022年8月7日

Few-shot Learning with Retrieval Augmented Language Model

Arxiv

0+阅读 · 2022年8月5日

DnS: Distill-and-Select for Efficient and Accurate Video Indexing and Retrieval

Arxiv

0+阅读 · 2022年8月5日

Self-Supervised Learning for Recommender Systems: A Survey

Arxiv

12+阅读 · 2022年3月29日

Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Arxiv

12+阅读 · 2020年6月23日

Meta Learning for End-to-End Low-Resource Speech Recognition

Meta Learning for End-to-End Low-Resource Speech Recognition

Arxiv

20+阅读 · 2019年10月26日

Label-aware Double Transfer Learning for Cross-Specialty Medical Named Entity Recognition

Arxiv

10+阅读 · 2018年4月28日

VIP会员

文章信息

相关主题

相关VIP内容

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

125+阅读 · 2022年4月21日

【Manning新书】迁移学习自然语言处理，266页pdf，Transfer Learning for NLP

【Manning新书】迁移学习自然语言处理，266页pdf，Transfer Learning for NLP

专知会员服务

137+阅读 · 2021年11月6日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《在多无人机系统中利用软件定义无线电共享位置信息》

851页！《潮涨之海：代数几何的基础》新书

《2025年空天与国防领域的新兴趋势：驾驭创新与韧性的新时代》最新40页报告

航天遥感大模型发展综述与产业化应用展望

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Industry Talk2

【ICIG2021】Latest News & Announcements of the Industry Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年7月29日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Continual Machine Reading Comprehension via Uncertainty-aware Fixed Memory and Adversarial Domain Adaptation

Arxiv

0+阅读 · 2022年8月10日

Few-shot Learning with Retrieval Augmented Language Models

Arxiv

0+阅读 · 2022年8月8日

MVDG: A Unified Multi-view Framework for Domain Generalization

Arxiv

0+阅读 · 2022年8月8日

DyTox: Transformers for Continual Learning with DYnamic TOken eXpansion

Arxiv

0+阅读 · 2022年8月7日

Few-shot Learning with Retrieval Augmented Language Model

Arxiv

0+阅读 · 2022年8月5日

DnS: Distill-and-Select for Efficient and Accurate Video Indexing and Retrieval

Arxiv

0+阅读 · 2022年8月5日

Self-Supervised Learning for Recommender Systems: A Survey

Arxiv

12+阅读 · 2022年3月29日

Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Arxiv

12+阅读 · 2020年6月23日

Meta Learning for End-to-End Low-Resource Speech Recognition

Meta Learning for End-to-End Low-Resource Speech Recognition

Arxiv

20+阅读 · 2019年10月26日

Label-aware Double Transfer Learning for Cross-Specialty Medical Named Entity Recognition

Arxiv

10+阅读 · 2018年4月28日

相关基金

孤独症儿童早期干预的同步TMS-EEG研究

国家自然科学基金

0+阅读 · 2017年12月31日

SATB2-Nanog-mTOR轴调控BMSCs衰老及其在颌骨增龄性骨量丢失中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

核苷肽类抗生素缩合机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

功能性遗传变异调控BARD1/BRCA1泛素化通路的机制及与儿童神经母细胞瘤的关联研究

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Legendre 级数多极边界元法理论研究

国家自然科学基金

0+阅读 · 2013年12月31日

PGC-1α调节骨骼肌脂肪酸代谢和胰岛素抵抗的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

精神分裂症超高危人群经颅磁刺激干预的磁共振研究

国家自然科学基金

0+阅读 · 2012年12月31日

暴露于电子垃圾污染物中的育龄人群基因组DNA甲基化状态的研究

国家自然科学基金

0+阅读 · 2011年12月31日

认知无线电系统中的非凸优化资源分配算法

国家自然科学基金

1+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员