学习强健的语音表现,配有动脉结构正规化的变异自动编码自动编码器 (Learning robust speech representation with an articulatory-regularized variational autoencoder)

It is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model (here a variational autoencoder) trained to encode and decode acoustic speech features. First we develop an articulatory model able to associate articulatory parameters describing the jaw, tongue, lips and velum configurations with vocal tract shapes and spectral features. Then we incorporate these articulatory parameters into a variational autoencoder applied on spectral features by using a regularization technique that constraints part of the latent space to follow articulatory trajectories. We show that this articulatory constraint improves model training by decreasing time to convergence and reconstruction loss at convergence, and yields better performance in a speech denoising task.

翻译：人们日益认识到,人类的言语感知和制作都依赖于动脉表征。在本文中,我们调查这种表征是否能够改善深层基因模型(这里是一个可变自动编码器)的性能,该模型经过了对声调特征进行编码和解码的培训。首先,我们开发了一种动脉模型,能够将描述下巴、舌头、嘴唇和排卵管配置的动脉参数与声道形状和光谱特征联系起来。然后,我们将这些动脉参数纳入对光谱特征应用的变异自动编码器中,使用一种正规化技术,限制潜在空间的一部分以跟踪动脉道轨迹。我们表明,这种动脉冲制约通过缩短时间以趋同和在趋同时重建损耗来改进模式培训,并在语言解析任务中产生更好的性能。

相关内容

自编码器

关注 140

自动编码器是一种人工神经网络，用于以无监督的方式学习有效的数据编码。自动编码器的目的是通过训练网络忽略信号“噪声”来学习一组数据的表示（编码），通常用于降维。与简化方面一起，学习了重构方面，在此，自动编码器尝试从简化编码中生成尽可能接近其原始输入的表示形式，从而得到其名称。基本模型存在几种变体，其目的是迫使学习的输入表示形式具有有用的属性。自动编码器可有效地解决许多应用问题，从面部识别到获取单词的语义。

多标签学习的新趋势（2020 Survey）

专知会员服务

43+阅读 · 2020年12月6日

【CVPR2020】在线深度聚类的无监督表示学习, Online Deep Clustering for Unsupervised Representation Learning

专知会员服务

69+阅读 · 2020年6月19日

强化学习的对比无监督表示，CURL: Contrastive Unsupervised Representations for Reinforcement Learning

专知会员服务

41+阅读 · 2020年4月11日

【InterSpeech2020】混合语音识别系统中的词汇扩展技术，Techniques for Vocabulary Expansion in Hybrid Speech Recognition Systems

专知会员服务

17+阅读 · 2020年3月23日