重现性缺乏正确性不足：测试NLP代码的重要性 (Reproducibility is Nothing without Correctness: The Importance of Testing Code in NLP) - 专知论文

会员服务 ·

0

正确性 · 重现性 · 潜在 · 代码 · Conformer ·

2023 年 3 月 28 日

Reproducibility is Nothing without Correctness: The Importance of Testing Code in NLP

翻译：重现性缺乏正确性不足：测试NLP代码的重要性

Sara Papi,Marco Gaido,Matteo Negri,Andrea Pilzer

Despite its pivotal role in research experiments, code correctness is often presumed only on the basis of the perceived quality of the results. This comes with the risk of erroneous outcomes and potentially misleading findings. To address this issue, we posit that the current focus on result reproducibility should go hand in hand with the emphasis on coding best practices. We bolster our call to the NLP community by presenting a case study, in which we identify (and correct) three bugs in widely used open-source implementations of the state-of-the-art Conformer architecture. Through comparative experiments on automatic speech recognition and translation in various language settings, we demonstrate that the existence of bugs does not prevent the achievement of good and reproducible results and can lead to incorrect conclusions that potentially misguide future research. In response to this, this study is a call to action toward the adoption of coding best practices aimed at fostering correctness and improving the quality of the developed software.

翻译：尽管其在研究实验中具有关键作用，但代码正确性通常只基于结果质量的感知被认定。这会带来错误结果和潜在的误导性发现的风险。为解决这个问题，我们认为当前对结果可重复性的关注应该与对编码最佳实践的强调相辅相成。通过案例研究，我们证明了这一点，我们在其中识别(和纠正)了当前广泛使用的最先进的Conformer结构的开源实现中的三个错误。通过对各种语言设置中的自动语音识别和翻译的比较实验，我们证明了错误的存在并不会阻止实现良好和可重复的结果，并且可能导致不正确的结论，从而潜在地误导未来的研究。为了应对这个问题，这项研究呼吁采用旨在促进正确性和提高开发软件质量的编码最佳实践。

0

相关内容

正确性

【AAAI 2022】机器学习模型的解释方法效果如何？MIT、微软学者为你解读，Do Feature Attribution Methods Correctly Attribute Features?

【AAAI 2022】机器学习模型的解释方法效果如何？MIT、微软学者为你解读，Do Feature Attribution Methods Correctly Attribute Features?

专知会员服务

31+阅读 · 2022年3月12日

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

专知会员服务

34+阅读 · 2022年3月5日

【干货书】机器学习设计模式，408页pdf，Machine Learning Design Patterns

【干货书】机器学习设计模式，408页pdf，Machine Learning Design Patterns

专知会员服务

138+阅读 · 2022年2月6日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

神经网络与形式语言综述，12页pdf，A Survey of Neural Networks and Formal Languages

神经网络与形式语言综述，12页pdf，A Survey of Neural Networks and Formal Languages

专知会员服务

21+阅读 · 2020年6月4日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【Google可解释人工智能白皮书】27页pdf，AI Explainability Whitepaper ，Introduction to AI Explanations for AI Platform

【Google可解释人工智能白皮书】27页pdf，AI Explainability Whitepaper ，Introduction to AI Explanations for AI Platform

专知会员服务

127+阅读 · 2019年12月13日

【Google论文】ALBERT:自我监督学习语言表达的精简BERT

【Google论文】ALBERT:自我监督学习语言表达的精简BERT

专知会员服务

24+阅读 · 2019年11月4日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

使用BERT做文本摘要

使用BERT做文本摘要

专知

23+阅读 · 2019年12月7日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

AINLP

40+阅读 · 2019年6月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【推荐】自然语言处理（NLP）指南

【推荐】自然语言处理（NLP）指南

机器学习研究会

35+阅读 · 2017年11月17日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

【推荐】用Tensorflow理解LSTM

【推荐】用Tensorflow理解LSTM

机器学习研究会

36+阅读 · 2017年9月11日

基于连续循环平移理论的Shearlet域稀疏表示SAR图像去噪算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

三维椭圆方程Cauchy问题的正则化方法

国家自然科学基金

0+阅读 · 2013年12月31日

分子束外延生长n型BaBiO3拓扑绝缘体的原位角分辨光电子能谱研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于BOTDR技术的岩溶塌陷监测预警试验研究

国家自然科学基金

0+阅读 · 2013年12月31日

LncRNA MEG3在糖尿病肾病足细胞损伤中的作用及机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

半参数回归分析的随机函数法及其高维情形

国家自然科学基金

2+阅读 · 2012年12月31日

支原体的必需基因识别与核心基因组分析

国家自然科学基金

0+阅读 · 2011年12月31日

基于格值逻辑的α-n(t)元归结动态自动推理研究

国家自然科学基金

0+阅读 · 2011年12月31日

适应纳米尺度CMOS集成电路DFM的ULTRA模型完善和偏差模拟技术研究

国家自然科学基金

0+阅读 · 2009年12月31日

洋葱新病原细菌Burkholderia cenocepacia的致病基因及检测技术研究

国家自然科学基金

0+阅读 · 2008年12月31日

Segment Anything Model for Medical Images?

Arxiv

0+阅读 · 2023年5月19日

A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks

Arxiv

0+阅读 · 2023年5月18日

An Empirical Study on the Language Modal in Visual Question Answering

Arxiv

0+阅读 · 2023年5月17日

Probing the Role of Positional Information in Vision-Language Models

Arxiv

0+阅读 · 2023年5月17日

Toward Falsifying Causal Graphs Using a Permutation-Based Test

Arxiv

0+阅读 · 2023年5月16日

Combining datasets to increase the number of samples and improve model fitting

Arxiv

0+阅读 · 2023年5月16日

On the Origins of Bias in NLP through the Lens of the Jim Code

Arxiv

0+阅读 · 2023年5月16日

Soft Prompt Decoding for Multilingual Dense Retrieval

Arxiv

0+阅读 · 2023年5月15日

InteractE: Improving Convolution-based Knowledge Graph Embeddings by Increasing Feature Interactions

InteractE: Improving Convolution-based Knowledge Graph Embeddings by Increasing Feature Interactions

Arxiv

13+阅读 · 2019年11月1日

Additive Margin Softmax for Face Verification

Arxiv

11+阅读 · 2018年1月18日

VIP会员

文章信息

相关主题

相关VIP内容

【AAAI 2022】机器学习模型的解释方法效果如何？MIT、微软学者为你解读，Do Feature Attribution Methods Correctly Attribute Features?

【AAAI 2022】机器学习模型的解释方法效果如何？MIT、微软学者为你解读，Do Feature Attribution Methods Correctly Attribute Features?

专知会员服务

31+阅读 · 2022年3月12日

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

专知会员服务

34+阅读 · 2022年3月5日

【干货书】机器学习设计模式，408页pdf，Machine Learning Design Patterns

【干货书】机器学习设计模式，408页pdf，Machine Learning Design Patterns

专知会员服务

138+阅读 · 2022年2月6日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

神经网络与形式语言综述，12页pdf，A Survey of Neural Networks and Formal Languages

神经网络与形式语言综述，12页pdf，A Survey of Neural Networks and Formal Languages

专知会员服务

21+阅读 · 2020年6月4日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【Google可解释人工智能白皮书】27页pdf，AI Explainability Whitepaper ，Introduction to AI Explanations for AI Platform

【Google可解释人工智能白皮书】27页pdf，AI Explainability Whitepaper ，Introduction to AI Explanations for AI Platform

专知会员服务

127+阅读 · 2019年12月13日

【Google论文】ALBERT:自我监督学习语言表达的精简BERT

【Google论文】ALBERT:自我监督学习语言表达的精简BERT

专知会员服务

24+阅读 · 2019年11月4日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《多智能体不确定环境追逃博弈研究》216页

美智库最新发布《解放军"人机编组协同作战"发展路径：理论与实践》53页

现代战争"杀伤区"理论：空间尺度与结构特征、控制手段与毁伤机制、生存策略与战线转移

《俄军无人机创新技术或已在乌克兰达成"战场空中封锁"作战效果》最新18页报告

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

使用BERT做文本摘要

使用BERT做文本摘要

专知

23+阅读 · 2019年12月7日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

AINLP

40+阅读 · 2019年6月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【推荐】自然语言处理（NLP）指南

【推荐】自然语言处理（NLP）指南

机器学习研究会

35+阅读 · 2017年11月17日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

【推荐】用Tensorflow理解LSTM

【推荐】用Tensorflow理解LSTM

机器学习研究会

36+阅读 · 2017年9月11日

相关论文

Segment Anything Model for Medical Images?

Arxiv

0+阅读 · 2023年5月19日

A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks

Arxiv

0+阅读 · 2023年5月18日

An Empirical Study on the Language Modal in Visual Question Answering

Arxiv

0+阅读 · 2023年5月17日

Probing the Role of Positional Information in Vision-Language Models

Arxiv

0+阅读 · 2023年5月17日

Toward Falsifying Causal Graphs Using a Permutation-Based Test

Arxiv

0+阅读 · 2023年5月16日

Combining datasets to increase the number of samples and improve model fitting

Arxiv

0+阅读 · 2023年5月16日

On the Origins of Bias in NLP through the Lens of the Jim Code

Arxiv

0+阅读 · 2023年5月16日

Soft Prompt Decoding for Multilingual Dense Retrieval

Arxiv

0+阅读 · 2023年5月15日

InteractE: Improving Convolution-based Knowledge Graph Embeddings by Increasing Feature Interactions

InteractE: Improving Convolution-based Knowledge Graph Embeddings by Increasing Feature Interactions

Arxiv

13+阅读 · 2019年11月1日

Additive Margin Softmax for Face Verification

Arxiv

11+阅读 · 2018年1月18日

相关基金

基于连续循环平移理论的Shearlet域稀疏表示SAR图像去噪算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

三维椭圆方程Cauchy问题的正则化方法

国家自然科学基金

0+阅读 · 2013年12月31日

分子束外延生长n型BaBiO3拓扑绝缘体的原位角分辨光电子能谱研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于BOTDR技术的岩溶塌陷监测预警试验研究

国家自然科学基金

0+阅读 · 2013年12月31日

LncRNA MEG3在糖尿病肾病足细胞损伤中的作用及机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

半参数回归分析的随机函数法及其高维情形

国家自然科学基金

2+阅读 · 2012年12月31日

支原体的必需基因识别与核心基因组分析

国家自然科学基金

0+阅读 · 2011年12月31日

基于格值逻辑的α-n(t)元归结动态自动推理研究

国家自然科学基金

0+阅读 · 2011年12月31日

适应纳米尺度CMOS集成电路DFM的ULTRA模型完善和偏差模拟技术研究

国家自然科学基金

0+阅读 · 2009年12月31日

洋葱新病原细菌Burkholderia cenocepacia的致病基因及检测技术研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员