多语种核对表:生成和评价 (Multilingual CheckList: Generation and Evaluation) - 专知论文

会员服务 ·

0

TEA · MoDELS · 缩放 · Machine Translation · 估计/估计量 ·

2022 年 10 月 12 日

Multilingual CheckList: Generation and Evaluation

翻译：多语种核对表:生成和评价

Karthikeyan K,Shaily Bhatt,Pankaj Singh,Somak Aditya,Sandipan Dandapat,Sunayana Sitaram,Monojit Choudhury

from arxiv, Accepted to Findings of AACL-IJCNLP 2022

Multilingual evaluation benchmarks usually contain limited high-resource languages and do not test models for specific linguistic capabilities. CheckList is a template-based evaluation approach that tests models for specific capabilities. The CheckList template creation process requires native speakers, posing a challenge in scaling to hundreds of languages. In this work, we explore multiple approaches to generate Multilingual CheckLists. We device an algorithm - Template Extraction Algorithm (TEA) for automatically extracting target language CheckList templates from machine translated instances of a source language templates. We compare the TEA CheckLists with CheckLists created with different levels of human intervention. We further introduce metrics along the dimensions of cost, diversity, utility, and correctness to compare the CheckLists. We thoroughly analyze different approaches to creating CheckLists in Hindi. Furthermore, we experiment with 9 more different languages. We find that TEA followed by human verification is ideal for scaling Checklist-based evaluation to multiple languages while TEA gives a good estimates of model performance.

翻译：多语言评价基准通常包含有限的高资源语言,并不测试特定语言能力的模式。核对列表是一种基于模板的评估方法,用于测试特定能力的模式。核对列表的创建过程需要本地语言者,这给推广到数百种语言带来了挑战。在这项工作中,我们探索了多种方法来生成多语言核对列表者。我们设置了一种算法 — 模板提取算法(TEA),用于自动从源语言模板的机器翻译实例中提取目标语言校验列表模板。我们比较了TEA核对列表者与以不同水平的人类干预生成的校验列表者。我们进一步引入了成本、多样性、实用性和正确性等层面的计量标准,以比较核对列表者。我们深入分析了在印地语中创建核对列表者的不同方法。此外,我们尝试了9种更不同的语言。我们发现,通过人类核查的计算方法是将核对列表评估范围扩大到多种语言的理想方法,而TEA则对模型性能做了良好的估计。

0

相关内容

TEA

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

114+阅读 · 2020年4月5日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

52+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

45+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

32+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

54+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

171+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

77+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

91+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

64+阅读 · 2019年10月9日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

26+阅读 · 2019年5月18日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

19+阅读 · 2017年12月17日

有限域上指数和与量子码的研究

国家自然科学基金

0+阅读 · 2014年12月31日

Anderson型多酸的不对称修饰及可控组装研究

国家自然科学基金

1+阅读 · 2014年12月31日

离散时间马氏链的泛函不等式及遍历性

国家自然科学基金

0+阅读 · 2014年12月31日

有限域上多项式的p-进与T-进指数和

国家自然科学基金

0+阅读 · 2013年12月31日

非高斯过程驱动系统的随机不变流形

国家自然科学基金

0+阅读 · 2013年12月31日

KIBRA及APOE基因多态性对人脑记忆功能调控机制的多模态MRI研究

国家自然科学基金

0+阅读 · 2013年12月31日

中国桦木属植物外生菌根真菌多样性及分布格局研究

国家自然科学基金

0+阅读 · 2013年12月31日

分子水平研究放射性Cs(I)、Sr(II)、Am(III)在高岭石/水界面的吸附形态

国家自然科学基金

0+阅读 · 2012年12月31日

随机环境中随机游动与分枝系统

国家自然科学基金

0+阅读 · 2011年12月31日

Asperger综合症情绪认知的神经心理调控机制研究

国家自然科学基金

0+阅读 · 2008年12月31日

Holistic Evaluation of Language Models

Arxiv

0+阅读 · 2022年11月16日

Reasoning Circuits: Few-shot Multihop Question Generation with Structured Rationales

Arxiv

0+阅读 · 2022年11月15日

Automatic Evaluation of Excavator Operators using Learned Reward Functions

Arxiv

0+阅读 · 2022年11月15日

Finding Skill Neurons in Pre-trained Transformer-based Language Models

Arxiv

0+阅读 · 2022年11月14日

Grafting Pre-trained Models for Multimodal Headline Generation

Arxiv

0+阅读 · 2022年11月14日

Shape-based Evaluation of Epidemic Forecasts

Arxiv

0+阅读 · 2022年11月11日

Few-shot Image Generation with Diffusion Models

Arxiv

0+阅读 · 2022年11月11日

HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation

HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation

Arxiv

0+阅读 · 2022年11月11日

A Survey on Generative Diffusion Model

Arxiv

44+阅读 · 2022年9月6日

Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Arxiv

11+阅读 · 2020年5月8日

VIP会员

文章信息

相关主题

Machine Translation

估计/估计量

相关VIP内容

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

114+阅读 · 2020年4月5日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

52+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

45+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

32+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

54+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

171+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

77+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

91+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

64+阅读 · 2019年10月9日

热门VIP内容

相关资讯

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium7

中国图象图形学学会CSIG

0+阅读 · 2021年11月15日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

灾难性遗忘问题新视角：迁移-干扰平衡

灾难性遗忘问题新视角：迁移-干扰平衡

CreateAMind

17+阅读 · 2019年7月6日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

26+阅读 · 2019年5月18日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

19+阅读 · 2017年12月17日

相关论文

Holistic Evaluation of Language Models

Arxiv

0+阅读 · 2022年11月16日

Reasoning Circuits: Few-shot Multihop Question Generation with Structured Rationales

Arxiv

0+阅读 · 2022年11月15日

Automatic Evaluation of Excavator Operators using Learned Reward Functions

Arxiv

0+阅读 · 2022年11月15日

Finding Skill Neurons in Pre-trained Transformer-based Language Models

Arxiv

0+阅读 · 2022年11月14日

Grafting Pre-trained Models for Multimodal Headline Generation

Arxiv

0+阅读 · 2022年11月14日

Shape-based Evaluation of Epidemic Forecasts

Arxiv

0+阅读 · 2022年11月11日

Few-shot Image Generation with Diffusion Models

Arxiv

0+阅读 · 2022年11月11日

HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation

HumanDiffusion: a Coarse-to-Fine Alignment Diffusion Framework for Controllable Text-Driven Person Image Generation

Arxiv

0+阅读 · 2022年11月11日

A Survey on Generative Diffusion Model

Arxiv

44+阅读 · 2022年9月6日

Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Arxiv

11+阅读 · 2020年5月8日

相关基金

有限域上指数和与量子码的研究

国家自然科学基金

0+阅读 · 2014年12月31日

Anderson型多酸的不对称修饰及可控组装研究

国家自然科学基金

1+阅读 · 2014年12月31日

离散时间马氏链的泛函不等式及遍历性

国家自然科学基金

0+阅读 · 2014年12月31日

有限域上多项式的p-进与T-进指数和

国家自然科学基金

0+阅读 · 2013年12月31日

非高斯过程驱动系统的随机不变流形

国家自然科学基金

0+阅读 · 2013年12月31日

KIBRA及APOE基因多态性对人脑记忆功能调控机制的多模态MRI研究

国家自然科学基金

0+阅读 · 2013年12月31日

中国桦木属植物外生菌根真菌多样性及分布格局研究

国家自然科学基金

0+阅读 · 2013年12月31日

分子水平研究放射性Cs(I)、Sr(II)、Am(III)在高岭石/水界面的吸附形态

国家自然科学基金

0+阅读 · 2012年12月31日

随机环境中随机游动与分枝系统

国家自然科学基金

0+阅读 · 2011年12月31日

Asperger综合症情绪认知的神经心理调控机制研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员