DuReader:来自现实世界应用的中国机器阅读综合数据集 (DuReader: a Chinese Machine Reading Comprehension Dataset from Real-world Applications)

In this paper, we introduce DuReader, a new large-scale, open-domain Chinese machine reading comprehension (MRC) dataset, aiming to tackle real-world MRC problems. In comparison to prior datasets, DuReader has the following characteristics: (a) the questions and the documents are all extracted from real application data, and the answers are human generated; (b) it provides rich annotations for question types, especially yes-no and opinion questions, which take a large proportion in real users' questions but have not been well studied before; (c) it provides multiple answers for each question. The first release of DuReader contains 200k questions, 1,000k documents, and 420k answers, which, to the best of our knowledge, is the largest Chinese MRC dataset so far. Experimental results show there exists big gap between the state-of-the-art baseline systems and human performance, which indicates DuReader is a challenging dataset that deserves future study. The dataset and the code of the baseline systems are publicly available now.

翻译：本文介绍DuReader(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader))(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(MRC)(MRC)(DuReader)(DuReader)(DuReader)(DuReader)(DuReader)(Dumber)(DuReader)(Dureader)(Dublead) (Dublead) (Dread) (Dragues) (Dravelop) (MRC(MRC) (MRC(MRC) (MRC) (MRC) (MRC) (MRC) (MRC) (D) (D) (D) (D) (D) (D) (D) (D) (DR) (D)) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) ) (D) (D) (D) (D) (D) (Dir) (D) (D) (D) (D)) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (Dr) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D) (D

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

【硬核书】金融数学C++编程，411页pdf，C++ for Financial Mathematics

专知会员服务

75+阅读 · 2020年4月6日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日