与短文本模型的高效长文本理解 (Efficient Long-Text Understanding with Short-Text Models) - 专知论文

会员服务 ·

0

可理解性 · MoDELS · Seven · INFORMS · 语言模型化 ·

2022 年 12 月 27 日

Efficient Long-Text Understanding with Short-Text Models

翻译：与短文本模型的高效长文本理解

Maor Ivgi,Uri Shaham,Jonathan Berant

from arxiv, Accepted for publication in Transactions of the Association for Computational Linguistics (TACL), 2023. Authors' final version (pre-MIT)

Transformer-based pretrained language models (LMs) are ubiquitous across natural language understanding, but cannot be applied to long sequences such as stories, scientific articles and long documents, due to their quadratic complexity. While a myriad of efficient transformer variants have been proposed, they are typically based on custom implementations that require expensive pretraining from scratch. In this work, we propose SLED: SLiding-Encoder and Decoder, a simple approach for processing long sequences that re-uses and leverages battle-tested short-text pretrained LMs. Specifically, we partition the input into overlapping chunks, encode each with a short-text LM encoder and use the pretrained decoder to fuse information across chunks (fusion-in-decoder). We illustrate through controlled experiments that SLED offers a viable strategy for long text understanding and evaluate our approach on SCROLLS, a benchmark with seven datasets across a wide range of language understanding tasks. We find that SLED is competitive with specialized models that are up to 50x larger and require a dedicated and expensive pretraining step.

翻译：在这项工作中,我们提议SLED:SLED:SLED:Lidedy-Encoder和Decoder,这是一个处理长序列的简单方法,可以重新使用并利用经过战斗测试的短文本预培训LMS。具体地说,我们把输入分解成重叠块,每个输入以短文本 LM 编码,并使用经过预先训练的解码器将各块信息(聚解解码器)连接起来。我们通过受控实验说明,SLED为长文本理解提供了可行的战略,并评估了我们在SCROLLS上的做法。 SLED是七套基准,跨越了广泛的语言理解任务。我们发现,SLED具有竞争力,专门模型多达50x大,需要专门和昂贵的培训前步骤。

0

相关内容

可理解性

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

48+阅读 · 2022年10月2日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

72+阅读 · 2022年3月15日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

123+阅读 · 2020年7月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

161+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

52+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

45+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

77+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

90+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

77+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

64+阅读 · 2019年10月9日

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

1+阅读 · 2022年11月2日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

23+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

26+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

41+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

16+阅读 · 2018年12月24日

【论文推荐】最新7篇视觉问答（VQA）相关论文—解释、读写记忆网络、逆视觉问答、视觉推理、可解释性、注意力机制、计数

【论文推荐】最新7篇视觉问答（VQA）相关论文—解释、读写记忆网络、逆视觉问答、视觉推理、可解释性、注意力机制、计数

专知

30+阅读 · 2018年3月22日

Erk调控Wnt通路影响间充质干细胞分化的机制及其在骨质疏松中的作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

新型半金属性Heusler合金及其半金属性的稳定性研究

国家自然科学基金

0+阅读 · 2012年12月31日

异基因MSC抑制TNF-α促进成骨分化治疗类风湿关节炎骨质疏松的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

NiCoMnIn/Mg智能复合材料研究

国家自然科学基金

0+阅读 · 2012年12月31日

高固溶度Mg-RE二元合金塑性变形机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

时效镁合金的沉淀析出与强韧化机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

多模态电磁分支电路振动控制理论与应用研究

国家自然科学基金

0+阅读 · 2011年12月31日

微量Zr、Mg等在Cu-Cr-Zr铜合金时效过程中的作用机理

国家自然科学基金

0+阅读 · 2011年12月31日

Mg-Ca-Sr合金的腐蚀降解及其降解产物的生物学效应

国家自然科学基金

0+阅读 · 2011年12月31日

miR146a抑制Smad4对骨髓间充质干细胞成骨分化调控的研究

国家自然科学基金

0+阅读 · 2009年12月31日

SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks

Arxiv

0+阅读 · 2023年2月27日

I-ViT: Integer-only Quantization for Efficient Vision Transformer Inference

Arxiv

0+阅读 · 2023年2月27日

Principled and Efficient Transfer Learning of Deep Models via Neural Collapse

Arxiv

0+阅读 · 2023年2月26日

Computing the Difference of Conjunctive Queries Efficiently

Arxiv

0+阅读 · 2023年2月25日

Less is More: Data Pruning for Faster Adversarial Training

Arxiv

0+阅读 · 2023年2月23日

Efficient Transformers: A Survey

Arxiv

35+阅读 · 2022年3月14日

Interpretable and Efficient Heterogeneous Graph Convolutional Network

Arxiv

15+阅读 · 2021年9月8日

Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions

Arxiv

20+阅读 · 2021年8月30日

Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better

Arxiv

27+阅读 · 2021年6月16日

Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Arxiv

11+阅读 · 2020年6月23日

VIP会员

文章信息

相关主题

语言模型化

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

48+阅读 · 2022年10月2日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

72+阅读 · 2022年3月15日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

123+阅读 · 2020年7月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

161+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

52+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

45+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

77+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

90+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

77+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

64+阅读 · 2019年10月9日

热门VIP内容

相关资讯

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

1+阅读 · 2022年11月2日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

23+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

26+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

41+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

16+阅读 · 2018年12月24日

【论文推荐】最新7篇视觉问答（VQA）相关论文—解释、读写记忆网络、逆视觉问答、视觉推理、可解释性、注意力机制、计数

【论文推荐】最新7篇视觉问答（VQA）相关论文—解释、读写记忆网络、逆视觉问答、视觉推理、可解释性、注意力机制、计数

专知

30+阅读 · 2018年3月22日

相关论文

SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks

Arxiv

0+阅读 · 2023年2月27日

I-ViT: Integer-only Quantization for Efficient Vision Transformer Inference

Arxiv

0+阅读 · 2023年2月27日

Principled and Efficient Transfer Learning of Deep Models via Neural Collapse

Arxiv

0+阅读 · 2023年2月26日

Computing the Difference of Conjunctive Queries Efficiently

Arxiv

0+阅读 · 2023年2月25日

Less is More: Data Pruning for Faster Adversarial Training

Arxiv

0+阅读 · 2023年2月23日

Efficient Transformers: A Survey

Arxiv

35+阅读 · 2022年3月14日

Interpretable and Efficient Heterogeneous Graph Convolutional Network

Arxiv

15+阅读 · 2021年9月8日

Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions

Arxiv

20+阅读 · 2021年8月30日

Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better

Arxiv

27+阅读 · 2021年6月16日

Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Arxiv

11+阅读 · 2020年6月23日

相关基金

Erk调控Wnt通路影响间充质干细胞分化的机制及其在骨质疏松中的作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

新型半金属性Heusler合金及其半金属性的稳定性研究

国家自然科学基金

0+阅读 · 2012年12月31日

异基因MSC抑制TNF-α促进成骨分化治疗类风湿关节炎骨质疏松的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

NiCoMnIn/Mg智能复合材料研究

国家自然科学基金

0+阅读 · 2012年12月31日

高固溶度Mg-RE二元合金塑性变形机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

时效镁合金的沉淀析出与强韧化机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

多模态电磁分支电路振动控制理论与应用研究

国家自然科学基金

0+阅读 · 2011年12月31日

微量Zr、Mg等在Cu-Cr-Zr铜合金时效过程中的作用机理

国家自然科学基金

0+阅读 · 2011年12月31日

Mg-Ca-Sr合金的腐蚀降解及其降解产物的生物学效应

国家自然科学基金

0+阅读 · 2011年12月31日

miR146a抑制Smad4对骨髓间充质干细胞成骨分化调控的研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员