变色龙: 基于大语言模型的即插即用组合推理 (Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models) - 专知论文

会员服务 ·

0

即插即用 · GPT-4 · 工具 · 语言模型 · 推断 ·

2023 年 4 月 19 日

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

翻译：变色龙: 基于大语言模型的即插即用组合推理

Pan Lu,Baolin Peng,Hao Cheng,Michel Galley,Kai-Wei Chang,Ying Nian Wu,Song-Chun Zhu,Jianfeng Gao

from arxiv, 25 pages, 10 figures. Project page: https://chameleon-llm.github.io

Large language models (LLMs) have achieved remarkable progress in various natural language processing tasks with emergent abilities. However, they face inherent limitations, such as an inability to access up-to-date information, utilize external tools, or perform precise mathematical reasoning. In this paper, we introduce Chameleon, a plug-and-play compositional reasoning framework that augments LLMs to help address these challenges. Chameleon synthesizes programs to compose various tools, including LLM models, off-the-shelf vision models, web search engines, Python functions, and rule-based modules tailored to user interests. Built on top of an LLM as a natural language planner, Chameleon infers the appropriate sequence of tools to compose and execute in order to generate a final response. We showcase the adaptability and effectiveness of Chameleon on two tasks: ScienceQA and TabMWP. Notably, Chameleon with GPT-4 achieves an 86.54% accuracy on ScienceQA, significantly improving upon the best published few-shot model by 11.37%; using GPT-4 as the underlying LLM, Chameleon achieves a 17.8% increase over the state-of-the-art model, leading to a 98.78% overall accuracy on TabMWP. Further studies suggest that using GPT-4 as a planner exhibits more consistent and rational tool selection and is able to infer potential constraints given the instructions, compared to other LLMs like ChatGPT.

翻译：大型语言模型（LLM）在各种自然语言处理任务中取得了显着进展，并具有新兴的能力。然而，它们面临固有的限制，如无法访问最新信息、利用外部工具或执行精确的数学推理。在本文中，我们介绍了变色龙，这是一个即插即用的组合推理框架，它增强了LLM以帮助解决这些挑战。变色龙综合各种工具，包括LLM模型、现成的视觉模型、网络搜索引擎、Python函数和针对用户兴趣量身定制的基于规则的模块。基于LLM作为自然语言计划器，变色龙推断出合适的工具序列以组合和执行，生成最终的响应。我们展示了变色龙在两个任务中的适应性和有效性: ScienceQA和TabMWP。值得注意的是，使用GPT-4的变色龙在ScienceQA上实现了86.54%的准确率，比最好的发布的少量训练模型提高了11.37%;使用GPT-4作为底层LLM，变色龙在TabMWP上实现了17.8%的增长，导致98.78%的整体准确度。进一步的研究表明，使用GPT-4作为计划器表现出更一致和合理的工具选择，并能够推断出可能的约束条件，给出指令，相比于其他LLM，如ChatGPT更加优秀。

1

相关内容

即插即用

【ICML2023】调整语言模型作为增强少样本学习的训练数据生成器

【ICML2023】调整语言模型作为增强少样本学习的训练数据生成器

专知会员服务

32+阅读 · 2023年5月19日

【CIKM2021】基于检索的个性化聊天机器人模型IMPChat

专知会员服务

17+阅读 · 2021年8月25日

NLP新范式-预训练，提示(Prompt)，预测！CMU刘鹏飞等论文综述预训练语言模型提示学习进展

NLP新范式-预训练，提示(Prompt)，预测！CMU刘鹏飞等论文综述预训练语言模型提示学习进展

专知会员服务

71+阅读 · 2021年7月31日

【神经自然语言处理进展：建模，学习，推理】Progress in Neural NLP: Modeling, Learning, and Reasoning

【神经自然语言处理进展：建模，学习，推理】Progress in Neural NLP: Modeling, Learning, and Reasoning

专知会员服务

78+阅读 · 2020年8月13日

【2020新书】自然语言处理Python与spaCy实践，216页pdf，NLP with Python

【2020新书】自然语言处理Python与spaCy实践，216页pdf，NLP with Python

专知会员服务

108+阅读 · 2020年5月1日

【论文翻译】NLP注意力机制综述论文翻译，Attention, please! A Critical Review of Neural Attention Models in Natural Language Processing

【论文翻译】NLP注意力机制综述论文翻译，Attention, please! A Critical Review of Neural Attention Models in Natural Language Processing

专知会员服务

96+阅读 · 2020年4月18日

【博士论文】CHAMELEON:新闻推荐系统的深度学习元架构，187页pdf，CHAMELEON: A Deep Learning Meta-Architecture for News Recommender Systems [Phd. Thesis]

【博士论文】CHAMELEON:新闻推荐系统的深度学习元架构，187页pdf，CHAMELEON: A Deep Learning Meta-Architecture for News Recommender Systems [Phd. Thesis]

专知会员服务

22+阅读 · 2020年1月15日

【Google论文强烈推荐】ALBERT:基于精简BERT的自我监督学习的语言表示，ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations

【Google论文强烈推荐】ALBERT:基于精简BERT的自我监督学习的语言表示，ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations

专知会员服务

24+阅读 · 2019年12月21日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

深度学习自然语言处理

18+阅读 · 2020年5月22日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

基于LSTM-CNN组合模型的Twitter情感分析（附代码）

基于LSTM-CNN组合模型的Twitter情感分析（附代码）

机器学习研究会

50+阅读 · 2018年2月21日

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

专知

15+阅读 · 2018年2月3日

【论文】图上的表示学习综述

【论文】图上的表示学习综述

机器学习研究会

15+阅读 · 2017年9月24日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

基于非独立同分布学习理论的图模型词义消歧及领域适应方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

宽窄带混合主动噪声控制系统高效稳健算法及应用研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于蛋白质/多肽自组装模板的新型介孔催化材料设计

国家自然科学基金

0+阅读 · 2014年12月31日

Anderson型多酸的不对称修饰及可控组装研究

国家自然科学基金

1+阅读 · 2014年12月31日

星座自主运行时间同步精度测评方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

多通道SAR地面运动目标自动检测与定位技术研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于非下采样Contourlet变换的多源影像自动匹配方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

三维线性阱离子囚禁的实验研究

国家自然科学基金

0+阅读 · 2009年12月31日

川芎嗪衍生物的化学合成与SAR研究

国家自然科学基金

0+阅读 · 2009年12月31日

Langmuir环流在上层海洋混合中的作用

国家自然科学基金

0+阅读 · 2008年12月31日

In-Context Analogical Reasoning with Pre-Trained Language Models

Arxiv

0+阅读 · 2023年6月5日

ThinkSum: Probabilistic reasoning over sets using large language models

Arxiv

0+阅读 · 2023年6月2日

Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker

Arxiv

0+阅读 · 2023年6月1日

Augmented Large Language Models with Parametric Knowledge Guiding

Arxiv

20+阅读 · 2023年5月8日

A Survey of Large Language Models

A Survey of Large Language Models

Arxiv

470+阅读 · 2023年3月31日

Towards Reasoning in Large Language Models: A Survey

Arxiv

34+阅读 · 2022年12月20日

Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey

Arxiv

31+阅读 · 2021年11月1日

K-AID: Enhancing Pre-trained Language Models with Domain Knowledge for Question Answering

Arxiv

15+阅读 · 2021年9月22日

Differentiable Reasoning on Large Knowledge Bases and Natural Language

Arxiv

12+阅读 · 2019年12月17日

An Attentive Survey of Attention Models

Arxiv

19+阅读 · 2019年4月5日

VIP会员

文章信息

相关主题

相关VIP内容

【ICML2023】调整语言模型作为增强少样本学习的训练数据生成器

【ICML2023】调整语言模型作为增强少样本学习的训练数据生成器

专知会员服务

32+阅读 · 2023年5月19日

【CIKM2021】基于检索的个性化聊天机器人模型IMPChat

专知会员服务

17+阅读 · 2021年8月25日

NLP新范式-预训练，提示(Prompt)，预测！CMU刘鹏飞等论文综述预训练语言模型提示学习进展

NLP新范式-预训练，提示(Prompt)，预测！CMU刘鹏飞等论文综述预训练语言模型提示学习进展

专知会员服务

71+阅读 · 2021年7月31日

【神经自然语言处理进展：建模，学习，推理】Progress in Neural NLP: Modeling, Learning, and Reasoning

【神经自然语言处理进展：建模，学习，推理】Progress in Neural NLP: Modeling, Learning, and Reasoning

专知会员服务

78+阅读 · 2020年8月13日

【2020新书】自然语言处理Python与spaCy实践，216页pdf，NLP with Python

【2020新书】自然语言处理Python与spaCy实践，216页pdf，NLP with Python

专知会员服务

108+阅读 · 2020年5月1日

【论文翻译】NLP注意力机制综述论文翻译，Attention, please! A Critical Review of Neural Attention Models in Natural Language Processing

【论文翻译】NLP注意力机制综述论文翻译，Attention, please! A Critical Review of Neural Attention Models in Natural Language Processing

专知会员服务

96+阅读 · 2020年4月18日

【博士论文】CHAMELEON:新闻推荐系统的深度学习元架构，187页pdf，CHAMELEON: A Deep Learning Meta-Architecture for News Recommender Systems [Phd. Thesis]

【博士论文】CHAMELEON:新闻推荐系统的深度学习元架构，187页pdf，CHAMELEON: A Deep Learning Meta-Architecture for News Recommender Systems [Phd. Thesis]

专知会员服务

22+阅读 · 2020年1月15日

【Google论文强烈推荐】ALBERT:基于精简BERT的自我监督学习的语言表示，ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations

【Google论文强烈推荐】ALBERT:基于精简BERT的自我监督学习的语言表示，ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations

专知会员服务

24+阅读 · 2019年12月21日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《多视角时空一致多模态感知目标检测的对抗鲁棒性研究》DARPA赞助最新96页技术报告

4300字《创新防御：生成式人工智能在美国军事演进中的角色》附原文

《主动式社会工程防御（ASED）项目》美空军24页项目报告

《军事域人工智能及其对国际影响：未来政策行动路线图》2025最新31页报告

相关资讯

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理

深度学习自然语言处理

18+阅读 · 2020年5月22日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

基于LSTM-CNN组合模型的Twitter情感分析（附代码）

基于LSTM-CNN组合模型的Twitter情感分析（附代码）

机器学习研究会

50+阅读 · 2018年2月21日

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

专知

15+阅读 · 2018年2月3日

【论文】图上的表示学习综述

【论文】图上的表示学习综述

机器学习研究会

15+阅读 · 2017年9月24日

【推荐】RNN/LSTM时序预测

【推荐】RNN/LSTM时序预测

机器学习研究会

25+阅读 · 2017年9月8日

相关论文

In-Context Analogical Reasoning with Pre-Trained Language Models

Arxiv

0+阅读 · 2023年6月5日

ThinkSum: Probabilistic reasoning over sets using large language models

Arxiv

0+阅读 · 2023年6月2日

Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker

Arxiv

0+阅读 · 2023年6月1日

Augmented Large Language Models with Parametric Knowledge Guiding

Arxiv

20+阅读 · 2023年5月8日

A Survey of Large Language Models

A Survey of Large Language Models

Arxiv

470+阅读 · 2023年3月31日

Towards Reasoning in Large Language Models: A Survey

Arxiv

34+阅读 · 2022年12月20日

Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey

Arxiv

31+阅读 · 2021年11月1日

K-AID: Enhancing Pre-trained Language Models with Domain Knowledge for Question Answering

Arxiv

15+阅读 · 2021年9月22日

Differentiable Reasoning on Large Knowledge Bases and Natural Language

Arxiv

12+阅读 · 2019年12月17日

An Attentive Survey of Attention Models

Arxiv

19+阅读 · 2019年4月5日

相关基金

基于非独立同分布学习理论的图模型词义消歧及领域适应方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

宽窄带混合主动噪声控制系统高效稳健算法及应用研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于蛋白质/多肽自组装模板的新型介孔催化材料设计

国家自然科学基金

0+阅读 · 2014年12月31日

Anderson型多酸的不对称修饰及可控组装研究

国家自然科学基金

1+阅读 · 2014年12月31日

星座自主运行时间同步精度测评方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

多通道SAR地面运动目标自动检测与定位技术研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于非下采样Contourlet变换的多源影像自动匹配方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

三维线性阱离子囚禁的实验研究

国家自然科学基金

0+阅读 · 2009年12月31日

川芎嗪衍生物的化学合成与SAR研究

国家自然科学基金

0+阅读 · 2009年12月31日

Langmuir环流在上层海洋混合中的作用

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员