以乘数启动装置为基础的探索 (Multiplier Bootstrap-based Exploration) - 专知论文

会员服务 ·

0

MBE · Extensibility · 赌博机/老虎机 · MoDELS · Bandits ·

2023 年 2 月 3 日

Multiplier Bootstrap-based Exploration

翻译：以乘数启动装置为基础的探索

Runzhe Wan,Haoyu Wei,Branislav Kveton,Rui Song

Despite the great interest in the bandit problem, designing efficient algorithms for complex models remains challenging, as there is typically no analytical way to quantify uncertainty. In this paper, we propose Multiplier Bootstrap-based Exploration (MBE), a novel exploration strategy that is applicable to any reward model amenable to weighted loss minimization. We prove both instance-dependent and instance-independent rate-optimal regret bounds for MBE in sub-Gaussian multi-armed bandits. With extensive simulation and real data experiments, we show the generality and adaptivity of MBE.

翻译：尽管对强盗问题非常感兴趣,但设计复杂模型的有效算法仍具有挑战性,因为通常没有分析方法来量化不确定性。在本文件中,我们提出“倍增诱杀陷阱探索”(MBE),这是一项适用于任何可加权减低损失的奖励模式的新探索战略。我们证明,在亚加盟多武装强盗中,MBE既依赖实例,又依赖实例的速率最佳遗憾。通过广泛的模拟和真实的数据实验,我们展示了MBE的普遍性和适应性。

0

相关内容

MBE

不可错过！700+ppt《因果推理》课程！杜克大学Fan Li教程

不可错过！700+ppt《因果推理》课程！杜克大学Fan Li教程

专知会员服务

72+阅读 · 2022年7月11日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

Berezin变换及相关的算子理论

国家自然科学基金

1+阅读 · 2014年12月31日

ZmEREB58转录因子在玉米虫害胁迫响应中的调控机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

实时安全关键系统的建模、仿真与验证

国家自然科学基金

1+阅读 · 2012年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

PhoBR双组份调控系统对胸膜肺炎放线杆菌致病性调控机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

Expert Kaplan--Meier estimation

Arxiv

0+阅读 · 2023年3月27日

Towards black-box parameter estimation

Arxiv

0+阅读 · 2023年3月27日

On the Convergence of Numerical Integration as a Finite Matrix Approximation to Multiplication Operator

Arxiv

0+阅读 · 2023年3月27日

Multi-agent Black-box Optimization using a Bayesian Approach to Alternating Direction Method of Multipliers

Arxiv

0+阅读 · 2023年3月25日

Weighted Pressure and Mode Matching for Sound Field Reproduction: Theoretical and Experimental Comparisons

Arxiv

0+阅读 · 2023年3月23日

VIP会员

文章信息

相关主题

赌博机/老虎机

相关VIP内容

不可错过！700+ppt《因果推理》课程！杜克大学Fan Li教程

不可错过！700+ppt《因果推理》课程！杜克大学Fan Li教程

专知会员服务

72+阅读 · 2022年7月11日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

前沿人工智能趋势报告（Frontier AI Trends Report）

【AAAI2026】善始则事半功倍：基于前缀优化的大语言模型推理强化学习

Andrej Karpathy：2025 年 LLM 年度回顾（2025 LLM Year in Review）

音退化问题：基于输入操控的鲁棒语音转换综述

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Expert Kaplan--Meier estimation

Arxiv

0+阅读 · 2023年3月27日

Towards black-box parameter estimation

Arxiv

0+阅读 · 2023年3月27日

On the Convergence of Numerical Integration as a Finite Matrix Approximation to Multiplication Operator

Arxiv

0+阅读 · 2023年3月27日

Multi-agent Black-box Optimization using a Bayesian Approach to Alternating Direction Method of Multipliers

Arxiv

0+阅读 · 2023年3月25日

Weighted Pressure and Mode Matching for Sound Field Reproduction: Theoretical and Experimental Comparisons

Arxiv

0+阅读 · 2023年3月23日

相关基金

Berezin变换及相关的算子理论

国家自然科学基金

1+阅读 · 2014年12月31日

ZmEREB58转录因子在玉米虫害胁迫响应中的调控机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

实时安全关键系统的建模、仿真与验证

国家自然科学基金

1+阅读 · 2012年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

PhoBR双组份调控系统对胸膜肺炎放线杆菌致病性调控机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员