减少的保守政策变化 (Variance-Reduced Conservative Policy Iteration) - 专知论文

会员服务 ·

0

策略迭代 · 样本复杂度 · 样本 · 可约的 · 经验风险 ·

2023 年 1 月 25 日

Variance-Reduced Conservative Policy Iteration

翻译：减少的保守政策变化

Naman Agarwal,Brian Bullins,Karan Singh

from arxiv, To appear in proceedings of ALT 2023; updated references

We study the sample complexity of reducing reinforcement learning to a sequence of empirical risk minimization problems over the policy space. Such reductions-based algorithms exhibit local convergence in the function space, as opposed to the parameter space for policy gradient algorithms, and thus are unaffected by the possibly non-linear or discontinuous parameterization of the policy class. We propose a variance-reduced variant of Conservative Policy Iteration that improves the sample complexity of producing a $\varepsilon$-functional local optimum from $O(\varepsilon^{-4})$ to $O(\varepsilon^{-3})$. Under state-coverage and policy-completeness assumptions, the algorithm enjoys $\varepsilon$-global optimality after sampling $O(\varepsilon^{-2})$ times, improving upon the previously established $O(\varepsilon^{-3})$ sample requirement.

翻译：我们研究了将强化学习降低到一系列在政策空间方面尽量减少风险的经验性问题的抽样复杂性。这种基于削减的算法在功能空间中表现出当地趋同,而不是政策梯度算法的参数空间,因此不受政策类别可能非线性或不连续参数化的影响。我们提议了一个减少差异的保守政策循环变式,将生产美元(varepsilon)-功能最佳当地价格的抽样复杂性从O(\ varepsilon)美元提高到O(\ varepsilon}-3})美元。根据国家覆盖和政策完整性假设,在取样美元( \ varepsilon}-2})之后,该算法享有瓦列普斯隆-全球最佳性,比以前确定的美元( varepsilon}-3}抽样要求有所改善。

0

相关内容

策略迭代

【干货书】数据分析优化，Optimization for Modern Data Analysis，117页pdf

【干货书】数据分析优化，Optimization for Modern Data Analysis，117页pdf

专知会员服务

60+阅读 · 2023年2月15日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

59+阅读 · 2022年4月22日

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

专知会员服务

29+阅读 · 2022年3月8日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

122+阅读 · 2020年11月20日

UC.Berkeley CS189讲义教材:《机器学习全面指南》，185页pdf

专知会员服务

158+阅读 · 2020年1月16日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

45+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

31+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

53+阅读 · 2019年10月17日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

90+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

64+阅读 · 2019年10月9日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

23+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

26+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

17+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

41+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

16+阅读 · 2018年12月24日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

Yang-Baxter矩阵方程解的研究与应用

国家自然科学基金

0+阅读 · 2015年12月31日

TMS1基因响应高温胁迫和ER Stress的分子机制

国家自然科学基金

0+阅读 · 2014年12月31日

S3AGA样本（Spitzer-SDSS Spectral Atlas of Galaxies and AGNs)及其AGN研究

国家自然科学基金

0+阅读 · 2014年12月31日

蛋白酶体抑制剂Bortezomib对肺动脉平滑肌细胞钙离子通路的作用研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于Fermi-LAT和AMS-02的暗物质理论研究

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

新型非甾体类AR拮抗剂的设计合成及生物活性评价

国家自然科学基金

0+阅读 · 2012年12月31日

猪带绦虫胰岛素与受体的相互作用及其对葡萄糖代谢的调节

国家自然科学基金

0+阅读 · 2012年12月31日

干旱、寒冷灌区土壤水盐耦合遥感监测敏感性研究

国家自然科学基金

0+阅读 · 2011年12月31日

模-相对Hochschild同调与上同调

国家自然科学基金

0+阅读 · 2011年12月31日

Dynamic Update-to-Data Ratio: Minimizing World Model Overfitting

Arxiv

0+阅读 · 2023年3月17日

A Policy Iteration Approach for Flock Motion Control

Arxiv

0+阅读 · 2023年3月17日

A New Policy Iteration Algorithm For Reinforcement Learning in Zero-Sum Markov Games

Arxiv

0+阅读 · 2023年3月17日

Enabling First-Order Gradient-Based Learning for Equilibrium Computation in Markets

Enabling First-Order Gradient-Based Learning for Equilibrium Computation in Markets

Arxiv

0+阅读 · 2023年3月16日

Time-marching based quantum solvers for time-dependent linear differential equations

Arxiv

0+阅读 · 2023年3月15日

Information-Theoretic Regret Bounds for Bandits with Fixed Expert Advice

Arxiv

0+阅读 · 2023年3月15日

Betty: An Automatic Differentiation Library for Multilevel Optimization

Arxiv

0+阅读 · 2023年3月15日

Policy learning "without'' overlap: Pessimism and generalized empirical Bernstein's inequality

Arxiv

0+阅读 · 2023年3月15日

Adaptive Testing for High-dimensional Data

Arxiv

0+阅读 · 2023年3月14日

Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy Optimization

Arxiv

12+阅读 · 2021年12月20日

VIP会员

文章信息

相关主题

样本复杂度

相关VIP内容

【干货书】数据分析优化，Optimization for Modern Data Analysis，117页pdf

【干货书】数据分析优化，Optimization for Modern Data Analysis，117页pdf

专知会员服务

60+阅读 · 2023年2月15日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

59+阅读 · 2022年4月22日

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

【超赞的#C++#速查&信息图】“hacking c++ - Cheat Sheets & Infographics”

专知会员服务

29+阅读 · 2022年3月8日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

122+阅读 · 2020年11月20日

UC.Berkeley CS189讲义教材:《机器学习全面指南》，185页pdf

专知会员服务

158+阅读 · 2020年1月16日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

45+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

31+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

53+阅读 · 2019年10月17日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

90+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

64+阅读 · 2019年10月9日

热门VIP内容

相关资讯

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

23+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

26+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

17+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

41+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

16+阅读 · 2018年12月24日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

相关论文

Dynamic Update-to-Data Ratio: Minimizing World Model Overfitting

Arxiv

0+阅读 · 2023年3月17日

A Policy Iteration Approach for Flock Motion Control

Arxiv

0+阅读 · 2023年3月17日

A New Policy Iteration Algorithm For Reinforcement Learning in Zero-Sum Markov Games

Arxiv

0+阅读 · 2023年3月17日

Enabling First-Order Gradient-Based Learning for Equilibrium Computation in Markets

Enabling First-Order Gradient-Based Learning for Equilibrium Computation in Markets

Arxiv

0+阅读 · 2023年3月16日

Time-marching based quantum solvers for time-dependent linear differential equations

Arxiv

0+阅读 · 2023年3月15日

Information-Theoretic Regret Bounds for Bandits with Fixed Expert Advice

Arxiv

0+阅读 · 2023年3月15日

Betty: An Automatic Differentiation Library for Multilevel Optimization

Arxiv

0+阅读 · 2023年3月15日

Policy learning "without'' overlap: Pessimism and generalized empirical Bernstein's inequality

Arxiv

0+阅读 · 2023年3月15日

Adaptive Testing for High-dimensional Data

Arxiv

0+阅读 · 2023年3月14日

Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy Optimization

Arxiv

12+阅读 · 2021年12月20日

相关基金

Yang-Baxter矩阵方程解的研究与应用

国家自然科学基金

0+阅读 · 2015年12月31日

TMS1基因响应高温胁迫和ER Stress的分子机制

国家自然科学基金

0+阅读 · 2014年12月31日

S3AGA样本（Spitzer-SDSS Spectral Atlas of Galaxies and AGNs)及其AGN研究

国家自然科学基金

0+阅读 · 2014年12月31日

蛋白酶体抑制剂Bortezomib对肺动脉平滑肌细胞钙离子通路的作用研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于Fermi-LAT和AMS-02的暗物质理论研究

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

新型非甾体类AR拮抗剂的设计合成及生物活性评价

国家自然科学基金

0+阅读 · 2012年12月31日

猪带绦虫胰岛素与受体的相互作用及其对葡萄糖代谢的调节

国家自然科学基金

0+阅读 · 2012年12月31日

干旱、寒冷灌区土壤水盐耦合遥感监测敏感性研究

国家自然科学基金

0+阅读 · 2011年12月31日

模-相对Hochschild同调与上同调

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员