首页 > 其他 > 详细

CS294-112深度增强学习课程（加州大学伯克利分校 2017）NO.4 Learning policies by imitating optimal controllers

时间：2018-05-23 19:29:35 阅读：207 评论：0 收藏：0 [点我收藏+]

技术分享图片

技术分享图片

There are some problems: mismatch of model and reality; gradient explosion

技术分享图片

技术分享图片

技术分享图片

so, the dynamics can be quite messy, and backpropogating can be quite problematic.

sudden change in velocity and so on. schochastic system. gradient descent can be tough.

技术分享图片

can we apply this trajectory optimization method to optimize policy?

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

GPS: guided policy search

技术分享图片

in this case, o_t is from the camera and the joint velocity

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

技术分享图片

CS294-112深度增强学习课程（加州大学伯克利分校 2017）NO.4 Learning policies by imitating optimal controllers

原文：https://www.cnblogs.com/ecoflex/p/9078801.html

踩

(0)

赞

(0)

举报

评论一句话评论（0）

分享档案

更多>

2021年09月23日 (328)
2021年09月24日 (313)
2021年09月17日 (191)
2021年09月15日 (369)
2021年09月16日 (411)
2021年09月13日 (439)
2021年09月11日 (398)
2021年09月12日 (393)
2021年09月10日 (160)
2021年09月08日 (222)

最新文章

更多>

教程昨日排行

更多>

友情链接

汇智网 PHP教程插件网

关于我们 - 联系我们 - 留言反馈 - 联系我们:wmxa8@hotmail.com

© 2014 bubuko.com 版权所有

打开技术之扣，分享程序人生！