08 2021 档案

摘要:Value-Based Reinforcement Learning 一、Deep Q-Network (DQN) 本质就是用神经网络近似$Q^*$函数,将 \(Q^{*}(s_t,a_t)\) 当作是一个先知,先知可以告诉你每个动作带来的平均回报,我们就应该听先知的话选平均回报最高的动作 Goal 阅读全文
posted @ 2021-08-29 00:11 TR_Goldfish 阅读(108) 评论(0) 推荐(0)
摘要:Monte Carlo Algorithms 一大类随机算法,根据随机样本来估计真实值 一、Integration We are given a function, e.g: Calculate the integral: If \(f(x)\) is very involved, there is 阅读全文
posted @ 2021-08-28 00:41 TR_Goldfish 阅读(49) 评论(0) 推荐(0)
摘要:Reinforcement Learning Basics 课程地址:https://www.bilibili.com/video/BV1rv41167yx 一、probability theory 1.1 Random Variable Random variable: unknown; its 阅读全文
posted @ 2021-08-26 22:28 TR_Goldfish 阅读(188) 评论(0) 推荐(0)
摘要:https://www.zhihu.com/question/49136398/answer/1654722335 阅读全文
posted @ 2021-08-20 11:56 TR_Goldfish 阅读(29) 评论(0) 推荐(0)
摘要:一、Markov Decision Process 马尔可夫决策过程(Markov Decision Process),即马尔可夫奖励过程的基础上加上action,即:Markov Chain + Reward + action。如果还用刚才的股票为例子的话,我们只能每天看到股票价格的上涨或者下降, 阅读全文
posted @ 2021-08-19 10:55 TR_Goldfish 阅读(219) 评论(0) 推荐(0)
摘要:1. Properties of Reinforcement Learning Reward delay reward角度 Agent’s actions affect the subsequent data it receives 环境角度 2. Policy-based Approach:Lea 阅读全文
posted @ 2021-08-18 22:13 TR_Goldfish 阅读(587) 评论(0) 推荐(0)
摘要:https://blog.csdn.net/qq_43062920/article/details/104436779?utm_medium=distribute.pc_relevant.none-task-blog-2%7Edefault%7EsearchFromBaidu%7Edefault-3 阅读全文
posted @ 2021-08-05 05:14 TR_Goldfish 阅读(489) 评论(0) 推荐(0)