Foundations of Deep Reinforcement Learning: Theory and Practice in PythonAddison-Wesley Professional, 2019 M11 20 - 416 pages The Contemporary Introduction to Deep Reinforcement Learning that Combines Theory and Practice Deep reinforcement learning (deep RL) combines deep learning and reinforcement learning, in which artificial agents learn to solve sequential decision-making problems. In the past decade deep RL has achieved remarkable results on a range of problems, from single and multiplayer games—such as Go, Atari games, and DotA 2—to robotics. Foundations of Deep Reinforcement Learning is an introduction to deep RL that uniquely combines both theory and implementation. It starts with intuition, then carefully explains the theory of deep RL algorithms, discusses implementations in its companion software library SLM Lab, and finishes with the practical details of getting deep RL to work. This guide is ideal for both computer science students and software engineers who are familiar with basic machine learning concepts and have a working understanding of Python.
|
Contents
Deep QNetworks DQN | 3-39 |
Combined Methods | 6-1 |
PolicyBased and ValueBased Algorithms | 6-2 |
Proximal Policy Optimization PPO | 18 |
Parallelization Methods | 34 |
Algorithm Summary | 47 |
Environment Design | 33 |
Actions | 78 |
SARSA | 3 |
Improving | 5 |
Network Architectures | 12 |
Epilogue | 19 |
B Example Environments | 25 |
References | 34 |
SLM | 41 |
Hardware | 49 |
Index | 40 |
Other editions - View all
Foundations of Deep Reinforcement Learning: Theory and Practice in Python Laura Graesser,Wah Loon Keng No preview available - 2020 |
Deep Reinforcement Learning in Python: A Hands-On Introduction Laura Harding Graesser,Keng Wah Loon No preview available - 2020 |
Common terms and phrases
a₁ action space actor Advantage Estimation advantage function algorithm Atari games Atari Pong batch calculate CartPole Chapter Click CNNs complex convolutional debugging Deep Reinforcement Learning deep RL algorithms discrete Double DQN DQN algorithm elements entropy environment episode example Figure frame rate frame skipping global network graphs grayscale hyperparameters implementation input layers learning rate max Q MDPs method MLPs n-step returns network architecture network parameters neural network numpy on-policy OpenAI output parameter updates performance pixels policy gradient policy loss POMDPs preprocessing priorities problem Q-function Q-learning Q-values Qtar r₁ reinforcement learning replay memory reward RNNs s₁ sampling SARSA Section session shown in Code shown in Equation SLM Lab spec file surrogate objective target network tensor training step trajectories trial unit tests view code image π π πο
