Search references for DEEP REINFORCEMENT-LEARNING. Phrases containing DEEP REINFORCEMENT-LEARNING
See searches and references containing DEEP REINFORCEMENT-LEARNING!DEEP REINFORCEMENT-LEARNING
Machine learning that combines deep learning and reinforcement learning
Deep reinforcement learning (deep RL) is a subfield of machine learning that combines reinforcement learning (RL) and deep learning. RL considers the
Deep_reinforcement_learning
Field of machine learning
In machine learning and optimal control, reinforcement learning (RL) is concerned with how an intelligent agent should take actions in a dynamic environment
Reinforcement_learning
Sub-field of reinforcement learning
Multi-agent reinforcement learning (MARL) is a sub-field of reinforcement learning. It focuses on studying the behavior of multiple learning agents that
Multi-agent reinforcement learning
Multi-agent_reinforcement_learning
Model-free reinforcement learning algorithm
Q-learning is a reinforcement learning algorithm that trains an agent to assign values to its possible actions based on its current state, without requiring
Q-learning
Machine learning technique
In machine learning, reinforcement learning from human feedback (RLHF) is a technique to align an intelligent agent with human preferences. It involves
Reinforcement learning from human feedback
Reinforcement_learning_from_human_feedback
losing. Reinforcement learning is used heavily in the field of machine learning and can be seen in methods such as Q-learning, policy search, Deep Q-networks
Machine learning in video games
Machine_learning_in_video_games
AI research laboratory
(Japanese chess) after a few days of play against itself using reinforcement learning. DeepMind has since trained models for game-playing (MuZero, AlphaStar)
Google_DeepMind
Computer scientist and researcher
London. From 2013 to 2026, Silver worked full-time at DeepMind, leading reinforcement learning research. He notably led the development of AlphaGo and
David Silver (computer scientist)
David_Silver_(computer_scientist)
Model-free reinforcement learning algorithm
is a reinforcement learning (RL) algorithm for training an intelligent agent. Specifically, it is a policy gradient method, often used for deep RL when
Proximal_policy_optimization
Class of reinforcement learning algorithm
In reinforcement learning (RL), a model-free algorithm is an algorithm which does not estimate the transition probability distribution (and the reward
Model-free (reinforcement learning)
Model-free_(reinforcement_learning)
Machine learning technique where agents learn from demonstrations
Imitation learning is a paradigm in reinforcement learning, where an agent learns to perform a task by supervised learning from expert demonstrations
Imitation_learning
Machine learning researcher at Berkeley
his cutting-edge research in robotics and machine learning, particularly in deep reinforcement learning. In 2021, he joined AIX Ventures as an Investment
Pieter_Abbeel
Ability of artificial intelligence to play different games
Starting in 2013, significant progress was made following the deep reinforcement learning approach, including the development of programs that can learn
General_game_playing
American computer scientist and academic
worked on robot learning algorithms from deep predictive models. She delivered a massive open online course on deep reinforcement learning. She was the first
Chelsea_Finn
Deep reinforcement learning method
AlphaChip is a deep reinforcement learning method for automated chip floorplanning. It was developed at Google and is now a portion of the offerings of
AlphaChip
Type of feedforward neural network
predictions. A deep Q-network (DQN) is a type of deep learning model that combines a deep neural network with Q-learning, a form of reinforcement learning. Unlike
Convolutional_neural_network
Artificial intelligence that plays Go
Furthermore, AlphaGo Zero performed better than standard deep reinforcement learning models (such as Deep Q-Network implementations) due to its integration of
AlphaGo_Zero
Technique in machine learning
Jian; Han, Jiawei (2018). Curriculum learning for heterogeneous star network embedding via deep reinforcement learning. pp. 468–476. doi:10.1145/3159652
Curriculum_learning
American AI safety researcher
co-authored the paper "Deep Reinforcement Learning from Human Preferences" (2017) and other works developing reinforcement learning from human feedback (RLHF)
Paul_Christiano
Reinforcement learning algorithms
The actor-critic algorithm (AC) is a family of reinforcement learning (RL) algorithms that combine policy-based RL algorithms such as policy gradient methods
Actor-critic_algorithm
Research field that lies at the intersection of machine learning and computer security
resembles Ridge regression. Adversarial deep reinforcement learning is an active area of research in reinforcement learning focusing on vulnerabilities of learned
Adversarial_machine_learning
Computer scientist
co‑authored Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels (Yarats, Kostrikov & Fergus, ICLR 2021), which introduced
Denis_Yarats
Computer scientist and professor
trains deep neural networks to execute complex robotic tasks. He contributed to end-to-end visuomotor policy learning, model-based reinforcement learning for
Sergey_Levine
and tools used for machine learning, deep learning, natural language processing, computer vision, reinforcement learning, artificial general intelligence
Lists of open-source artificial intelligence software
Lists_of_open-source_artificial_intelligence_software
Intelligence of machines
four of the world's best Gran Turismo drivers using deep reinforcement learning. In 2024, Google DeepMind introduced SIMA, a type of AI capable of autonomously
Artificial_intelligence
Subset of artificial intelligence
in the field of deep learning have allowed neural networks, a class of statistical algorithms, to surpass many previous machine learning approaches in performance
Machine_learning
Branch of machine learning
In machine learning, deep learning (DL) focuses on utilizing multilayered neural networks to perform tasks such as classification, regression, and representation
Deep_learning
Computational model used in machine learning
Alternative to Reinforcement Learning". arXiv:1703.03864 [stat.ML]. Such FP, Madhavan V, Conti E, Lehman J, Stanley KO, Clune J (20 April 2018). "Deep Neuroevolution:
Neural network (machine learning)
Neural_network_(machine_learning)
Canadian researcher
for exploiting deep neural networks in the context of reinforcement learning, and new recurrent memory architectures for one-shot learning. His numerous
Timothy_Lillicrap
computer-generated music for stress and pain relief. The Watson Beat uses reinforcement learning and deep belief networks to compose music on a simple seed input melody
Applications of artificial intelligence
Applications_of_artificial_intelligence
Electricity generation by nuclear fusion
address fusion heating, measurement, and power production. A deep reinforcement learning system has been used to control a tokamak-based reactor. The
Fusion_power
Blueprint for intelligent agents
Wierstra, Daan; Riedmiller, Martin (2013). "Playing Atari with Deep Reinforcement Learning". arXiv:1312.5602 [cs.LG]. Mnih, Volodymyr; Kavukcuoglu, Koray;
Cognitive_architecture
Algorithm for modelling sequential data
In deep learning, the transformer is a family of artificial neural network architectures based on the multi-head attention mechanism, in which input data
Transformer_(deep_learning)
Concept in artificial intelligence
or the robot. In 2017, OpenAI and DeepMind applied deep learning to the cooperative inverse reinforcement learning in simple domains such as Atari games
Apprenticeship_learning
American scholar of computer science
peer-to-peer networks, Internet privacy, social networks, and deep reinforcement learning. He is the Dean of Engineering and Computer Science at NYU Shanghai
Keith_W._Ross
Subfield of machine learning
classification benchmarks and to policy-gradient-based reinforcement learning. Variational Bayes-Adaptive Deep RL (VariBAD) was introduced in 2019. While MAML
Meta-learning (computer science)
Meta-learning_(computer_science)
Artificial intelligence field of study
in Deep Reinforcement Learning". Proceedings of the 39th International Conference on Machine Learning. International Conference on Machine Learning. PMLR
AI_safety
Concept in decision-making
context of machine learning, the exploration–exploitation tradeoff is fundamental in reinforcement learning (RL), a type of machine learning that involves
Exploration–exploitation dilemma
Exploration–exploitation_dilemma
breakthroughs, most notably through London-based DeepMind's achievements in deep reinforcement learning and protein structure prediction. The UK AI sector
Artificial intelligence industry in the United Kingdom
Artificial_intelligence_industry_in_the_United_Kingdom
Robot that performs behaviors or tasks with a high degree of autonomy
Al; Najjaran, Homayoun (2024-02-08). "Learning team-based navigation: a review of deep reinforcement learning techniques for multi-agent pathfinding"
Autonomous_robot
Deep learning artificial intelligence research team
Google Brain was a deep learning artificial intelligence research team that served as the sole AI branch of Google before being incorporated under the
Google_Brain
British AI researcher (born 1976)
made significant advances in deep learning and reinforcement learning, and pioneered the field of deep reinforcement learning which combines these two methods
Demis_Hassabis
AI model that developer a super-human sorting algorithm
intelligence system developed by Google DeepMind to discover enhanced computer science algorithms using reinforcement learning. AlphaDev is based on AlphaZero
AlphaDev
Language models designed for reasoning tasks
LLMs via Reinforcement Learning". arXiv:2501.12948 [cs.CL]. DeepSeek 支持"深度思考+联网检索"能力 [DeepSeek adds a search feature supporting simultaneous deep thinking
Reasoning_model
Jewish American scientist
focusing on deep reinforcement learning techniques with neural networks in virtual environments. These networks underwent trial-and-error learning in VR before
Vladlen_Koltun
Artificial intelligence concept
hacking or specification gaming occurs when an AI trained with reinforcement learning optimizes an objective function—achieving the literal, formal specification
Reward_hacking
Conformance of AI to intended objectives
in Deep Reinforcement Learning". Proceedings of the 39th International Conference on Machine Learning. International Conference on Machine Learning. PMLR
AI_alignment
Machine learning strategy
Mainini, https://arxiv.org/abs/2303.01560v2 Learning how to Active Learn: A Deep Reinforcement Learning Approach, Meng Fang, Yuan Li, Trevor Cohn, https://arxiv
Active learning (machine learning)
Active_learning_(machine_learning)
Computer programming concept
Temporal difference (TD) learning refers to a class of model-free reinforcement learning methods which learn by bootstrapping from the current estimate
Temporal_difference_learning
Method of executing orders
pivotal shift in algorithmic trading as machine learning was adopted. Specifically deep reinforcement learning (DRL) which allows systems to dynamically adapt
Algorithmic_trading
Series of actions for bettering effective usage
distillation performance by combining artificial intelligence, deep reinforcement learning and real-time crude oil analysis.. Changes in crude oil properties
Process_optimization
Use of artificial intelligence in the automation of electronic design
from Google researchers between 2020 and 2021. They created a deep reinforcement learning method for planning the layout of a chip, known as floorplanning
AI-driven_design_automation
Decentralized machine learning
Guo, Weisi; Nallanathan, Arumugam; Wu, Qihui (2021). "Green Deep Reinforcement Learning for Radio Resource Management: Architecture, Algorithm Compression
Federated_learning
Interdisciplinary research area
Xiaoli; Goan, Hsi-Sheng (2020). "Variational Quantum Circuits for Deep Reinforcement Learning". IEEE Access. 8: 141007–141024. arXiv:1907.00397. Bibcode:2020IEEEA
Quantum_machine_learning
Artificial production of media by automated means
social media platforms through tactics such as astroturfing. Deep reinforcement learning-based natural-language generators could potentially be used to
Synthetic_media
on reinforcement learning, marked by breakthroughs such as generative AI models from Krutrim, Sarvam, CoRover, OpenAI and Alphafold by Google DeepMind
Artificial intelligence in India
Artificial_intelligence_in_India
Overview of and topical guide to machine learning
unlabeled data Reinforcement learning, where the model learns to make decisions by receiving rewards or penalties. Applications of machine learning Bioinformatics
Outline_of_machine_learning
PMC 346238. PMID 6953413. Bozinovski, S. (1982). "A self-learning system using secondary reinforcement". In Trappl, Robert (ed.). Cybernetics and Systems Research:
Timeline_of_machine_learning
skilled human dogfighter. Heron Systems corporation wrote a deep reinforcement learning software tool that bested the human pilot by a score of 5–0.
DARPA_AlphaDogfight
Chinese artificial intelligence company
optimization (DPO). DPO is also known as reinforcement learning from human feedback.[original research?] DeepSeek-MoE models (Base and Chat), each have
DeepSeek
Artificial intelligence company
floorplanning, placement, routing, and verification. AlphaChip is a deep reinforcement learning method for automated chip floorplanning. The basic ideas were
Ricursive
Machine learning paradigm
of fully self-contained autoencoder training. In reinforcement learning, self-supervising learning from a combination of losses can create abstract representations
Self-supervised_learning
Aircraft without any human pilot on board
operating system such as Linux with relaxed time constraints. Deep reinforcement learning has also been investigated for UAV flight control, particularly
Unmanned_aerial_vehicle
Deep learning method
unsupervised learning, GANs have also proved useful for semi-supervised learning, fully supervised learning, and reinforcement learning. The core idea
Generative adversarial network
Generative_adversarial_network
Class of reinforcement learning algorithms
Policy gradient methods are a class of reinforcement learning algorithms and a sub-class of policy optimization methods. Unlike value-based methods which
Policy_gradient_method
Coordination of multiple robots as a system
Multi-Robot Autonomous Exploration in Unknown Environments via Deep Reinforcement Learning" IEEE Transactions on Vehicular Technology, 2020. Hu, J.; Turgut
Swarm_robotics
Canadian technology company
generation. Maluuba published a research paper learning dialogue policies with deep reinforcement learning. In 2016, Maluuba also freely released the Frames
Maluuba
Computer backgammon program (1992)
as an early success of reinforcement learning and neural networks, and was cited in, for example, papers for deep Q-learning and AlphaGo. During play
TD-Gammon
Initial estimate or framework to the solution of a mathematical problem
; Prati, E. (2019). "Coherent transport of quantum states by deep reinforcement learning". Communications Physics. 2 (1): 61. arXiv:1901.06603. Bibcode:2019CmPhy
Ansatz
Suite of reinforcement learning algorithms
Critic (DSAC) is a suite of model-free off-policy reinforcement learning algorithms, tailored for learning decision-making or control policies in complex
Distributional Soft Actor Critic
Distributional_Soft_Actor_Critic
Academic conference in machine learning
machine learning conferences NeurIPS and ICLR, ICML traditionally features more content on statistical learning theory, reinforcement learning and robotics
International Conference on Machine Learning
International_Conference_on_Machine_Learning
Software program
Neural Networks Through Deep Visualization. Deep Learning Workshop, International Conference on Machine Learning (ICML) Deep Learning Workshop. arXiv:1506
DeepDream
Computer scientist
establish the Reinforcement Learning and Artificial Intelligence Laboratory. In 2017 he became a distinguished research scientist with Google DeepMind and helped
Richard_S._Sutton
include neural networks in their evaluation function. Yet the deep reinforcement learning used for AlphaZero remains uncommon in top engines. Computer
History_of_chess_engines
Machine learning technique
In deep learning, fine-tuning is the process of adapting a computational model trained for one task (the upstream task) to perform a different, usually
Fine-tuning_(deep_learning)
Hyperparameter optimization framework
machine learning models. It was first introduced in 2018 by Preferred Networks, a Japanese startup that works on practical applications of deep learning in
Optuna
Problem setup in machine learning
Zero-shot learning (ZSL) is a problem setup in machine learning where, at test time, a learner observes samples from classes which were not observed during
Zero-shot_learning
Computer scientist
February 2022). "Magnetic control of tokamak plasmas through deep reinforcement learning". Nature. 602 (7897): 414–419. Bibcode:2022Natur.602..414D. doi:10
Pushmeet_Kohli
AI that generates content
et al. (June 2023). "Faster sorting algorithms discovered using deep reinforcement learning". Nature. 618 (7964): 257–263. Bibcode:2023Natur.618..257M. doi:10
Generative_AI
Cross-platform video game and simulation engine
researchers in the field of deep reinforcement learning to train agents inside Unity-created environments. Unity Machine Learning Agents can act as virtual
Unity_(game_engine)
Artificial intelligence system for discovering matrix multiplication algorithms
intelligence system developed by DeepMind for discovering efficient matrix multiplication algorithms using reinforcement learning. Introduced in 2022, the system
AlphaTensor
Capital budgeting analysis term
data-driven Markov decision process, and uses advanced machine learning like deep reinforcement learning to evaluate a wide range of possible real option and design
Real_options_valuation
Type of large language model
and audio. Additionally, GPT models like o3 and DeepSeek R1 have been trained with reinforcement learning to generate multi-step chain-of-thought reasoning
Generative pre-trained transformer
Generative_pre-trained_transformer
Machine learning technique
"Self-organizing maps for storage and transfer of knowledge in reinforcement learning". Adaptive Behavior. 27 (2): 111–126. arXiv:1811.08318. doi:10
Transfer_learning
Physics engine
Jorge Pena; Westerlund, Tomi (2020). "Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: A Survey". 2020 IEEE Symposium Series on Computational
MuJoCo
Algorithms in constraint satisfaction
David; Kavukcuoglu, Koray (2016-06-16). "Asynchronous Methods for Deep Reinforcement Learning". arXiv:1602.01783v2 [cs.LG]. A.K. Mackworth. Consistency in
AC-3_algorithm
Computer program for poker
Bakhtin, Anton; Lerer, Adam; Gong, Qucheng (2020). "Combining deep reinforcement learning and search for imperfect-information games". Advances in Neural
DeepStack
Sequence of operations for a task
et al. (June 2023). "Faster sorting algorithms discovered using deep reinforcement learning". Nature. 618 (7964): 257–263. doi:10.1038/s41586-023-06004-9
Algorithm
Computer hardware and software capable of playing chess
some engines use deep neural networks in their evaluation function. Neural networks are usually trained using some reinforcement learning algorithm, in conjunction
Computer_chess
Research Institute in Chennai, Tamil Nadu, India
one of the country's largest groups in network analytics and deep reinforcement learning. Google has granted IIT Madras $1 million for setting up India's
IIT_Madras
Artificial intelligence control techniques
supposed to capture the dynamics of a system. For the control part, deep reinforcement learning has shown its ability to control complex systems. Bayesian probability
Intelligent_control
Local boundary-limited electrical grid
Fonteneau, Raphael. Deep reinforcement learning solutions for energy microgrids management. European Workshop on Reinforcement Learning (EWRL 2016). hdl:2268/203831
Microgrid
1984 video game
Petersen, Stig (February 2015). "Human-level control through deep reinforcement learning". Nature. 518 (7540): 529–533. Bibcode:2015Natur.518..529M. doi:10
Montezuma's Revenge (video game)
Montezuma's_Revenge_(video_game)
Fictional alien race
Pieter (2018). "Modular Architecture for StarCraft II with Deep Reinforcement Learning". arXiv:1811.03555 [cs.AI]. Han, Lei; Xiong, Jiechao; Sun, Peng;
Zerg
Fictional character from My Little Pony
"Rainbow Dash" that successfully taught itself to walk using deep reinforcement learning. The robot demonstrated the ability to learn to walk backward
Rainbow_Dash
Reinforcement learning technique
reinforcement learning agents. Intuitively, agents learn to improve their performance by playing "against themselves". In multi-agent reinforcement learning
Self-play
Ability of some flying animals
2023). "Exploring storm petrel pattering and sea-anchoring using deep reinforcement learning". Bioinspiration & Biomimetics. 18 (6). University of Portland
Hover_(behaviour)
British recruitment business
Conference". Retrieved 27 November 2019. "Step into the AI Era: Deep Reinforcement Learning Workshop". Retrieved 27 November 2019. "UX Sessions". Retrieved
InterQuest_Group_Ltd
Chinese tile-based game
Hsiao-Wuen (31 March 2020). "Suphx: Mastering Mahjong with Deep Reinforcement Learning". arXiv:2003.13590 [cs.AI]. "新年打麻雀4大風水秘訣". Lam, Desmond. "Chinese
Mahjong
Technology made by American organization
included many projects focused on reinforcement learning (RL). OpenAI has been viewed as an important competitor to DeepMind. Announced in 2016, Gym was
Products and applications of OpenAI
Products_and_applications_of_OpenAI
DEEP REINFORCEMENT-LEARNING
DEEP REINFORCEMENT-LEARNING
DEEP REINFORCEMENT-LEARNING
DEEP REINFORCEMENT-LEARNING
DEEP REINFORCEMENT-LEARNING
DEEP REINFORCEMENT-LEARNING
DEEP REINFORCEMENT-LEARNING
DEEP REINFORCEMENT-LEARNING
DEEP REINFORCEMENT-LEARNING