IEEE Internet of Things Journal · 2018 · 319 citations · 36 references
EngineeringEducationReinforcement Learning (Educational Psychology)Cloud-based InternetJoint OptimizationService LatencyReinforcement Learning (Computer Engineering)Fog ComputingInternet Of ThingsMobile Data OffloadingComputer EngineeringRadio ResourcesComputer ScienceMobile ComputingDeep Neural NetworkDeep Reinforcement LearningEdge ComputingCloud ComputingMulti-access Edge Computing
Fog computing and caching promise lower latency and reduced backhaul traffic for rapidly growing cloud‑based IoT, yet their performance hinges on intelligent coordination of caching, computing, and communication resources. This work proposes a joint optimization framework that simultaneously determines content caching, computation offloading, and radio resource allocation for fog‑enabled IoT. An actor‑critic deep reinforcement learning algorithm with dual neural networks and a natural policy gradient is employed to solve the joint decision problem and minimize average end‑to‑end delay. Simulation results confirm the algorithm’s learning capability and demonstrate reduced end‑to‑end service latency.
The cloud-based Internet of Things (IoT) develops rapidly but suffer from large latency and backhaul bandwidth requirement, the technology of fog computing and caching has emerged as a promising paradigm for IoT to provide proximity services, and thus reduce service latency and save backhaul bandwidth. However, the performance of the fog-enabled IoT depends on the intelligent and efficient management of various network resources, and consequently the synergy of caching, computing, and communications becomes the big challenge. This paper simultaneously tackles the issues of content caching strategy, computation offloading policy, and radio resource allocation, and propose a joint optimization solution for the fog-enabled IoT. Since wireless signals and service requests have stochastic properties, we use the actor-critic reinforcement learning framework to solve the joint decision-making problem with the objective of minimizing the average end-to-end delay. The deep neural network (DNN) is employed as the function approximator to estimate the value functions in the critic part due to the extremely large state and action space in our problem. The actor part uses another DNN to represent a parameterized stochastic policy and improves the policy with the help of the critic. Furthermore, the Natural policy gradient method is used to avoid converging to the local maximum. Using the numerical simulations, we demonstrate the learning capacity of the proposed algorithm and analyze the end-to-end service latency.
36
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al. · Nature · 2015 · 28.8K citations
Artificial Intelligence, Engineering, Deep Reinforcement Learning +3
Continuous control with deep reinforcement learning
Timothy Lillicrap, Jonathan J. Hunt, Alexander Pritzel et al. · arXiv (Cornell University) · 2016 · 6.8K citations · Full text
Natural Gradient Works Efficiently in Learning
Neural Computation · 1998 · 2.7K citations