Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)

Due to the safety risks and training sample inefficiency, it is often preferred to develop controllers in simulation. However, minor differences between the simulation and the real world can cause a significant sim-to-real gap. This gap can reduce the effectiveness of the developed controller. In this paper, we examine a case study of transferring an octorotor reinforcement learning controller from simulation to the real world. First, we quantify the effectiveness of the real-world transfer by examining safety metrics. We find that although there is a noticeable (around 100%) increase in deviation in real flights, this deviation may not be considered unsafe, as it will be within > 2m safety corridors. Then, we estimate the densities of the measurement distributions and compare the Jensen-Shannon divergences of simulated and real measurements. From this, we show that the vehicle’s orientation is significantly different between simulated and real flights. We attribute this to a different flight mode in real flights where the vehicle turns to face the next waypoint. We also find that the reinforcement learning controller actions appear to correctly counteract disturbance forces. Then, we analyze the errors of a measurement autoencoder and state transition model neural network applied to real data. We find that these models further reinforce the difference between the simulated and real attitude control, showing the errors directly on the flight paths. Finally, we discuss important lessons learned in the sim-to-real transfer of our controller.

Approximately OptimalApproximate…Approximately Optimal Approximate Reinforcement LearningMuJoCo: A physics enginefor model-based controlMuJoCo: A physics engine for model-based controlAdam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationHuman-level controlthrough deep…Human-level control through deep reinforcement learningTrust Region PolicyOptimizationTrust Region Policy OptimizationOpenAI GymOpenAI GymAsynchronous Methods forDeep Reinforcement…Asynchronous Methods for Deep Reinforcement LearningHigh-DimensionalContinuous Control Usin…High-Dimensional Continuous Control Using Generalized Advantage EstimationBenchmarking DeepReinforcement Learning…Benchmarking Deep Reinforcement Learning for Continuous ControlSample EfficientActor-Critic with…Sample Efficient Actor-Critic with Experience ReplayEmergence of LocomotionBehaviours in Rich…Emergence of Locomotion Behaviours in Rich EnvironmentsLIFT: ReinforcementLearning in Computer…LIFT: Reinforcement Learning in Computer Systems by Learning From DemonstrationsComputation Offloadingin Multi-Access Edge…Computation Offloading in Multi-Access Edge Computing Using a Deep Sequential Model Based on Reinforcement LearningModel-based LookaheadReinforcement LearningModel-based Lookahead Reinforcement LearningFinRL: A DeepReinforcement Learning…FinRL: A Deep Reinforcement Learning Library for Automated Stock Trading in Quantitative FinanceA Hybrid Learning Methodfor System…A Hybrid Learning Method for System Identification and Optimal ControlReconfigurableIntelligent Surfaces…Reconfigurable Intelligent Surfaces: Principles and OpportunitiesReal-Time Scheduling forDynamic Partial-No-Wait…Real-Time Scheduling for Dynamic Partial-No-Wait Multiobjective Flexible Job Shop by Deep Reinforcement LearningCoordinate-wise ControlVariates for Deep Polic…Coordinate-wise Control Variates for Deep Policy GradientsColO-RAN: DevelopingMachine Learning-Based…ColO-RAN: Developing Machine Learning-Based xApps for Open RAN Closed-Loop Control on Programmable Experimental PlatformsLook where you look!Saliency-guided…Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement LearningJaxMARL: Multi-Agent RLEnvironments and…JaxMARL: Multi-Agent RL Environments and Algorithms in JAXContrastive PreferenceOptimization: Pushing…Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationDiagnosingNon-Intermittent…Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。