Jose Lavariega, advised by Chenhao Li & Victor Klemm, Robotics Systems Lab, ETH Zürich
Abstract:
In this work, we learn an Uncertainty-Aware Model of Forward Dynamics for a walking quadruped. Current methods for learning forward dynamics rely on deterministic models, which have the downside of limited interpretability when presented with observations outside of the training distribution. By learning an uncertainty-aware model, we can calculate and use the epistemic uncertainty of the model as an indicator to determine out-of-distribution (OOD) observations. We show that epistemic uncertainty is a reliable metric for distinguishing between additive noises to the observations, as well as for quantifying when a model is given observations outside of the training regime. An example of the latter is shown with proprioceptive data from quadruped locomotion on flat terrains versus locomotion on rugged terrain. Our approach is implemented in three network architectures and tested through three network training paradigms. The model dynamics accuracy in In-Distribution and OOD data is evaluated on observations taken from an ANYmal-D platform in locomotion. Finally, we set the foundations for extension into model-based RL to aid exploration of the state-action space the model learns.
In-Distribution Training Environment
Out of Distribution Deployment Environment
How can we strengthen robot performance in the presence of Out of Distribution (OOD) Observations?
In real-world deployments and in the field, we often encounter OOD observations across terrains, visual features or planning scenarios. When a learned model-free policy encounters such a situation, we have no notion of distinguishing an incorrect
We prepare for such scenarios by creating an uncertainty-aware world model, which allows us to explicitly model how reliable predictions are through the epistemic uncertainty (from limited data).
Since we model the next motion step and the reliability, we augment the autoregressive error and identify regions where predictions become unreliable due to OOD observations. We can then use the epistemic uncertainty to switch to a safe locomotion policy, guide policy training to diversify the dataset, and prepare robots to enter an uncertain field deployment.
Method
Estimating Uncertainty
We estimate uncertainty through an ensemble network. Each ensemble member takes predictions from a base set of layers to predict a Gaussian distribution over subsequent observations. We then obtain the aleatoric and epistemic uncertainties directly: the variance within each ensemble member prediction captures aleatoric uncertainty, or the uncertainty arising from statistical variability or noise.
The variance across the ensemble means captures the epistemic uncertainty, or uncertainty that arises from lack or biased data. In other words, a high epistemic uncertainty indicates an OOD observation. For any environment, epistemic uncertainty decreases as more data is observed.
Autoregressive Error
We utilize an autoregressive error to capture accuracy over long horizon prediction rollouts. Model-based RL approaches avoid hallucinated dynamics by limiting rollouts to short horizons, requiring accuracy over single-step predictions. However real locomotion is not representative of short-horizon scenarios, and uncertainties and error accrue over long horizons. Hence, we use a long-horizon autoregressive error, and experiment on base layer architectures, to capture the dynamics of locomotion seconds into the future, instead of milliseconds. With our epistemic estimate, we track where predictions degrade as rollouts project over hundreds of steps.
Base Layers
Our ensemble heads are composed of a shallow MLPs that take in the Latent Space output from the Base Layers. We experiment with three types, each assuming inherent structure to be captured by the uncertainty-aware dynamics world model:
1) MLP (no assumed structure)
2) Gated Recurrent Units (consequentiality of recent dynamics on future dynamics)
3) Fourier Latent Dynamics (periodicity and consequentiality)
We evaluate three training methods to assess how best to design a training pipeline for the uncertainty aware world model:
In Offline training we train the dynamics network through autoregressive loss on the captured trajectories only.
Captured trajectories are obtained from a fully trained PPO policy on flat and rough terrains, with no significant obstacles.
In Policy-aided self-supervised training, we train the dynamics network alongside a PPO Policy. The trajectories are obtained directly from the state-action buffer, which is updated with trajectories from the next iteration step.
Our dynamics model then captures and learns to identify failures, falling over, and early interactions as our robot learns to walk on rough and flat terrains.
In Policy-aided Exploration Training, we place a reward on the dynamics model's epistemic uncertainty output, and use the epistemic uncertainty as a reward to aid which regions the MOPO-PPO selects for the locomotion training step.
Our dynamics model learns alongside the locomotion policy, and can accurately predict failures, falling over, and also has a richer dataset on terrains with high epistemic uncertainty, i.e. extreme versions of the rough terrains the locomotion policy trains on and edge cases.
We then evaluate each uncertainty-aware world model on the OOD Trajectories, answering: Can we identify OOD observations and does an uncertainty-aware world model provide higher prediction accuracy?
Results
Learned Dynamics Model
Model Prediction (Left)
Ground Truth (Right)
GRU - In Distribution Epistemic Uncertainty
GRU - OOD Epistemic Uncertainty
Across the flat terrain in-distribution trajectory dataset and the rough terrain out of distribution terrain dataset, our Uncertainty Aware model distinguishes in-distribution and OOD joint space configurations and dynamics through the epistemic uncertainty.
Across our three base layers the GRU, which assumes consequentiality across dynamic steps was able to most accurately model the dynamics, up to 400 timesteps into the future (8 seconds), from a moving horizon of 32 timesteps. The modeled dynamics in the picture show an quadruped walking over uneven but uniformly-noised terrain.
Epistemic uncertainties (black, second from bottom) show spikes at the foot impacts during the gait.
True (Solid) vs Predicted(Dashed) forward dynamics predictions across proprioceptive measurements. GRU
Cumulative Error - OOD Observations
Cumulative Error - OOD Observations
Cumulative Error - OOD Observations
Across our tested locomotion policies, the reward on epistemic uncertainty exploration alone was not enough to visualize any differences across performance on Out of Distribution Rough Terrains, (including large obstacles, more challenging versions of trained terrains, and different noise profiles). We observed minor evidence that policy-aided exploration training with epistemic uncertainty cleared these OOD obstacles faster.
As Future Work, we propose the integration of Uncertainty-Aware models to guide exploration of the task space during training, and showcase practical implementations for using Uncertainty-Aware models in deployments. A particular example is to use epistemic uncertainty as an indicator for policy switching to a safe policy under OOD observation, thus conserving performance guarantees. We are motivated by the prospect of using uncertainty-aware models in On-Deployment learning across hardware.
Presentation - February 2025
Manuscript - December 2024