Writing University College London Research (UCLR): U butterfly, C front crawl, L backstroke, R breaststroke, all in still water.
Challenging conditions: U without lower legs, C in a 0.51 m/s wave, L in honey, R pushed with 0.5 mg every 6 s.
Deep reinforcement learning (RL) has driven rapid progress in physically-based motion generation, yet synthesizing robust motion policies in dense fluid environments (e.g. swimming) remains unsolved. Unlike land motion, where environmental impact is sparse and can be coarsely modeled (e.g. gravity, normal reaction), swimming requires continuous, full-body coordination under pressure and flow forces across the entire body surface; fully-coupled rigid–fluid simulation is far too slow for the millions of interactions RL requires. We propose SWIM, an RL framework for physically-based human swimming learned from a single reference motion. Its core is a compact, low-dimensional environment representation: a dual-branch graph-convolutional VQ-VAE that tokenizes per-link body–water forces and torques, informative enough for control yet robust to the rapidly changing force exchanges that would otherwise destabilize training. We pair this with a GPU-based Lagrangian solver with rigid–fluid coupling that yields 10–15× faster training. Across hundreds of training runs and several hundred zero-shot conditions, SWIM generalizes to unseen goals and trajectories, fluids with varying density, external flows and perturbations, altered body morphologies, and all four competitive swimming styles. Against imitation-learning baselines (MimicKit-DeepMimic, MimicKit-AMP, ADD) and alternative force models (a PINN world model, MuJoCo's inertial fluid model, and an underwater-robotics simulator), SWIM achieves better stability, goal satisfaction, and physical realism: a 63.9% success rate on held-out generalization versus 36.8% for the strongest baseline.
Videos play while you hover over them. On a touch screen, tap a video to play it. Each section shows one example; See more results opens every stroke, task and camera view for that section.
Below is a visual comparison of SWIM with three baseline imitation methods (AMP, DeepMimic, ADD), three surrogate models for simulation (CFC particle PINN, OceanSim, MuJoCo), and one high-resolution simulator (High-resolution SPH). We provide results across four swimming styles and eight control trajectories. For each motion, we provide both top and side views.
Videos play while you hover over them. On a touch screen, tap a video to play it.
We report generalization results on exhaustive and challenging combinations of conditions in tasks, environments, perturbation, and body geometry and morphologies.
We evaluate our method under various water flow scenarios across four flow directions: downstream (0°), diagonal (45°), crossflow (90°), and upstream (180°), and eight different speeds. We present representative video results for two example flow directions at two flow speeds.
Videos play while you hover over them. On a touch screen, tap a video to play it.
We evaluate our method on a larger pool than the training setup across various combinations. Below, we provide representative example videos in 5 m and 8 m pools with various trajectories.
Videos play while you hover over them. On a touch screen, tap a video to play it.
We evaluate our method on fluids with different viscosities across various trajectories. Below, we provide example results for glycerol, which has a viscosity about 1,100 times that of water.
Videos play while you hover over them. On a touch screen, tap a video to play it.
We also test our method against perturbations directed backward along the swimming path, as well as sideways and downward perturbations. Representative video results are presented below.
Videos play while you hover over them. On a touch screen, tap a video to play it.
Our method also generalizes to various morphological alterations. Specifically, we evaluated adding swim fins of varying lengths to the feet, as well as partial and full amputations of an arm or a leg. Representative video results are presented below.
Videos play while you hover over them. On a touch screen, tap a video to play it.
Given the large number of experiments in Ablation and limited space, we only show results of the top three configurations: Whole-body raw, Per-body token, and SWIM.
Videos play while you hover over them. On a touch screen, tap a video to play it.