SWIM

Compact Environment Representation for
Reinforcement Learning Human Swimming

Writing University College London Research (UCLR): U butterfly, C front crawl, L backstroke, R breaststroke, all in still water.

SWIM: Compact Environment Representation for Reinforcement Learning Human Swimming

1University College London 2University of Glasgow *Corresponding author
SWIM overview
Trained on a single goal-reaching task in still water (left), SWIM learns one policy (middle) that generalizes to unseen tasks, flows, fluids, perturbations and body geometries (right).

Abstract

Deep reinforcement learning (RL) has driven rapid progress in physically-based motion generation, yet synthesizing robust motion policies in dense fluid environments (e.g. swimming) remains unsolved. Unlike land motion, where environmental impact is sparse and can be coarsely modeled (e.g. gravity, normal reaction), swimming requires continuous, full-body coordination under pressure and flow forces across the entire body surface; fully-coupled rigid–fluid simulation is far too slow for the millions of interactions RL requires. We propose SWIM, an RL framework for physically-based human swimming learned from a single reference motion. Its core is a compact, low-dimensional environment representation: a dual-branch graph-convolutional VQ-VAE that tokenizes per-link body–water forces and torques, informative enough for control yet robust to the rapidly changing force exchanges that would otherwise destabilize training. We pair this with a GPU-based Lagrangian solver with rigid–fluid coupling that yields 10–15× faster training. Across hundreds of training runs and several hundred zero-shot conditions, SWIM generalizes to unseen goals and trajectories, fluids with varying density, external flows and perturbations, altered body morphologies, and all four competitive swimming styles. Against imitation-learning baselines (MimicKit-DeepMimic, MimicKit-AMP, ADD) and alternative force models (a PINN world model, MuJoCo's inertial fluid model, and an underwater-robotics simulator), SWIM achieves better stability, goal satisfaction, and physical realism: a 63.9% success rate on held-out generalization versus 36.8% for the strongest baseline.

Results

Videos play while you hover over them. On a touch screen, tap a video to play it. Each section shows one example; See more results opens every stroke, task and camera view for that section.

1 Baseline Comparison

Below is a visual comparison of SWIM with three baseline imitation methods (AMP, DeepMimic, ADD), three surrogate models for simulation (CFC particle PINN, OceanSim, MuJoCo), and one high-resolution simulator (High-resolution SPH). We provide results across four swimming styles and eight control trajectories. For each motion, we provide both top and side views.

2 Generalization

We report generalization results on exhaustive and challenging combinations of conditions in tasks, environments, perturbation, and body geometry and morphologies.

2.1 Moving Water

We evaluate our method under various water flow scenarios across four flow directions: downstream (0°), diagonal (45°), crossflow (90°), and upstream (180°), and eight different speeds. We present representative video results for two example flow directions at two flow speeds.

2.2 Larger Pool

We evaluate our method on a larger pool than the training setup across various combinations. Below, we provide representative example videos in 5 m and 8 m pools with various trajectories.

2.3 Unseen Fluid

We evaluate our method on fluids with different viscosities across various trajectories. Below, we provide example results for glycerol, which has a viscosity about 1,100 times that of water.

2.4 External Perturbations

We also test our method against perturbations directed backward along the swimming path, as well as sideways and downward perturbations. Representative video results are presented below.

2.5 Morphological Alterations

Our method also generalizes to various morphological alterations. Specifically, we evaluated adding swim fins of varying lengths to the feet, as well as partial and full amputations of an arm or a leg. Representative video results are presented below.

3 Ablation

Given the large number of experiments in Ablation and limited space, we only show results of the top three configurations: Whole-body raw, Per-body token, and SWIM.