Action-space noise
Exploration is attached to the executed action, avoiding dilution.
Flow-based vision-language-action models are powerful robot policies, but their usual reinforcement-learning exploration lives inside a multi-step denoising chain. That makes it difficult to control the actions that are ultimately executed.
We identify structured noise dilution: temporal correlations and action-group scales injected partway through the chain can be weakened by the remaining denoising steps. StructRL instead places exploration directly at the action output, where its intended structure reaches the robot intact.
StructRL decodes a clean action chunk with a deterministic ODE, then adds noise directly in action space. Fixed AR(1) correlation smooths the noise over the chunk. The decoder's terminal features predict a scale for each action dimension and time step, within bounds set for position, rotation, and gripper controls.
Terminal replay supplies a tractable policy-gradient signal to the flow decoder without assigning likelihoods to intermediate denoising states. The default setting replays the last denoising step.
Exploration is attached to the executed action, avoiding dilution.
Smooth noise matches the temporal structure of action chunks.
State-conditioned scales are bounded separately for position, rotation, and gripper controls.
The final denoising steps provide an efficient learning signal.
We evaluate StructRL on LIBERO, ManiSkill, and CALVIN across three flow-based VLA backbones. Green marks StructRL; bold marks the best result in each column, including ties.
| Backbone | Method | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|---|
| GR00T N1.5 | Few-shot SFT | 41.4 | 58.6 | 48.2 | 61.9 | 52.5 |
| πRL (Flow-SDE + PPO) | 96.6 | 100.0 | 93.8 | 95.6 | 96.5 | |
| Baseline | 96.2 | 99.4 | 91.8 | 95.6 | 95.8 | |
| StructRL | 99.2 | 100.0 | 97.6 | 99.0 | 99.0 | |
| π0 | Few-shot SFT | 65.3 | 64.4 | 49.8 | 51.2 | 57.6 |
| πRL (Flow-SDE + PPO) | 98.4 | 99.4 | 96.2 | 90.2 | 96.0 | |
| π-StepNFT | 93.5 | 98.0 | 83.7 | 86.7 | 90.5 | |
| Baseline | 98.8 | 99.0 | 97.2 | 90.2 | 96.3 | |
| StructRL | 98.8 | 99.8 | 98.4 | 91.8 | 97.2 | |
| π0.5 | Few-shot SFT | 84.6 | 95.4 | 84.6 | 43.9 | 77.1 |
| πRL (Flow-SDE + PPO) | 99.6 | 100.0 | 98.8 | 93.0 | 97.9 | |
| π-StepNFT | 97.8 | 100.0 | 98.2 | 79.8 | 94.0 | |
| Baseline | 99.4 | 100.0 | 93.8 | 86.0 | 94.8 | |
| StructRL | 99.6 | 100.0 | 98.8 | 90.2 | 97.2 |
| Backbone | Method | IND | Vision | Semantic | Execution | OOD Avg. |
|---|---|---|---|---|---|---|
| π0 | SFT | 38.4 | 32.6 | 8.4 | 13.2 | 18.1 |
| πRL (Flow-SDE + PPO) | 78.8 | 61.1 | 25.4 | 31.5 | 39.3 | |
| π-StepNFT | 79.2 | 69.1 | 49.1 | 33.1 | 50.4 | |
| Baseline | 84.1 | 79.4 | 63.3 | 59.1 | 67.3 | |
| StructRL | 87.8 | 80.0 | 64.2 | 56.2 | 66.8 | |
| π0.5 | SFT | 40.1 | 38.8 | 16.6 | 22.3 | 25.9 |
| πRL (Flow-SDE + PPO) | 90.9 | 68.0 | 34.5 | 45.4 | 49.3 | |
| π-StepNFT | 85.4 | 76.9 | 56.6 | 45.1 | 59.5 | |
| Baseline | 85.9 | 76.2 | 62.2 | 59.4 | 65.9 | |
| StructRL | 90.9 | 78.3 | 68.4 | 64.2 | 70.3 |
| Method | Len1 | Len2 | Len3 | Len4 | Len5 | Avg. length |
|---|---|---|---|---|---|---|
| SFT | 0.927 | 0.843 | 0.767 | 0.688 | 0.613 | 3.838 |
| πRL (Flow-SDE + PPO) | 0.997 | 0.982 | 0.958 | 0.910 | 0.870 | 4.717 |
| πRL (Flow-Noise + PPO) | 0.996 | 0.976 | 0.939 | 0.896 | 0.845 | 4.652 |
| Baseline | 0.998 | 0.990 | 0.964 | 0.929 | 0.868 | 4.749 |
| StructRL | 1.000 | 0.993 | 0.968 | 0.934 | 0.880 | 4.775 |




We initialize each policy with SFT on 10 demonstrations, with 0% success over 25 trials before online RL. We then continue training with StructRL + AWAC and compare against Gaussian action-space exploration.
Task 01
After 60 minutes, StructRL succeeds in 21/25 trials (84%); the Gaussian baseline succeeds in 14/25 (56%).
Task 02
The SFT policy starts at 0% success. StructRL reaches stable, fully successful behavior after about 30 minutes; the Gaussian baseline takes about 35 minutes.