PEARS turns physical priors in VLMs into a compass for RL post-training.
PEARS at a glance. Tactile-aware diffusion steering reinforcement learning (DSRL) and physics-guided force reasoning (PFR) jointly accelerate online adaptation.
The idea
Abstract
Pretrained robotic policies can suffer substantial performance degradation under out-of-distribution (OOD) conditions encountered during deployment, motivating post-training through real-world interaction. However, reinforcement-learning (RL)-based post-training typically requires substantial environment interactions, a burden that is especially significant in manipulation, where each trial can be slow, costly, or destructive. Therefore, we present PEARS, a physics-prior-guided hybrid RL framework for sample-efficient online adaptation of pretrained policies with tactile feedback. After each episode, its physics-guided force reasoning (PFR) module uses physical priors encoded in a vision-language model (VLM) to diagnose failures from the visual outcome and tactile interaction history and update task-appropriate contact-force bounds. A high-frequency hybrid force-position controller then enforces these bounds during contact. Complementarily, tactile-conditioned diffusion steering reinforcement learning adjusts the latent noise of the frozen flow-matching policy to correct errors in free-space motion and contact timing without updating the base model. In simulation, PEARS improves success rates by 12.4–37.4 percentage points over the strongest per-task baselines. PEARS also reduces the number of interaction episodes required for a certain success threshold by up to 53.2% relative to the fastest baseline. In real-world experiments, PEARS achieves success rates of 95% on Whiteboard Erasing and 90% on Pipette Liquid Aspiration. These results show that combining the PFR module with policy steering can accelerate adaptation while reducing costly interactions.
Method
Learn from failures. Refine motion. Regulate contact.
PEARS combines two complementary adaptation mechanisms around a frozen pretrained policy.
01
Reason about interaction forces
After each episode, PFR interprets the final visual outcome, measured force trajectory, and interaction history. It distinguishes insufficient or excessive force from positioning and trajectory errors, then updates the relevant force bounds.
02
Steer the frozen policy
A tactile-conditioned RL actor adapts the latent noise supplied to the flow-matching policy. This corrects motion alignment and contact timing while preserving the pretrained policy weights.
03
Control force during contact
A stage-aware hybrid force-position controller regulates the selected force or torque axes within the inferred bounds. The remaining directions follow the policy, which also controls free-space motion.
Stage-aware force control. During whiteboard wiping, the controller regulates normal contact force while the policy drives tangential motion.
Physics-guided force reasoning
Turn each interaction into a better force estimate.
Force bounds remain fixed within an episode and are updated between episodes. PFR uses the accumulated history to refine its estimates; it retains the current bounds when a failure is attributed to motion or the evidence is inconclusive.
Refining initially inaccurate bounds. Example force and torque intervals predicted by PFR during online interaction.
Evaluation
More successful adaptation, fewer interactions.
Evaluation spans three simulated contact-rich tasks and two tasks on a physical robot.
12.4–37.4 pp
Higher simulation success than the strongest per-task baselines
53.2%
Fewer episodes to reach 80% online success on Bottle Cap Twisting
95% / 90%
Real-world success on Whiteboard Erasing / Pipette Liquid Aspiration
Simulation Tasks
Thin Sheet TransferBottle Cap TwistingFragile Fruit Picking
Online adaptation in simulation
Mean trailing-10 online success rates over three 100-episode runs. Shaded bands indicate one sample standard deviation. Complete windows end at episodes 10–100.
Final-policy success and online adaptation efficiency
Simulation task
PEARS success
Best baseline success
Episodes to 80%
Thin Sheet Transfer
94.7 ± 1.5%
82.3 ± 4.5% DSRL
10
Bottle Cap Twisting
87.7 ± 1.5%
68.0 ± 20.3% Residual RL
22
Fragile Fruit Picking
77.7 ± 1.2%
40.3 ± 4.2% DSRL
25
Final-policy success: mean ± sample standard deviation across three runs, with 100 test trials per run. Episodes to 80% refers to PEARS' mean trailing-10 online success. The 53.2% reduction compares 22 episodes for PEARS with 47 for the fastest baseline, Residual RL.
Real-world deployment
Precise contact, from wiping to liquid handling.
Each adaptive method receives 40 episodes of online interaction, followed by 20 test trials per task.
Real-world rollouts
For each task, a failed behavior-cloned (BC) rollout before adaptation is followed by a successful rollout after PEARS training.
Pipette Liquid Aspiration
BC rollout · Failure
Incorrect positioning. The pipette is not properly aligned, causing the task to fail.
PEARS · Success
After PEARS training, the robot successfully completes liquid aspiration and transfer. PEARS succeeds in 18 of 20 trials.
Whiteboard Erasing
BC rollout · Failure
Inappropriate force magnitude. The applied force is unsuitable for the task, resulting in unsuccessful erasing.
PEARS · Success
After PEARS training, the robot applies appropriate contact force to erase the board while keeping the eraser securely grasped. PEARS succeeds in 19 of 20 trials.
Liquid transfer · Close-up comparison
More appropriate force, less liquid leakage.
After RL fine-tuning, PEARS applies more appropriate grasping force to the pipette, reducing liquid leakage during transfer compared with the behavior-cloned (BC) policy.
BC
Before RL fine-tuning
PEARS
After RL fine-tuning
Real-world performance: successful trials / total trials
Task
Base policy
DSRL
PEARS
Whiteboard Erasing
11 / 20
15 / 20
19 / 20
Pipette Liquid Aspiration
8 / 20
12 / 20
18 / 20
An interactive explanation
A little physics goes a long way.
Too gentle, and the marks remain. Too forceful, and the eraser slips. Follow three rollouts to see how failure reasoning points toward a better contact force.
Whiteboard lab
Changing the force starts a fresh rollout. The force target stays fixed during each attempt.
Rollout 01 / 03Ready
Contact target2 N
Visual outcomeReady to wipe
Rollout 1 · 0%
2 N · too gentle9 N · too forceful6 N · just right
Beyond this simplified demo. Real-world failures can involve more complex combinations of force, positioning, and contact timing. To address this, we propose PEARS: a framework that uses a VLM to reason with physical priors and the complete history of interaction outcomes and tactile feedback, diagnosing failures and inferring suitable contact-force ranges for the next interaction.
The 5–8 N window, forces, and outcomes above are illustrative. In PEARS, force bounds are updated between episodes, while RL separately refines motion and contact timing.