Seeing Through the Displaced Frame

Privileged Noise Distillation for Vision-Force Precision Assembly

Supplementary video

Simulation and hardware, 2 min 50 s. No audio.

Abstract

Pose error in precision assembly can corrupt not only what a robot observes but also the coordinate frame in which it acts. On the FORGE benchmark, the official state-based policy succeeds in 97% to 99% of episodes with the true pose but only 32% to 60% at the benchmark's σ = 5 mm pose-noise setting. The same estimated pose enters the observation and anchors the action frame, making the offset unidentifiable from proprioceptive state alone before contact. We supply this missing information during training in two ways. A privileged teacher observes the offset in simulation, while clean demonstrations can instead be relabelled into the displaced frame in closed form. The deployed student is trained with behaviour cloning followed by one DAgger round and receives only noisy state, a raw wrench window, and two RGB cameras at test time. On the unmodified FORGE tasks, the teacher-route student maintains 92% to 99% success across σ = 0 to 5 mm, while six non-privileged baselines fall to 2% to 80%. A matched behaviour-cloning experiment isolates the source of this robustness. With the same student architecture, data budget, and training procedure, demonstrations generated without offset access yield only 31.5% success at σ = 5 mm, whereas privileged and relabelled demonstrations reach 88.7% and 95.8%. Deployed zero-shot on a Franka, the student reaches 83.3% pooled success at σ = 5 mm against 34.4% for the strongest state-based policy. Deployable sensing alone is insufficient. Robustness requires supervision that encodes compensation for the latent frame offset.

The displaced action frame

Why a pose estimate that is 5 mm off is worse than it sounds.

The estimated fixture pose does two jobs at once. It enters the observation, and it anchors the frame the actions are expressed in. A policy that sees only proprioceptive state therefore cannot identify the offset before contact: every signal it has is already expressed in the displaced frame, so nothing in its own state reveals that the frame has moved. The peg's hole leaves 0.114 mm of clearance, and at σ = 5 mm the believed hole barely overlaps the real one.

Diagram: the true hole at p-star and the believed hole at p-hat, separated by epsilon. The plug descends in the believed frame and misses.
The offset ε displaces both the observation and the commanded descent. The robot acts correctly in a frame that is in the wrong place.

Method

Put the offset into the training signal, then take it away at test time.

Two routes to the same supervision

  • A privileged teacher observes ε in simulation and acts through it.
  • Or clean demonstrations are relabelled into the displaced frame in closed form.

Either route leaves the compensation present in the training data.

What the deployed student gets

  • Noisy state, no access to ε
  • A raw wrench window
  • Two RGB cameras, fused into one action at 15 Hz

Behaviour cloning followed by one DAgger round. No privileged input at test time.

Architecture: privileged teacher and closed-form relabelling produce demonstrations; the student fuses noisy state, a force window and two RGB cameras into an action at 15 Hz; the same network deploys on a Franka.
The same network and the same sensing run on a Franka Panda, zero-shot, with no real-world data and no fine-tuning.

Results

Three unmodified FORGE tasks in simulation, then zero-shot on hardware.

92–99%
student success in simulation, σ = 0 to 5 mm
32–60%
official state-based policy at σ = 5 mm
83.3%
hardware, pooled, σ = 5 mm
34.4%
strongest state-based policy, same trials
Simulation. Success rate (%) across the noise spectrum, three seeds, n = 256 per cell. Cells report dynamics randomisation on / off. T-B is the privileged oracle, which the student is not given at test time.
TaskPolicy0 mm1 mm2.5 mm5 mm
pegT-A official baseline98.7 / 98.297.3 / 94.773.7 / 74.531.8 / 33.3
T-A+ noise-augmented84.1 / 77.378.0 / 75.158.7 / 56.229.9 / 25.8
T-B privileged (oracle)99.3 / 96.699.3 / 96.799.7 / 97.599.3 / 96.6
student (ours)98.3 / 96.298.6 / 96.298.3 / 95.496.6 / 91.8
gearT-A official baseline98.6 / 100.097.5 / 99.286.9 / 90.649.9 / 53.8
T-A+ noise-augmented94.9 / 99.794.7 / 99.792.1 / 99.377.9 / 90.0
T-B privileged (oracle)98.7 / 99.799.0 / 100.098.2 / 99.798.2 / 98.8
student (ours)98.8 / 98.499.2 / 97.899.3 / 96.598.2 / 96.5
nutT-A official baseline97.4 / 97.896.1 / 96.987.4 / 87.960.4 / 64.3
T-A+ noise-augmented89.7 / 88.488.8 / 88.383.7 / 85.273.0 / 71.2
T-B privileged (oracle)99.2 / 99.799.5 / 99.999.9 / 99.699.2 / 99.9
student (ours)99.1 / 98.899.0 / 98.899.0 / 98.696.5 / 95.1

At 5 mm the improvement over the official baseline reaches +64.8 points on peg, +48.3 on gear and +36.1 on nut. Measured as the fraction of baseline failures eliminated, the student removes 91.1% to 96.4% of them. The recovered oracle gap is 96.0% on peg, 100% on gear and 93.0% on nut.

Hardware. Franka Panda, printed parts, dark work surface, 30 trials per cell. Transfer is zero-shot: the simulation checkpoint runs unchanged. T-A+ is the noise-augmented state baseline without cameras. Pooled rows combine all three tasks (n = 90).
TaskPolicy0 mm1 mm2.5 mm5 mm
pegT-A+86.783.363.320.0
student (ours)90.093.386.780.0
gearT-A+90.090.073.336.7
student (ours)93.390.083.386.7
nutT-A+90.083.370.046.7
student (ours)93.390.086.783.3
pooledT-A+88.985.668.934.4
student (ours)92.291.185.683.3

The margin grows with pose error: 17 points at 2.5 mm (Fisher's exact test, p = 0.012) and 49 points at 5 mm (Wilson 95% intervals [74, 90] versus [25, 45], p < 10−10), with the largest task-level gap on peg, 80.0% against 20.0%.

Initial and goal states of peg insertion, gear meshing and nut threading, in simulation and on hardware.
Initial and goal states of the three tasks, in simulation and on the Franka.

Try it yourself

Pick a task, a noise level and a policy. Every clip is a recorded rollout.

task
pose noise σ
policy

Where the robustness comes from

A matched behaviour-cloning experiment, same architecture and data budget.

31.5%
demonstrations generated without offset access
88.7%
privileged demonstrations
95.8%
closed-form relabelled demonstrations

All three share the student's architecture, data budget and training procedure, and all three are evaluated at σ = 5 mm. The only difference is whether the demonstrations encode compensation for the offset. Adding cameras and force to a policy is not what produces the robustness. The supervision is.

What this does not solve

  • Unseen geometries. On 70 AutoMate plug and socket pairs in the unmodified peg environment, at 2.5 mm of pose error the student succeeds in 4.6% of trials against 18.4% for the noise-augmented state baseline. Robustness to pose error does not imply cross-geometry generalisation.
  • Scene composition. On peg, replacing the dark work surface with a light one takes hardware success from 93% to 25%. Photometric augmentation narrows the appearance gap, it does not close it.
  • Camera calibration. Increasing camera extrinsic error from 5 mm and 1 degree to 10 mm and 2 degrees roughly halves gear success.
  • Sample size. The main results are simulated. The hardware sweep is 30 trials per cell, and several cells elsewhere are single-seed, marked as such in the paper.

BibTeX

Anonymous entry while the paper is under review. Update on acceptance.

@misc{anonymous_displaced_frame,
  title  = {Seeing Through the Displaced Frame: Privileged Noise
            Distillation for Vision-Force Precision Assembly},
  author = {Anonymous},
  note   = {Under review}
}