Learning from Human Corrections

Seunghwan Um

Sungkyunkwan University

Advisor: Prof. Hyouk Ryeol Choi

01 Dataset collection

Even for a human operator, repeating the same insertion consistently is difficult, making demonstration collection challenging.

02 Base policy

It can insert the box. Contact can still stop it.

BASE-POLICY TRAINING

Absolute-action demonstrations. Collected at the Original position only.

Wrist camera Same rollout · recorded RGB
Insertion 3× approach → 1× insertion
Robot trajectory Video-derived view
Measured EE pathBase target
0.0 / 15.0 s

Same rollout · ep03 · timestamp-aligned

Synchronization details

All three views use ep_00003. The original video-to-log offset is +109.864 s, followed by the same 3× approach / 1× insertion edit. Wrist RGB and trajectory use the same nearest logged sample; maximum sampling mismatch is 64 ms, plus source-video frame quantization. The path shows the logged end-effector position, not an estimated fingertip. Its viewing direction comes from the original video projection matrix.

Contact problem

03 Residual policy

Correct alignment before insertion.

Contact failure Base only · concept illustration
Observations feed a frozen base policy and a residual policy trained on human corrections; their actions are added into the executed action.
Residual policy
Related paper: CR-DAgger ↗
Residual correction Base + residual · concept illustration

04 Human guidance

Cyan: human correction vector

05 Evaluation

Orange: residual correction toward −x

BACK −3 CM · OUTSIDE THE BASE TRAINING SET

The base policy learned absolute actions at the Original position only.
With the target shifted back by 3 cm, the base policy alone fails.

OriginalTraining position · 0 cm
Base + residual · Original position
BackUnseen target shift: −3 cm
Box displacement relative to fixed black tapeThe tape stays fixed. A dashed circle marks the Original box corner; a blue arrow points toward the Back corner. Minus 3 centimeters is the stated experiment condition, not a distance measured from these pixels. Fixed tape Back −3 cm
Base + residual · No base-training demos at this offset

Dashed: Original corner · Blue: shifted position

Separate autonomous rollouts

Position cues are visual references; −3 cm is the experiment setting.