Sammanfattning

Like human touch, tactile sensing can provide an important additional source of information for contact-rich precision tasks, especially when vision is impaired, unavailable, or insufficient. This thesis introduces a vision and force Diffusion Policy framework that fuses low-rate vision and robot state observations with a longer, high-rate force and torque (F/T) history to generate robot action sequences. The multimodal policy is evaluated against vision-based and F/T- based variants in a precision peg-in-hole task using a pair of compliant grippers that deform upon interaction. The grippers are adapted from the universal manipulation interface (UMI) design. In the multi-location evaluation, referred to as Experiment 1.5, the vision and F/T fusion model achieved a full-task success rate of 22%, compared with 5% for the vision-based policy and 3% for the F/T-based policy. In the same experiment, the fusion model achieved an insertion success rate of 57%, compared with 32% for the vision-based policy and 25% for the F/T-based policy. Among the trials that achieved full-task success, the fusion model also completed the task faster than both unimodal policies and had a lower mean summed absolute force metric than the F/T-based model. These results suggest that combining vision and F/T sensing can improve performance in contact-rich precision manipulation with compliant grippers.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.