Search arXiv⌕ Search

arXiv subjects

Bastian Orthmann

Publications and source records attributed to Bastian Orthmann.

3 recordsLinked to original sources

Hybrid Imitation Learning: Teleoperation Augmentation Primitives that Policies Learn to Trigger

What an operator can demonstrate bounds what imitation learning can learn. Teleoperation interfaces map the human body to the robot, so motions that are hard for a human, such as holding an exact orientation, returning to the same viewpoint, or turning a wrist joint several full revolutions, are difficult to demonstrate on most teleoperation interfaces, even when they are trivial for the robot. We introduce Teleoperation Augmentation Primitives (TAPs): axis locks, perching waypoints, pose anchors, and embodiment-specific routines that the operator triggers during a demonstration by speech, an AR menu, or simply with a controller button. TAPs are themselves recorded in the demonstration and can therefore also be learned and invoked by the policy itself. In simulation, where benchmarks provide human demonstrations, we show that augmenting them with TAPs after the fact yields better policies on a wristcamera-only peg insertion (0.372 to 0.531 success) and on a mug-cleanup task with a learned "remember this pose" anchor (0.203 to 0.323). On a real robot, three tasks show the same pattern: primitives that help the operator do not hurt the policy, and a policy triggering an unscrewing routine succeeds where a plain policy cannot resolve the multi-turn motion from images alone (67% vs. 38% success). We close with an industrial proof-of-concept case leveraging this idea. Project website: https://hybrid-imitation.github.io/

cs.RO↗

Gesture First, LLM-Assisted Voice Complement: Exploring Multimodal Robot 'Puppeteer' Teleoperation Via Virtual Counterpart in Augmented Reality

Robot teleoperation via augmented reality (AR) offers a promising path toward more intuitive human-robot interaction (HRI). We present a head-mounted AR 'puppeteer' system in which users control a physical robot by interacting with its virtual counterpart robot using large language model (LLM)-assisted voice commands and hand-gesture interaction on the Meta Quest 3. In a within-subject user study with 42 participants performing an AR-based robotic pick-and-place pattern-matching task, we empirically compare two interaction conditions: gesture-only (GO) and combined voice+gesture (VG) on performance and user experience (UX). In VG, voice and gesture operate in a sequential role-allocated manner, with voice handling high-level navigation and gesture handling fine manipulation. Our results show that GO currently provides more reliable and efficient control for this time-critical task, while VG introduces additional flexibility but also latency and recognition issues that can increase workload. We additionally analyze how prior robotics expertise differentiates performance and UX across conditions. Based on these findings, we distill a set of design guidelines for AR 'puppeteer' metaphoric robot teleoperation, framing multimodality as an adaptive strategy that must balance efficiency, robustness, and user expertise rather than assuming that additional modalities are universally beneficial.

cs.HC↗

LLM-Driven Augmented Reality Puppeteer: Controller-Free Voice-Commanded Robot Teleoperation

The integration of robotics and augmented reality (AR) presents transformative opportunities for advancing human-robot interaction (HRI) by improving usability, intuitiveness, and accessibility. This work introduces a controller-free, LLM-driven voice-commanded AR puppeteering system, enabling users to teleoperate a robot by manipulating its virtual counterpart in real time. By leveraging natural language processing (NLP) and AR technologies, our system -- prototyped using Meta Quest 3 -- eliminates the need for physical controllers, enhancing ease of use while minimizing potential safety risks associated with direct robot operation. A preliminary user demonstration successfully validated the system's functionality, demonstrating its potential for safer, more intuitive, and immersive robotic control.

cs.HC↗