Human-Robot Interaction: Directing Embodied Agents
How people instruct robots and agents, and how those agents show what they understood
When an agent has a body and shares physical space with us, instructing it becomes a spatial and social problem. We build interaction techniques that let people tell robots and embodied agents what to do. The techniques borrow from how people already direct each other’s attention and effort, and we evaluate them in lab studies against the alternatives people would otherwise use.
Telling an agent what to do
People allocate tasks to each other with a mixture of speech, pointing, and gaze, and they rarely say everything out loud. In Grip-that-there [Mahadevan, K., Sousa, M., Tang, A., and Grossman, T. (2021). "Grip-that-there": An Investigation of Explicit and Implicit Task Allocation Techniques for Human-Robot Collaboration. In CHI 2021: Proceedings of the 2021 SIGCHI Conference on Human Factors in Computing Systems.] we built explicit and implicit techniques for allocating tasks to a robot and compared them in a 16-participant block-stacking study. Mimic [Mahadevan, K., Chen, Y., Cakmak, M., Tang, A., and Grossman, T. (2022). Mimic: In-Situ Recording and Re-Use of Demonstrations to Support Robot Teleoperation. In UIST 2022: The 35th Annual ACM Symposium on User Interface Software and Technology.] lets a teleoperator save a demonstrated trajectory as a template and re-use it in new situations; we evaluated it against direct control in a simulated environment.
Instruction does not have to be verbal. In ImageInThat [Mahadevan, K., Lewis, B., Li, J., Mutlu, B., Tang, A., and Grossman, T. (2025). ImageInThat: Manipulating Images to Convey User Instructions to Robots. In Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction, 757–766.] people manipulate images of a scene directly to convey what they want a robot to do, so the image itself becomes the shared referent. We compared it against typed natural-language instruction in a user study. Stargazer [Li, J., Sousa, M., Mahadevan, K., Wang, B., Aoyagui, P., Yu, N., Yang, A., Balakrishnan, R., Tang, A., and Grossman, T. (2023). Stargazer: An Interactive Camera Robot for Capturing How-To Videos Based on Subtle Instructor Cues. In CHI 2023: Proceedings of the 2023 SIGCHI Conference on Human Factors in Computing Systems.] inverts the arrangement: rather than waiting to be instructed, a camera robot reads a person’s unintentional cues while they work and repositions itself accordingly. We studied it with six instructors, each filming a tutorial for a different skill.
How embodiment changes prompting
We built an embodied VR agent that signals its understanding by turning its head and highlighting the objects being referred to. A Wizard of Oz study examined how embodiment and multimodal signalling change the way people prompt [Zhang, T., Au Yeung, C., Aurelia, E., Onishi, Y., Chulpongsatorn, N., Li, J., and Tang, A. (2025). Prompting an Embodied AI Agent: How Embodiment and Multimodal Signaling Affects Prompting Behaviour. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 60.] . A companion Wizard of Oz elicitation study with 22 participants looked at what people implicitly expect when they verbally prompt an agent to build an interactive VR scene [Manesh, S., Zhang, T., Onishi, Y., Hara, K., Bateman, S., Li, J., and Tang, A. (2024). How People Prompt Generative AI to Create Interactive VR Scenes. In DIS 2024: Conference on Designing Interactive Systems 2024.] .
Proxemics as an interaction resource
People use position and orientation to manage their interactions with each other, and a system that can sense those cues can act on them too. We have used proxemics to mediate interaction with robots [Mahadevan, K., Sousa, M., Tang, A., and Grossman, T. (2021). "Grip-that-there": An Investigation of Explicit and Implicit Task Allocation Techniques for Human-Robot Collaboration. In CHI 2021: Proceedings of the 2021 SIGCHI Conference on Human Factors in Computing Systems.] , and to decide when a VR headset should reveal a nearby bystander who may want to interact. That second question we approached with a user study alongside simulation across VR content of varying interactivity [Kudo, Y., Tang, A., Fujita, K., Endo, I., Takashima, K., and Kitamura, Y. (2021). Towards Balancing VR Immersion and Bystander Awareness. In Proceedings of the ACM on Human-Computer Interaction (PACMHCI).] .
Showing what the agent understood
When an agent misreads a gesture or a phrase, the person usually discovers it only after the robot has acted. We are working on interfaces that make an agent’s interpretation visible before it commits: what it took the instruction to mean, and what it intends to do next. The aim is to let a person catch and correct a misreading early.
Publications
Acceptance: 26.3% - 749/2844. Best Paper Nominee (top 5% of submissions)
Acceptance: 24% - 268/1124. Best Paper Award (Top 1% of all submissions)
Acceptance: 27.0% - 1198/4444. Honourable Mention Award (Top 5% of all submissions)
Acceptance: 190/695 - 27.4%.
Best Paper award