← Search

Yen-Ling Kuo

19 accepted papers

2025

CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation

CoRL 2025poster

Recent advances in Behavior Cloning (BC) have led to strong performance in robotic manipulation, driven by expressive models, sequence modeling of actions, and large-scale demonstration data. However, BC faces significant challenges when applied to heterogeneous datasets, such as visual shift with d…

Cited by 0SourceScholar
2025

Diff-Dagger: Uncertainty Estimation With Diffusion Policy for Robotic Manipulation

ICRA 2025

Recently, diffusion policy has shown impressive results in handling multi-modal tasks in robotic manipulation. However, it has fundamental limitations in out-of-distribution failures that persist due to compounding errors and its limited capability to extrapolate. One way to address these limitation

Cited by 20SourcecodeScholar
2025

Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foraging

CVPR 2025poster

Imagine searching a collection of coins for quarters (0.25), dimes (0.10), nickels (0.05), and pennies (0.01)--a hybrid foraging task where observers search for multiple instances of multiple target types. In such tasks, how do target values and their prevalence influence foraging and eye movement b…

2025

MuMA-ToM: Multi-modal Multi-Agent Theory of Mind

AAAI 2025technical

Understanding people's social interactions in complex real-world scenarios often relies on intricate mental reasoning. To truly understand how and why people interact with one another, we must infer the underlying mental states that give rise to the social interactions, i.e., Theory of Mind reasonin…

2025

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation

CoRL 2025poster

Grounding object affordance is fundamental to robotic manipulation as it establishes the critical link between perception and action among interacting objects. However, prior works predominantly focus on predicting single-object affordance, overlooking the fact that most real-world interactions invo…

Cited by 0SourceScholar
2024

MMToM-QA: Multimodal Theory of Mind Question Answering

ACL 2024long

Theory of Mind (ToM), the ability to understand people’s mental states, is an essential ingredient for developing machines with human-level social intelligence. Recent machine learning models, particularly large language models, seem to show some aspects of ToM understanding. However, existing ToM b…

2024

Neural Amortized Inference for Nested Multi-Agent Reasoning

AAAI 2024technical

Multi-agent interactions, such as communication, teaching, and bluffing, often rely on higher-order social inference, i.e., understanding how others infer oneself. Such intricate reasoning can be effectively modeled through nested multi-agent reasoning. Nonetheless, the computational complexity esca…

2024

Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation

CVPR 2024poster

We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects coined "action context". We propose TransFusion a multimodal transformer-based architecture for short-term object interaction anticipati…

Cited by 5SourcePDFScholar
2023

Zero-Shot Linear Combinations of Grounded Social Interactions with Linear Social MDPs

AAAI 2023technical

Humans and animals engage in rich social interactions. It is often theorized that a relatively small number of basic social interactions give rise to the full range of behavior observed. But no computational theory explaining how social interactions combine together has been proposed before. We do s…

Cited by 1SourcePDFScholar
2022

Incorporating Rich Social Interactions Into MDPs

ICRA 2022poster

Much of what we do as humans is engage socially with other agents, a skill that robots must also eventually possess. We demonstrate that a rich theory of social interactions originating from microsociology can be formalized by extending a nested MDP where agents reason about arbitrary functions of e…

Cited by 10SourceScholar
2022

Trajectory Prediction with Linguistic Representations

ICRA 2022poster

Language allows humans to build mental models that interpret what is happening around them resulting in more accurate long-term predictions. We present a novel trajectory prediction model that uses linguistic intermediate representations to forecast trajectories, and is trained using trajectory samp…

Cited by 22SourceScholar
2021

Compositional Networks Enable Systematic Generalization for Grounded Language Understanding

EMNLP 2021finding

Humans are remarkably flexible when understanding new sentences that include combinations of concepts they have never encountered before. Recent work has shown that while deep networks can mimic some human language abilities when presented with novel sentences, systematic variation uncovers the limi…

2020

Encoding formulas as deep networks: Reinforcement learning for zero-shot execution of LTL formulas

IROS 2020poster

We demonstrate a reinforcement learning agent which uses a compositional recurrent neural network that takes as input an LTL formula and determines satisfying actions. The input LTL formulas have never been seen before, yet the network performs zero-shot generalization to satisfy them. This is a nov…

Cited by 57SourcecodeScholar
2020

Learning a natural-language to LTL executable semantic parser for grounded robotics

CoRL 2020

Children acquire their native language with apparent ease by observing how language is used in context and attempting to use it themselves. They do so without laborious annotations, negative examples, or even direct corrections. We take a step toward robots that can do the same by training a grounde

Cited by 0SourcePDFScholar