← Search

Yixin Zhu

93 accepted papers

2026

B*: Efficient and Optimal Base Placement for Fixed-Base Manipulators

ICRA 2026poster

B* is a novel optimization framework that addresses a critical challenge in fixed-base manipulator robotics: optimal base placement. Current methods rely on pre-computed kinematics databases generated through sampling to search for solutions. However, they face an inherent trade-off between solution…

2026

Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis

CVPR 2026

Clinical video diagnosis, in which physicians assess dynamic tissue responses across procedural stages, is critical for detecting diseases such as cervical and colorectal cancers. Recent spatiotemporal models map visual progressions directly to diagnostic outputs, yet overlook two hallmarks of exper

Cited by 0SourceScholar
2026

IntrinsicWeather: Controllable Weather Editing in Intrinsic Space

CVPR 2026

We present IntrinsicWeather, a diffusion-based framework for controllable weather editing in intrinsic space. Our framework includes two components based on diffusion priors: an inverse renderer that estimates material properties, scene geometry, and lighting as intrinsic maps from an input image, a

Cited by 0SourceScholar
2026

Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force Fields

ICLR 2026poster

Predicting physical dynamics from raw visual data remains a major challenge in AI. While recent video generation models have achieved impressive visual quality, they still cannot consistently generate physically plausible videos due to a lack of modeling of physical laws. Recent approaches combining…

Cited by 0SourcecodeScholar
2026

MotionMaster: Generalizable Text-Driven Motion Generation and Editing

CVPR 2026

Synthesizing realistic human motion from natural language holds transformative potential for animation, robotics, and virtual reality. Recent methods handle single-action sequences and simple textual instructions, yet multi-action compositions and precise editing remain elusive due to limited data d

Cited by 0SourcecodeScholar
2026

Neural Force Field: Few-shot Learning of Generalized Physical Reasoning

ICLR 2026poster

Physical reasoning is a remarkable human ability that enables rapid learning and generalization from limited experience. Current AI models, despite extensive training, still struggle to achieve similar generalization, especially in Out-of-distribution (OOD) settings. This limitation stems from their…

Cited by 0SourcecodeScholar
2026

Scalable Trajectory Generation for Whole-Body Mobile Manipulation

CVPR 2026

Robots deployed in unstructured environments must coordinate whole-body motion---simultaneously moving a mobile base and arm---to interact with the physical world. This coupled mobility and dexterity yields a state space that grows combinatorially with scene and object diversity, demanding datasets

Cited by 0SourcecodeScholar
2026

Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation

RA-L 2026

Robotic manipulation requires both rich multimodal perception and effective learning frameworks to handle complex real-world tasks. See-Through-Skin (STS) sensors, which combine tactile and visual perception, offer promising sensing capabilities, while modern imitation learning provides powerful too

Cited by 5SourceScholar
2026

Vi-TacMan: Articulated Object Manipulation Via Vision and Touch

ICRA 2026poster

Autonomous manipulation of articulated objects represents a basic skill for robots deployed in human environments. Current vision-based methods can infer object hidden kinematics, but their estimates are sometimes imprecise in driving reliable actions, especially on previously unseen objects. Tactil…

2025

Ag2x2: Robust Agent-Agnostic Visual Representations for Zero-Shot Bimanual Manipulation

IROS 2025

Bimanual manipulation, fundamental to human daily activities, remains a challenging task due to its inherent complexity of coordinated control. Recent advances have enabled zero-shot learning of single-arm manipulation skills through agent-agnostic visual representations derived from human videos; h

Cited by 0SourceScholar
2025

CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks

CoRL 2025poster

Humanoid robot teleoperation plays a vital role in demonstrating and collecting data for complex interactions. Current methods suffer from two key limitations: (1) restricted controllability due to decoupled upper- and lower-body control, and (2) severe drift caused by open-loop execution. These iss…

Cited by 0SourceScholar
2025

DrivAerStar: An Industrial-Grade CFD Dataset for Vehicle Aerodynamic Optimization

NeurIPS 2025poster

Vehicle aerodynamics optimization has become critical for automotive electrification, where drag reduction directly determines electric vehicle range and energy efficiency. Traditional approaches face an intractable trade-off: computationally expensive Computational Fluid Dynamics (CFD) simulations…

Cited by 0SourcecodeScholar
2025

Dynamic Motion Blending for Versatile Motion Editing

CVPR 2025poster

Text-guided motion editing enables high-level semantic control and iterative modifications beyond traditional keyframe animation. Existing methods rely on limited pre-collected training triplets (original motion, edited motion, and instruction), which severely hinders their versatility in diverse ed…

Cited by 0SourcePDFScholar
2025

GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill

CVPR 2025poster

Learning open-vocabulary physical skills for simulated agents presents a significant challenge in artificial intelligence. Current reinforcement learning approaches face critical limitations: manually designed rewards lack scalability across diverse tasks, while demonstration-based methods struggle…

Cited by 2SourcePDFScholar
2025

GlobalTomo: A global dataset for physics-ML seismic wavefield modeling and FWI

NeurIPS 2025poster

Global seismic tomography, taking advantage of seismic waves from natural earthquakes, provides essential insights into the earth's internal dynamics. Advanced Full-Waveform Inversion (FWI) techniques, whose aim is to meticulously interpret every detail in seismograms, confront formidable computatio…

Cited by 0SourcecodeScholar
2025

Heterogeneous Adversarial Play in Interactive Environments

NeurIPS 2025poster

Self-play constitutes a fundamental paradigm for autonomous skill acquisition, whereby agents iteratively enhance their capabilities through self-directed environmental exploration. Conventional self-play frameworks exploit agent symmetry within zero-sum competitive settings, yet this approach prove…

Cited by 0SourceScholar
2025

PrimHOI: Compositional Human-Object Interaction via Reusable Primitives

ICCV 2025accepted

Synthesizing realistic Human-Object Interaction (HOI) motions is essential for creating believable digital characters and intelligent robots. Existing approaches rely on data-intensive learning models that struggle with the compositional structure of daily HOI motions, particularly for complex multi…

Cited by 0SourcePDFScholar
2025

Taccel: Scaling Up Vision-based Tactile Robotics via High-performance GPU Simulation

NeurIPS 2025spotlight

Tactile sensing is crucial for achieving human-level robotic capabilities in manipulation tasks. As a promising solution, Vision-based Tactile Sensors (VBTSs) offer high spatial resolution and cost-effectiveness, but present unique challenges in robotics for their complex physical characteristics an…

Cited by 0SourcecodeScholar
2024

Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations

IROS 2024poster

Autonomous robotic systems capable of learning novel manipulation tasks are poised to transform industries from manufacturing to service automation. However, current methods (e.g., VIP and R3M) still face significant hurdles, notably the domain gap among robotic embodiments and the sparsity of succe…

Cited by 15SourcecodeScholar
2024

AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents

CVPR 2024poster

Traditional approaches in physics-based motion generation centered around imitation learning and reward shaping often struggle to adapt to new scenarios. To tackle this limitation we propose AnySkill a novel hierarchical method that learns physically plausible interactions following open-vocabulary…

Cited by 20SourcePDFScholar
2024

MiniTac: An Ultra-Compact 8 mm Vision-Based Tactile Sensor for Enhanced Palpation in Robot-Assisted Minimally Invasive Surgery

RA-L 2024

Robot-assisted minimally invasive surgery (RAMIS) provides substantial benefits over traditional open and laparoscopic methods. However, a significant limitation of robot-assisted minimally invasive surgery (RAMIS) is the surgeon's inability to palpate tissues, a crucial technique for examining tiss

Cited by 12SourceScholar
2024

Move as You Say Interact as You Can: Language-guided Human Motion Generation with Scene Affordance

CVPR 2024highlight

Despite significant advancements in text-to-motion synthesis generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models capable of jointly modeling natural language 3D scenes and huma…

2024

Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping

CoRL 2024poster

One-shot transfer of dexterous grasps to novel scenes with object and context variations has been a challenging problem. While distilled feature fields from large vision models have enabled semantic correspondences across 3D scenes, their features are point-based and restricted to object surfaces, l…

Cited by 2SourceScholar
2024

Neural-Symbolic Recursive Machine for Systematic Generalization

ICLR 2024poster

Current learning models often struggle with human-like systematic generalization, particularly in learning compositional rules from limited data and extrapolating them to novel combinations. We introduce the Neural-Symbolic Recursive Ma- chine ( NSR), whose core is a Grounded Symbol System ( GSS), a…

Cited by 9SourcePDFScholar
2024

PhyRecon: Physically Plausible Neural Scene Reconstruction

NeurIPS 2024poster

We address the issue of physical implausibility in multi-view neural reconstruction. While implicit representations have gained popularity in multi-view 3D reconstruction, previous work struggles to yield physically plausible results, limiting their utility in domains requiring rigorous physical acc…

Cited by 10SourcePDFScholar
2024

PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments

IROS 2024poster

Robotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are limited in their adaptability across different object categorie…

Cited by 4SourceScholar
2024

Scaling Up Dynamic Human-Scene Interaction Modeling

CVPR 2024highlight

Confronting the challenges of data scarcity and advanced motion synthesis in human-scene interaction modeling we introduce the TRUMANS dataset alongside a novel HSI motion synthesis method. TRUMANS stands as the most comprehensive motion-captured HSI dataset currently available encompassing over 15…

Cited by 54SourcePDFScholar
2024

SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation

ICLR 2024poster

Humans demonstrate remarkable skill in transferring manipulation abilities across objects of varying shapes, poses, and appearances, a capability rooted in their understanding of semantic correspondences between different instances. To equip robots with a similar high-level comprehension, we present…

Cited by 16SourcePDFScholar
2024

Temporal Spiking Neural Networks with Synaptic Delay for Graph Reasoning

ICML 2024poster

Spiking neural networks (SNNs) are investigated as biologically inspired models of neural computation, distinguished by their computational capability and energy efficiency due to precise spiking times and sparse spikes with event-driven computation. A significant question is how SNNs can emulate hu…

2024

Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation

AAAI 2024technical

New Natural Langauge Process~(NLP) benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present Xiezhi, the most comprehensive evaluation suite designed to assess holistic domain knowledge.Xiezhi comprises multiple-choice questions across 516 diverse…

2023

A Minimalist Dataset for Systematic Generalization of Perception, Syntax, and Semantics

ICLR 2023top-25%

Inspired by humans' exceptional ability to master arithmetic and generalize to new problems, we present a new dataset, HINT, to examine machines' capability of learning generalizable concepts at three levels: perception, syntax, and semantics. In HINT, machines are tasked with learning how concepts…

Cited by 6SourcePDFScholar
2023

ChimpACT: A Longitudinal Dataset for Understanding Chimpanzee Behaviors

NeurIPS 2023poster

Understanding the behavior of non-human primates is crucial for improving animal welfare, modeling social behavior, and gaining insights into distinctively human and phylogenetically shared behaviors. However, the lack of datasets on non-human primate behavior hinders in-depth exploration of primate…

2023

Diffusion-Based Generation, Optimization, and Planning in 3D Scenes

CVPR 2023poster

We introduce SceneDiffuser, a conditional generative model for 3D scene understanding. SceneDiffuser provides a unified model for solving scene-conditioned generation, optimization, and planning. In contrast to prior works, SceneDiffuser is intrinsically scene-aware, physics-based, and goal-oriented…

2023

Evaluating and Inducing Personality in Pre-trained Language Models

NeurIPS 2023spotlight

Standardized and quantified evaluation of machine behaviors is a crux of understanding LLMs. In this study, we draw inspiration from psychometric studies by leveraging human personality theory as a tool for studying machine behaviors. Originating as a philosophical quest for human behaviors, the stu…

Cited by 143SourcePDFScholar
2023

Full-Body Articulated Human-Object Interaction

ICCV 2023poster

Fine-grained capture of 3D Human-Object Interactions (HOIs) boosts human activity understanding and facilitates various downstream visual tasks. Prior models mostly assume that humans interact with rigid objects using only a few body parts, limiting their scope. In this paper, we address the challen…

Cited by 59PDFcodeScholar
2023

GenDexGrasp: Generalizable Dexterous Grasping

ICRA 2023poster

Generating dexterous grasping has been a long-standing and challenging robotic task. Despite recent progress, existing methods primarily suffer from two issues. First, most prior art focuses on a specific type of robot hand, lacking generalizable capability of handling unseen ones. Second, prior art…

Cited by 80SourcecodeScholar
2023

Learning a Causal Transition Model for Object Cutting

IROS 2023poster

Cutting objects into desired fragments is challenging for robots due to the spatially unstructured nature of fragments and the complex one-to-many object fragmentation caused by actions. We present a novel approach to model object fragmentation using an attributed stochastic grammar. This grammar ab…

Cited by 2SourceScholar
2023

MEWL: Few-shot multimodal word learning with referential uncertainty

ICML 2023poster

Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning…

2023

On the Complexity of Bayesian Generalization

ICML 2023poster

We examine concept generalization at a large scale in the natural visual spectrum. Established computational modes (*i.e.*, rule-based or similarity-based) are primarily studied isolated, focusing on confined and abstract problem spaces. In this work, we study these two modes when the *problem space…

2023

Part-level Scene Reconstruction Affords Robot Interaction

IROS 2023poster

Existing methods for reconstructing interactive scenes primarily focus on replacing reconstructed objects with CAD models retrieved from a limited database, resulting in significant discrepancies between the reconstructed and observed scenes. To address this issue, our work introduces a part-level r…

Cited by 9SourceScholar
2023

ProBio: A Protocol-guided Multimodal Dataset for Molecular Biology Lab

NeurIPS 2023poster

The challenge of replicating research results has posed a significant impediment to the field of molecular biology. The advent of modern intelligent systems has led to notable progress in various domains. Consequently, we embarked on an investigation of intelligent monitoring systems as a means of t…

Cited by 3SourcePDFScholar
2023

Rearrange Indoor Scenes for Human-Robot Co-Activity

ICRA 2023poster

We present an optimization-based framework for rearranging indoor furniture to accommodate human-robot co-activities better. The rearrangement aims to afford sufficient accessible space for robot activities without compromising everyday human activities. To retain human activities, our algorithm pre…

Cited by 7SourceScholar
2023

Sequential Manipulation Planning for Over-Actuated Unmanned Aerial Manipulators

IROS 2023poster

We investigate the sequential manipulation planning problem for unmanned aerial manipulators (UAMs). Unlike prior work that primarily focuses on one-step manipulation tasks, sequential manipulations require coordinated motions of a UAM's floating base, the manipulator, and the object being manipulat…

Cited by 17SourceScholar
2023

Understanding Embodied Reference with Touch-Line Transformer

ICLR 2023poster

We study embodied reference understanding, the task of locating referents using embodied gestural signals and language references. Human studies have revealed that, contrary to popular belief, objects referred to or pointed to do not lie on the elbow-wrist line, but rather on the so-called virtual t…

2023

X-VoE: Measuring eXplanatory Violation of Expectation in Physical Events

ICCV 2023oral

Intuitive physics is pivotal for human understanding of the physical world, enabling prediction and interpretation of events even in infancy. Nonetheless, replicating this level of intuitive physics in artificial intelligence (AI) remains a formidable challenge. This study introduces X-VoE, a compre…

Cited by 4PDFcodeScholar
2022

Combined Magnetic Field Decoupling and Disturbance Rejection Control of Microrobots Based on Extended State Observer

RA-L 2022

Magnetic microrobots are potentially used in various biomedical applications due to their distinguished properties in bio-related manipulation such as minimally invasive and accessible in complex bio-environment. Precise path tracking control of magnetic microrobots in complex interference is an imp

Cited by 12SourceScholar
2022

Downwash-aware Control Allocation for Over-actuated UAV Platforms

IROS 2022poster

Tracking position and orientation independently affords more agile maneuver for over-actuated multirotor Unmanned Aerial Vehicles (UAVs) while introducing undesired downwash effects; downwash flows generated by thrust generators may counteract others due to close proximity, which significantly threa…

Cited by 17SourceScholar
2022

Emergent Graphical Conventions in a Visual Communication Game

NeurIPS 2022accept

Humans communicate with graphical sketches apart from symbolic languages. Primarily focusing on the latter, recent studies of emergent communication overlook the sketches; they do not account for the evolution process through which symbolic sign systems emerge in the trade-off between iconicity and…

Cited by 19SourcePDFScholar
2022

HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes

NeurIPS 2022accept

Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characters of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality and lack semantics. To fill in the gap, we propose a large-scale an…

2022

Latent Diffusion Energy-Based Model for Interpretable Text Modelling

ICML 2022spotlight

Latent space Energy-Based Models (EBMs), also known as energy-based priors, have drawn growing interests in generative modeling. Fueled by its flexibility in the formulation and strong modeling power of the latent space, recent works built upon it have made interesting attempts aiming at the interpr…

2022

Learning Algebraic Representation for Systematic Generalization in Abstract Reasoning

ECCV 2022poster

"Is intelligence realized by connectionist or classicist? While connectionist approaches have achieved superhuman performance, there has been growing evidence that such task-specific superiority is particularly fragile in systematic generalization. This observation lies in the central debate between…

Cited by 37SourcePDFScholar
2022

Sequential Manipulation Planning on Scene Graph

IROS 2022poster

We devise a 3D scene graph representation, contact graph+ (cg+), for efficient sequential manipulation planning. Augmented with predicate-like attributes, this contact graph-based representation abstracts scene layouts with succinct geometric information and valid robot-scene interactions. Goal conf…

Cited by 35SourcecodeScholar
2022

Synthesizing Diverse and Physically Stable Grasps With Arbitrary Hand Structures Using Differentiable Force Closure Estimator

RA-L 2022

Existing grasp synthesis methods are either analytical or data-driven. The former one is oftentimes limited to specific application scope. The latter one depends heavily on demonstrations, thus suffers from generalization issues; <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="ht

Cited by 155SourceScholar
2021

Abstract Spatial-Temporal Reasoning via Probabilistic Abduction and Execution

CVPR 2021poster

Spatial-temporal reasoning is a challenging task in Artificial Intelligence (AI) due to its demanding but unique nature: a theoretic requirement on representing and reasoning based on spatial-temporal knowledge in mind, and an applied requirement on a high-level cognitive system capable of navigatin…

Cited by 75PDFScholar
2021

Communicative Learning with Natural Gestures for Embodied Navigation Agents with Human-in-the-Scene

IROS 2021poster

Human-robot collaboration is an essential re-search topic in artificial intelligence (AI), enabling researchers to devise cognitive AI systems and affords an intuitive means for users to interact with the robot. Of note, communication plays a central role. To date, prior studies in embodied agent na…

Cited by 24SourceScholar
2021

Congestion-aware Multi-agent Trajectory Prediction for Collision Avoidance

ICRA 2021poster

Predicting agents’ future trajectories plays a crucial role in modern AI systems, yet it is challenging due to intricate interactions exhibited in multi-agent systems, especially when it comes to collision avoidance. To address this challenge, we propose to learn congestion patterns as contextual cu…

Cited by 51SourcecodeScholar
2021

Consolidating Kinematic Models to Promote Coordinated Mobile Manipulations

IROS 2021poster

We construct a Virtual Kinematic Chain (VKC) that readily consolidates the kinematics of the mobile base, the arm, and the object to be manipulated in mobile manipulations. Accordingly, a mobile manipulation task is represented by altering the state of the constructed VKC, which can be converted to…

Cited by 21SourcecodeScholar
2021

Efficient Task Planning for Mobile Manipulation: a Virtual Kinematic Chain Perspective

IROS 2021poster

We present a Virtual Kinematic Chain (VKC) perspective, a simple yet effective method, to improve task planning efficacy for mobile manipulation. By consolidating the kinematics of the mobile base, the arm, and the object being manipulated collectively as a whole, this novel VKC perspective naturall…

Cited by 21SourcecodeScholar
2021

Learning Triadic Belief Dynamics in Nonverbal Communication From Videos

CVPR 2021poster

Humans possess a unique social cognition capability; nonverbal communication can convey rich social information among agents. In contrast, such crucial social characteristics are mostly missing in the existing scene understanding literature. In this paper, we incorporate different nonverbal communic…

Cited by 27PDFcodeScholar
2021

Reconstructing Interactive 3D Scenes by Panoptic Mapping and CAD Model Alignments

ICRA 2021poster

In this paper, we rethink the problem of scene reconstruction from an embodied agent’s perspective: While the classic view focuses on the reconstruction accuracy, our new perspective emphasizes the underlying functions and constraints such that the reconstructed scenes provide actionable information…

Cited by 32SourcecodeScholar
2021

Spatio-Temporal Self-Supervised Representation Learning for 3D Point Clouds

ICCV 2021poster

To date, various 3D scene understanding tasks still lack practical and generalizable pre-trained models, primarily due to the intricate nature of 3D scene understanding tasks and their immerse variations due to camera views, lighting, occlusions, etc. In this paper, we tackle this immanent challenge…

Cited by 248PDFcodeScholar
2021

Unsupervised Foreground Extraction via Deep Region Competition

NeurIPS 2021poster

We present Deep Region Competition (DRC), an algorithm designed to extract foreground objects from images in a fully unsupervised manner. Foreground extraction can be viewed as a special case of generic image segmentation that focuses on identifying and disentangling objects from the background. In…

Cited by 41SourcePDFScholar
2021

YouRefIt: Embodied Reference Understanding With Language and Gesture

ICCV 2021poster

We study the machine's understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding multimodal cues with perspective-taking to identify which object is being…

Cited by 46PDFScholar
2020

Congestion-aware Evacuation Routing using Augmented Reality Devices

ICRA 2020poster

We present a congestion-aware routing solution for indoor evacuation, which produces real-time individual-customized evacuation routes among multiple destinations while keeping tracks of all evacuees’ locations. A population density map, obtained on-the-fly by aggregating locations of evacuees from…

Cited by 17SourceScholar
2020

Graph-based Hierarchical Knowledge Representation for Robot Task Transfer from Virtual to Physical World

IROS 2020poster

We study the hierarchical knowledge transfer problem using a cloth-folding task, wherein the agent is first given a set of human demonstrations in the virtual world using an Oculus Headset, and later transferred and validated on a physical Baxter robot. We argue that such an intricate robot task tra…

Cited by 23SourceScholar
2020

Human-Robot Interaction in a Shared Augmented Reality Workspace

IROS 2020poster

We design and develop a new shared Augmented Reality (AR) workspace for Human-Robot Interaction (HRI), which establishes a bi-directional communication between human agents and robots. In a prototype system, the shared AR workspace enables a shared perception, so that a physical robot not only perce…

Cited by 39SourceScholar
2020

Joint Inference of States, Robot Knowledge, and Human (False-)Beliefs

ICRA 2020poster

Aiming to understand how human (false-)belief— a core socio-cognitive ability—would affect human interactions with robots, this paper proposes to adopt a graphical model to unify the representation of object states, robot knowledge, and human (false-)beliefs. Specifically, a parse graph (pg) is lear…

Cited by 27SourceScholar
2020

LEMMA: A Multi-view Dataset for LEarning Multi-agent Multi-task Activities

ECCV 2020poster

The ability to understand and interpret human actions is a long-standing challenge and a critical indicator of perception in artificial intelligence. However, a few imperative components of daily human activities are largely missed in prior literature, including the goal-directed actions, concurrent…

2019

High-Fidelity Grasping in Virtual Reality using a Glove-based System

ICRA 2019poster

This paper presents a design that jointly provides hand pose sensing, hand localization, and haptic feedback to facilitate real-time stable grasps in Virtual Reality (VR). The design is based on an easy-to-replicate glove-based system that can reliably perform (i) a high-fidelity hand pose sensing i…

Cited by 81SourceScholar
2019

Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical Commonsense

ICCV 2019poster

We propose a new 3D holistic++ scene understanding problem, which jointly tackles two tasks from a single-view image: (i) holistic scene parsing and reconstruction---3D estimations of object bounding boxes, camera pose, and room layout, and (ii) 3D human pose estimation. The intuition behind is to l…

Cited by 145PDFScholar
2019

Learning Perceptual Inference by Contrasting

NeurIPS 2019spotlight

“Thinking in pictures,” [1] i.e., spatial-temporal reasoning, effortless and instantaneous for humans, is believed to be a significant ability to perform logical induction and a crucial factor in the intellectual history of technology development. Modern Artificial Intelligence (AI), fueled by massi…

2019

Learning Virtual Grasp with Failed Demonstrations via Bayesian Inverse Reinforcement Learning

IROS 2019poster

We propose Bayesian Inverse Reinforcement Learning with Failure (BIRLF), which makes use of failed demonstrations that were often ignored or filtered in previous methods due to the difficulties to incorporate them in addition to the successful ones. Specifically, we leverage halfspaces derived from…

Cited by 27SourceScholar
2019

PerspectiveNet: 3D Object Detection from a Single RGB Image via Perspective Points

NeurIPS 2019poster

Detecting 3D objects from a single RGB image is intrinsically ambiguous, thus requiring appropriate prior knowledge and intermediate representations as constraints to reduce the uncertainties and improve the consistencies between the 2D image plane and the 3D world coordinate. To address this challe…

2019

RAVEN: A Dataset for Relational and Analogical Visual REasoNing

CVPR 2019poster

Dramatic progress has been witnessed in basic vision tasks involving low-level perception, such as object recognition, detection, and tracking. Unfortunately, there is still enormous performance gap between artificial vision systems and human intelligence in terms of higher-level vision problems, es…

Cited by 357PDFScholar
2019

Self-Supervised Incremental Learning for Sound Source Localization in Complex Indoor Environment

ICRA 2019poster

This paper presents an incremental learning framework for mobile robots localizing the human sound source using a microphone array in a complex indoor environment consisting of multiple rooms. In contrast to conventional approaches that leverage direction-of-arrival (DOA) estimation, the framework a…

Cited by 14SourceScholar
2018

Cooperative Holistic Scene Understanding: Unifying 3D Object, Layout, and Camera Pose Estimation

NeurIPS 2018poster

Holistic 3D indoor scene understanding refers to jointly recovering the i) object bounding boxes, ii) room layout, and iii) camera pose, all in 3D. The existing methods either are ineffective or only tackle the problem partially. In this paper, we propose an end-to-end model that simultaneously solv…

2018

Holistic 3D Scene Parsing and Reconstruction from a Single RGB Image

ECCV 2018poster

We propose a computational framework to jointly parse a single RGB image and reconstruct a holistic 3D configuration composed by a set of CAD models using a stochastic grammar model. Specifically, we introduce a Holistic Scene Grammar (HSG) to represent the 3D scene structure, which characterizes a…

Cited by 171SourcePDFScholar
2018

Human-Centric Indoor Scene Synthesis Using Stochastic Grammar

CVPR 2018poster

We present a human-centric method to sample and synthesize 3D room layouts and 2D images thereof, for the purpose of obtaining large-scale 2D/3D image data with the perfect per-pixel ground truth. An attributed spatial And-Or graph (S-AOG) is proposed to represent indoor scenes. The S-AOG is a proba…

2018

Interactive Robot Knowledge Patching Using Augmented Reality

ICRA 2018poster

We present a novel Augmented Reality (AR) approach, through Microsoft HoloLens, to address the challenging problems of diagnosing, teaching, and patching interpretable knowledge of a robot. A Temporal And-Or graph (T-AOG) of opening bottles is learned from human demonstration and programmed to the r…

Cited by 83SourceScholar
2018

Unsupervised Learning of Hierarchical Models for Hand-Object Interactions

ICRA 2018poster

Contact forces of the hand are visually unobservable, but play a crucial role in understanding hand-object interactions. In this paper, we propose an unsupervised learning approach for manipulation event segmentation and manipulation event parsing. The proposed framework incorporates hand pose kinem…

Cited by 15SourceScholar
2017

A glove-based system for studying hand-object manipulation via joint pose and force sensing

IROS 2017poster

We present a design of an easy-to-replicate glove-based system that can reliably perform simultaneous hand pose and force sensing in real time, for the purpose of collecting human hand data during fine manipulative actions. The design consists of a sensory glove that is capable of jointly collecting…

Cited by 70SourceScholar
2017

Feeling the force: Integrating force and pose for fluent discovery through imitation learning to open medicine bottles

IROS 2017poster

Learning complex robot manipulation policies for real-world objects is challenging, often requiring significant tuning within controlled environments. In this paper, we learn a manipulation model to execute tasks with multiple stages and variable structure, which typically are not suitable for most…

Cited by 78SourceScholar
2016

Inferring Forces and Learning Human Utilities From Videos

CVPR 2016oral

We propose a notion of affordance that takes into account physical quantities generated when the human body interacts with real-world objects, and introduce a learning framework that incorporates the concept of human utilities, which in our opinion provides a deeper and finer-grained account not onl…

Cited by 113PDFScholar
2015

Understanding Tools: Task-Oriented Object Modeling, Learning and Recognition

CVPR 2015poster

In this paper, we present a new framework - task-oriented modeling, learning and recognition which aims at understanding the underlying functions, physics and causality in using objects as "tools". Given a task, such as, cracking a nut or painting a wall, we represent each object, e.g. a hammer or…

Cited by 224SourcePDFScholar