← Search

Song-Chun Zhu

158 accepted papers

2026

Aegis: Automated Error Generation and Identification for Multi-Agent Systems

ICLR 2026poster

Large language model based multi-agent systems (MAS) have unlocked significant advancements in tackling complex problems, but their increasing capability introduces a structural fragility that makes them difficult to debug. A key obstacle to improving their reliability is the severe scarcity of larg…

Cited by 0SourceScholar
2026

G4Splat: Geometry-Guided Gaussian Splatting with Generative Prior

ICLR 2026poster

Despite recent advances in leveraging generative prior from pre-trained diffusion models for 3D scene reconstruction, existing methods still face two critical limitations. First, due to the lack of reliable geometric supervision, they struggle to produce high-quality reconstructions even in observed…

Cited by 0SourcecodeScholar
2026

Learning AND–OR Templates for Compositional Representation in Art and Design

ICLR 2026poster

This work proposes a compositional AND–OR template for art and design that encodes the part–relation–geometry organization of images in a structured and interpretable form. Within a maximum-entropy log-linear model, we define a unified consistency score as log-likelihood gain against a reference dis…

Cited by 0SourceScholar
2026

Lifting Unlabeled Internet-level Data for 3D Scene Understanding

CVPR 2026

Annotated 3D scene data is scarce and expensive to acquire, while abundant unlabeled videos are readily available on the internet. In this paper, we demonstrate that carefully designed data engines can leverage web-curated, unlabeled videos to automatically generate training data, to facilitate end-

Cited by 0SourcecodeScholar
2026

M3Bench: Benchmarking Whole-Body Motion Generation for Mobile Manipulation in 3D Scenes

ICRA 2026poster

We propose M3Bench, a new benchmark for whole-body motion generation in mobile manipulation tasks. Given a 3D scene context, M3Bench requires an embodied agent to reason about its configuration, environmental constraints, and task objectives to generate coordinated whole-body motion trajectories for…

2026

MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning

ICLR 2026poster

Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image generation either restrict reasoning to a single modality (image or text) or rely on high-quality reasoning data for fine-t…

Cited by 0SourcecodeScholar
2026

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

ICML 2026poster

We introduce **Native Parallel Reasoner (NPR)**, a teacher-free framework that enables Large Language Models (LLMs) to self-evolve genuine parallel reasoning capabilities. NPR transforms the model from sequential emulation to native parallel cognition through three key innovations: 1) a **self-disti…

Cited by 0SourceScholar
2026

TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents

AAAI 2026technical

Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. However, a major challenge in developing generalized GUI agents is the lack of sufficient trajectory data across various o

Cited by 0SourcePDFScholar
2026

Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation

AAAI 2026technical

Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal an

Cited by 0SourcePDFScholar
2025

Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting

ICLR 2025poster

Building interactable replicas of articulated objects is a key challenge in computer vision. Existing methods often fail to effectively integrate information across different object states, limiting the accuracy of part-mesh reconstruction and part dynamics modeling, particularly for complex multi-p…

Cited by 0SourcePDFScholar
2025

ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models

CoRL 2025poster

Learning real-world robotic manipulation is challenging, particularly when limited demonstrations are available. Existing methods for few-shot manipulation often rely on simulation-augmented data or pre-built modules like grasping and pose estimation, which struggle with sim-to-real gaps and lack ex…

Cited by 0SourceScholar
2025

Decompositional Neural Scene Reconstruction with Generative Diffusion Prior

CVPR 2025poster

Decompositional reconstruction of 3D scenes, with complete shapes and detailed texture of all objects within, is intriguing for downstream applications but remains challenging, particularly with sparse views as input. Recent approaches incorporate semantic or geometric regularization to address this…

2025

Differentiable Information Enhanced Model-Based Reinforcement Learning

AAAI 2025technical

Differentiable environments have heralded new possibilities for learning control policies by offering rich differentiable information that facilitates gradient-based methods. In comparison to prevailing model-free reinforcement learning approaches, model-based reinforcement learning (MBRL) methods e…

Cited by 0SourcePDFScholar
2025

Enhancing LLM-Based Social Bot via an Adversarial Learning Framework

EMNLP 2025

Developing Large Language Model (LLM) agents that exhibit human-like behavior, encompassing not only individual heterogeneity rooted in unique user profiles but also adaptive response to socially connected neighbors, is a significant research challenge. Social media platforms, with their diverse use

2025

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning

NeurIPS 2025poster

Multimodal agents, which integrate a controller (e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks. Existing approaches for training these agents, both supervised fine-tuning and reinforcement learning, depend on extens…

Cited by 0SourceScholar
2025

M${}{3}$Bench: Benchmarking Whole-Body Motion Generation for Mobile Manipulation in 3D Scenes

RA-L 2025

We propose M <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">${}^{3}$</tex-math></inline-formula> Bench, a new benchmark for whole-body motion generation in mobile manipulation tasks. Given a 3D scene context, M <in

Cited by 4SourceScholar
2025

Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

ICLR 2025spotlight

The advancement of large language models (LLMs) prompts the development of multi-modal agents, which are used as a controller to call external tools, providing a feasible way to solve practical tasks. In this paper, we propose a multi-modal agent tuning method that automatically generates multi-moda…

Cited by 5SourcePDFScholar
2025

ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection

ACL 2025finding

We present a novel pipeline, ReflectEvo, to demonstrate that small language models (SLMs) can enhance meta introspection through reflection learning. This process iteratively generates self-reflection for self-training, fostering a continuous and self-evolving process. Leveraging this pipeline, we c…

Cited by 0SourcePDFScholar
2025

Social World Model-Augmented Mechanism Design Policy Learning

NeurIPS 2025poster

Designing adaptive mechanisms to align individual and collective interests remains a central challenge in artificial social intelligence. Existing methods often struggle with modeling heterogeneous agents possessing persistent latent traits (e.g., skills, preferences) and dealing with complex multi-…

Cited by 0SourceScholar
2025

Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis

CVPR 2025poster

Existing 3D vision-language (3D-VL) benchmarks fall short in evaluating 3D-VL models, creating a "mist" that obscures rigorous insights into model capabilities and 3D-VL tasks. This mist persists due to three key limitations. First, flawed test data, like ambiguous referential text in the grounding…

2025

World Models Should Prioritize the Unification of Physical and Social Dynamics

NeurIPS 2025poster

World models, which explicitly learn environmental dynamics to lay the foundation for planning, reasoning, and decision-making, are rapidly advancing in predicting both physical dynamics and aspects of social behavior, yet predominantly in separate silos. This division results in a systemic failure…

Cited by 0SourceScholar
2024

AdaSociety: An Adaptive Environment with Social Structures for Multi-Agent Decision-Making

NeurIPS 2024poster

Traditional interactive environments limit agents' intelligence growth with fixed tasks. Recently, single-agent environments address this by generating new tasks based on agent actions, enhancing task diversity. We consider the decision-making problem in multi-agent settings, where tasks are further…

2024

Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations

IROS 2024poster

Autonomous robotic systems capable of learning novel manipulation tasks are poised to transform industries from manufacturing to service automation. However, current methods (e.g., VIP and R3M) still face significant hurdles, notably the domain gap among robotic embodiments and the sparsity of succe…

Cited by 15SourcecodeScholar
2024

An Embodied Generalist Agent in 3D World

ICML 2024poster

Leveraging massive knowledge from large language models (LLMs), recent machine learning models show notable successes in general-purpose task solving in diverse domains such as computer vision and robotics. However, several significant challenges remain: (i) most of these models rely on 2D images ye…

2024

Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real World

ICLR 2024poster

We introduce Bongard-OpenWorld, a new benchmark for evaluating real-world few-shot reasoning for machine vision. It originates from the classical Bongard Problems (BPs): Given two sets of images (positive and negative), the model needs to identify the set that query images belong to by inducing the…

2024

CLOVA: A Closed-LOop Visual Assistant with Tool Usage and Update

CVPR 2024poster

Utilizing large language models (LLMs) to compose off-the-shelf visual tools represents a promising avenue of research for developing robust visual assistants capable of addressing diverse visual tasks. However these methods often overlook the potential for continual learning typically by freezing t…

Cited by 29SourcePDFScholar
2024

CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making Agents

ICLR 2024spotlight

The generalization of decision-making agents encompasses two fundamental elements: learning from past experiences and reasoning in novel contexts. However, the predominant emphasis in most interactive environments is on learning, often at the expense of complexity in reasoning. In this paper, we int…

2024

Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning

ICML 2024poster

Despite the recent successes of multi-agent reinforcement learning (MARL) algorithms, efficiently adapting to co-players in mixed-motive environments remains a significant challenge. One feasible approach is to hierarchically model co-players' behavior based on inferring their characteristics. Howev…

Cited by 1SourcePDFScholar
2024

FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models

NeurIPS 2024poster

Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spon…

Cited by 4SourcePDFScholar
2024

Fast Peer Adaptation with Context-aware Exploration

ICML 2024poster

Fast adapting to unknown peers (partners or opponents) with different strategies is a key challenge in multi-agent games. To do so, it is crucial for the agent to probe and identify the peer’s strategy efficiently, as this is the prerequisite for carrying out the best response in adaptation. However…

Cited by 2SourcePDFScholar
2024

INTERPRET: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning

RSS 2024poster

Learning abstract state representations and knowledge is crucial for long-horizon robot planning. We present InterPreT, an LLM-powered framework for robots to learn symbolic predicates from language feedback of human non-experts during embodied interaction. The learned predicates provide relational…

2024

LLM3: Large Language Model-based Task and Motion Planning with Motion Failure Reasoning

IROS 2024poster

Conventional Task and Motion Planning (TAMP) approaches rely on manually designed interfaces connecting symbolic task planning with continuous motion generation. These domain-specific and labor-intensive modules are limited in addressing emerging tasks in real-world settings. Here, we present LLM3,…

Cited by 45SourcecodeScholar
2024

LangSuit·E: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments

ACL 2024findings

Recent advances in Large Language Models (LLMs) have shown inspiring achievements in constructing autonomous agents that rely onlanguage descriptions as inputs. However, it remains unclear how well LLMs can function as few-shot or zero-shot embodied agents in dynamic interactive environments. To add…

2024

Learning to Balance Altruism and Self-interest Based on Empathy in Mixed-Motive Games

NeurIPS 2024poster

Real-world multi-agent scenarios often involve mixed motives, demanding altruistic agents capable of self-protection against potential exploitation. However, existing approaches often struggle to achieve both objectives. In this paper, based on that empathic responses are modulated by learned social…

Cited by 0SourcePDFScholar
2024

Mars: Situated Inductive Reasoning in an Open-World Environment

NeurIPS 2024poster

Large Language Models (LLMs) trained on massive corpora have shown remarkable success in knowledge-intensive tasks. Yet, most of them rely on pre-stored knowledge. Inducing new general knowledge from a specific environment and performing reasoning with the acquired knowledge—situated inductive reaso…

Cited by 1SourcePDFScholar
2024

Neural-Symbolic Recursive Machine for Systematic Generalization

ICLR 2024poster

Current learning models often struggle with human-like systematic generalization, particularly in learning compositional rules from limited data and extrapolating them to novel combinations. We introduce the Neural-Symbolic Recursive Ma- chine ( NSR), whose core is a Grounded Symbol System ( GSS), a…

Cited by 9SourcePDFScholar
2024

PhyRecon: Physically Plausible Neural Scene Reconstruction

NeurIPS 2024poster

We address the issue of physical implausibility in multi-view neural reconstruction. While implicit representations have gained popularity in multi-view 3D reconstruction, previous work struggles to yield physically plausible results, limiting their utility in domains requiring rigorous physical acc…

Cited by 10SourcePDFScholar
2024

ProAgent: Building Proactive Cooperative Agents with Large Language Models

AAAI 2024technical

Building agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, whose policy generalization depends heavily on the diversity of teammates they int…

2024

RulE: Knowledge Graph Reasoning with Rule Embedding

ACL 2024findings

Knowledge graph reasoning is an important problem for knowledge graphs. In this paper, we propose a novel and principled framework called RulE (stands for Rule Embedding) to effectively leverage logical rules to enhance KG reasoning. Unlike knowledge graph embedding methods, RulE learns rule embeddi…

2023

A Minimalist Dataset for Systematic Generalization of Perception, Syntax, and Semantics

ICLR 2023top-25%

Inspired by humans' exceptional ability to master arithmetic and generalize to new problems, we present a new dataset, HINT, to examine machines' capability of learning generalizable concepts at three levels: perception, syntax, and semantics. In HINT, machines are tasked with learning how concepts…

Cited by 6SourcePDFScholar
2023

ARNOLD: A Benchmark for Language-Grounded Task Learning with Continuous States in Realistic 3D Scenes

ICCV 2023poster

Understanding the continuous states of objects is essential for task learning and planning in the real world. However, most existing task learning benchmarks assume discrete (e.g., binary) object states, which poses challenges for learning complex tasks and transferring learned policy from the simul…

Cited by 29PDFcodeScholar
2023

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

NeurIPS 2023poster

Large language models (LLMs) have achieved remarkable progress in solving various natural language processing tasks due to emergent reasoning abilities. However, LLMs have inherent limitations as they are incapable of accessing up-to-date information (stored on the Web or in task-specific knowledge…

Cited by 456SourcePDFScholar
2023

Diffusion-Based Generation, Optimization, and Planning in 3D Scenes

CVPR 2023poster

We introduce SceneDiffuser, a conditional generative model for 3D scene understanding. SceneDiffuser provides a unified model for solving scene-conditioned generation, optimization, and planning. In contrast to prior works, SceneDiffuser is intrinsically scene-aware, physics-based, and goal-oriented…

2023

Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

ICLR 2023poster

Mathematical reasoning, a core ability of human intelligence, presents unique challenges for machines in abstract thinking and logical reasoning. Recent large pre-trained language models such as GPT-3 have achieved remarkable progress on mathematical reasoning tasks written in text form, such as mat…

2023

Evaluating and Inducing Personality in Pre-trained Language Models

NeurIPS 2023spotlight

Standardized and quantified evaluation of machine behaviors is a crux of understanding LLMs. In this study, we draw inspiration from psychometric studies by leveraging human personality theory as a tool for studying machine behaviors. Originating as a philosophical quest for human behaviors, the stu…

Cited by 143SourcePDFScholar
2023

Learning Energy-Based Prior Model with Diffusion-Amortized MCMC

NeurIPS 2023poster

Latent space EBMs, also known as energy-based priors, have drawn growing interests in the field of generative modeling due to its flexibility in the formulation and strong modeling power of the latent space. However, the common practice of learning latent space EBMs with non-convergent short-run MCM…

2023

Learning a Causal Transition Model for Object Cutting

IROS 2023poster

Cutting objects into desired fragments is challenging for robots due to the spatially unstructured nature of fragments and the complex one-to-many object fragmentation caused by actions. We present a novel approach to model object fragmentation using an attributed stochastic grammar. This grammar ab…

Cited by 2SourceScholar
2023

Learning non-Markovian Decision-Making from State-only Sequences

NeurIPS 2023poster

Conventional imitation learning assumes access to the actions of demonstrators, but these motor signals are often non-observable in naturalistic settings. Additionally, sequential decision-making behaviors in these settings can deviate from the assumptions of a standard Markov Decision Process (MDP)…

Cited by 9SourcePDFScholar
2023

On the Complexity of Bayesian Generalization

ICML 2023poster

We examine concept generalization at a large scale in the natural visual spectrum. Established computational modes (*i.e.*, rule-based or similarity-based) are primarily studied isolated, focusing on confined and abstract problem spaces. In this work, we study these two modes when the *problem space…

2023

Part-level Scene Reconstruction Affords Robot Interaction

IROS 2023poster

Existing methods for reconstructing interactive scenes primarily focus on replacing reconstructed objects with CAD models retrieved from a limited database, resulting in significant discrepancies between the reconstructed and observed scenes. To address this issue, our work introduces a part-level r…

Cited by 9SourceScholar
2023

Rearrange Indoor Scenes for Human-Robot Co-Activity

ICRA 2023poster

We present an optimization-based framework for rearranging indoor furniture to accommodate human-robot co-activities better. The rearrangement aims to afford sufficient accessible space for robot activities without compromising everyday human activities. To retain human activities, our algorithm pre…

Cited by 7SourceScholar
2023

SQA3D: Situated Question Answering in 3D Scenes

ICLR 2023poster

We propose a new task to benchmark scene understanding of embodied agents: Situated Question Answering in 3D Scenes (SQA3D). Given a scene context (e.g., 3D scan), SQA3D requires the tested agent to first understand its situation (position, orientation, etc.) in the 3D scene as described by text, th…

2023

X-VoE: Measuring eXplanatory Violation of Expectation in Physical Events

ICCV 2023oral

Intuitive physics is pivotal for human understanding of the physical world, enabling prediction and interpretation of events even in infancy. Nonetheless, replicating this level of intuitive physics in artificial intelligence (AI) remains a formidable challenge. This study introduces X-VoE, a compre…

Cited by 4PDFcodeScholar
2022

COAT: Measuring Object Compositionality in Emergent Representations

ICML 2022spotlight

Learning representations that can decompose a multi-object scene into its constituent objects and recompose them flexibly is desirable for object-oriented reasoning and planning. Built upon object masks in the pixel space, existing metrics for objectness can only evaluate generative models with an o…

Cited by 9SourcePDFScholar
2022

EgoTaskQA: Understanding Human Tasks in Egocentric Videos

NeurIPS 2022accept

Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, their effects on object states (\ie, state changes), and their causal dependencie…

2022

Emergent Graphical Conventions in a Visual Communication Game

NeurIPS 2022accept

Humans communicate with graphical sketches apart from symbolic languages. Primarily focusing on the latter, recent studies of emergent communication overlook the sketches; they do not account for the evolution process through which symbolic sign systems emerge in the trade-off between iconicity and…

Cited by 19SourcePDFScholar
2022

Latent Diffusion Energy-Based Model for Interpretable Text Modelling

ICML 2022spotlight

Latent space Energy-Based Models (EBMs), also known as energy-based priors, have drawn growing interests in generative modeling. Fueled by its flexibility in the formulation and strong modeling power of the latent space, recent works built upon it have made interesting attempts aiming at the interpr…

2022

Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

NeurIPS 2022accept

When answering a question, humans utilize the information available across different modalities to synthesize a consistent and complete chain of thought (CoT). This process is normally a black box in the case of deep learning models like large-scale language models. Recently, science question benchm…

2022

Learning Algebraic Representation for Systematic Generalization in Abstract Reasoning

ECCV 2022poster

"Is intelligence realized by connectionist or classicist? While connectionist approaches have achieved superhuman performance, there has been growing evidence that such task-specific superiority is particularly fragile in systematic generalization. This observation lies in the central debate between…

Cited by 37SourcePDFScholar
2022

Learning Probabilistic Models from Generator Latent Spaces with Hat EBM

NeurIPS 2022accept

This work proposes a method for using any generator network as the foundation of an Energy-Based Model (EBM). Our formulation posits that observed images are the sum of unobserved latent variables passed through the generator network and a residual random variable that spans the gap between the gene…

2022

Learning V1 Simple Cells with Vector Representation of Local Content and Matrix Representation of Local Motion

AAAI 2022technical

This paper proposes a representational model for image pairs such as consecutive video frames that are related by local pixel displacements, in the hope that the model may shed light on motion perception in primary visual cortex (V1). The model couples the following two components: (1) the vector re…

Cited by 2SourcePDFScholar
2022

Learning from the Tangram to Solve Mini Visual Tasks

AAAI 2022technical

Current pre-training methods in computer vision focus on natural images in the daily-life context. However, abstract diagrams such as icons and symbols are common and important in the real world. We are inspired by Tangram, a game that requires replicating an abstract pattern from seven dissected sh…

2022

MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control

NeurIPS 2022accept

We introduce the Multi-Agent Tracking Environment (MATE), a novel multi-agent environment simulates the target coverage control problems in the real world. MATE hosts an asymmetric cooperative-competitive game consisting of two groups of learning agents--"cameras" and "targets"--with opposing intere…

2022

MCMC Should Mix: Learning Energy-Based Model with Neural Transport Latent Space MCMC

ICLR 2022poster

Learning energy-based model (EBM) requires MCMC sampling of the learned model as an inner loop of the learning algorithm. However, MCMC sampling of EBMs in high-dimensional data space is generally not mixing, because the energy function, which is usually parametrized by deep network, is highly multi…

Cited by 31SourcePDFScholar
2022

RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning

ICLR 2022poster

Reasoning about visual relationships is central to how humans interpret the visual world. This task remains challenging for current deep learning algorithms since it requires addressing three key technical problems jointly: 1) identifying object entities and their properties, 2) inferring semantic r…

2022

Sequential Manipulation Planning on Scene Graph

IROS 2022poster

We devise a 3D scene graph representation, contact graph+ (cg+), for efficient sequential manipulation planning. Augmented with predicate-like attributes, this contact graph-based representation abstracts scene layouts with succinct geometric information and valid robot-scene interactions. Goal conf…

Cited by 35SourcecodeScholar
2022

Show Me What You Can Do: Capability Calibration on Reachable Workspace for Human-Robot Collaboration

RA-L 2022

Aligning humans’ assessment of what a robot can do with its true capability is crucial for establishing a common ground between human and robot partners when they collaborate on a joint task. In this work, we propose an approach to calibrate humans’ estimate of a robot’s reachable workspace through

Cited by 4SourceScholar
2022

Synthesizing Diverse and Physically Stable Grasps With Arbitrary Hand Structures Using Differentiable Force Closure Estimator

RA-L 2022

Existing grasp synthesis methods are either analytical or data-driven. The former one is oftentimes limited to specific application scope. The latter one depends heavily on demonstrations, thus suffers from generalization issues; <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="ht

Cited by 155SourceScholar
2022

Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning

NeurIPS 2022accept

Achieving human-level dexterity is an important open problem in robotics. However, tasks of dexterous hand manipulation even at the baby level are challenging to solve through reinforcement learning (RL). The difficulty lies in the high degrees of freedom and the required cooperation among heterogen…

2022

ValueNet: A New Dataset for Human Value Driven Dialogue System

AAAI 2022technical

Building a socially intelligent agent involves many challenges, one of which is to teach the agent to speak guided by its value like a human. However, value-driven chatbots are still understudied in the area of dialogue systems. Most existing datasets focus on commonsense reasoning or social norm mo…

2021

Abstract Spatial-Temporal Reasoning via Probabilistic Abduction and Execution

CVPR 2021poster

Spatial-temporal reasoning is a challenging task in Artificial Intelligence (AI) due to its demanding but unique nature: a theoretic requirement on representing and reasoning based on spatial-temporal knowledge in mind, and an applied requirement on a high-level cognitive system capable of navigatin…

Cited by 75PDFScholar
2021

Congestion-aware Multi-agent Trajectory Prediction for Collision Avoidance

ICRA 2021poster

Predicting agents’ future trajectories plays a crucial role in modern AI systems, yet it is challenging due to intricate interactions exhibited in multi-agent systems, especially when it comes to collision avoidance. To address this challenge, we propose to learn congestion patterns as contextual cu…

Cited by 51SourcecodeScholar
2021

Consolidating Kinematic Models to Promote Coordinated Mobile Manipulations

IROS 2021poster

We construct a Virtual Kinematic Chain (VKC) that readily consolidates the kinematics of the mobile base, the arm, and the object to be manipulated in mobile manipulations. Accordingly, a mobile manipulation task is represented by altering the state of the constructed VKC, which can be converted to…

Cited by 21SourcecodeScholar
2021

CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA Generalization

EMNLP 2021main

One challenge in evaluating visual question answering (VQA) models in the cross-dataset adaptation setting is that the distribution shifts are multi-modal, making it difficult to identify if it is the shifts in visual or language features that play a key role. In this paper, we propose a semi-automa…

Cited by 30SourcePDFScholar
2021

Efficient Task Planning for Mobile Manipulation: a Virtual Kinematic Chain Perspective

IROS 2021poster

We present a Virtual Kinematic Chain (VKC) perspective, a simple yet effective method, to improve task planning efficacy for mobile manipulation. By consolidating the kinematics of the mobile base, the arm, and the object being manipulated collectively as a whole, this novel VKC perspective naturall…

Cited by 21SourcecodeScholar
2021

Generative PointNet: Deep Energy-Based Learning on Unordered Point Sets for 3D Generation, Reconstruction and Classification

CVPR 2021poster

We propose a generative model of unordered point sets, such as point clouds, in the forms of an energy-based model, where the energy function is parameterized by an input-permutation-invariant bottom-up neural network. The energy function learns a coordinate encoding of each point and then aggregate…

Cited by 91PDFcodeScholar
2021

IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

NeurIPS 2021poster

Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images. However, aside from natural images, abstract diagrams with semantic richness are still understudied in visual understanding and reasoning research. In this work, we introduce a new c…

Cited by 204SourcecodeScholar
2021

Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

ACL 2021long

Geometry problem solving has attracted much attention in the NLP community recently. The task is challenging as it requires abstract problem understanding and symbolic reasoning with axiomatic knowledge. However, current datasets are either small in scale or not publicly available. Thus, we construc…

2021

Iterative Teacher-Aware Learning

NeurIPS 2021poster

In human pedagogy, teachers and students can interact adaptively to maximize communication efficiency. The teacher adjusts her teaching method for different students, and the student, after getting familiar with the teacher’s instruction mechanism, can infer the teacher’s intention to learn faster.…

Cited by 15SourcePDFScholar
2021

Learning Cycle-Consistent Cooperative Networks via Alternating MCMC Teaching for Unsupervised Cross-Domain Translation

AAAI 2021technical

This paper studies the unsupervised cross-domain translation problem by proposing a generative framework, in which the probability distribution of each domain is represented by a generative cooperative network that consists of an energy-based model and a latent variable model. The use of generative…

Cited by 15SourcePDFScholar
2021

Learning Neural Representation of Camera Pose with Matrix Representation of Pose Shift via View Synthesis

CVPR 2021poster

How to efficiently represent camera pose is an essential problem in 3D computer vision, especially in tasks like camera pose regression and novel view synthesis. Traditionally, 3D position of the camera is represented by Cartesian coordinate and the orientation is represented by Euler angle or quate…

Cited by 9PDFcodeScholar
2021

Learning Triadic Belief Dynamics in Nonverbal Communication From Videos

CVPR 2021poster

Humans possess a unique social cognition capability; nonverbal communication can convey rich social information among agents. In contrast, such crucial social characteristics are mostly missing in the existing scene understanding literature. In this paper, we incorporate different nonverbal communic…

Cited by 27PDFcodeScholar
2021

Learning by Fixing: Solving Math Word Problems with Weak Supervision

AAAI 2021technical

Previous neural solvers of math word problems (MWPs) are learned with full supervision and fail to generate diverse solutions. In this paper, we address this issue by introducing a weakly-supervised paradigm for learning MWPs. Our method only requires the annotations of the final answers and can gen…

2021

Mind the Context: The Impact of Contextualization in Neural Module Networks for Grounding Visual Referring Expressions

EMNLP 2021main

Neural module networks (NMN) are a popular approach for grounding visual referring expressions. Prior implementations of NMN use pre-defined and fixed textual inputs in their module instantiation. This necessitates a large number of modules as they lack the ability to share weights and exploit assoc…

2021

On Path Integration of Grid Cells: Group Representation and Isotropic Scaling

NeurIPS 2021poster

Understanding how grid cells perform path integration calculations remains a fundamental problem. In this paper, we conduct theoretical analysis of a general representation model of path integration by grid cells, where the 2D self-position is encoded as a higher dimensional vector, and the 2D self-…

2021

Reconstructing Interactive 3D Scenes by Panoptic Mapping and CAD Model Alignments

ICRA 2021poster

In this paper, we rethink the problem of scene reconstruction from an embodied agent’s perspective: While the classic view focuses on the reconstruction accuracy, our new perspective emphasizes the underlying functions and constraints such that the reconstructed scenes provide actionable information…

Cited by 32SourcecodeScholar
2021

Robust Visual Reasoning via Language Guided Neural Module Networks

NeurIPS 2021poster

Neural module networks (NMN) are a popular approach for solving multi-modal tasks such as visual question answering (VQA) and visual referring expression recognition (REF). A key limitation in prior implementations of NMN is that the neural modules do not effectively capture the association between…

Cited by 27SourcePDFScholar
2021

SMART: A Situation Model for Algebra Story Problems via Attributed Grammar

AAAI 2021technical

Solving algebra story problems remains a challenging task in artificial intelligence, which requires a detailed understanding of real-world situations and a strong mathematical reasoning capability. Previous neural solvers of math word problems directly translate problem texts into equations, lackin…

Cited by 37SourcePDFScholar
2021

SocAoG: Incremental Graph Parsing for Social Relation Inference in Dialogues

ACL 2021long

Inferring social relations from dialogues is vital for building emotionally intelligent robots to interpret human language better and act accordingly. We model the social network as an And-or Graph, named SocAoG, for the consistency of relations among a group and leveraging attributes as inference c…

2021

Spatio-Temporal Self-Supervised Representation Learning for 3D Point Clouds

ICCV 2021poster

To date, various 3D scene understanding tasks still lack practical and generalizable pre-trained models, primarily due to the intricate nature of 3D scene understanding tasks and their immerse variations due to camera views, lighting, occlusions, etc. In this paper, we tackle this immanent challenge…

Cited by 248PDFcodeScholar
2021

Stochastic Security: Adversarial Defense Using Long-Run Dynamics of Energy-Based Models

ICLR 2021poster

The vulnerability of deep networks to adversarial attacks is a central problem for deep learning from the perspective of both cognition and security. The current most successful defense method is to train a classifier using adversarial images created during learning. Another defense approach involve…

2021

Unsupervised Foreground Extraction via Deep Region Competition

NeurIPS 2021poster

We present Deep Region Competition (DRC), an algorithm designed to extract foreground objects from images in a fully unsupervised manner. Foreground extraction can be viewed as a special case of generic image segmentation that focuses on identifying and disentangling objects from the background. In…

Cited by 41SourcePDFScholar
2021

YouRefIt: Embodied Reference Understanding With Language and Gesture

ICCV 2021poster

We study the machine's understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding multimodal cues with perspective-taking to identify which object is being…

Cited by 46PDFScholar
2020

A Competence-aware Curriculum for Visual Concepts Learning via Question Answering

ECCV 2020poster

Humans can progressively learn visual concepts from easy to hard questions. To mimic this efficient learning ability, we propose a competence-aware curriculum for visual concept learning in a question-answering manner. Specifically, we design a neural-symbolic concept learner for learning the visual…

Cited by 40SourcePDFScholar
2020

Closed Loop Neural-Symbolic Learning via Integrating Neural Perception, Grammar Parsing, and Symbolic Reasoning

ICML 2020poster

The goal of neural-symbolic computation is to integrate the connectionist and symbolist paradigms. Prior methods learn the neural-symbolic models using reinforcement learning (RL) approaches, which ignore the error propagation in the symbolic reasoning module and thus converge slowly with sparse rew…

2020

Congestion-aware Evacuation Routing using Augmented Reality Devices

ICRA 2020poster

We present a congestion-aware routing solution for indoor evacuation, which produces real-time individual-customized evacuation routes among multiple destinations while keeping tracks of all evacuees’ locations. A population density map, obtained on-the-fly by aggregating locations of evacuees from…

Cited by 17SourceScholar
2020

Graph-based Hierarchical Knowledge Representation for Robot Task Transfer from Virtual to Physical World

IROS 2020poster

We study the hierarchical knowledge transfer problem using a cloth-folding task, wherein the agent is first given a set of human demonstrations in the virtual world using an Oculus Headset, and later transferred and validated on a physical Baxter robot. We argue that such an intricate robot task tra…

Cited by 23SourceScholar
2020

Human-Robot Interaction in a Shared Augmented Reality Workspace

IROS 2020poster

We design and develop a new shared Augmented Reality (AR) workspace for Human-Robot Interaction (HRI), which establishes a bi-directional communication between human agents and robots. In a prototype system, the shared AR workspace enables a shared perception, so that a physical robot not only perce…

Cited by 39SourceScholar
2020

Inducing Hierarchical Compositional Model by Sparsifying Generator Network

CVPR 2020poster

This paper proposes to learn hierarchical compositional AND-OR model for interpretable image synthesis by sparsifying the generator network. The proposed method adopts the scene-objects-parts-subparts-primitives hierarchy in image representation. A scene has different types (i.e., OR) each of which…

Cited by 9PDFScholar
2020

Joint Inference of States, Robot Knowledge, and Human (False-)Beliefs

ICRA 2020poster

Aiming to understand how human (false-)belief— a core socio-cognitive ability—would affect human interactions with robots, this paper proposes to adopt a graphical model to unify the representation of object states, robot knowledge, and human (false-)beliefs. Specifically, a parse graph (pg) is lear…

Cited by 27SourceScholar
2020

Joint Training of Variational Auto-Encoder and Latent Energy-Based Model

CVPR 2020poster

This paper proposes a joint training method to learn both the variational auto-encoder (VAE) and the latent energy-based model (EBM). The joint training of VAE and latent EBM are based on an objective function that consists of three Kullback-Leibler divergences between three joint distributions on t…

Cited by 57PDFScholar
2020

LEMMA: A Multi-view Dataset for LEarning Multi-agent Multi-task Activities

ECCV 2020poster

The ability to understand and interpret human actions is a long-standing challenge and a critical indicator of perception in artificial intelligence. However, a few imperative components of daily human activities are largely missed in prior literature, including the goal-directed actions, concurrent…

2020

Learning Multi-layer Latent Variable Model via Variational Optimization of Short Run MCMC for Approximate Inference

ECCV 2020poster

This paper studies the fundamental problem of learning deep generative models that consist of multiple layers of latent variables organized in top-down architectures. Such models have high expressivity and allow for learning hierarchical representations. Learning such a generative model requires inf…

Cited by 56SourcePDFScholar
2019

Divergence Triangle for Joint Training of Generator Model, Energy-Based Model, and Inferential Model

CVPR 2019oral

This paper proposes the divergence triangle as a framework for joint training of a generator model, energy-based model and inference model. The divergence triangle is a compact and symmetric (anti-symmetric) objective function that seamlessly integrates variational learning, adversarial learning, wa…

Cited by 77PDFcodeScholar
2019

High-Fidelity Grasping in Virtual Reality using a Glove-based System

ICRA 2019poster

This paper presents a design that jointly provides hand pose sensing, hand localization, and haptic feedback to facilitate real-time stable grasps in Virtual Reality (VR). The design is based on an easy-to-replicate glove-based system that can reliably perform (i) a high-fidelity hand pose sensing i…

Cited by 81SourceScholar
2019

Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical Commonsense

ICCV 2019poster

We propose a new 3D holistic++ scene understanding problem, which jointly tackles two tasks from a single-view image: (i) holistic scene parsing and reconstruction---3D estimations of object bounding boxes, camera pose, and room layout, and (ii) 3D human pose estimation. The intuition behind is to l…

Cited by 145PDFScholar
2019

Learning Grid Cells as Vector Representation of Self-Position Coupled with Matrix Representation of Self-Motion

ICLR 2019poster

This paper proposes a representational model for grid cells. In this model, the 2D self-position of the agent is represented by a high-dimensional vector, and the 2D self-motion or displacement of the agent is represented by a matrix that transforms the vector. Each component of the vector is a unit…

2019

Learning Non-Convergent Non-Persistent Short-Run MCMC Toward Energy-Based Model

NeurIPS 2019poster

This paper studies a curious phenomenon in learning energy-based model (EBM) using MCMC. In each learning iteration, we generate synthesized examples by running a non-convergent, non-mixing, and non-persistent short-run MCMC toward the current model, always starting from the same initial distributio…

Cited by 267SourcePDFScholar
2019

Learning Perceptual Inference by Contrasting

NeurIPS 2019spotlight

“Thinking in pictures,” [1] i.e., spatial-temporal reasoning, effortless and instantaneous for humans, is believed to be a significant ability to perform logical induction and a crucial factor in the intellectual history of technology development. Modern Artificial Intelligence (AI), fueled by massi…

2019

Learning Virtual Grasp with Failed Demonstrations via Bayesian Inverse Reinforcement Learning

IROS 2019poster

We propose Bayesian Inverse Reinforcement Learning with Failure (BIRLF), which makes use of failed demonstrations that were often ignored or filtered in previous methods due to the difficulties to incorporate them in addition to the successful ones. Specifically, we leverage halfspaces derived from…

Cited by 27SourceScholar
2019

PerspectiveNet: 3D Object Detection from a Single RGB Image via Perspective Points

NeurIPS 2019poster

Detecting 3D objects from a single RGB image is intrinsically ambiguous, thus requiring appropriate prior knowledge and intermediate representations as constraints to reduce the uncertainties and improve the consistencies between the 2D image plane and the 3D world coordinate. To address this challe…

2019

RAVEN: A Dataset for Relational and Analogical Visual REasoNing

CVPR 2019poster

Dramatic progress has been witnessed in basic vision tasks involving low-level perception, such as object recognition, detection, and tracking. Unfortunately, there is still enormous performance gap between artificial vision systems and human intelligence in terms of higher-level vision problems, es…

Cited by 357PDFScholar
2019

Reasoning Visual Dialogs With Structural and Partial Observations

CVPR 2019oral

We propose a novel model to address the task of Visual Dialog which exhibits complex dialog structures. To obtain a reasonable answer based on the current question and the dialog history, the underlying semantic dependencies between dialog entities are essential. In this paper, we explicitly formali…

Cited by 143PDFcodeScholar
2019

Self-Supervised Incremental Learning for Sound Source Localization in Complex Indoor Environment

ICRA 2019poster

This paper presents an incremental learning framework for mobile robots localizing the human sound source using a microphone array in a complex indoor environment consisting of multiple rooms. In contrast to conventional approaches that leverage direction-of-arrival (DOA) estimation, the framework a…

Cited by 14SourceScholar
2019

Understanding Human Gaze Communication by Spatio-Temporal Graph Reasoning

ICCV 2019poster

This paper addresses a new problem of understanding human gaze communication in social videos from both atomic-level and event-level, which is significant for studying human social interactions. To tackle this novel and challenging problem, we contribute a large-scale video dataset, VACATION, which…

Cited by 145PDFcodeScholar
2019

Unsupervised Disentangling of Appearance and Geometry by Deformable Generator Network

CVPR 2019poster

We present a deformable generator model to disentangle the appearance and geometric information in purely unsupervised manner. The appearance generator models the appearance related information, including color, illumination, identity or category, of an image, while the geometric generator performs…

Cited by 33PDFScholar
2018

A Causal And-Or Graph Model for Visibility Fluent Reasoning in Tracking Interacting Objects

CVPR 2018poster

Tracking humans that are interacting with the other subjects or environment remains unsolved in visual tracking, because the visibility of the human of interests in videos is unknown and might vary over time. In particular, it is still difficult for state-of-the-art human trackers to recover complet…

Cited by 36SourcePDFScholar
2018

Attentive Fashion Grammar Network for Fashion Landmark Detection and Clothing Category Classification

CVPR 2018poster

This paper proposes a knowledge-guided fashion network to solve the problem of visual fashion analysis, e.g., fashion landmark localization and clothing category classification. The suggested fashion model is leveraged with high-level human knowledge in this domain. We propose two important fashion…

Cited by 307SourcePDFScholar
2018

Cooperative Holistic Scene Understanding: Unifying 3D Object, Layout, and Camera Pose Estimation

NeurIPS 2018poster

Holistic 3D indoor scene understanding refers to jointly recovering the i) object bounding boxes, ii) room layout, and iii) camera pose, all in 3D. The existing methods either are ineffective or only tackle the problem partially. In this paper, we propose an end-to-end model that simultaneously solv…

2018

Generalized Earley Parser: Bridging Symbolic Grammars and Sequence Data for Future Prediction

ICML 2018oral

Future predictions on sequence data (e.g., videos or audios) require the algorithms to capture non-Markovian and compositional properties of high-level semantics. Context-free grammars are natural choices to capture such properties, but traditional grammar parsers (e.g., Earley parser) only take sym…

Cited by 40SourcePDFScholar
2018

Holistic 3D Scene Parsing and Reconstruction from a Single RGB Image

ECCV 2018poster

We propose a computational framework to jointly parse a single RGB image and reconstruct a holistic 3D configuration composed by a set of CAD models using a stochastic grammar model. Specifically, we introduce a Holistic Scene Grammar (HSG) to represent the 3D scene structure, which characterizes a…

Cited by 171SourcePDFScholar
2018

Human-Centric Indoor Scene Synthesis Using Stochastic Grammar

CVPR 2018poster

We present a human-centric method to sample and synthesize 3D room layouts and 2D images thereof, for the purpose of obtaining large-scale 2D/3D image data with the perfect per-pixel ground truth. An attributed spatial And-Or graph (S-AOG) is proposed to represent indoor scenes. The S-AOG is a proba…

2018

Interactive Robot Knowledge Patching Using Augmented Reality

ICRA 2018poster

We present a novel Augmented Reality (AR) approach, through Microsoft HoloLens, to address the challenging problems of diagnosing, teaching, and patching interpretable knowledge of a robot. A Temporal And-Or graph (T-AOG) of opening bottles is learned from human demonstration and programmed to the r…

Cited by 83SourceScholar
2018

Learning Descriptor Networks for 3D Shape Synthesis and Analysis

CVPR 2018poster

This paper proposes a 3D shape descriptor network, which is a deep convolutional energy-based model, for modeling volumetric shape patterns. The maximum likelihood training of the model follows an "analysis by synthesis" scheme and can be interpreted as a mode seeking and mode shifting process. The…

2018

Learning Generative ConvNets via Multi-Grid Modeling and Sampling

CVPR 2018poster

This paper proposes a multi-grid method for learning energy-based generative ConvNet models of images. For each grid, we learn an energy-based probabilistic model where the energy function is defined by a bottom-up convolutional neural network (ConvNet or CNN). Learning such a model requires generat…

Cited by 93SourcePDFScholar
2018

Learning Human-Object Interactions by Graph Parsing Neural Networks

ECCV 2018poster

This paper addresses the task of detecting and recognizing human-object interactions (HOI) in images and videos. We introduce the Graph Parsing Neural Network (GPNN), a framework that incorporates structural knowledge while being differentiable end-to-end. For a given scene, GPNN infers a parse grap…

2018

Unsupervised Learning of Hierarchical Models for Hand-Object Interactions

ICRA 2018poster

Contact forces of the hand are visually unobservable, but play a crucial role in understanding hand-object interactions. In this paper, we propose an unsupervised learning approach for manipulation event segmentation and manipulation event parsing. The proposed framework incorporates hand pose kinem…

Cited by 15SourceScholar
2018

Where and Why Are They Looking? Jointly Inferring Human Attention and Intentions in Complex Tasks

CVPR 2018poster

This paper addresses a new problem - jointly inferring human attention, intentions, and tasks from videos. Given an RGB-D video where a human performs a task, we answer three questions simultaneously: 1) where the human is looking - attention prediction; 2) why the human is looking there - intention…

Cited by 82SourcePDFScholar
2017

A glove-based system for studying hand-object manipulation via joint pose and force sensing

IROS 2017poster

We present a design of an easy-to-replicate glove-based system that can reliably perform simultaneous hand pose and force sensing in real time, for the purpose of collecting human hand data during fine manipulative actions. The design consists of a sensory glove that is capable of jointly collecting…

Cited by 70SourceScholar
2017

CERN: Confidence-Energy Recurrent Network for Group Activity Recognition

CVPR 2017poster

This work is about recognizing human activities occurring in videos at distinct semantic levels, including individual actions, interactions, and group activities. The recognition is realized using a two-level hierarchy of Long Short-Term Memory (LSTM) networks, forming a feed-forward deep architectu…

Cited by 228PDFcodeScholar
2017

Feeling the force: Integrating force and pose for fluent discovery through imitation learning to open medicine bottles

IROS 2017poster

Learning complex robot manipulation policies for real-world objects is challenging, often requiring significant tuning within controlled environments. In this paper, we learn a manipulation model to execute tasks with multiple stages and variable structure, which typically are not suitable for most…

Cited by 78SourceScholar
2017

Learning Human Utility from Video Demonstrations for Deductive Planning in Robotics

CoRL 2017

We uncouple three components of autonomous behavior (utilitarian value, causal reasoning, and fine motion control) to design an interpretable model of tasks from video demonstrations. Utilitarian value is learned from aggregating human preferences to understand the implicit goal of a task, explainin

2017

Learning social affordance grammar from videos: Transferring human interactions to human-robot interactions

ICRA 2017poster

In this paper, we present a general framework for learning social affordance grammar as a spatiotemporal AND-OR graph (ST-AOG) from RGB-D videos of human interactions, and transfer the grammar to humanoids to enable a real-time motion inference for human-robot interaction (HRI). Based on Gibbs sampl…

Cited by 54SourceScholar
2016

Inferring Forces and Learning Human Utilities From Videos

CVPR 2016oral

We propose a notion of affordance that takes into account physical quantities generated when the human body interacts with real-world objects, and introduce a learning framework that incorporates the concept of human utilities, which in our opinion provides a deeper and finer-grained account not onl…

Cited by 113PDFScholar
2016

Inferring human intent from video by sampling hierarchical plans

IROS 2016poster

This paper presents a method which allows robots to infer a human's hierarchical intent from partially observed RGBD videos by imagining how the human will behave in the future. This capability is critical for creating robots which can interact socially or collaboratively with humans. We represent i…

Cited by 43SourceScholar
2015

Automated Facial Trait Judgment and Election Outcome Prediction: Social Dimensions of Face

ICCV 2015poster

The human face is a primary medium of human communication and a prominent source of information used to infer various attributes. In this paper, we study a fully automated system that can infer the perceived traits of a person from his face -- social dimensions, such as "intelligence," "honesty," an…

Cited by 92PDFScholar