← Search

Chao Tang

17 accepted papers

2025

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

CVPR 2025poster

Story visualization, the task of creating visual narratives from textual descriptions, has seen progress with text-to-image generation models. However, these models often lack effective control over character appearances and interactions, particularly in multi-character scenes. To address these limi…

Cited by 4SourcePDFScholar
2025

FlowPlan: Zero-Shot Task Planning with LLM Flow Engineering for Robotic Instruction Following

IROS 2025

Robotic instruction following tasks require seamless integration of visual perception, task planning, target localization, and motion execution. However, existing task planning methods for instruction following are either data-driven or underperform in zero-shot scenarios due to difficulties in grou

Cited by 1SourcecodeScholar
2025

HGDiffuser: Efficient Task-Oriented Grasp Generation via Human-Guided Grasp Diffusion Models

IROS 2025

Task-oriented grasping (TOG) is essential for robots to perform manipulation tasks, requiring grasps that are both stable and compliant with task-specific constraints. Humans naturally grasp objects in a task-oriented manner to facilitate subsequent manipulation tasks. By leveraging human grasp demo

Cited by 5SourceScholar
2025

MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence

CoRL 2025poster

Imitating tool manipulation from human videos offers an intuitive approach to teaching robots, while also providing a promising and scalable alternative to labor-intensive teleoperation data collection for visuomotor policy learning. While humans can mimic tool manipulation behavior by observing oth…

Cited by 0SourceScholar
2025

RTAGrasp: Learning Task-Oriented Grasping from Human Videos via Retrieval, Transfer, and Alignment

ICRA 2025

Task-oriented grasping (TOG) is crucial for robots to accomplish manipulation tasks, requiring the determination of TOG positions and directions. Existing methods either rely on costly manual TOG annotations or only extract coarse grasping positions or regions from human demonstrations, limiting the

Cited by 12SourceScholar
2024

Estimating before Debiasing: A Bayesian Approach to Detaching Prior Bias in Federated Semi-Supervised Learning

IJCAI 2024poster

Federated Semi-Supervised Learning (FSSL) leverages both labeled and unlabeled data on clients to collaboratively train a model. In FSSL, the heterogeneous data can introduce prediction bias into the model, causing the model's prediction to skew towards some certain classes. Existing FSSL method…

2023

GraspGPT: Leveraging Semantic Knowledge From a Large Language Model for Task-Oriented Grasping

RA-L 2023

Task-oriented grasping (TOG) refers to the problem of predicting grasps on an object that enable subsequent manipulation tasks. To model the complex relationships between objects, tasks, and grasps, existing methods incorporate semantic knowledge as priors into TOG pipelines. However, the existing s

Cited by 122SourceScholar
2023

Task-Oriented Grasp Prediction with Visual-Language Inputs

IROS 2023poster

To perform household tasks, assistive robots receive commands in the form of user language instructions for tool manipulation. The initial stage involves selecting the intended tool (i.e., object grounding) and grasping it in a task-oriented manner (i.e., task grounding). Nevertheless, prior researc…

Cited by 40SourceScholar
2022

Keyframe Selection with Information Occupancy Grid Model for Long-term Data Association

IROS 2022poster

As the basics of Visual Simultaneous Localization And Mapping (VSLAM), keyframes play an essential role. In previous works, keyframes are selected according to a series of view change-based strategies for short-term data association (STDA). However, the texture enrichment of frames is always ignored…

Cited by 2SourceScholar
2022

Relationship Oriented Semantic Scene Understanding for Daily Manipulation Tasks

IROS 2022poster

Assistive robot systems have been developed to help people accomplish daily manipulation tasks especially for those with disabilities, where scene understanding plays a crucial role in enabling robots to interpret the surroundings and behave accordingly. Most of the current systems approach scene un…

Cited by 5SourceScholar
2022

Untethered Robotic Millipede Driven by Low-Pressure Microfluidic Actuators for Multi-Terrain Exploration

RA-L 2022

Mobile robots that can adapt to an extensive range of terrains play essential roles in many applications. Millipedes are one of the most terrain-adaptive creatures in nature due to their multi-legged locomotion and flexible body. Inspired by natural millipedes, we report an untethered robotic millip

Cited by 20SourceScholar
2021

Increasing the Payload and Terrain Adaptivity of an Untethered Crawling Robot Via Soft-Rigid Coupled Linear Actuators

RA-L 2021

Fluidic Elastomer Actuators (FEAs) provide new opportunities for developing agile, adaptive, and strong mobile robots for field explorations. In this letter, we propose a soft-rigid coupled linear FEA with a bio-comparable energy density of 10.9 J/kg at 50 kPa. This actuator achieves an increased bl

Cited by 25SourceScholar
2020

Using Synthetic Data and Deep Networks to Recognize Primitive Shapes for Object Grasping

ICRA 2020poster

A segmentation-based architecture is proposed to decompose objects into multiple primitive shapes from monocular depth input for robotic manipulation. The backbone deep network is trained on synthetic data with 6 classes of primitive shapes generated by a simulation engine. Each primitive shape is d…

Cited by 54SourceScholar