← Search

Fuchun Sun

107 accepted papers

2026

Inference-Stage Adaptation-Projection Strategy Adapts Diffusion Policy to Cross-Manipulators Scenarios

ICRA 2026poster

Diffusion policies are powerful visuomotor models for robotic manipulation, yet they often fail to generalize to manipulators or end-effectors unseen during training and struggle to accommodate new task requirements at inference time. Addressing this typically requires costly data recollection and p…

2026

SceneTransporter: Optimal Transport-Guided Compositional Latent Diffusion for Single-Image Structured 3D Scene Generation

ICLR 2026poster

We introduce SceneTransporter, an end-to-end framework for structured 3D scene generation from a single image. While existing methods generate part-level 3D objects, they often fail to organize these parts into distinct instances in open-world scenes. Through a debiased clustering probe, we reveal a…

Cited by 0SourcecodeScholar
2026

Test-Time Perturbation Tuning with Delayed Feedback for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object pose. We attribute this brittleness to trajectory overfitting, where VLAs over-attend to the spurious correlation betwe

Cited by 0SourcecodeScholar
2025

A Bionic Robotic Hand Designed with Multiple Grasping Modes and Magnetic-tactile Perception

IROS 2025

This paper presents a novel multi-mode bionic robotic hand. Its bionic finger (BIF) ingeniously combines a magnetic-silica-gel skin with a rigid skeletal framework and integrates a vacuum suction cup at the fingertip. This design enables the bionic manipulator to execute multiple grasping modes, nam

Cited by 0SourceScholar
2025

Adversarial Locomotion and Motion Imitation for Humanoid Policy Learning

NeurIPS 2025poster

Humans exhibit diverse and expressive whole-body movements. However, attaining human-like whole-body coordination in humanoid robots remains challenging, as conventional approaches that mimic whole-body motions often neglect the distinct roles of upper and lower body. This oversight leads to computa…

Cited by 0SourcecodeScholar
2025

Efficiently Kinematic-Constraint-Coupled State Estimation for Integrated Aerial Platforms in GPS-Denied Environments

RA-L 2025

Small-scale autonomous aerial vehicles (AAVs) are widely used in various fields. However, their underactuated design limits their ability to perform complex tasks that require physical interaction with environments. The fully-actuated Integrated Aerial Platforms (IAPs), where multiple AAVs are conne

Cited by 0SourceScholar
2025

Evolutionary Reinforcement Learning with Parameterized Action Primitives for Diverse Manipulation Tasks

AAAI 2025technical

Reinforcement learning (RL) has shown promising performance in tackling robotic manipulation tasks (RMTs), which require learning a prolonged sequence of manipulation actions to control robots efficiently. However, most RL algorithms often suffer from two problems when solving RMTs: inefficient expl…

Cited by 0SourcePDFScholar
2025

Multi-Segment Soft Robot Control Via Deep Koopman-Based Model Predictive Control

ICRA 2025

Soft robots, compared to regular rigid robots, as their multiple segments with soft materials bring flexibility and compliance, have the advantages of safe interaction and dexterous operation in the environment. However, due to its characteristics of high dimensional, nonlinearity, time-varying natu

Cited by 0SourcecodeScholar
2025

Slimming the Fat-Tail: Morphing-Flow for Adaptive Time Series Modeling

ICML 2025poster

Temporal sequences, even after stationarization, often exhibit leptokurtic distributions with fat tails and persistent distribution shifts. These properties destabilize feature dynamics, amplify model variance, and hinder model convergence in time series forecasting. To address this, we propose Morp…

Cited by 0SourcePDFScholar
2025

Soft Growing Robot Explore Unknown Environments Through Obstacle Interaction

RA-L 2025

In low-light, unstructured, and confined environments, performing Simultaneous Localization and Mapping (SLAM) with conventional methods presents significant challenges. Soft growing robots, characterized by their compliance and extensibility, interact safely with the environment, making them well-s

Cited by 5SourceScholar
2025

Towards the Causal Complete Cause of Multi-Modal Representation Learning

ICML 2025poster

Multi-Modal Learning (MML) aims to learn effective representations across modalities for accurate predictions. Existing methods typically focus on modality consistency and specificity to learn effective representations. However, from a causal perspective, they may lead to representations that contai…

Cited by 0SourcePDFScholar
2024

ACE: Off-Policy Actor-Critic with Causality-Aware Entropy Regularization

ICML 2024oral

The varying significance of distinct primitive behaviors during the policy learning process has been overlooked by prior model-free RL algorithms. Leveraging this insight, we explore the causal relationship between different action dimensions and rewards to evaluate the significance of various primi…

2024

BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain Abstraction

ICLR 2024poster

As a novel and effective fine-tuning paradigm based on large-scale pre-trained language models (PLMs), prompt-tuning aims to reduce the gap between downstream tasks and pre-training objectives. While prompt-tuning has yielded continuous advancements in various tasks, such an approach still remains a…

2024

Bionic Soft Fingers with Hybrid Variable Stiffness Mechanisms for Multimode Grasping

ICRA 2024poster

This paper presents a novel Bionic Soft Finger (BSF) that aims to overcome the limitations of conventional rigid manipulators in terms of adaptability and safety, as well as the challenges faced by soft hands regarding carrying capacity and stability. The BSF design uses a hybrid variable stiffness…

Cited by 1SourceScholar
2024

Hierarchical Topology Isomorphism Expertise Embedded Graph Contrastive Learning

AAAI 2024technical

Graph contrastive learning (GCL) aims to align the positive features while differentiating the negative features in the latent space by minimizing a pair-wise contrastive loss. As the embodiment of an outstanding discriminative unsupervised graph representation learning approach, GCL achieves impres…

2024

Hybrid Robot for Percutaneous Needle Intervention Procedures: Mechanism Design and Experiment Verification

ICRA 2024poster

This paper presents a 6-DOF hybrid robot for percutaneous needle intervention procedures. The new robot combines the advantages of both serial robots and parallel robots, featuring compactness, high accuracy, and small footprint while overcoming the problems of the high cost of serial robots and the…

Cited by 0SourceScholar
2024

IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation

ICML 2024poster

Recent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual features without employing selective attention mechanisms, which is in sharp contras…

2024

InstanceVO: Self-Supervised Semantic Visual Odometry by Using Metric Learning to Incorporate Geometrical Priors in Instance Objects

RA-L 2024

Visual odometry is one of the key technologies for unmanned ground vehicles. To improve the robustness of the systems and enable intelligent tasks, researchers introduced learning-based recognition modules into visual odometry systems, but didn't realize tight coupling between visual odometry system

Cited by 2SourceScholar
2024

Modeling and Control of PADUAV: a Passively Articulated Dual UAVs Platform for Aerial Manipulation*

ICRA 2024poster

In this paper, we introduce PADUAV, a novel 5-DOF aerial platform designed to overcome the limitations of traditional tiltrotor vehicles. PADUAV features a unique mechanical design that incorporates two off-the-shelf quadrotors passively articulated to a rigid frame. This innovation enables free pit…

Cited by 2SourceScholar
2024

OMPO: A Unified Framework for RL under Policy and Dynamics Shifts

ICML 2024oral

Training reinforcement learning policies using environment interaction data collected from varying policies or dynamics presents a fundamental challenge. Existing works often overlook the distribution discrepancies induced by policy or dynamics shifts, or rely on specialized algorithms with task pri…

2024

Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL

ICML 2024poster

Off-policy reinforcement learning (RL) has achieved notable success in tackling many complex real-world tasks, by leveraging previously collected data for policy learning. However, most existing off-policy RL algorithms fail to maximally exploit the information in the replay buffer, limiting sample…

2024

Radardiff: Improving Sea Clutter Suppression Using Diffusion Models for Radar Images

ICASSP 2024accepted

Marine radar is employed across multiple fields, notably in navigation, meteorology, defense, and security. Marine radar images are highly sensitive to sea clutter, highlighting the crucial importance of sea clutter suppression in radar image processing. However, existing algorithms for sea clutter…

Cited by 0SourceScholar
2024

Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-Critic

ICML 2024poster

Learning high-quality $Q$-value functions plays a key role in the success of many modern off-policy deep reinforcement learning (RL) algorithms. Previous works primarily focus on addressing the value overestimation issue, an outcome of adopting function approximators and off-policy learning. Deviati…

2024

Smooth Computation without Input Delay: Robust Tube-Based Model Predictive Control for Robot Manipulator Planning

ICRA 2024poster

Model Predictive Control (MPC) has exhibited remarkable capabilities in optimizing objectives and meeting constraints. However, the substantial computational burden associated with solving the Optimal Control Problem (OCP) at each triggering instant introduces significant delays between state sampli…

Cited by 2SourceScholar
2024

Soft Magnetic Skin With Motion and Contact Sensing for Anthropomorphic Robotic Finger

RA-L 2024

Drawing inspiration from human fine tactile and proprioceptive kinaesthetic sensing pathways, we propose a soft magnetic skin (m-skin) with multimodal sensing functions integrated into the anthropomorphic robotic finger. This paper mainly explores the magnetic tactile sensor's structural design, per

Cited by 8SourceScholar
2024

Subequivariant Reinforcement Learning in 3D Multi-Entity Physical Environments

ICML 2024poster

Learning policies for multi-entity systems in 3D environments is far more complicated against single-entity scenarios, due to the exponential expansion of the global state space as the number of entities increases. One potential solution of alleviating the exponential complexity is dividing the glob…

Cited by 0SourcePDFScholar
2024

Tight Fusion of Odometry and Kinematic Constraints for Multiple Aerial Vehicles in Physical Interconnection

ICRA 2024poster

Integrated aerial Platforms (IAPs), comprising multiple aircrafts, are typically fully actuated and hold significant potential for aerial manipulation tasks. Differing from a multiple aerial swarm, the aircrafts within the IAP are interconnected, presenting promising opportunities for enhancing loca…

Cited by 3SourceScholar
2024

Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels

NeurIPS 2024poster

Video generative models are receiving particular attention given their ability to generate realistic and imaginative frames. Besides, these models are also observed to exhibit strong 3D consistency, significantly enhancing their potential to act as world simulators. In this work, we present Vidu4D,…

Cited by 18SourcePDFScholar
2023

Acquisition and Prediction of High-Density Tactile Field Data for Rigid and Flexible Objects

IROS 2023poster

Obtaining high-density tactile field information is a critical aspect of research in the field of robotic haptics, as it plays a decisive role in determining the precision of robot manipulations. Vision-based tactile sensors have unique high-resolution features, which make them promising for related…

Cited by 3SourceScholar
2023

Compacting Binary Neural Networks by Sparse Kernel Selection

CVPR 2023poster

Binary Neural Network (BNN) represents convolution weights with 1-bit values, which enhances the efficiency of storage and computation. This paper is motivated by a previously revealed phenomenon that the binary kernels in successful BNNs are nearly power-law distributed: their values are mostly clu…

Cited by 7SourcePDFScholar
2023

Disentangle and Remerge: Interventional Knowledge Distillation for Few-Shot Object Detection from a Conditional Causal Perspective

AAAI 2023technical

Few-shot learning models learn representations with limited human annotations, and such a learning paradigm demonstrates practicability in various tasks, e.g., image classification, object detection, etc. However, few-shot object detection methods suffer from an intrinsic defect that the limited tra…

2023

Embodied Referring Expression for Manipulation Question Answering in Interactive Environment

ICRA 2023poster

Embodied agents are expected to perform more complicated tasks in an interactive environment, with the progress of Embodied AI in recent years. Existing embodied tasks including Embodied Referring Expression (ERE) and other QA-form tasks mainly focuses on interaction in term of linguistic instructio…

Cited by 7SourceScholar
2023

Implementation and Optimization of Grasping Learning with Dual-modal Soft Gripper

ICRA 2023poster

Robust and efficient grasping of different objects is still an open problem due to the difficulty of integrating multidisciplinary knowledge such as gripper ontology design, perception, control, and learning. In recent years, learning-based methods have achieved excellent results in grasping various…

Cited by 4SourceScholar
2023

Measuring Acoustics with Collaborative Multiple Agents

IJCAI 2023poster

As humans, we hear sound every second of our life. The sound we hear is often affected by the acoustics of the environment surrounding us. For example, a spacious hall leads to more reverberation. Room Impulse Responses (RIR) are commonly used to characterize environment acoustics as a function of t…

Cited by 3SourcePDFScholar
2023

Robust Causal Graph Representation Learning against Confounding Effects

AAAI 2023technical

The prevailing graph neural network models have achieved significant progress in graph representation learning. However, in this paper, we uncover an ever-overlooked phenomenon: the pre-trained graph representation learning model tested with full graphs underperforms the model tested with well-prune…

2023

Root Pose Decomposition Towards Generic Non-rigid 3D Reconstruction with Monocular Videos

ICCV 2023poster

This work focuses on the 3D reconstruction of non-rigid objects based on monocular RGB video sequences. Concretely, we aim at building high-fidelity models for generic object categories and casually captured scenes. To this end, we do not assume known root poses of objects, and do not utilize catego…

Cited by 9PDFcodeScholar
2023

SRTNET: Time Domain Speech Enhancement via Stochastic Refinement

ICASSP 2023accepted

Diffusion model, as a new generative model which is very popular in image generation and audio synthesis, is rarely used in speech enhancement. In this paper, we use the diffusion model as a module for stochastic refinement. We propose SRTNet, a novel method for speech enhancement via Stochastic Ref…

Cited by 0SourceScholar
2023

Subequivariant Graph Reinforcement Learning in 3D Environments

ICML 2023oral

Learning a shared policy that guides the locomotion of different agents is of core interest in Reinforcement Learning (RL), which leads to the study of morphology-agnostic RL. However, existing benchmarks are highly restrictive in the choice of starting point and target point, constraining the movem…

2023

TIRgel: A Visuo-Tactile Sensor With Total Internal Reflection Mechanism for External Observation and Contact Detection

RA-L 2023

This letter proposes a vision-based tactile sensor named TIRgel, leveraging visual integration to simplify the sensing system. First, the sensor achieves conversion of visual and tactile modality via focus adjustment. Under far-focus imaging, the camera can observe external environments; Under near-

Cited by 36SourceScholar
2023

Tacchi: A Pluggable and Low Computational Cost Elastomer Deformation Simulator for Optical Tactile Sensors

RA-L 2023

Simulation is widely applied in robotics research to save time and resources. There have been several works to simulate optical tactile sensors that leverage either a smoothing method or Finite Element Method (FEM). However, elastomer deformation physics is not considered in the former method, where

Cited by 51SourcecodeScholar
2023

Timestamp-Supervised Action Segmentation from the Perspective of Clustering

IJCAI 2023poster

Video action segmentation under timestamp supervision has recently received much attention due to lower annotation costs. Most existing methods generate pseudo-labels for all frames in each video to train the segmentation model. However, these methods suffer from incorrect pseudo-labels, especially…

2023

Variable Admittance Interaction Control of UAVs via Deep Reinforcement Learning

ICRA 2023poster

A compliant control model based on reinforcement learning (RL) is proposed to allow robots to interact with the environment more effectively and autonomously execute force control tasks. The admittance model learns an optimal adjustment policy for interactions with the external environment using RL…

Cited by 10SourceScholar
2022

Adversarial Texture for Fooling Person Detectors in the Physical World

CVPR 2022oral

Nowadays, cameras equipped with AI systems can capture and analyze images to detect people automatically. However, the AI system can make mistakes when receiving deliberately designed patterns in the real world, i.e., physical adversarial examples. Prior works have shown that it is possible to print…

Cited by 147PDFcodeScholar
2022

Audio-Visual Grounding Referring Expression for Robotic Manipulation

ICRA 2022poster

Referring expressions are commonly used when referring to a specific target in people's daily dialogue. In this paper, we develop a novel task of audio-visual grounding referring expression for robotic manipulation. The robot leverages both the audio and visual information to understand the referrin…

Cited by 19SourceScholar
2022

Bootstrapping Informative Graph Augmentation via A Meta Learning Approach

IJCAI 2022poster

Recent works explore learning graph representations in a self-supervised manner. In graph contrastive learning, benchmark methods apply various graph augmentation approaches. However, most of the augmentation methods are non-learnable, which causes the issue of generating unbeneficial augmented grap…

2022

Bridged Transformer for Vision and Point Cloud 3D Object Detection

CVPR 2022poster

3D object detection is a crucial research topic in computer vision, which usually uses 3D point clouds as input in conventional setups. Recently, there is a trend of leveraging multiple sources of input data, such as complementing the 3D point cloud with 2D images that often have richer color and fe…

Cited by 53PDFScholar
2022

C-Shaped Bidirectional Stiffness Joint Design For Anthropomorphic Hand

RA-L 2022

In this letter, we propose a C-shaped bidirectional stiffness joint for an anthropomorphic hand. The spring steel piece is used to connect the knuckles with inner concave C-shape as a rotational joint. The finger is bent by the tendon driven with low stiffness and reset by its own elasticity. The re

Cited by 11SourceScholar
2022

Depth-Aware Vision-and-Language Navigation using Scene Query Attention Network

ICRA 2022poster

Vision-and-language navigation (VLN) has been an important task in the field of Robotics and Computer Vision. However, most existing vision-and-language navigation models only use features extracted from RGB observation as input, while robots can utilize depth sensors in the real world. Existing res…

Cited by 4SourceScholar
2022

Embodied Multi-Agent Task Planning from Ambiguous Instruction

RSS 2022poster

In human-robots collaboration scenarios, a human would give robots an instruction that is intuitive for the human himself to accomplish. However, the instruction given to robots is likely ambiguous for them to understand as some information is implicit in the instruction. Therefore, it is necessary…

Cited by 26SourcePDFScholar
2022

Equivariant Graph Mechanics Networks with Constraints

ICLR 2022poster

Learning to reason about relations and dynamics over multiple interacting objects is a challenging topic in machine learning. The challenges mainly stem from that the interacting systems are exponentially-compositional, symmetrical, and commonly geometrically-constrained. Current methods, particular…

2022

Learning 6-DoF Task-oriented Grasp Detection via Implicit Estimation and Visual Affordance

IROS 2022poster

Currently, task-oriented grasp detection approaches are mostly based on pixel-level affordance detection and semantic segmentation. These pixel-level approaches heavily rely on the accuracy of a 2D affordance mask, and the generated grasp candidates are restricted to a small workspace. To mitigate t…

Cited by 24SourceScholar
2022

Multifingered Grasping Based on Multimodal Reinforcement Learning

RA-L 2022

In this work, we tackle the challenging problem of grasping novel objects using a high-DoF anthropomorphic hand-arm system. Combining fingertip tactile sensing, joint torques and proprioception, a multimodal agent is trained in simulation to learn the finger motions and to determine when to lift an

Cited by 34SourceScholar
2022

Multimodal Token Fusion for Vision Transformers

CVPR 2022poster

Many adaptations of transformers have emerged to address the single-modal vision tasks, where self-attention modules are stacked to handle input sources like images. Intuitively, feeding multiple modalities of data to vision transformers could improve the performance, yet the inner-modal attentive w…

Cited by 217PDFcodeScholar
2022

Non-destructive Fruit Firmness Evaluation Using Vision-Based Tactile Information

ICRA 2022poster

During postharvest storage, fruit firmness usually decreases due to respiration and bruise, the former of which indicates the fruit ripeness while the latter negatively influence consumers' taste preference. This paper presents a portable and low-cost device using vision-based tactile information to…

Cited by 27SourceScholar
2022

SNAKE: Shape-aware Neural 3D Keypoint Field

NeurIPS 2022accept

Detecting 3D keypoints from point clouds is important for shape reconstruction, while this work investigates the dual question: can shape reconstruction benefit 3D keypoint detection? Existing methods either seek salient features according to statistics of different orders or learn to predict keypoi…

2022

Sim2Real Object-Centric Keypoint Detection and Description

AAAI 2022technical

Keypoint detection and description play a central role in computer vision. Most existing methods are in the form of scene-level prediction, without returning the object classes of different keypoints. In this paper, we propose the object-centric formulation, which, beyond the conventional setting, r…

Cited by 9SourcePDFScholar
2022

Sound Adversarial Audio-Visual Navigation

ICLR 2022poster

Audio-visual navigation task requires an agent to find a sound source in a realistic, unmapped 3D environment by utilizing egocentric audio-visual observations. Existing audio-visual navigation works assume a clean environment that solely contains the target sound, which, however, would not be suita…

2022

When to Update Your Model: Constrained Model-based Reinforcement Learning

NeurIPS 2022accept

Designing and analyzing model-based RL (MBRL) algorithms with guaranteed monotonic improvement has been challenging, mainly due to the interdependence between policy optimization and model learning. Existing discrepancy bounds generally ignore the impacts of model shifts, and their corresponding alg…

2021

Adversarial Option-Aware Hierarchical Imitation Learning

ICML 2021spotlight

It has been a challenge to learning skills for an agent from long-horizon unannotated demonstrations. Existing approaches like Hierarchical Imitation Learning(HIL) are prone to compounding errors or suboptimal solutions. In this paper, we propose Option-GAIL, a novel method to learn skills at long h…

2021

Evaluations of the Gap between Supervised and Reinforcement Lifelong Learning on Robotic Manipulation Tasks

CoRL 2021poster

Overcoming catastrophic forgetting is of great importance for deep learning and robotics. Recent lifelong learning research has great advances in supervised learning. However, little work focuses on reinforcement learning(RL). We focus on evaluating the performances of state-of-the-art lifelong lear…

Cited by 13SourceScholar
2021

Sub-Bit Neural Networks: Learning To Compress and Accelerate Binary Neural Networks

ICCV 2021poster

In the low-bit quantization field, training Binarized Neural Networks (BNNs) is the extreme solution to ease the deployment of deep models on resource-constrained devices, having the lowest storage cost and significantly cheaper bit-wise operations compared to 32-bit floating-point counterparts. In…

Cited by 20PDFcodeScholar
2020

A Boundary Based Out-of-Distribution Classifier for Generalized Zero-Shot Learning

ECCV 2020poster

Generalized Zero-Shot Learning (GZSL) is a challenging topic that has promising prospects in many realistic scenarios. Using a gating mechanism that discriminates the unseen samples from the seen samples can decompose the GZSL problem to a conventional Zero-Shot Learning (ZSL) problem and a supervis…

Cited by 107SourcePDFScholar
2020

A Mobile Robot Hand-Arm Teleoperation System by Vision and IMU

IROS 2020poster

In this paper, we present a multimodal mobile teleoperation system that consists of a novel vision-based hand pose regression network (Transteleop) and an IMU (inertial measurement units)-based arm tracking method. Transteleop observes the human hand through a low-cost depth camera and generates not…

Cited by 75SourceScholar
2020

Deep Multimodal Fusion by Channel Exchanging

NeurIPS 2020poster

Deep multimodal fusion by using multiple sources of data for classification or regression has exhibited a clear advantage over the unimodal counterpart on various applications. Yet, current methods including aggregation-based and alignment-based fusion are still inadequate in balancing the trade-off…

2020

Multi-Agent Embodied Question Answering in Interactive Environments

ECCV 2020poster

We investigate a new AI task --- Multi-Agent Interactive Question Answering --- where several agents explore the scene jointly in interactive environments to answer a question. To cooperate efficiently and answer accurately, agents must be well-organized to have balanced work division and share know…

Cited by 38SourcePDFScholar
2020

Resolution Switchable Networks for Runtime Efficient Image Recognition

ECCV 2020poster

We propose a general method to train a single convolutional neural network which is capable of switching image resolutions at inference. Thus the running speed can be selected to meet various computational resource limits. Networks trained with the proposed method are named Resolution Switchable Net…

2020

Reusing Discriminators for Encoding: Towards Unsupervised Image-to-Image Translation

CVPR 2020poster

Unsupervised image-to-image translation is a central task in computer vision. Current translation frameworks will abandon the discriminator once the training process is completed. This paper contends a novel role of the discriminator by reusing it for encoding the images of the target domain. The pr…

Cited by 259PDFcodeScholar
2020

Robust Robotic Pouring using Audition and Haptics

IROS 2020poster

Robust and accurate estimation of liquid height lies as an essential part of pouring tasks for service robots. However, vision-based methods often fail in occluded conditions while audio-based methods cannot work well in a noisy environment. We instead propose a multimodal pouring network (MP-Net) t…

Cited by 24SourcecodeScholar
2020

Self-Supervised Learning for Alignment of Objects and Sound

ICRA 2020poster

The sound source separation problem has many useful applications in the field of robotics, such as human-robot interaction, scene understanding, etc. However, it remains a very challenging problem. In this paper, we utilize both visual and audio information of videos to perform the sound source sepa…

Cited by 5SourceScholar
2019

Attention-based Transfer Learning for Brain-computer Interface

ICASSP 2019accepted

Different functional areas of the human brain play different roles in brain activity, which has not been paid sufficient research attention in the brain-computer interface (BCI) field. This paper presents a new approach for electroencephalography (EEG) classification that applies attention-based tra…

Cited by 0SourceScholar
2019

Deep Reinforcement Learning for Robotic Pushing and Picking in Cluttered Environment

IROS 2019poster

In this paper, a novel robotic grasping system is established to automatically pick up objects in cluttered scenes. A composite robotic hand composed of a suction cup and a gripper is designed for grasping the object stably. The suction cup is used for lifting the object from the clutter first and t…

Cited by 103SourceScholar
2019

Imitation Learning from Observations by Minimizing Inverse Dynamics Disagreement

NeurIPS 2019spotlight

This paper studies Learning from Observations (LfO) for imitation learning with access to state-only demonstrations. In contrast to Learning from Demonstration (LfD) that involves both action and state supervisions, LfO is more practical in leveraging previously inapplicable resources (e.g., videos)…

Cited by 90SourcePDFScholar
2019

Making Sense of Audio Vibration for Liquid Height Estimation in Robotic Pouring

IROS 2019poster

In this paper, we focus on the challenging perception problem in robotic pouring. Most of the existing approaches either leverage visual or haptic information. However, these techniques may suffer from poor generalization performances on opaque containers or concerning measuring precision. To tackle…

Cited by 43SourceScholar
2019

PointNetGPD: Detecting Grasp Configurations from Point Sets

ICRA 2019poster

In this paper, we propose an end-to-end grasp evaluation model to address the challenging problem of localizing robot grasp configurations directly from the point cloud. Compared to recent grasp evaluation metrics that are based on handcrafted depth features and a convolutional neural network (CNN),…

Cited by 444SourcecodeScholar
2019

Vision-based Teleoperation of Shadow Dexterous Hand using End-to-End Deep Neural Network

ICRA 2019poster

In this paper, we present TeachNet, a novel neural network architecture for intuitive and markerless vision-based teleoperation of dexterous robotic hands. Robot joint angles are directly generated from depth images of the human hand that produce visually similar robot hand poses in an end-to-end fa…

Cited by 120SourceScholar
2018

A Dual-Modal Vision-Based Tactile Sensor for Robotic Hand Grasping

ICRA 2018poster

Humans' fingertips can perceive not only the magnitude and the direction of force but also the texture of object. When we grasp an object, the surface texture sensing of the fingertip helps us recognize the object and the force feeling that is parallel to the skin helps us grasp stably. Focusing on…

Cited by 64SourceScholar
2018

Deep Feature Pyramid Reconfiguration for Object Detection

ECCV 2018poster

State-of-the-art object detectors usually learn multi-scale representations to get better results by employing feature pyramids. However, the current designs for feature pyramids are still inefficient to integrate the semantic information over different scales. In this paper, we begin by investigati…

2017

Efficient Optimization for Linear Dynamical Systems with Applications to Clustering and Sparse Coding

NeurIPS 2017poster

Linear Dynamical Systems (LDSs) are fundamental tools for modeling spatio-temporal data in various disciplines. Though rich in modeling, analyzing LDSs is not free of difficulty, mainly because LDSs do not comply with Euclidean geometry and hence conventional learning techniques can not be applied d…

Cited by 12SourcePDFScholar
2017

RON: Reverse Connection With Objectness Prior Networks for Object Detection

CVPR 2017poster

We present RON, an efficient and effective framework for generic object detection. Our motivation is to smartly associate the best of the region-based (e.g., Faster R-CNN) and region-free (e.g., SSD) methodologies. Under fully convolutional architecture, RON mainly focuses on two fundamental problem…

Cited by 539PDFScholar
2016

HyperNet: Towards Accurate Region Proposal Generation and Joint Object Detection

CVPR 2016spotlight

Almost all of the current top-performing object detection networks employ region proposals to guide the search for object instances. State-of-the-art region proposal methods usually need several thousand proposals to get high recall, thus hurting the detection efficiency. Although the latest Region…

Cited by 1150PDFScholar
2016

Learning Cooperative Primitives with physical Human-Robot Interaction for a HUman-powered Lower EXoskeleton

IROS 2016poster

Human-powered lower exoskeletons have gained considerable interests from both academia and industry over the past few decades, and thus have seen increasing applications in areas of human locomotion assistance and strength augmentation. One of the most important aspects in those applications is to a…

Cited by 20SourceScholar
2016

Sparse Coding and Dictionary Learning With Linear Dynamical Systems

CVPR 2016oral

Linear Dynamical Systems (LDSs) are the fundamental tools for encoding spatio-temporal data in various disciplines. To enhance the performance of LDSs, in this paper, we address the challenging issue of performing sparse coding on the space of LDSs, where both data and dictionary atoms are LDSs. Rat…

Cited by 38PDFScholar
2015

Grasp planning by human experience on a variety of objects with complex geometry

IROS 2015poster

We present an effective method of identifying the graspable components of a variety of complex objects for grasp planning based on human experience. Instead of focusing on individual objects, our method identifies graspable components on the category level under the assumption that geometrically ali…

Cited by 10SourceScholar
2015

Transmissive optical pretouch sensing for robotic grasping

IROS 2015poster

Robotic grasping has been hindered by the inability of robots to perceive unstructured environments. Because these environments can be complex or dynamic, it is important to obtain additional and precise sensing information just before grasping. This paper expands upon the pretouch modality by intro…

Cited by 20SourceScholar