← Search

Xiangyu Chen

63 accepted papers

2026

Adaptor: Advancing Assistive Teleoperation with Few-Shot Learning and Cross-Operator Generalization

ICRA 2026poster

Assistive teleoperation enhances efficiency via shared control, yet inter-operator variability, stemming from diverse habits and expertise, induces highly heterogeneous trajectory distributions that undermine intent recognition stability. We present Adaptor, a few-shot framework for robust cross-ope…

2026

ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed Graphs

ICML 2026poster

Text-attributed Graphs (TAGs) incorporate textual node attributes with graph structures to describe rich relational semantics. Recent efforts to integrate Graph Neural Networks (GNNs) and Large Language Models (LLMs) have shown promise for learning on TAGs, yet achieving well-aligned representations…

Cited by 0SourceScholar
2026

FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration

CVPR 2026

All-in-One Image Restoration (AIO-IR) aims to develop a unified model that can handle multiple degradations under complex conditions. However, existing methods often rely on task-specific designs or latent routing strategies, making it hard to adapt to real-world scenarios with various degradations.

Cited by 0SourcecodeScholar
2026

LatentLLM: Activation-Aware Transform to Multi-Head Latent Attention

AAAI 2026technical

Modern foundation models such as large language models (LLMs) require a massive amount of computational and memory resources. We propose a new framework to convert such LLMs into a reduced-dimension latent structure. Our method extends a local activation-aware tensor decomposition to a global attent

Cited by 0SourcePDFScholar
2026

RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation

ICRA 2026poster

Precise robot manipulation is critical for fine-grained applications such as chemical and biological experiments, where even small errors (e.g., reagent spillage) can invalidate an entire task. Existing approaches often rely on pre-collected expert demonstrations and train policies via imitation lea…

2026

Scan Clusters, Not Pixels: A Cluster-Centric Paradigm for Efficient Ultra-high-definition Image Restoration

CVPR 2026

Ultra-High-Definition (UHD) image restoration is trapped in a scalability crisis: existing models, bound to pixel-wise operations, demand unsustainable computation. While state space models (SSMs) like Mamba promise linear complexity, their pixel-serial scanning remains a fundamental bottleneck for

Cited by 0SourcecodeScholar
2026

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

ICML 2026spotlight

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains limited. In this work, we present UniPercept-Bench, a unified fr…

Cited by 0SourceScholar
2025

Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image Compression

AAAI 2025technical

Neural image compression often faces a challenging trade-off among rate, distortion and perception. While most existing methods typically focus on either achieving high pixel-level fidelity or optimizing for perceptual metrics, we propose a novel approach that simultaneously addresses both aspects f…

Cited by 1SourcePDFScholar
2025

DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex Degradations

ICCV 2025poster

Diffusion models have demonstrated exceptional capabilities in image restoration, yet their application to video super-resolution (VSR) faces significant challenges in balancing fidelity with temporal consistency. Our evaluation reveals a critical gap: existing approaches consistently fail on severe…

Cited by 0SourcePDFScholar
2025

Exploiting Task Relationships in Continual Learning via Transferability-Aware Task Embeddings

NeurIPS 2025poster

Continual learning (CL) has been a critical topic in contemporary deep neural network applications, where higher levels of both forward and backward transfer are desirable for an effective CL performance. Existing CL strategies primarily focus on task models — either by regularizing model updates or…

Cited by 0SourcecodeScholar
2025

Learning 3D Perception from Others' Predictions

ICLR 2025poster

Accurate 3D object detection in real-world environments requires a huge amount of annotated data with high quality. Acquiring such data is tedious and expensive, and often needs repeated effort when a new sensor is adopted or when the detector is deployed in a new environment. We investigate a new s…

Cited by 1SourcePDFScholar
2025

WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning

ICLR 2025poster

The Earth's weather system involves intricate weather data modalities and diverse weather understanding tasks, which hold significant value to human life. Existing data-driven models focus on single weather understanding tasks (e.g., weather forecasting). While these models have achieved promising…

2024

Block-Map-Based Localization in Large-Scale Environment

ICRA 2024poster

Accurate localization is an essential technology for the flexible navigation of robots in large-scale environments. Both SLAM-based and map-based localization will increase the computing load due to the increase in map size, which will affect downstream tasks such as robot navigation and services. T…

Cited by 4SourcecodeScholar
2024

DiffuBox: Refining 3D Object Detection with Point Diffusion

NeurIPS 2024poster

Ensuring robust 3D object detection and localization is crucial for many applications in robotics and autonomous driving. Recent models, however, face difficulties in maintaining high performance when applied to domains with differing sensor setups or geographic locations, often resulting in poor lo…

2024

Direction-Aware Video Demoiréing with Temporal-Guided Bilateral Learning

AAAI 2024technical

Moiré patterns occur when capturing images or videos on screens, severely degrading the quality of the captured images or videos. Despite the recent progresses, existing video demoiréing methods neglect the physical characteristics and formation process of moiré patterns, significantly limiting the…

Cited by 10SourcePDFScholar
2024

GaussianGrasper: 3D Language Gaussian Splatting for Open-Vocabulary Robotic Grasping

RA-L 2024

Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit in the domain of robotics, which facilitates robots in executing object manipulations based on human language directives. To achieve this, some research efforts have been dedicated to the development o

Cited by 102SourcecodeScholar
2024

Image Demoireing in RAW and sRGB Domains

ECCV 2024poster

"Moiré patterns frequently appear when capturing screens with smartphones or cameras, potentially compromising image quality. Previous studies suggest that moiré pattern elimination in the RAW domain offers greater effectiveness compared to demoiréing in the sRGB domain. Nevertheless, relying sol…

2024

SEAL: A Framework for Systematic Evaluation of Real-World Super-Resolution

ICLR 2024spotlight

Real-world Super-Resolution (Real-SR) methods focus on dealing with diverse real-world images and have attracted increasing attention in recent years. The key idea is to use a complex and high-order degradation model to mimic real-world degradations. Although they have achieved impressive results i…

2024

Unifying Image Processing as Visual Prompting Question Answering

ICML 2024poster

Image processing is a fundamental task in computer vision, which aims at enhancing image quality and extracting essential features for subsequent vision applications. Traditionally, task-specific models are developed for individual tasks and designing such models requires distinct expertise. Buildin…

2023

Activating More Pixels in Image Super-Resolution Transformer

CVPR 2023poster

Transformer-based methods have shown impressive performance in low-level vision tasks, such as image super-resolution. However, we find that these networks can only utilize a limited spatial range of input information through attribution analysis. This implies that the potential of Transformer is st…

2023

DeSRA: Detect and Delete the Artifacts of GAN-based Real-World Super-Resolution Models

ICML 2023poster

Image super-resolution (SR) with generative adversarial networks (GAN) has achieved great success in restoring realistic details. However, it is notorious that GAN-based SR models will inevitably produce unpleasant and undesirable artifacts, especially in practical scenarios. Previous works typicall…

2023

Effective Ambiguity Attack Against Passport-Based DNN Intellectual Property Protection Schemes Through Fully Connected Layer Substitution

CVPR 2023poster

Since training a deep neural network (DNN) is costly, the well-trained deep models can be regarded as valuable intellectual property (IP) assets. The IP protection associated with deep models has been receiving increasing attentions in recent years. Passport-based method, which replaces normalizatio…

Cited by 16SourcePDFScholar
2023

Learning Iterative Neural Optimizers for Image Steganography

ICLR 2023poster

Image steganography is the process of concealing secret information in images through imperceptible changes. Recent work has formulated this task as a classic constrained optimization problem. In this paper, we argue that image steganography is inherently performed on the (elusive) manifold of natu…

2023

Learning To Invert: Simple Adaptive Attacks for Gradient Inversion in Federated Learning

UAI 2023poster

Gradient inversion attack enables recovery of training samples from model gradients in federated learning (FL), and constitutes a serious threat to data privacy. To mitigate this vulnerability, prior work proposed both principled defenses based on differential privacy, as well as heuristic defenses…

2023

Low-Light Video Enhancement with Synthetic Event Guidance

AAAI 2023technical

Low-light video enhancement (LLVE) is an important yet challenging task with many applications such as photographing and autonomous driving. Unlike single image low-light enhancement, most LLVE methods utilize temporal information from adjacent frames to restore the color and remove the noise of the…

Cited by 29SourcePDFScholar
2023

Real-World Image Super-Resolution as Multi-Task Learning

NeurIPS 2023poster

In this paper, we take a new look at real-world image super-resolution (real-SR) from a multi-task learning perspective. We demonstrate that the conventional formulation of real-SR can be viewed as solving multiple distinct degradation tasks using a single shared model. This poses a challenge known…

2023

Reward Finetuning for Faster and More Accurate Unsupervised Object Discovery

NeurIPS 2023poster

Recent advances in machine learning have shown that Reinforcement Learning from Human Feedback (RLHF) can improve machine learning models and align them with human preferences. Although very successful for Large Language Models (LLMs), these advancements have not had a comparable impact in research…

2023

Waymax: An Accelerated, Data-Driven Simulator for Large-Scale Autonomous Driving Research

NeurIPS 2023poster

Simulation is an essential tool to develop and benchmark autonomous vehicle planning software in a safe and cost-effective manner. However, realistic simulation requires accurate modeling of multi-agent interactive behaviors to be trustworthy, behaviors which can be highly nuanced and complex. To ad…

Cited by 116SourcePDFScholar
2022

Dilated Continuous Random Field for Semantic Segmentation

ICRA 2022poster

Mean field approximation methodology has laid the foundation of modern Continuous Random Field (CRF) based solutions for the refinement of semantic segmentation. In this paper, we propose to relax the hard constraint of mean field approximation - minimizing the energy term of each node from probabil…

Cited by 1SourcecodeScholar
2022

Fixed Neural Network Steganography: Train the images, not the network

ICLR 2022poster

Recent attempts at image steganography make use of advances in deep learning to train an encoder-decoder network pair to hide and retrieve secret messages in images. These methods are able to hide large amounts of data, but they also incur high decoding error rates (around 20%). In this paper, we pr…

2022

Hindsight is 20/20: Leveraging Past Traversals to Aid 3D Perception

ICLR 2022poster

Self-driving cars must detect vehicles, pedestrians, and other traffic participants accurately to operate safely. Small, far-away, or highly occluded objects are particularly challenging because there is limited information in the LiDAR point clouds for detecting them. To address this challenge, we l…

2022

Ithaca365: Dataset and Driving Perception Under Repeated and Challenging Weather Conditions

CVPR 2022poster

Advances in perception for self-driving cars have accelerated in recent years due to the availability of large-scale datasets, typically collected at specific locations and under nice weather conditions. Yet, to achieve the high safety requirement, these perceptual systems must operate robustly unde…

Cited by 53PDFScholar
2022

Low-drift LiDAR-only Odometry and Mapping for UGVs in Environments with Non-level Roads

IROS 2022poster

This study focuses on localization and mapping for UGVs when they are deployed in environments with non-level roads. In these scenarios, the vehicles need to travel through flat but not necessarily level grounds, i.e., ascent or descent, which may cause drifts of the robot pose and distortion of the…

Cited by 2SourceScholar
2022

Soft Capacitive Force Sensors With Low Hysteresis Based on Folded and Rolled Structures

RA-L 2022

Force sensors made of a polymer material with soft characteristics have application potential in the fields of soft robotics, exoskeletons, and human motion measurement. However, the hysteresis of soft force sensors is generally large because their sensing materials are rubbers with large dynamic vi

Cited by 9SourceScholar
2022

TAPE: Task-Agnostic Prior Embedding for Image Restoration

ECCV 2022poster

"Learning a generalized prior for natural image restoration is an important yet challenging task. Early methods mostly involved handcrafted priors including normalized sparsity, â„“0 gradients, dark channel priors, etc.. Recently, deep neural networks have been used to learn various image priors but…

Cited by 65SourcePDFScholar
2021

Low-Precision Reinforcement Learning: Running Soft Actor-Critic in Half Precision

ICML 2021spotlight

Low-precision training has become a popular approach to reduce compute requirements, memory footprint, and energy consumption in supervised learning. In contrast, this promising approach has not yet enjoyed similarly widespread adoption within the reinforcement learning (RL) community, partly becaus…

Cited by 30SourcePDFScholar
2021

Run Like a Dog: Learning Based Whole-Body Control Framework for Quadruped Gait Style Transfer

IROS 2021poster

In this paper, a learning-based whole-body loco-motion controller is proposed, which enables quadruped robots to perform running in the style of real animals. We use a low-level controller based on multi-rigid body dynamics to calculate desired torques for each joint, while the high-level neural net…

Cited by 8SourceScholar
2020

Combating Noisy Labels by Agreement: A Joint Training Method with Co-Regularization

CVPR 2020poster

Deep Learning with noisy labels is a practically challenging problem in weakly-supervised learning. The state-of-the-art approaches "Decoupling" and "Co-teaching+" claim that the "disagreement" strategy is crucial for alleviating the problem of learning with noisy labels. In this paper, we start fro…

Cited by 706PDFcodeScholar
2020

Gain Scheduled Controller Design for Balancing an Autonomous Bicycle

IROS 2020poster

In this paper, the gain scheduling technique is applied to design a balance controller for an autonomous bicycle with an inertia wheel. Previously, two different balance controllers are needed depending on whether the bicycle is stationary or dynamic. The switch between the two different controllers…

Cited by 16SourceScholar
2020

Nonlinear Balance Control of an Unmanned Bicycle: Design and Experiments

IROS 2020poster

In this paper, nonlinear control techniques are exploited to balance an unmanned bicycle with enlarged stability domain. We consider two cases. For the first case when the autonomous bicycle is balanced by the flywheel, the steering angle is set to zero, and the torque of the flywheel is used as the…

Cited by 26SourceScholar
2020

Train in Germany, Test in the USA: Making 3D Object Detectors Generalize

CVPR 2020poster

In the domain of autonomous driving, deep learning has substantially improved the 3D object detection accuracy for LiDAR and stereo camera data alike. While deep networks are great at generalization, they are also notorious to overfit to all kinds of spurious artifacts, such as brightness, car sizes…

Cited by 215PDFcodeScholar
2020

Transferable Active Grasping and Real Embodied Dataset

ICRA 2020poster

Grasping in cluttered scenes is challenging for robot vision systems, as detection accuracy can be hindered by partial occlusion of objects. We adopt a reinforcement learning (RL) framework and 3D vision architectures to search for feasible viewpoints for grasping by the use of hand-mounted RGB-D ca…

Cited by 26SourcecodeScholar
2019

Achievement of Online Agile Manipulation Task for Aerial Transformable Multilink Robot

IROS 2019poster

Transformable aerial robots are favorable in aerial manipulation tasks for their flexible ability to change configuration during the flight. By assuming robot keeping in the mild motion, the previous researches sacrifice aerial agility to simplify the complex non-linear system into a single rigid bo…

Cited by 6SourceScholar
2019

External Wrench Estimation for Multilink Aerial Robot by Center of Mass Estimator Based on Distributed IMU System

ICRA 2019poster

External wrench estimation is very helpful for aerial exploration and manipulation tasks. During the exploration, there might be unseen obstacles to cause dangerous collisions. The estimation of the external force and torque is also beneficial in aerial manipulation tasks. In this paper, we present…

Cited by 20SourceScholar
2019

Semantic Predictive Control for Explainable and Efficient Policy Learning

ICRA 2019poster

Visual anticipation of ego and object motion over a short time horizons is a key feature of human-level performance in complex environments. We propose a driving policy learning framework that predicts feature representations of future visual inputs; our predictive model infers not only future event…

Cited by 16SourceScholar
2019

TendencyRL: Multi-stage Discriminative Hints for Efficient Goal-Oriented Reverse Curriculum Learning

IROS 2019poster

Deep reinforcement learning algorithms have been proven successful in a variety of simulation tasks with dense reward feedback. However, real-world RL applications, e.g. robotic manipulation, remain challenging as most of them are multi-stage and a positive reward can only be received when the final…

Cited by 4SourceScholar
2019

Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection

ICCV 2019poster

Multispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not st…

Cited by 241PDFcodeScholar
2018

Design, Modeling, and Control of an Aerial Robot DRAGON: A Dual-Rotor-Embedded Multilink Robot With the Ability of Multi-Degree-of-Freedom Aerial Transformation

RA-L 2018

In this letter, we introduce a novel transformable aerial robot called DRAGON, which is a dual-rotor-embedded multilink robot with the ability of multi-degree-of-freedom (DoF) aerial transformation. The new aerial robot can control the full pose in SE(3) regarding the center of gravity (CoG) of mult

Cited by 160SourceScholar
2018

Flight Motion of Passing Through Small Opening by DRAGON: Transformable Multilinked Aerial Robot

IROS 2018poster

In this paper, we introduce the achievement of the flight motion to pass through small opening by the multilinked and transformable aerial robot. Previous works about such motion are based on under-actuated multirotors, indicating that aggressive maneuvering is necessary condition. This involves two…

Cited by 21SourceScholar
2018

Learning to Segment Generic Handheld Objects Using Class-Agnostic Deep Comparison and Segmentation Network

RA-L 2018

Learning unknown objects in the environment is important for detection and manipulation tasks. Prior to learning the unknown objects the ground-truth labels have to be provided. The data annotation or labeling can be achieved in a number of ways but the most widely used method is still manual annota

Cited by 7SourceScholar
2018

Predicting Part Affordances of Objects Using Two-Stream Fully Convolutional Network with Multimodal Inputs

IROS 2018poster

For a robot to manipulate an object, it has to understand the functions and the actions that can be subjected to the object. This set of information is known as affordance of the object. Affordances are generally defined by the geometrical structures and physical properties of the objects. In this p…

Cited by 15SourceScholar
2017

Multilinked multirotor with internal communication system for multiple objects transportation based on form optimization method

IROS 2017poster

In this paper, we show the achievement of a transformable aerial robot with internal communication system for multiple objects transportation. As it is not easy to make the flight endurance of an aerial robot longer, we study the problem to transport multiple objects at the same time to improve the…

Cited by 33SourceScholar
2017

Robust real-time visual tracking using dual-frame deep comparison network integrated with correlation filters

IROS 2017poster

In recent years, applications of visual tracking algorithms has seen a substantial growth with deployments in intelligent robots such as drones for human tracking. The algorithms for such tasks has to be efficient in terms of computational cost while been robust, accurate and fast. Object tracking a…

Cited by 21SourceScholar
2017

Whole-body aerial manipulation by transformable multirotor with two-dimensional multilinks

ICRA 2017poster

In this paper, we introduce the achievement of the aerial manipulation by using the whole body of a transformable aerial robot, instead of attaching an additional manipulator. The aerial robot in our work is composed by two-dimensional multilinks which enable a stable aerial transformation and can b…

Cited by 113SourceScholar
2016

Development of a low-cost ultra-tiny line laser range sensor

IROS 2016poster

To enable robotic sensing for tasks with requirements on weight, size, and cost, we develop an ultra-tiny line laser range sensor based on the Time-of-Flight (TOF) principle. With delicate circuit design and optical attachments, we create a sensor as small as 35[mm] × 27[mm] × 30[mm] and as light as…

Cited by 11SourceScholar
2015

Reasoning-based vision recognition for agricultural humanoid robot toward tomato harvesting

IROS 2015poster

We present a vision cognition framework for tomato harvesting humanoid robot based on geometrical and physical reasoning. Inspired from the natural human harvesting behaviour, our goal is to build a humanoid robot to pick tomatoes autonomously or with minimal human efforts. The proposed vision appro…

Cited by 52SourceScholar