← Search

Ruigang Yang

43 accepted papers

2026

DART: Distribution-Aware Adaptive Relational Transfer for Adversarial Attacks against Closed-Source MLLMs

ICML 2026poster

This paper studies the critical problem of targeted adversarial attacks against closed-source MLLMs, which aim to generate highly transferable adversarial samples with open-source MLLMs. Previous approaches typically focus on maximizing the similarity of latent representations between adversarial sa…

Cited by 0SourceScholar
2026

Real Garment Benchmark (RGBench): A Comprehensive Benchmark for Robotic Garment Manipulation Featuring a High-Fidelity Scalable Simulator

AAAI 2026technical

While there has been significant progress to use simulated data to learn robotic manipulation of rigid objects, applying its success to deformable objects has been hindered by the lack of both deformable object models and realistic non-rigid body simulators. In this paper, we present Real Garment Be

Cited by 0SourcePDFScholar
2025

Fuel-Optimal Operational Speed Planning for Autonomous Trucking on Highways

ICRA 2025

The rapid advancement of autonomous driving technology, particularly in autonomous trucking on highways, shows great value for enhancing efficiency and reducing costs in the logistics industry. In this work, we define the full-trip speed planning problem for autonomous trucks under delivery time and

Cited by 0SourceScholar
2025

OLiDM: Object-aware LiDAR Diffusion Models for Autonomous Driving

AAAI 2025technical

To enhance autonomous driving, innovative approaches have been proposed to generate simulated LiDAR data. However, these methods often face challenges in producing high-quality and controllable foreground objects. To cater to the needs of object-aware tasks in 3D perception, we introduce OLiDM, a no…

Cited by 1SourcePDFScholar
2024

DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection

AAAI 2024technical

Vehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent e…

2024

ESP: Extro-Spective Prediction for Long-term Behavior Reasoning in Emergency Scenarios

ICRA 2024poster

Emergent-scene safety is the key milestone for fully autonomous driving, and reliable on-time prediction is essential to maintain safety in emergency scenarios. However, these emergency scenarios are long-tailed and hard to collect, which restricts the system from getting reliable predictions. In th…

Cited by 1SourcecodeScholar
2024

IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection

CVPR 2024highlight

Bird's eye view (BEV) representation has emerged as a dominant solution for describing 3D space in autonomous driving scenarios. However objects in the BEV representation typically exhibit small sizes and the associated point cloud context is inherently sparse which leads to great challenges for rel…

2024

NPC: Neural Predictive Control for Fuel-Efficient Autonomous Trucks

ICRA 2024poster

Fuel efficiency is a crucial aspect of long-distance cargo transportation by oil-powered trucks that economize on costs and decrease carbon emissions. Current predictive control methods depend on an accurate model of vehicle dynamics and engine, including weight, drag coefficient, and the Brake-spec…

Cited by 0SourceScholar
2023

LWSIS: LiDAR-Guided Weakly Supervised Instance Segmentation for Autonomous Driving

AAAI 2023technical

Image instance segmentation is a fundamental research topic in autonomous driving, which is crucial for scene understanding and road safety. Advanced learning-based approaches often rely on the costly 2D mask annotations for training. In this paper, we present a more artful framework, LiDAR-guided…

2023

SSDA3D: Semi-supervised Domain Adaptation for 3D Object Detection from Point Cloud

AAAI 2023technical

LiDAR-based 3D object detection is an indispensable task in advanced autonomous driving systems. Though impressive detection results have been achieved by superior 3D detectors, they suffer from significant performance degeneration when facing unseen domains, such as different LiDAR configurations,…

2023

Transformation-Equivariant 3D Object Detection for Autonomous Driving

AAAI 2023technical

3D object detection received increasing attention in autonomous driving recently. Objects in 3D scenes are distributed with diverse orientations. Ordinary detectors do not explicitly model the variations of rotation and reflection transformations. Consequently, large networks and extensive data augm…

2022

STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded Scenes

CVPR 2022poster

Accurately detecting and tracking pedestrians in 3D space is challenging due to large variations in rotations, poses and scales. The situation becomes even worse for dense crowds with severe occlusions. However, existing benchmarks either only provide 2D annotations, or have limited 3D annotations w…

Cited by 49PDFcodeScholar
2022

TraEDITS: Diversity and Irregularity-Aware Traffic Trajectory Editing

RA-L 2022

We present TraEDITS, a novel traffic trajectory editing framework for autonomous vehicle testing, which can generate new traffic behaviors by controlling each vehicle interactively to increase the diversity or irregularity of traffic testing data. Given a traffic flow with its original trajectories,

Cited by 4SourceScholar
2021

Adaptive Surface Normal Constraint for Depth Estimation

ICCV 2021poster

We present a novel method for single image depth estimation using surface normal constraints. Existing depth estimation methods either suffer from the lack of geometric constraints, or are limited to the difficulty of reliably capturing geometric context, which leads to a bottleneck of depth estimat…

Cited by 72PDFcodeScholar
2020

3D Part Guided Image Editing for Fine-Grained Object Understanding

CVPR 2020poster

Holistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving ve…

Cited by 14PDFcodeScholar
2020

A Unified Object Motion and Affinity Model for Online Multi-Object Tracking

CVPR 2020poster

Current popular online multi-object tracking (MOT) solutions apply single object trackers (SOTs) to capture object motions, while often requiring an extra affinity network to associate objects, especially for the occluded ones. This brings extra computational overhead due to repetitive feature extra…

Cited by 139PDFcodeScholar
2020

AutoTrajectory: Label-free Trajectory Extraction and Prediction from Videos using Dynamic Points

ECCV 2020poster

Current methods for trajectory prediction operate in supervised manners, and therefore require vast quantities of corresponding ground truth data for training. In this paper, we present a novel, label-free algorithm, AutoTrajectory, for trajectory extraction and prediction to use raw videos directly…

2020

Channel Attention Based Iterative Residual Learning for Depth Map Super-Resolution

CVPR 2020poster

Despite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what…

Cited by 103PDFScholar
2020

DVI: Depth Guided Video Inpainting for Autonomous Driving

ECCV 2020poster

To get clear street-view and photo-realistic simulation in autonomous driving, we present an automatic video inpainting algorithm that can remove traffic agents from videos and synthesize missing regions with the guidance of depth/point cloud. By building a dense 3D map from stitched point clouds, f…

2020

Domain-invariant Stereo Matching Networks

ECCV 2020poster

State-of-the-art stereo matching networks have difficulties in generalizing to new unseen environments due to significant domain differences, such as color, illumination, contrast, and texture. In this paper, we aim at designing a domain-invariant stereo matching network (DSMNet) that generalizes we…

2020

FaceScape: A Large-Scale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction

CVPR 2020poster

In this paper, we present a large-scale detailed 3D face dataset, FaceScape, and propose a novel algorithm that is able to predict elaborate riggable 3D face models from a single image input. FaceScape dataset provides 18,760 textured 3D faces, captured from 938 subjects and each with 20 specific ex…

Cited by 361PDFcodeScholar
2020

Instance Segmentation of LiDAR Point Clouds

ICRA 2020poster

We propose a robust baseline method for instance segmentation which are specially designed for large-scale outdoor LiDAR point clouds. Our method includes a novel dense feature encoding technique, allowing the localization and segmentation of small, far-away objects, a simple but effective solution…

Cited by 73SourcecodeScholar
2020

Joint 3D Instance Segmentation and Object Detection for Autonomous Driving

CVPR 2020poster

Currently, in Autonomous Driving (AD), most of the 3D object detection frameworks (either anchor- or anchor-free-based) consider the detection as a Bounding Box (BBox) regression problem. However, this compact representation is not sufficient to explore all the information of the objects. To tackle…

Cited by 132PDFScholar
2020

Learning Resilient Behaviors for Navigation Under Uncertainty

ICRA 2020poster

Deep reinforcement learning has great potential to acquire complex, adaptive behaviors for autonomous agents automatically. However, the underlying neural network polices have not been widely deployed in real-world applications, especially in these safety-critical tasks (e.g., autonomous driving). O…

Cited by 28SourceScholar
2020

LiDAR-Based Online 3D Video Object Detection With Graph-Based Message Passing and Spatiotemporal Transformer Attention

CVPR 2020poster

Existing LiDAR-based 3D object detectors usually focus on the single-frame detection, while ignoring the spatiotemporal information in consecutive point cloud frames. In this paper, we propose an end-to-end online 3D video object detector that operates on point cloud sequences. The proposed model co…

Cited by 184PDFcodeScholar
2019

ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving

CVPR 2019poster

Autonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision communit…

Cited by 224PDFcodeScholar
2019

Compact Reachability Map for Excavator Motion Planning

IROS 2019poster

In this paper, we propose a novel compact reachability map representation for excavator motion planning. The constructed reachability map can concisely encode the bucket’s reachable pose and the translation capability limited by excavator’s kinematic structure. By explicitly exploiting the property…

Cited by 24SourceScholar
2019

Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh Deformation

CVPR 2019oral

This paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template…

Cited by 173PDFcodeScholar
2019

GA-Net: Guided Aggregation Net for End-To-End Stereo Matching

CVPR 2019oral

In the stereo matching task, matching cost aggregation is crucial in both traditional methods and deep neural network models in order to accurately estimate disparities. We propose two novel neural net layers, aimed at capturing local and the whole-image cost dependencies respectively. The first is…

Cited by 915PDFcodeScholar
2019

Getting Robots Unfrozen and Unlost in Dense Pedestrian Crowds

RA-L 2019

Our goal is to navigate a mobile robot to navigate through environments with dense crowds, e.g., shopping malls, canteens, train stations, or airport terminals. In these challenging environments, existing approaches suffer from two common problems: the robot may get frozen and cannot make any progre

Cited by 67SourceScholar
2018

DeLS-3D: Deep Localization and Segmentation With a 3D Semantic Map

CVPR 2018poster

For applications such as augmented reality, autonomous driving, self-localization/camera pose estimation and scene parsing are crucial technologies. In this paper, we propose a unified framework to tackle these two problems simultaneously. The uniqueness of our design is a sensor fusion scheme which…

2018

Depth Estimation via Affinity Learned with Convolutional Spatial Propagation Network

ECCV 2018poster

Depth estimation from a single image is a fundamental problem in computer vision. In this paper, we propose a simple yet effective convolutional spatial propagation network (CSPN) to learn the affinity matrix for depth prediction. Specifically, we adopt an efficient linear propagation model, where t…

2018

Learning Warped Guidance for Blind Face Restoration

ECCV 2018poster

This paper studies the problem of blind face restoration from an unconstrained blurry, noisy, low-resolution, or compressed image (i.e., degraded observation). For better recovery of fine facial details, we modify the problem setting by taking both the degraded observation and a high-quality guided…

2017

A generative human-robot motion retargeting approach using a single depth sensor

ICRA 2017poster

The goal of human-robot motion retargeting is to let a robot follow the movements performed by a human subject. This is traditionally achieved by applying the estimated poses from a human pose tracking system to a robot via explicit joint mapping strategies. In this paper, we present a novel approac…

Cited by 24SourceScholar
2017

Detailed Surface Geometry and Albedo Recovery From RGB-D Video Under Natural Illumination

ICCV 2017poster

In this paper we present a novel approach for depth map enhancement from an RGB-D video sequence. The basic idea is to exploit the photometric information in the color sequence. Instead of making any assumption about surface albedo or controlled object motion and lighting, we use the lighting variat…

Cited by 15PDFScholar
2015

3D Reconstruction in the Presence of Glasses by Acoustic and Stereo Fusion

CVPR 2015poster

We present a practical and inexpensive method to reconstruct 3D scenes that include piece-wise planar transparent objects. Our work is motivated by the need for automatically generating 3D models of interior scenes, in which glass structures are common. These large structures are often invisible to…

Cited by 44SourcePDFScholar
2015

Interactive Visual Hull Refinement for Specular and Transparent Object Surface Reconstruction

ICCV 2015poster

In this paper we present a method of using standard multi-view images for 3D surface reconstruction of non-Lambertian objects. We extend the original visual hull concept to incorporate 3D cues presented by internal occluding contours, i.e., occluding contours that are inside the object's silhouettes…

Cited by 23PDFScholar
2015

Simultaneous Time-of-Flight Sensing and Photometric Stereo With a Single ToF Sensor

CVPR 2015poster

We present a novel system which incorporates photometric stereo with the Time-of-Flight depth sensor. Adding to the classic ToF, the system utilizes multiple point light sources that enable the capturing of a normal field whilst taking depth images. Two calibration methods are proposed to determine…

Cited by 28SourcePDFScholar