← Search

Hongkai Yu

24 accepted papers

2026

DarkDriving: A Real-World Day and Night Aligned Dataset for Autonomous Driving in the Dark Environment

ICRA 2026poster

The low-light conditions are challenging to the vision-centric perception systems for autonomous driving in the dark environment. In this paper, we propose a new benchmark dataset (named DarkDriving) to investigate the low-light enhancement for autonomous driving. The existing real-world low-light e…

2026

Unsupervised Multi-agent and Single-agent Perception from Cooperative Views

CVPR 2026

The LiDAR-based multi-agent and single-agent perception has shown promising performance in environmental understanding for robots and automated vehicles. However, there is no existing method that simultaneously solves both multi-agent and single-agent perception in an unsupervised way. By sharing se

Cited by 0SourceScholar
2025

An Efficient and Accurate Dynamic Sparse Training Framework Based on Parameter-Freezing

AAAI 2025technical

Federated learning is a decentralized machine learning approach that consists of servers and clients. It protects data privacy during model training by keeping the training data locally in each client. However, the requirement for the server and clients to frequently synchronize the parameters of th…

2025

Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis

ICCV 2025poster

Egocentricly comprehending the causes and effects of car accidents is crucial for the safety of self-driving cars, and synthesizing causal-entity reflected accident videos can facilitate the capability test to respond to unaffordable accidents in reality. However, incorporating causal relations as s…

Cited by 0SourcePDFScholar
2025

Co-Fix3D: Enhancing 3D Object Detection With Collaborative Refinement

RA-L 2025

3D object detection in driving scenarios is particularly challenging due to factors such as sensor noise, occlusions, and the inherent sparsity of LiDAR point clouds, which can lead to the loss or incompleteness of key features, in turn affecting perception performance. To address these challenges,

Cited by 0SourcecodeScholar
2025

CoMamba: Real-time Cooperative Perception Unlocked with State-Space Models

IROS 2025

Cooperative perception systems play a vital role in enhancing the safety and efficiency of vehicular autonomy. Although recent studies have highlighted the efficacy of vehicle-to-everything (V2X) communication techniques in autonomous driving, a significant challenge persists: how to efficiently int

Cited by 7SourceScholar
2025

Robust Multi-task Adversarial Attacks Using Min-max Optimization

ICASSP 2025accepted

Deep neural networks have achieved exceptional performance across a wide range of applications but remain susceptible to adversarial attacks. While most prior research has focused on single-task scenarios, increasing attention is being directed toward adversarial attacks targeting multiple tasks sim…

Cited by 0SourceScholar
2025

V2X-DG: Domain Generalization for Vehicle-to-Everything Cooperative Perception

ICRA 2025

LiDAR-based Vehicle-to-Everything (V2X) cooperative perception has demonstrated its impact on the safety and effectiveness of autonomous driving. Since current cooperative perception algorithms are trained and tested on the same dataset, the generalization ability of cooperative perception systems r

Cited by 3SourceScholar
2025

V2X-DGW: Domain Generalization for Multi-Agent Perception Under Adverse Weather Conditions

ICRA 2025

Current LiDAR-based Vehicle-to-Everything (V2X) multi-agent perception systems have shown the significant success on 3D object detection. While these models perform well in the trained clean weather, they struggle in unseen adverse weather conditions with the domain gap. In this paper, we propose a

Cited by 19SourcecodeScholar
2024

Abductive Ego-View Accident Video Understanding for Safe Driving Perception

CVPR 2024highlight

We present MM-AU a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11727 in-the-wild ego-view accident videos each with temporally aligned text descriptions. We annotate over 2.23 million object boxes and 58650 pairs of video-based accident reasons covering 58 accident cat…

Cited by 12SourcePDFScholar
2024

AdvGPS: Adversarial GPS for Multi-Agent Perception Attack

ICRA 2024poster

The multi-agent perception system collects visual data from sensors located on various agents and leverages their relative poses determined by GPS signals to effectively fuse information, mitigating the limitations of single-agent sensing, such as occlusion. However, the precision of GPS signals can…

Cited by 6SourcecodeScholar
2024

Breaking Data Silos: Cross-Domain Learning for Multi-Agent Perception from Independent Private Sources

ICRA 2024poster

The diverse agents in multi-agent perception systems may be from different companies. Each company might use the identical classic neural network architecture based encoder for feature extraction. However, the data source to train the various agents is independent and private in each company, leadin…

Cited by 7SourcecodeScholar
2024

Light the Night: A Multi-Condition Diffusion Framework for Unpaired Low-Light Enhancement in Autonomous Driving

CVPR 2024poster

Vision-centric perception systems for autonomous driving have gained considerable attention recently due to their cost-effectiveness and scalability especially compared to LiDAR-based systems. However these systems often struggle in low-light conditions potentially compromising their performance and…

Cited by 24SourcePDFScholar
2024

S2R-ViT for Multi-Agent Cooperative Perception: Bridging the Gap from Simulation to Reality

ICRA 2024poster

Due to the lack of enough real multi-agent data and time-consuming of labeling, existing multi-agent cooperative perception algorithms usually select the simulated sensor data for training and validating. However, the perception performance is degraded when these simulation-trained models are deploy…

Cited by 21SourceScholar
2024

SQLdepth: Generalizable Self-Supervised Fine-Structured Monocular Depth Estimation

AAAI 2024technical

Recently, self-supervised monocular depth estimation has gained popularity with numerous applications in autonomous driving and robotics. However, existing solutions primarily seek to estimate depth from immediate visual features, and struggle to recover fine-grained scene details. In this paper, we…

2023

V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception

CVPR 2023highlight

Modern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception…

2022

Can You Spot the Chameleon? Adversarially Camouflaging Images From Co-Salient Object Detection

CVPR 2022poster

Co-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods…

Cited by 25PDFcodeScholar
2021

Auto-Exposure Fusion for Single-Image Shadow Removal

CVPR 2021poster

Shadow removal is still a challenging task due to its inherent background-dependent and spatial-variant properties, leading to unknown and diverse shadow patterns. Even powerful deep neural networks could hardly recover traceless shadow-removed background. This paper proposes a new solution for this…

Cited by 174PDFcodeScholar
2019

Visual Attention Consistency Under Image Transforms for Multi-Label Image Classification

CVPR 2019poster

Human visual perception shows good consistency for many multi-label image classification tasks under certain spatial transforms, such as scaling, rotation, flipping and translation. This has motivated the data augmentation strategy widely used in CNN classifier training -- transformed images are inc…

Cited by 307PDFScholar
2017

Learning View-Invariant Features for Person Identification in Temporally Synchronized Videos Taken by Wearable Cameras

ICCV 2017poster

In this paper, we study the problem of Cross-View Person Identification (CVPI), which aims at identifying the same person from temporally synchronized videos taken by different wearable cameras. Our basic idea is to utilize the human motion consistency for CVPI, where human motion can be computed by…

Cited by 28PDFScholar
2016

Groupwise Tracking of Crowded Similar-Appearance Targets From Low-Continuity Image Sequences

CVPR 2016spotlight

Automatic tracking of large-scale crowded targets are of particular importance in many applications, such as crowded people/vehicle tracking in video surveillance, fiber tracking in materials science, and cell tracking in biomedical imaging. This problem becomes very challenging when the targets sho…

Cited by 37PDFScholar
2015

Co-Interest Person Detection From Multiple Wearable Camera Videos

ICCV 2015poster

Wearable cameras, such as Google Glass and Go Pro, enable video data collection over larger areas and from different views. In this paper, we tackle a new problem of locating the co-interest person (CIP), i.e., the one who draws attention from most camera wearers, from temporally synchronized videos…

Cited by 29PDFScholar