← Search

Dengxin Dai

66 accepted papers

2024

FreePoint: Unsupervised Point Cloud Instance Segmentation

CVPR 2024poster

Instance segmentation of point clouds is a crucial task in 3D field with numerous applications that involve localizing and segmenting objects in a scene. However achieving satisfactory results requires a large number of manual annotations which is time-consuming and expensive. To alleviate dependenc…

2024

MICDrop: Masking Image and Depth Features via Complementary Dropout for Domain-Adaptive Semantic Segmentation

ECCV 2024poster

"Unsupervised Domain Adaptation (UDA) is the task of bridging the domain gap between a labeled source domain, e.g., synthetic data, and an unlabeled target domain. We observe that current UDA methods show inferior results on fine structures and tend to oversegment objects with ambiguous appearance.…

2024

Object-centric Cross-modal Feature Distillation for Event-based Object Detection

ICRA 2024poster

Event cameras are gaining popularity due to their unique properties, such as their low latency and high dynamic range. One task where these benefits can be crucial is real-time object detection. However, RGB detectors still outperform event-based detectors due to the sparsity of the event data and m…

Cited by 9SourceScholar
2024

U-BEV: Height-aware Bird’s-Eye-View Segmentation and Neural Map-based Relocalization

IROS 2024poster

Efficient relocalization is essential for intelligent vehicles when GPS reception is insufficient or sensor-based localization fails. Recent advances in Bird’s-Eye-View (BEV) segmentation allow for accurate estimation of local scene appearance and in turn, can benefit the relocalization of the vehic…

Cited by 11SourceScholar
2023

Continuous Pseudo-Label Rectified Domain Adaptive Semantic Segmentation With Implicit Neural Representations

CVPR 2023poster

Unsupervised domain adaptation (UDA) for semantic segmentation aims at improving the model performance on the unlabeled target domain by leveraging a labeled source domain. Existing approaches have achieved impressive progress by utilizing pseudo-labels on the unlabeled target-domain images. Yet the…

2023

EDAPS: Enhanced Domain-Adaptive Panoptic Segmentation

ICCV 2023poster

With autonomous industries on the rise, domain adaptation of the visual perception stack is an important research direction due to the cost savings promise. Much prior art was dedicated to domain-adaptive semantic segmentation in the synthetic-to-real context. Despite being a crucial output of the p…

Cited by 14PDFcodeScholar
2023

Federated Incremental Semantic Segmentation

CVPR 2023poster

Federated learning-based semantic segmentation (FSS) has drawn widespread attention via decentralized training on local clients. However, most FSS models assume categories are fxed in advance, thus heavily undergoing forgetting on old categories in practical applications where local clients receive…

2023

HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic Segmentation

CVPR 2023poster

Current semantic segmentation models have achieved great success under the independent and identically distributed (i.i.d.) condition. However, in real-world applications, test data might come from a different domain than training data. Therefore, it is important to improve model robustness against…

2023

MIC: Masked Image Consistency for Context-Enhanced Domain Adaptation

CVPR 2023poster

In unsupervised domain adaptation (UDA), a model trained on source data (e.g. synthetic) is adapted to target data (e.g. real-world) without access to target annotation. Most previous UDA methods struggle with classes that have a similar visual appearance on the target domain as no ground truth is a…

2023

SSB: Simple but Strong Baseline for Boosting Performance of Open-Set Semi-Supervised Learning

ICCV 2023poster

Semi-supervised learning (SSL) methods effectively leverage unlabeled data to improve model generalization. However, SSL models often underperform in open-set scenarios, where unlabeled data contain outliers from novel categories that do not appear in the labeled set. In this paper, we study the cha…

Cited by 15PDFcodeScholar
2023

Self-Supervised Pre-Training With Masked Shape Prediction for 3D Scene Understanding

CVPR 2023poster

Masked signal modeling has greatly advanced self-supervised pre-training for language and 2D images. However, it is still not fully explored in 3D scene understanding. Thus, this paper introduces Masked Shape Prediction (MSP), a new framework to conduct masked signal modeling in 3D scenes. MSP uses…

2023

Towards Robust Object Detection Invariant to Real-World Domain Shifts

ICLR 2023poster

Safety-critical applications such as autonomous driving require robust object detection invariant to real-world domain shifts. Such shifts can be regarded as different domain styles, which can vary substantially due to environment changes and sensor noises, but deep models only know the training dom…

Cited by 37SourcePDFScholar
2023

TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction

ICRA 2023poster

Data-driven simulation has become a favorable way to train and test autonomous driving algorithms. The idea of replacing the actual environment with a learned simulator has also been explored in model-based reinforcement learning in the context of world models. In this work, we show data-driven traf…

Cited by 48SourceScholar
2023

Weakly-Supervised Domain Adaptive Semantic Segmentation With Prototypical Contrastive Learning

CVPR 2023poster

There has been a lot of effort in improving the performance of unsupervised domain adaptation for semantic segmentation task, however there is still a huge gap in performance when compared with supervised learning. In this work, we propose a common framework to use different weak labels, e.g. image,…

2022

Adiabatic Quantum Computing for Multi Object Tracking

CVPR 2022poster

Multi-Object Tracking (MOT) is most often approached in the tracking-by-detection paradigm, where object detections are associated through time. The association step naturally leads to discrete optimization problems. As these optimization problems are often NP-hard, they can only be solved exactly f…

Cited by 34PDFScholar
2022

Bi-Level Alignment for Cross-Domain Crowd Counting

CVPR 2022poster

Recently, crowd density estimation has received increasing attention. The main challenge for this task is to achieve high-quality manual annotations on a large amount of training data. To avoid reliance on such annotations, previous works apply unsupervised domain adaptation (UDA) techniques by tran…

Cited by 40PDFcodeScholar
2022

Both Style and Fog Matter: Cumulative Domain Adaptation for Semantic Foggy Scene Understanding

CVPR 2022oral

Although considerable progress has been made in semantic scene understanding under clear weather, it is still a tough problem under adverse weather conditions, such as dense fog, due to the uncertainty caused by imperfect observations. Besides, difficulties in collecting and labeling foggy images hi…

Cited by 64PDFScholar
2022

Class-Agnostic Object Counting Robust to Intraclass Diversity

ECCV 2022poster

"Most previous works on object counting are limited to pre-defined categories. In this paper, we focus on classagnostic counting, i.e., counting object instances in an image by simply specifying a few exemplar boxes of interest. We start with an analysis on intraclass diversity and point out three f…

2022

CoSSL: Co-Learning of Representation and Classifier for Imbalanced Semi-Supervised Learning

CVPR 2022poster

Standard semi-supervised learning (SSL) using class-balanced datasets has shown great progress to leverage unlabeled data effectively. However, the more realistic setting of class-imbalanced data - called imbalanced SSL - is largely underexplored and standard SSL tends to underperform. In this paper…

Cited by 69PDFcodeScholar
2022

DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic Segmentation

CVPR 2022poster

As acquiring pixel-wise annotations of real-world images for semantic segmentation is a costly process, a model can instead be trained with more accessible synthetic data and adapted to real images without requiring their annotations. This process is studied in unsupervised domain adaptation (UDA).…

Cited by 617PDFcodeScholar
2022

End-to-End Optimization of LiDAR Beam Configuration for 3D Object Detection and Localization

RA-L 2022

Existing learning methods for LiDAR-based applications use 3D points scanned under a pre-determined beam configuration, e.g., the elevation angles of beams are often evenly distributed. Those fixed configurations are task-agnostic, so simply using them can lead to sub-optimal performance. In this wo

Cited by 17SourcecodeScholar
2022

HRDA: Context-Aware High-Resolution Domain-Adaptive Semantic Segmentation

ECCV 2022poster

"Unsupervised domain adaptation (UDA) aims to adapt a model trained on the source domain (e.g. synthetic data) to the target domain (e.g. real-world data) without requiring further annotations on the target domain. This work focuses on UDA for semantic segmentation as real-world pixel-wise annotatio…

2022

Learnable Online Graph Representations for 3D Multi-Object Tracking

RA-L 2022

Autonomous systems that operate in dynamic environments require robust object tracking in 3D as one of their key components. Most recent approaches for 3D multi-object tracking (MOT) from LIDAR use object dynamics together with a set of handcrafted features to match detections of objects across mult

Cited by 78SourceScholar
2022

LiDAR Snowfall Simulation for Robust 3D Object Detection

CVPR 2022oral

3D object detection is a central task for applications such as autonomous driving, in which the system needs to localize and classify surrounding traffic agents, even in the presence of adverse weather. In this paper, we address the problem of LiDAR-based 3D object detection under snowfall. Due to t…

Cited by 147PDFcodeScholar
2022

Motion Transformer with Global Intention Localization and Local Movement Refinement

NeurIPS 2022accept

Predicting multimodal future behavior of traffic participants is essential for robotic vehicles to make safe decisions. Existing works explore to directly predict future trajectories based on latent features or utilize dense goal candidates to identify agent's destinations, where the former strategy…

2022

Multi-Scale Interaction for Real-Time LiDAR Data Segmentation on an Embedded Platform

RA-L 2022

Real-time semantic segmentation of LiDAR data is crucial for autonomously driving vehicles and robots, which are usually equipped with an embedded platform and have limited computational resources. Approaches that operate directly on the point cloud use complex spatial aggregation operations, which

Cited by 101SourcecodeScholar
2022

Pix2NeRF: Unsupervised Conditional p-GAN for Single Image to Neural Radiance Fields Translation

CVPR 2022poster

We propose a pipeline to generate Neural Radiance Fields (NeRF) of an object or a scene of a specific class, conditioned on a single input image. This is a challenging task, as training NeRF requires multiple views of the same scene, coupled with corresponding poses, which are hard to obtain. Our me…

Cited by 105PDFcodeScholar
2022

TACS: Taxonomy Adaptive Cross-Domain Semantic Segmentation

ECCV 2022poster

"Traditional domain adaptive semantic segmentation addresses the task of adapting a model to a novel target domain under limited or no additional supervision. While tackling the input domain gap, the standard domain adaptation settings assume no domain change in the output space. In semantic predict…

2021

ACDC: The Adverse Conditions Dataset With Correspondences for Semantic Driving Scene Understanding

ICCV 2021poster

Level 5 autonomy for self-driving cars requires a robust visual perception system that can parse input images under any visual condition. However, existing semantic segmentation datasets are either dominated by images captured under normal conditions or are small in scale. To address this, we introd…

Cited by 705PDFScholar
2021

Analogical Image Translation for Fog Generation

AAAI 2021technical

Image-to-image translation is to map images from a given style to another given style. While exceptionally successful, current methods assume the availability of training images in both source and target domains, which does not always hold in practice. Inspired by humans' reasoning capability of ana…

Cited by 18SourcePDFScholar
2021

Cluster, Split, Fuse, and Update: Meta-Learning for Open Compound Domain Adaptive Semantic Segmentation

CVPR 2021poster

Open compound domain adaptation (OCDA) is a domain adaptation setting, where target domain is modeled as a compound of multiple unknown homogeneous domains, which brings the advantage of improved generalization to unseen domains. In this work, we propose a principled meta-learning based approach to…

Cited by 45PDFScholar
2021

Domain Adaptive Semantic Segmentation With Self-Supervised Depth Estimation

ICCV 2021poster

Domain adaptation for semantic segmentation aims to improve the model performance in the presence of a distribution shift between source and target domain. Leveraging the supervision from auxiliary tasks (such as depth estimation) has the potential to heal this shift because many visual tasks are cl…

Cited by 169PDFcodeScholar
2021

End-to-End Urban Driving by Imitating a Reinforcement Learning Coach

ICCV 2021poster

End-to-end approaches to autonomous driving commonly rely on expert demonstrations. Although humans are good drivers, they are not good coaches for end-to-end algorithms that demand dense on-policy supervision. On the contrary, automated experts that leverage privileged information can efficiently g…

Cited by 232PDFcodeScholar
2021

Fog Simulation on Real LiDAR Point Clouds for 3D Object Detection in Adverse Weather

ICCV 2021poster

This work addresses the challenging task of LiDAR-based 3D object detection in foggy weather. Collecting and annotating data in such a scenario is very time, labor and cost intensive. In this paper, we tackle this problem by simulating physically accurate fog into clear-weather scenes, so that the a…

Cited by 185PDFcodeScholar
2021

Self-Aligned Video Deraining With Transmission-Depth Consistency

CVPR 2021poster

In this paper, we address the problems of rain streaks and rain accumulation removal in video, by developing a self-aligned network with transmission-depth consistency. Existing video based deraining method focus only on rain streak removal, and commonly use optical flow to align the rain video fram…

Cited by 40PDFcodeScholar
2021

Spectral Tensor Train Parameterization of Deep Learning Layers

AISTATS 2021poster

We study low-rank parameterizations of weight matrices with embedded spectral properties in the Deep Learning context. The low-rank property leads to parameter efficiency and permits taking computational shortcuts when computing mappings. Spectral properties are often subject to constraints in optim…

2021

Task Switching Network for Multi-Task Learning

ICCV 2021poster

We introduce Task Switching Networks (TSNs), a task-conditioned architecture with a single unified encoder/decoder for efficient multi-task learning. Multiple tasks are performed by switching between them, performing one task at a time. TSNs have a constant number of parameters irrespective of the n…

Cited by 64PDFScholar
2021

Three Ways To Improve Semantic Segmentation With Self-Supervised Depth Estimation

CVPR 2021poster

Training deep networks for semantic segmentation requires large amounts of labeled training data, which presents a major challenge in practice, as labeling segmentation masks is a highly labor-intensive process. To address this issue, we present a framework for semi-supervised semantic segmentation,…

Cited by 115PDFcodeScholar
2021

mDALU: Multi-Source Domain Adaptation and Label Unification With Partial Datasets

ICCV 2021poster

One challenge of object recognition is to generalize to new domains, to more classes and/or to new modalities. This necessitates methods to combine and reuse existing datasets that may belong to different domains, have partial annotations, and/or have different data modalities. This paper formulates…

Cited by 28PDFScholar
2020

Action Sequence Predictions of Vehicles in Urban Environments using Map and Social Context

IROS 2020poster

This work studies the problem of predicting the sequence of future actions for surrounding vehicles in real-world driving scenarios. To this aim, we make three main contributions. The first contribution is an automatic method to convert the trajectories recorded in real-world driving scenarios to ac…

Cited by 13SourceScholar
2020

Don't Forget The Past: Recurrent Depth Estimation from Monocular Video

RA-L 2020

Autonomous cars need continuously updated depth information. Thus far, depth is mostly estimated independently for a single frame at a time, even if the method starts from video input. Our method produces a time series of depth maps, which makes it an ideal candidate for online learning approaches.

Cited by 152SourceScholar
2020

Learning Accurate and Human-Like Driving using Semantic Maps and Attention

IROS 2020poster

This paper investigates how end-to-end driving models can be improved to drive more accurately and human-like. To tackle the first issue we exploit semantic and visual maps from HERE Technologies and augment the existing Drive360 dataset with such. The maps are used in an attention mechanism that pr…

Cited by 26SourceScholar
2020

Nighttime Defogging Using High-Low Frequency Decomposition and Grayscale-Color Networks

ECCV 2020poster

In this paper, we address the problem of nighttime defogging from a single image. We propose a framework consisting of two main modules: grayscale and color modules. Given an RGB foggy nighttime image, our grayscale module takes the grayscale version of the image as input, and decomposes it into hig…

Cited by 58SourcePDFScholar
2020

Off-Policy Reinforcement Learning for Efficient and Effective GAN Architecture Search

ECCV 2020poster

In this paper, we introduce a new reinforcement learning (RL) based neural architecture search (NAS) methodology for effective and efficient generative adversarial network (GAN) architecture search. The key idea is to formulate the GAN architecture search problem as a Markov decision process (MDP) f…

2020

Semantic Object Prediction and Spatial Sound Super-Resolution with Binaural Sounds

ECCV 2020poster

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic labelling of sound-making objects, purely based on binaural sound…

Cited by 55SourcePDFScholar
2020

T-Basis: a Compact Representation for Neural Networks

ICML 2020poster

We introduce T-Basis, a novel concept for a compact representation of a set of tensors, each of an arbitrary shape, which is often seen in Neural Networks. Each of the tensors in the set is modeled using Tensor Rings, though the concept applies to other Tensor Networks. Owing its name to the T-shape…

2020

Weakly Supervised 3D Object Detection from Lidar Point Cloud

ECCV 2020poster

It is laborious to manually label point cloud data for training high-quality 3D object detectors. This work proposes a weakly supervised approach for 3D object detection, only requiring a small set of weakly annotated scenes, associated with a few precisely labeled object instances. This is achieved…

2019

Guided Curriculum Model Adaptation and Uncertainty-Aware Evaluation for Semantic Nighttime Image Segmentation

ICCV 2019poster

Most progress in semantic segmentation reports on daytime images taken under favorable illumination conditions. We instead address the problem of semantic segmentation of nighttime images and improve the state-of-the-art, by adapting daytime models to nighttime without using nighttime annotations. M…

Cited by 311PDFcodeScholar
2018

Domain Adaptive Faster R-CNN for Object Detection in the Wild

CVPR 2018poster

Object detection typically assumes that training and test data are drawn from an identical distribution, which, however, does not always hold in practice. Such a distribution mismatch will lead to a significant performance drop. In this work, we aim to improve the cross-domain robustness of object d…

2018

End-to-End Learning of Driving Models with Surround-View Cameras and Route Planners

ECCV 2018poster

For human drivers, having rear and side-view mirrors is vital for safe driving. They deliver a more complete view of what is happening around the car. Human drivers also heavily exploit their mental map for navigation. Nonetheless, several methods have been published that learn driving models with o…

Cited by 213SourcePDFScholar
2018

Model Adaptation with Synthetic and Real Data for Semantic Dense Foggy Scene Understanding

ECCV 2018poster

This work addresses the problem of semantic scene understanding under dense fog. Although considerable progress has been made in semantic scene understanding, it is mainly related to clear-weather scenes. Extending recognition methods to adverse weather conditions such as fog is crucial for outdoor…

Cited by 285SourcePDFScholar
2015

Metric Imitation by Manifold Transfer for Efficient Vision Applications

CVPR 2015poster

Metric learning has proved very successful. However, human annotations are necessary. In this paper, we propose an unsupervised method, dubbed Metric Imitation (MI), where metrics over one cheap feature (target features, TFs) are learned by imitating the standard metrics over another sophisticated,…

Cited by 0SourcePDFScholar