← Search

Yuenan Hou

23 accepted papers

2026

RaCFusion: Improving Camera-Based 3D Object Detection via Radar-Assisted Hierarchical Refinement

RA-L 2026

Cameras and radar sensors are complementary in 3D object detection in that cameras specialize in capturing an object's visual information while radar provides spatial information and velocity hints. Existing radar-camera fusion methods often employ a symmetrical architecture that processes inputs fr

Cited by 0SourcecodeScholar
2026

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

AAAI 2026technical

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research

Cited by 0SourcePDFScholar
2025

OccMamba: Semantic Occupancy Prediction with State Space Models

CVPR 2025poster

Training deep learning models for semantic occupancy prediction is challenging due to factors such as a large number of occupancy cells, severe occlusion, limited visual cues, complicated driving scenarios, etc. Recent methods often adopt transformer-based architectures given their strong capability…

2024

Frozen CLIP Transformer Is an Efficient Point Cloud Encoder

AAAI 2024technical

The pretrain-finetune paradigm has achieved great success in NLP and 2D image fields because of the high-quality representation ability and transferability of their pretrained models. However, pretraining such a strong model is difficult in the 3D point cloud field due to the limited amount of point…

2024

Point Cloud Pre-training with Diffusion Models

CVPR 2024poster

Pre-training a model and then fine-tuning it on downstream tasks has demonstrated significant success in the 2D image and NLP domains. However due to the unordered and non-uniform density characteristics of point clouds it is non-trivial to explore the prior knowledge of point clouds and pre-train a…

2024

Semi-supervised 3D Object Detection with PatchTeacher and PillarMix

AAAI 2024technical

Semi-supervised learning aims to leverage numerous unlabeled data to improve the model performance. Current semi-supervised 3D object detection methods typically use a teacher to generate pseudo labels for a student, and the quality of the pseudo labels is essential for the final performance. In thi…

2024

TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation

CVPR 2024poster

Training deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the sparsity problem as it makes the input signal denser. However previous multi-frame fusion algorithms fall short in utilizing suffi…

2024

WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language

ECCV 2024poster

"We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a novel method, dubbed WildRefer, for this task by fully utilizing the rich appe…

2023

CLIP2Scene: Towards Label-Efficient 3D Scene Understanding by CLIP

CVPR 2023poster

Contrastive Language-Image Pre-training (CLIP) achieves promising results in 2D zero-shot and few-shot learning. Despite the impressive performance in 2D, applying CLIP to help the learning in 3D scene understanding has yet to be explored. In this paper, we make the first attempt to investigate how…

2023

CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection

NeurIPS 2023poster

Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution l…

Cited by 6SourcePDFScholar
2023

Human-centric Scene Understanding for 3D Large-scale Scenarios

ICCV 2023poster

Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc. In this paper, we present a large-scale multi-modal dataset…

Cited by 26PDFcodeScholar
2023

LoGoNet: Towards Accurate 3D Object Detection With Local-to-Global Cross-Modal Fusion

CVPR 2023poster

LiDAR-camera fusion methods have shown impressive performance in 3D object detection. Recent advanced multi-modal methods mainly perform global fusion, where image features and point cloud features are fused across the whole scene. Such practice lacks fine-grained region-level information, yielding…

2023

RangePerception: Taming LiDAR Range View for Efficient and Accurate 3D Object Detection

NeurIPS 2023poster

LiDAR-based 3D detection methods currently use bird's-eye view (BEV) or range view (RV) as their primary basis. The former relies on voxelization and 3D convolutions, resulting in inefficient training and inference processes. Conversely, RV-based methods demonstrate higher efficiency due to their co…

Cited by 8SourcePDFScholar
2023

Rethinking Range View Representation for LiDAR Segmentation

ICCV 2023poster

LiDAR segmentation is crucial for autonomous driving perception. Recent trends favor point- or voxel-based methods as they often yield better performance than the traditional range view representation. In this work, we unveil several key factors in building powerful range view models. We observe tha…

Cited by 173PDFScholar
2023

SCPNet: Semantic Scene Completion on Point Cloud

CVPR 2023highlight

Training deep models for semantic scene completion is challenging due to the sparse and incomplete input, a large quantity of objects of diverse scales as well as the inherent label noise for moving objects. To address the above-mentioned problems, we propose the following three solutions: 1) Redesi…

Cited by 95SourcePDFScholar
2023

See More and Know More: Zero-shot Point Cloud Segmentation via Multi-modal Visual Data

ICCV 2023poster

Zero-shot point cloud segmentation aims to make deep models capable of recognizing novel objects in point cloud that are unseen in the training phase. Recent trends favor the pipeline which transfers knowledge from seen classes with labels to unseen classes without labels. They typically align visua…

Cited by 33PDFScholar
2023

UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase

ICCV 2023poster

Point-, voxel-, and range-views are three representative forms of point clouds. All of them have accurate 3D measurements but lack color and texture information. RGB images are a natural complement to these point cloud views and fully utilizing the comprehensive information of them benefits more rob…

Cited by 46PDFcodeScholar
2022

Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection

ECCV 2022poster

"Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D points and dense 2D pixels. Recent approaches either fuse the image features with the point cloud features that are pr…

2022

Point-to-Voxel Knowledge Distillation for LiDAR Semantic Segmentation

CVPR 2022poster

This article addresses the problem of distilling knowledge from a large teacher model to a slim student network for LiDAR semantic segmentation. Directly employing previous distillation approaches yields inferior results due to the intrinsic challenges of point cloud, i.e., sparsity, randomness and…

Cited by 215PDFcodeScholar
2022

STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded Scenes

CVPR 2022poster

Accurately detecting and tracking pedestrians in 3D space is challenging due to large variations in rotations, poses and scales. The situation becomes even worse for dense crowds with severe occlusions. However, existing benchmarks either only provide 2D annotations, or have limited 3D annotations w…

Cited by 49PDFcodeScholar
2020

Inter-Region Affinity Distillation for Road Marking Segmentation

CVPR 2020poster

We study the problem of distilling knowledge from a large deep teacher network to a much smaller student network for the task of road marking segmentation. In this work, we explore a novel knowledge distillation (KD) approach that can transfer 'knowledge' on scene structure more effectively from a t…

Cited by 156PDFcodeScholar
2019

Learning Lightweight Lane Detection CNNs by Self Attention Distillation

ICCV 2019poster

Training deep models for lane detection is challenging due to the very subtle and sparse supervisory signals inherent in lane annotations. Without learning from much richer context, these models often fail in challenging scenarios, e.g., severe occlusion, ambiguous lanes, and poor lighting condition…

Cited by 824PDFcodeScholar