← Search

Shuangjie Xu

10 accepted papers

2024

Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection

ECCV 2024poster

"The Bird’s-Eye-View (BEV) representation is a critical factor that directly impacts the 3D object detection performance, but the traditional BEV grid representation induces quadratic computational cost as the spatial resolution grows. To address this limitation, we present a new camera-based 3D obj…

2024

PPAD: Iterative Interactions of Prediction and Planning for End-to-end Autonomous Driving

ECCV 2024poster

"We present a new interaction mechanism of prediction and planning for end-to-end autonomous driving, called PPAD (Iterative Interaction of Prediction and Planning Autonomous Driving), which considers the timestep-wise interaction to better integrate prediction and planning. An ego vehicle performs…

2023

SVQNet: Sparse Voxel-Adjacent Query Network for 4D Spatio-Temporal LiDAR Semantic Segmentation

ICCV 2023poster

LiDAR-based semantic perception tasks are critical yet challenging for autonomous driving. Due to the motion of objects and static/dynamic occlusion, temporal information plays an essential role in reinforcing perception by enhancing and completing single-frame knowledge. Previous approaches either…

Cited by 10PDFScholar
2022

Efficient Point Cloud Segmentation with Geometry-Aware Sparse Networks

ECCV 2022poster

"In point cloud learning, sparsity and geometry are two core properties. Recently, many approaches have been proposed through single or multiple representations to improve the performance of point cloud semantic segmentation. However, these works fail to maintain the balance among performance, effic…

Cited by 26SourcePDFScholar
2022

Sparse Cross-Scale Attention Network for Efficient LiDAR Panoptic Segmentation

AAAI 2022technical

Two major challenges of 3D LiDAR Panoptic Segmentation (PS) are that point clouds of an object are surface-aggregated and thus hard to model the long-range dependency especially for large instances, and that objects are too close to separate each other. Recent literature addresses these problems by…

Cited by 46SourcePDFScholar
2021

DRINet: A Dual-Representation Iterative Learning Network for Point Cloud Segmentation

ICCV 2021poster

We present a novel and flexible architecture for point cloud segmentation with dual-representation iterative learning. In point cloud processing, different representations have their own pros and cons. Thus, finding suitable ways to represent point cloud data structure while keeping its own internal…

Cited by 50PDFScholar
2021

Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object Segmentation

AAAI 2021technical

This paper addresses the task of segmenting class-agnostic objects in semi-supervised setting. Although previous detection based methods achieve relatively good performance, these approaches extract the best proposal by a greedy strategy, which may lose the local patch details outside the chosen can…

Cited by 27SourcePDFScholar
2019

MHP-VOS: Multiple Hypotheses Propagation for Video Object Segmentation

CVPR 2019oral

We address the problem of semi-supervised video object segmentation (VOS), where the masks of objects of interests are given in the first frame of an input video. To deal with challenging cases where objects are occluded or missing, previous work relies on greedy data association strategies that mak…

Cited by 67PDFcodeScholar
2017

Jointly Attentive Spatial-Temporal Pooling Networks for Video-Based Person Re-Identification

ICCV 2017poster

Person Re-Identification (person re-id) is a crucial task as its applications in visual surveillance and human-computer interaction. In this work, we present a novel joint Spatial and Temporal Attention Pooling Network (ASTPN) for video-based person re-identification, which enables the feature extra…

Cited by 333PDFcodeScholar