← Search

Lingting Ge

3 accepted papers

2026

SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection

CVPR 2026

Vision Transformers (ViTs) enable strong multi-view 3D detection but are limited by high inference latency from dense token and query processing across multiple views and large 3D regions. Existing sparsity methods, designed mainly for 2D vision, prune or merge image tokens but do not extend to full

Cited by 0SourceScholar
2024

LPFormer: LiDAR Pose Estimation Transformer with Multi-Task Network

ICRA 2024poster

Due to the difficulty of acquiring large-scale 3D human keypoint annotation, previous methods for 3D human pose estimation (HPE) have often relied on 2D image features and sequential 2D annotations. Furthermore, the training of these networks typically assumes the prediction of a human bounding box…

Cited by 11SourceScholar
2024

Multi-Granular Transformer for Motion Prediction with LiDAR

ICRA 2024poster

Motion prediction has been an essential component of autonomous driving systems since it handles highly uncertain and complex scenarios involving moving agents of different types. In this paper, we propose a Multi-Granular TRansformer (MGTR) framework, an encoder-decoder network that exploits contex…

Cited by 11SourceScholar