← Search

Maosheng Ye

12 accepted papers

2025

End-to-End HOI Reconstruction Transformer with Graph-based Encoding

CVPR 2025highlight

Human-object interaction (HOI) reconstruction has garnered significant attention due to its diverse applications and the success of capturing human meshes. Existing HOI reconstruction methods often rely on explicitly modeling interactions between humans and objects. However, such a way leads to a na…

Cited by 0SourcePDFScholar
2025

Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving

ICCV 2025poster

In light of the dynamic nature of autonomous driving environments and stringent safety requirements, general MLLMs combined with CLIP alone often struggle to accurately represent driving-specific scenarios, particularly in complex interactions and long-tail cases. To address this, we propose the Hin…

Cited by 0SourcePDFScholar
2024

Cross-Cluster Shifting for Efficient and Effective 3D Object Detection in Autonomous Driving

ICRA 2024poster

We present a new 3D point-based detector model, named Shift-SSD, for precise 3D object detection in autonomous driving. Traditional point-based 3D object detectors often employ architectures that rely on a progressive downsampling of points. While this method effectively reduces computational demand…

Cited by 1SourceScholar
2024

Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection

ECCV 2024poster

"The Bird’s-Eye-View (BEV) representation is a critical factor that directly impacts the 3D object detection performance, but the traditional BEV grid representation induces quadratic computational cost as the spatial resolution grows. To address this limitation, we present a new camera-based 3D obj…

2024

PPAD: Iterative Interactions of Prediction and Planning for End-to-end Autonomous Driving

ECCV 2024poster

"We present a new interaction mechanism of prediction and planning for end-to-end autonomous driving, called PPAD (Iterative Interaction of Prediction and Planning Autonomous Driving), which considers the timestep-wise interaction to better integrate prediction and planning. An ego vehicle performs…

2023

Bootstrap Motion Forecasting With Self-Consistent Constraints

ICCV 2023poster

We present a novel framework to bootstrap Motion forecasting with Self-consistent Constraints (MISC). The motion forecasting task aims at predicting future trajectories of vehicles by incorporating spatial and temporal information from the past. A key design of MISC is the proposed Dual Consistency…

Cited by 19PDFScholar
2022

Efficient Point Cloud Segmentation with Geometry-Aware Sparse Networks

ECCV 2022poster

"In point cloud learning, sparsity and geometry are two core properties. Recently, many approaches have been proposed through single or multiple representations to improve the performance of point cloud semantic segmentation. However, these works fail to maintain the balance among performance, effic…

Cited by 26SourcePDFScholar
2022

Sparse Cross-Scale Attention Network for Efficient LiDAR Panoptic Segmentation

AAAI 2022technical

Two major challenges of 3D LiDAR Panoptic Segmentation (PS) are that point clouds of an object are surface-aggregated and thus hard to model the long-range dependency especially for large instances, and that objects are too close to separate each other. Recent literature addresses these problems by…

Cited by 46SourcePDFScholar
2021

DRINet: A Dual-Representation Iterative Learning Network for Point Cloud Segmentation

ICCV 2021poster

We present a novel and flexible architecture for point cloud segmentation with dual-representation iterative learning. In point cloud processing, different representations have their own pros and cons. Thus, finding suitable ways to represent point cloud data structure while keeping its own internal…

Cited by 50PDFScholar
2020

GOSMatch: Graph-of-Semantics Matching for Detecting Loop Closures in 3D LiDAR data

IROS 2020poster

Detecting loop closures in 3D Light Detection and Ranging (LiDAR) data is a challenging task since point-level methods always suffer from instability. This paper presents a semantic-level approach named GOSMatch to perform reliable place recognition. Our method leverages novel descriptors, which are…

Cited by 74SourceScholar