← Search

Yunsong Zhou

11 accepted papers

2026

ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation

CVPR 2026

Embodied intelligence for contact-rich manipulation has predominantly relied on position control, while explicit awareness and regulation of interaction forces remain under-explored, limiting stability, precision, and robustness in real-world tasks. We propose ForceVLA2, an end-to-end vision-languag

Cited by 0SourceScholar
2026

SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-Body Manipulation

ICML 2026poster

Simulating deformable objects under rich interactions remains a fundamental challenge for real-to-sim robot manipulation, with dynamics jointly driven by environmental effects and robot actions. Existing simulators rely on predefined physics or data-driven dynamics without robot-conditioned control,…

Cited by 0SourceScholar
2025

Decoupled Diffusion Sparks Adaptive Scene Generation

ICCV 2025poster

Controllable scene generation could reduce the cost of diverse data collection substantially for autonomous driving. Prior works formulate the traffic layout generation as a predictive progress, either by denoising entire sequences at once or by iteratively predicting the next frame. However, full s…

Cited by 0SourcePDFScholar
2024

Extend Your Own Correspondences: Unsupervised Distant Point Cloud Registration by Progressive Distance Extension

CVPR 2024poster

Registration of point clouds collected from a pair of distant vehicles provides a comprehensive and accurate 3D view of the driving scenario which is vital for driving safety related applications yet existing literature suffers from the expensive pose label acquisition and the deficiency to generali…

2024

SimGen: Simulator-conditioned Driving Scene Generation

NeurIPS 2024poster

Controllable synthetic data generation can substantially lower the annotation cost of training data. Prior works use diffusion models to generate driving images conditioned on the 3D object layout. However, those models are trained on small-scale datasets like nuScenes, which lack appearance and lay…

Cited by 9SourcePDFScholar
2023

APR: Online Distant Point Cloud Registration through Aggregated Point Cloud Reconstruction

IJCAI 2023poster

For many driving safety applications, it is of great importance to accurately register LiDAR point clouds generated on distant moving vehicles. However, such point clouds have extremely different point density and sensor perspective on the same object, making registration on such point clouds very h…

2023

Density-invariant Features for Distant Point Cloud Registration

ICCV 2023poster

Registration of distant outdoor LiDAR point clouds is crucial to extending the 3D vision of collaborative autonomous vehicles, and yet is challenging due to small overlapping area and a huge disparity between observed point densities. In this paper, we propose Group-wise Contrastive Learning (GCL) s…

Cited by 22PDFcodeScholar
2023

MonoATT: Online Monocular 3D Object Detection With Adaptive Token Transformer

CVPR 2023poster

Mobile monocular 3D object detection (Mono3D) (e.g., on a vehicle, a drone, or a robot) is an important yet challenging task. Existing transformer-based offline Mono3D models adopt grid-based vision tokens, which is suboptimal when using coarse tokens due to the limited available computational power…

Cited by 29SourcePDFScholar
2022

MoGDE: Boosting Mobile Monocular 3D Object Detection with Ground Depth Estimation

NeurIPS 2022accept

Monocular 3D object detection (Mono3D) in mobile settings (e.g., on a vehicle, a drone, or a robot) is an important yet challenging task. Due to the near-far disparity phenomenon of monocular vision and the ever-changing camera pose, it is hard to acquire high detection accuracy, especially for far…

Cited by 18SourcePDFScholar
2021

Monocular 3D Object Detection: An Extrinsic Parameter Free Approach

CVPR 2021poster

Monocular 3D object detection is an important task in autonomous driving. It can be easily intractable where there exists ego-car pose change w.r.t. ground plane. This is common due to the slight fluctuation of road smoothness and slope. Due to the lack of insight in industrial application, existing…

Cited by 113PDFScholar
2021

TempNet: Online Semantic Segmentation on Large-Scale Point Cloud Series

ICCV 2021poster

Online semantic segmentation on a time series of point cloud frames is an essential task in autonomous driving. Existing models focus on single-frame segmentation, which cannot achieve satisfactory segmentation accuracy and offer unstably flicker among frames. In this paper, we propose a light-weigh…

Cited by 6PDFScholar