← Search

Yonggen Ling

19 accepted papers

2026

DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

CVPR 2026

Reliable 4D object detection, which refers to 3D object detection in streaming video, is crucial for perceiving and understanding the real world. Existing open-set 4D object detection methods typically make predictions on a frame-by-frame basis without modeling temporal consistency, or rely on compl

Cited by 0SourcecodeScholar
2026

EdgeGrasp: Enhancing Edge Perception for 7-DoF Grasping Pose Estimation in Cluttered Scenes

ICRA 2026poster

Estimating 7-DoF grasping poses (6-DoF with gripper width) in cluttered scenes is a critical challenge for robotic manipulation. In such environments, object edges often contain many promising grasp candidates, but relying solely on incomplete single-view point cloud to infer them is difficult. Whil…

Cited by 0Scholar
2026

UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo Pair

CVPR 2026

Perceiving and reconstructing objects from images are critical for real-to-sim transfer tasks, which are widely used in the robotics community.Existing methods rely on multiple submodules such as detection, segmentation, shape reconstruction, and pose estimation to complete the pipeline.However, suc

Cited by 0SourceScholar
2026

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

RSS 2026poster

Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA…

Cited by 0SourceScholar
2025

Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation

AAAI 2025technical

Aerial Vision-Dialog Navigation (AVDN) is a new task that requires drones to navigate to a target location based on human-robot dialog history. This paper focuses on the critical fine-grained cross-modal alignment problem in AVDN, requiring the drone to align language entities with visual landmarks…

2025

Leveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo Images

CVPR 2025poster

Stereo-based category-level shape and 6D pose estimation methods have the potential to generalize to a wider range of materials than RGBD methods, which often suffer from depth measurement errors. However, without explicit depth from two views, parameters to be estimated can become inherently entan…

Cited by 0SourcePDFScholar
2024

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

ICML 2024poster

This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the "Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision,…

2023

A Miniaturised Camera-based Multi-Modal Tactile Sensor

ICRA 2023poster

In conjunction with huge recent progress in cam-era and computer vision technology, camera-based sensors have increasingly shown considerable promise in relation to tactile sensing. In comparison to competing technologies (be they resistive, capacitive or magnetic based), they offer super-high-resol…

Cited by 10SourceScholar
2023

Mx2M: Masked Cross-Modality Modeling in Domain Adaptation for 3D Semantic Segmentation

AAAI 2023technical

Existing methods of cross-modal domain adaptation for 3D semantic segmentation predict results only via 2D-3D complementarity that is obtained by cross-modal feature matching. However, as lacking supervision in the target domain, the complementarity is not always reliable. The results are not ideal…

Cited by 6SourcePDFScholar
2022

"HVC-Net: Unifying Homography, Visibility, and Confidence Learning for Planar Object Tracking"

ECCV 2022poster

"Robust and accurate planar tracking over a whole video sequence is vitally important for many vision applications. The key to planar object tracking is to find object correspondences, modeled by homography, between the reference image and the tracked image. Existing methods tend to obtain wrong cor…

Cited by 9SourcePDFScholar
2022

E2EK: End-to-End Regression Network Based on Keypoint for 6D Pose Estimation

RA-L 2022

The methods based on deep learning are the mainstream of 6D object pose estimation, which mainly include direct regression and two-stage pipelines. The former are keen by many scholars at first due to their simplicity and differentiability to poses, but they usually lack in accuracy when compared wi

Cited by 41SourceScholar
2022

Multi-fingered Tactile Servoing for Grasping Adjustment under Partial Observation

IROS 2022poster

Grasping of objects using multi-fingered robotic hands often fails due to small uncertainties in the hand motion control and the object's pose estimation. To tackle this problem, we propose a grasping adjustment strategy based on tactile seroving. Our technique employs feedback from a sensorized mul…

Cited by 12SourceScholar
2019

MVF-Net: Multi-View 3D Face Morphable Model Regression

CVPR 2019poster

We address the problem of recovering the 3D geometry of a human face from a set of facial images in multiple views. While recent studies have shown impressive progress in 3D Morphable Model (3DMM) based facial reconstruction, the settings are mostly restricted to a single view. There is an inherent…

Cited by 144PDFScholar
2018

Left-Right Comparative Recurrent Model for Stereo Matching

CVPR 2018poster

Leveraging the disparity information from both left and right views is crucial for stereo disparity estimation. Left-right consistency check is an effective way to enhance the disparity estimation by referring to the information from the opposite view. However, the conventional left-right consisten…

Cited by 115SourcePDFScholar
2018

Modeling Varying Camera-IMU Time Offset in Optimization-Based Visual-Inertial Odometry

ECCV 2018poster

Combining cameras and inertial measurement units (IMUs) has been proven effective in motion tracking, as these two sensing modalities offer complementary characteristics that are suitable for fusion. While most works focus on global-shutter cameras and synchronized sensor measurements, consumer-grad…

Cited by 25SourcePDFScholar