← Search

Yiming Zeng

22 accepted papers

2026

DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation

CVPR 2026

Synthesis of diverse driving scenes serves as a crucial data augmentation technique for validating the robustness and generalizability of autonomous driving systems. Current methods aggregate high-definition (HD) maps and 3D bounding boxes as geometric conditions in diffusion models for conditional

Cited by 0SourceScholar
2026

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

ICML 2026poster

Diffusion models are effective for waypoint prediction in visual navigation, but standard sampling and test time guidance can produce unsafe or inefficient trajectories when updates drift off the training manifold. We propose Fisher Preserving Guidance with Outer Product Span Projection, a training-…

Cited by 0SourceScholar
2026

Gaussian Splatting-based Low-Rank Tensor Representation for Multi-Dimensional Image Recovery

CVPR 2026

Tensor singular value decomposition (t-SVD) is a promising tool for multi-dimensional image representation, which decomposes a multi-dimensional image into a latent tensor and an accompanying transform matrix. However, two critical limitations of t-SVD methods persist: (1) the approximation of the l

Cited by 1SourceScholar
2026

OpenPyRo-A1: An Open Python-Based Low-Cost Bimanual Robot for Embodied AI

RA-L 2026

Many real-world tasks, such as assembly, cooking, and object handovers, require bi-manual coordination. Learning such skills via imitation remains challenging due to dataset scarcity, mainly caused by the high cost of bi-manual robotic platforms and barriers to entry in robotics software. To address

Cited by 1SourceScholar
2026

OpenPyRo-A1: An Open Python-Based Low-Cost Bimanual Robot for Embodied AI

ICRA 2026poster

Many real-world tasks, such as assembly, cooking, and object handovers, require bi-manual coordination. Learning such skills via imitation remains challenging due to dataset scarcity, mainly caused by the high cost of bi-manual robotic platforms and barriers to entry in robotics software. To address…

Cited by 0SourceScholar
2026

Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry

ICLR 2026poster

Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this “LLM-as-a-Judge” paradigm is costly, opaque, and sensitive to prompt design. In this work, we investigate whether smaller models can serve as efficient evaluators by leveraging internal representations…

Cited by 0SourcecodeScholar
2026

SR-Planner: Sampling-Based Path Planning With Feasibility-Aware Focus Regions and Robust Trajectory Optimization for Mobile Manipulator

RA-L 2026

High-quality motion planning for mobile manipulators remains a challenging task due to the high dimensionality and complex constraints involved. While existing methods perform well in specific scenarios, their efficiency and the quality of the resulting paths and trajectories often degrade in dense

Cited by 0SourceScholar
2026

STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation

CVPR 2026

Visual navigation requires the robot to reach a specified goal such as an image, based on a sequence of first-person visual observations. While recent learning-based approaches have made significant progress, they often focus on improving policy heads or decision strategies while relying on simplist

Cited by 0SourcecodeScholar
2026

VLION: Vision-Language Guided Interactive Object Navigation with Mobile Manipulation

ICRA 2026poster

Object navigation for mobile robots typically assumes that targets are visible and paths are unobstructed. However, real-world scenarios often involve occluded targets like objects hidden behind doors or inside containers. Such scenarios require interactive navigation and manipulation by mobile mani…

Cited by 0Scholar
2025

Bridging the Editing Gap in LLMs: FineEdit for Precise and Targeted Text Modifications

EMNLP 2025

Large Language Models (LLMs) have significantly advanced natural language processing, demonstrating strong capabilities in tasks such as text generation, summarization, and reasoning. Recently, their potential for automating precise text editing tasks across specialized domains, such as programming

2025

NaviDiffusor: Cost-Guided Diffusion Model for Visual Navigation

ICRA 2025

Visual navigation, a fundamental challenge in mobile robotics, demands versatile policies to handle diverse environments. Classical methods leverage geometric solutions to minimize specific costs, offering adaptability to new scenarios but are prone to system errors due to their multi-modular design

Cited by 18SourcecodeScholar
2025

Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge Models

CVPR 2025poster

Recent advancements in diffusion-based imitation learning, which shows impressive performance in modeling multimodal distributions and training stability, have led to substantial progress in various robot learning tasks. In visual navigation, previous diffusion-based policies typically generate acti…

2025

RAPID Hand: Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platfrom for Embodied Intelligence

NeurIPS 2025poster

This paper addresses the scarcity of low-cost but high-dexterity platforms for collecting real-world multi-fingered robot manipulation data towards generalist robot autonomy. To achieve it, we propose the RAPID Hand, a co-optimized hardware and software platform where the compact 20-DoF hand, robus…

Cited by 0SourceScholar
2024

LVDiffusor: Distilling Functional Rearrangement Priors From Large Models Into Diffusor

RA-L 2024

Object rearrangement, a fundamental challenge in robotics, demands versatile strategies to handle diverse objects, configurations, and functional needs. To achieve this, the AI robot needs to learn functional rearrangement priors to specify precise goals that meet the functional requirements. Previo

Cited by 12SourceScholar
2024

OPG-Policy: Occluded Push-Grasp Policy Learning with Amodal Segmentation

IROS 2024poster

Goal-oriented grasping in dense clutter, a fundamental challenge in robotics, demands an adaptive policy to handle occluded target objects and diverse configurations. Previous methods typically learn policies based on partially observable segments of the occluded target to generate motions. However,…

Cited by 1SourceScholar
2022

IDEA-Net: Dynamic 3D Point Cloud Interpolation via Deep Embedding Alignment

CVPR 2022poster

This paper investigates the problem of temporally interpolating dynamic 3D point clouds with large non-rigid deformation. We formulate the problem as estimation of point-wise trajectories (i.e., smooth curves) and further reason that temporal irregularity and under-sampling are two major challenges.…

Cited by 23PDFcodeScholar
2022

WarpingGAN: Warping Multiple Uniform Priors for Adversarial 3D Point Cloud Generation

CVPR 2022poster

We propose WarpingGAN, an effective and efficient 3D point cloud generation network. Unlike existing methods that generate point clouds by directly learning the mapping functions between latent codes and 3D shapes, WarpingGAN learns a unified local-warping function to warp multiple identical pre-def…

Cited by 26PDFcodeScholar
2021

CorrNet3D: Unsupervised End-to-End Learning of Dense Correspondence for 3D Point Clouds

CVPR 2021poster

Motivated by the intuition that one can transform two aligned point clouds to each other more easily and meaningfully than a misaligned pair, we propose CorrNet3D -the first unsupervised and end-to-end deep learning-based framework - to drive the learning of dense correspondence between 3D shapes by…

Cited by 97PDFcodeScholar
2018

RT3D: Real-Time 3-D Vehicle Detection in LiDAR Point Cloud for Autonomous Driving

RA-L 2018

For autonomous driving, vehicle detection is the prerequisite for many tasks like collision avoidance and path planning. In this letter, we present a real-time three-dimensional (RT3D) vehicle detection method that utilizes pure LiDAR point cloud to predict the location, orientation, and size of veh

Cited by 175SourceScholar
2018

See and Think: Disentangling Semantic Scene Completion

NeurIPS 2018poster

Semantic scene completion predicts volumetric occupancy and object category of a 3D scene, which helps intelligent agents to understand and interact with the surroundings. In this work, we propose a disentangled framework, sequentially carrying out 2D semantic segmentation, 2D-3D reprojection and 3D…

2018

VarNet: Exploring Variations for Unsupervised Video Prediction

IROS 2018poster

Unsupervised video prediction is a very challenging task due to the complexity and diversity in natural scenes. Prior works directly predicting pixels or optical flows either have the blurring problem or require additional assumptions. We highlight that the crux for video frame prediction lies in pr…

Cited by 39SourcecodeScholar
2017

GeoCueDepth: Exploiting geometric structure cues to estimate depth from a single image

IROS 2017poster

Depth estimation from a single image is very challenging due to the inherent ambiguity of mapping a color image to a depth map. Previous work tackles this problem by exploiting various levels of features with multi-scale deep convolutional neural networks. However, most of the local geometric struct…

Cited by 7SourceScholar