← Search

Peng Jia

19 accepted papers

2026

A Type of Actuator with Large Deformation and Load Capacity: Design and Modeling

ICRA 2026poster

Flexible actuators have garnered extensive attention due to their flexibility and versatility. However, they still exhibit significant limitations in load capacity and structural stiffness. We have developed a multifunctional rigid-flexible coupled actuator with large deformation and high load capac…

Cited by 0Scholar
2026

From Manuals to Actions: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation

CVPR 2026

Vision-Language-Action (VLA) models have recently emerged, demonstrating strong generalization in robotic scene understanding and manipulation. However, when confronted with long-horizon tasks that require defined goal states, such as LEGO assembly or object rearrangement, existing VLA models still

Cited by 0SourceScholar
2026

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

ICRA 2026poster

Recent advancements in world models have revolutionized dynamic environment simulation, allowing systems to foresee future states and assess potential actions. In autonomous driving, these capabilities help vehicles anticipate the behavior of other road users, perform risk-aware planning, accelerate…

2026

TLDIFFGAN: A LATENT DIFFUSION-GAN FRAMEWORK WITH TEMPORAL INFORMATION FUSION FOR ANOMALOUS SOUND DETECTION

ICASSP 2026poster

Existing generative models for unsupervised anomalous sound detection are limited by their inability to fully capture the complex feature distribution of normal sounds, while the potential of powerful diffusion models in this domain remains largely unexplored. To address this challenge, we propose a…

Cited by 0SourcePDFScholar
2026

The Better You Learn, the Smarter You Prune: Towards Efficient Vision-Language-Action Models Via Differentiable Token Pruning

ICRA 2026poster

We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in executing real-world robotic tasks, their deployment on resource-constrained platforms is often bottlenecked by the heavy a…

2026

TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-Modal Representation for End-To-End Autonomous Driving

ICRA 2026poster

In recent years, diffusion models have demonstrated remarkable potential across diverse domains, from vision generation to language modeling. Transferring its generative capabilities to modern end-to-end autonomous driving systems has also emerged as a promising direction. However, existing diffusio…

2025

BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving

AAAI 2025technical

The rapid development of the autonomous driving industry has led to a significant accumulation of autonomous driving data. Consequently, there comes a growing demand for retrieving data to provide specialized optimization. However, directly applying previous image retrieval methods faces several cha…

Cited by 2SourcePDFScholar
2025

Cross-Scenario End-to-End Motion Planning in Off-Road Environment: A Lifelong Learning Perspective

RA-L 2025

Motion planning in off-road scenarios is particularly challenging due to diverse terrain features, surface characteristics, and environmental factors. Consequently, rule-based or fixed-parameter motion planning methods often fail to maintain optimal performance, especially in cross-scenario applicat

Cited by 2SourceScholar
2025

Generalizing Motion Planners with Mixture of Experts for Autonomous Driving

ICRA 2025

Large real-world driving datasets have sparked significant research into various aspects of learning-based motion planners for autonomous driving. These include data augmentation, model architecture, reward design, training strategies, and planner pipelines. In this paper, we review and benchmark pr

Cited by 23SourcecodeScholar
2025

HiNeuS: High-fidelity Neural Surface Mitigating Low-texture and Reflective Ambiguity

ICCV 2025poster

Neural surface reconstruction faces persistent challenges in reconciling geometric fidelity with photometric consistency under complex scene conditions. We present HiNeuS, a unified framework that holistically addresses three core limitations in existing approaches: multi-view radiance inconsistency…

2025

PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth

IROS 2025

Recent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and challenging driving conditions. However, a key challenge remains: precise and flexible camera pose control, which is crucia

Cited by 3SourceScholar
2025

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration

CVPR 2025poster

Closed-loop simulation is crucial for end-to-end autonomous driving. Existing sensor simulation methods (e.g., NeRF and 3DGS) reconstruct driving scenes based on conditions that closely mirror training data distributions. However, these methods struggle with rendering novel trajectories, such as lan…

Cited by 11SourcePDFScholar
2025

RoboPearls: Editable Video Simulation for Robot Manipulation

ICCV 2025poster

The development of generalist robot manipulation policies has seen significant progress, driven by large-scale demonstration data across diverse environments. However, the high cost and inefficiency of collecting real-world demonstrations hinder the scalability of data acquisition. While existing si…

Cited by 0SourcePDFScholar
2025

S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object Tracking

ICML 2025poster

3D multiple object tracking (MOT) plays a crucial role in autonomous driving perception. Recent end-to-end query-based trackers simultaneously detect and track objects, which have shown promising potential for the 3D MOT task. However, existing methods are still in the early stages of development an…

Cited by 0SourcePDFScholar
2025

World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

ICCV 2025poster

End-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information. A critical research challenge arises: constructing an informative driving world model to enable perception annotation-fr…

2024

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

CoRL 2024poster

A primary hurdle of autonomous driving in urban environments is understanding complex and long-tail scenarios, such as challenging road conditions and delicate human behaviors. We introduce DriveVLM, an autonomous driving system leveraging Vision-Language Models (VLMs) for enhanced scene understandi…

Cited by 190SourceScholar
2024

Enhancing Joint Dynamics Modeling for Underwater Robotics Through Stochastic Extension

RA-L 2024

Accurate joint dynamics models are essential for the compliance and robustness of robot control, especially for robots operating in complex underwater environments. To improve the precision of joint dynamics models, much research focuses on refining specific parameters or incorporating previously ov

Cited by 1SourceScholar
2024

TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

ECCV 2024poster

"3D dense captioning stands as a cornerstone in achieving a comprehensive understanding of 3D scenes through natural language. It has recently witnessed remarkable achievements, particularly in indoor settings. However, the exploration of 3D dense captioning in outdoor scenes is hindered by two majo…