← Search

Jieqi Shi

18 accepted papers

2026

AGiLe: Learning Robust Long-Horizon Manipulation via Affordance-Grounded Bidirectional Latent Planning

CVPR 2026

The robust execution of long-horizon manipulation tasks remains a central challenge in embodied intelligence, necessitating both coherent high-level planning and reliable low-level control. Existing approaches often encounter two critical limitations: the accumulation of prediction errors in subgoal

Cited by 0SourcecodeScholar
2026

DeCo: Task Decomposition and Skill Composition for Zero-Shot Generalization in Long-Horizon 3D Manipulation

RA-L 2026

Generalizing language-conditioned multi-task imitation learning (IL) models to novel long-horizon 3D manipulation tasks is challenging. To address this, we propose <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DeCo</b> (<italic xmlns:mml="http://www.

Cited by 11SourcecodeScholar
2026

LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments

ICRA 2026poster

Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to navigate unseen environments based on natural language instructions without any prior training. Current methods face a critical trade-off: either rely on environment-specific waypoint predictors that li…

2026

ManiLong-Shot: Interaction-Aware One-Shot Imitation Learning for Long-Horizon Manipulation

AAAI 2026technical

One-shot imitation learning (OSIL) offers a promising way to teach robots new skills without large-scale data collection. However, current OSIL methods are primarily limited to short-horizon tasks, thus limiting their applicability to complex, long-horizon manipulations. To address this limitation,

Cited by 0SourcePDFScholar
2026

RoTri-Diff: A Spatial Robot–Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation

ICRA 2026poster

Bimanual manipulation is a fundamental robotic skill that requires continuous and precise coordination between two arms. While imitation learning (IL) is the dominant paradigm for acquiring this capability, existing approaches, whether robot-centric or object-centric, often overlook the dynamic geom…

2026

SEPT: Standard-Definition Map Enhanced Scene Perception and Topology Reasoning for Autonomous Driving

ICRA 2026poster

Online scene perception and topology reasoning are critical for autonomous vehicles to understand their driving environment, particularly for mapless driving systems that endeavor to reduce reliance on costly High-Definition (HD) maps. However, recent advances in online scene understanding still fac…

2026

SG-Reg: Generalizable and Efficient Scene Graph Registration

ICRA 2026poster

This paper addresses the challenge of registering two rigid semantic scene graphs, an essential capability for autonomous agents to align with remote agents or prior maps. Traditional methods rely on hand-crafted descriptors or ground-truth annotations, limiting their applicability in real-world sce…

2025

GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation

ICLR 2025poster

Robots' ability to follow language instructions and execute diverse 3D manipulation tasks is vital in robot learning. Traditional imitation learning-based methods perform well on seen tasks but struggle with novel, unseen ones due to variability. Recent approaches leverage large foundation models to…

2025

SEPT: Standard-Definition Map Enhanced Scene Perception and Topology Reasoning for Autonomous Driving

RA-L 2025

Online scene perception and topology reasoning are critical for autonomous vehicles to understand their driving environments, particularly for mapless driving systems that endeavor to reduce reliance on costly High-Definition (HD) maps. However, recent advances in online scene understanding still fa

Cited by 6SourceScholar
2025

VDG: Vision-Only Dynamic Gaussian for Driving Simulation

RA-L 2025

Recent advances in dynamic Gaussian splatting have significantly improved scene reconstruction and novel-view synthesis. However, existing methods often rely on pre-computed camera poses and Gaussian initialization using Structure from Motion (SfM) or other costly sensors, limiting their scalability

Cited by 23SourceScholar
2024

FM-Fusion: Instance-Aware Semantic Mapping Boosted by Vision-Language Foundation Models

RA-L 2024

Semantic mapping based on the supervised object detectors is sensitive to image distribution. In real-world environments, the object detection and segmentation performance can lead to a major drop, preventing the use of semantic mapping in a wider domain. On the other hand, the development of vision

Cited by 11SourcecodeScholar
2023

Are All Point Clouds Suitable for Completion? Weakly Supervised Quality Evaluation Network for Point Cloud Completion

ICRA 2023poster

In the practical application of point cloud completion tasks, real data quality is usually much worse than the CAD datasets used for training. A small amount of noisy data will usually significantly impact the overall system's accuracy. In this paper, we propose a quality evaluation network to score…

Cited by 2SourceScholar
2023

You Only Label Once: 3D Box Adaptation From Point Cloud to Image With Semi-Supervised Learning

RA-L 2023

The image-based 3D object detection task expects that the predicted 3D bounding box has a “tightness” projection (also referred to as cuboid) to facilitate 2D-based training, which fits the object contour well on the image while remaining reasonable on the 3D space. These requirements bring signific

Cited by 1SourceScholar
2018

An Efficient Volumetric Mesh Representation for Real-Time Scene Reconstruction Using Spatial Hashing

ICRA 2018poster

Mesh plays an indispensable role in dense realtime reconstruction essential in robotics. Efforts have been made to maintain flexible data structures for 3D data fusion, yet an efficient incremental framework specifically designed for online mesh storage and manipulation is missing. We propose a nove…

Cited by 18SourceScholar