← Search

Long ZENG

30 accepted papers

2026

AssemMate: Graph-Based LLM for Robotic Assembly Assistance

ICRA 2026poster

Large Language Model (LLM)-based robotic assembly assistance has gained significant research attention. It requires the injection of domain-specific knowledge to guide the assembly process through natural language interaction with humans. Despite some progress, existing methods represent knowledge i…

2026

GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models

ICLR 2026poster

Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for Autonomous Driving, Embodied AI and General Artificial Intelligence. Existing spatial-temporal benchmarks mainly focus on egocentric perspective reasoning with images/video…

Cited by 0SourcecodeScholar
2026

PureCC: Pure Learning for Text-to-Image Concept Customization

CVPR 2026

Existing concept customization methods have achieved remarkable outcomes in high-fidelity and multi-concept customization. However, they often neglect the influence on the original model's behavior and capabilities when learning new personalized concepts. To address this issue, we propose PureCC. Pu

Cited by 0SourcecodeScholar
2026

S3LAM: Surfel Splatting SLAM for Geometrically Accurate Tracking and Mapping

ICRA 2026poster

We propose S3LAM, a novel RGB-D SLAM system that leverages 2D surfel splatting to achieve geometrically accurate scene representations for simultaneous tracking and mapping. Unlike existing 3DGS-based SLAM approaches that rely on 3D Gaussian ellipsoids, we utilize 2D Gaussian surfels as primitives f…

Cited by 0codeScholar
2026

SculptDrug: A Spatial Condition-Aware Bayesian Flow Model for Structure-based Drug Design

AAAI 2026technical

Structure-Based Drug Design (SBDD) has emerged as a popular approach in drug discovery, leveraging three-dimensional protein structures to generate drug ligands. However, existing generative models encounter several key challenges: (1) Incorporating boundary condition constraints, (2) Integrating hi

Cited by 0SourcePDFScholar
2026

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

ICLR 2026poster

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data, which lack accurate spatial awareness and hinder their abili…

Cited by 0SourcecodeScholar
2025

Constraint-Aware Feature Learning for Parametric Point Cloud

ICCV 2025poster

Parametric point clouds are sampled from CAD shapes and are becoming increasingly common in industrial manufacturing. Most CAD-specific deep learning methods focus on geometric features, while overlooking constraints inherent in CAD shapes. This limits their ability to discern CAD shapes with simila…

Cited by 0SourcePDFScholar
2025

DVS: Dynamic Virtual-Real Simulation Platform for Mobile Robotic Tasks

RSS 2025poster

With the development of Embodied AI, robotic research has increasingly focused on complex tasks. Existing simulation platforms, however, are often limited to idealized environments, single-task scenarios and lack data interoperability. This restricts task decomposition and multi-task learning. Addit…

Cited by 0PDFScholar
2025

Diffusion Suction Grasping with Large-Scale Parcel Dataset

IROS 2025

While recent advances in suction grasping have shown remarkable progress, significant challenges persist particularly in cluttered and complex parcel handling scenarios. Current approaches are limited by (1) the lack of comprehensive parcel-specific suction grasp datasets and (2) poor adaptability t

Cited by 1SourcecodeScholar
2025

DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder

AAAI 2025technical

Diffusion models for garment-centric human generation from text or image prompts have garnered emerging attention for their great application potential. However, existing methods often face a dilemma: lightweight approaches, such as adapters, are prone to generate inconsistent textures; while finetu…

Cited by 2SourcePDFScholar
2025

GaussianRoom: Improving 3D Gaussian Splatting with SDF Guidance and Monocular Cues for Indoor Scene Reconstruction

ICRA 2025

Embodied intelligence requires precise reconstruction and rendering to simulate large-scale real-world data. Although 3D Gaussian Splatting (3DGS) has recently demonstrated high-quality results with real-time performance, it still faces challenges in indoor scenes with large, textureless regions, re

Cited by 38SourcecodeScholar
2025

MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models

EMNLP 2025

As large language models (LLMs) become widely adopted, ensuring their alignment with human values is crucial to prevent jailbreaks where adversaries manipulate models to produce harmful content. While most defenses target single-turn attacks, real-world usage often involves multi-turn dialogues, exp

2025

Training-Free Point Cloud Recognition Based on Geometric and Semantic Information Fusion

ICASSP 2025accepted

The trend of employing training-free methods for point cloud recognition is becoming increasingly popular due to its significant reduction in computational resources and time costs. However, existing approaches are limited as they typically extract either geometric or semantic features. To address t…

Cited by 0SourceScholar
2024

A Two-Stage Reinforcement Learning Approach for Robot Navigation in Long-range Indoor Dense Crowd Environments

IROS 2024poster

Safe and efficient mobility is vital for mobile robots navigating long-range indoor crowd environments, such as supermarkets, restaurants, and railway stations. Traditional path planning methods are challenged because of the high dynamics of pedestrians and constrained feasible regions. Existing lon…

Cited by 4SourceScholar
2024

Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis

EMNLP 2024finding

In recent years, the rapid increase in scientific papers has overwhelmed traditional review mechanisms, resulting in varying quality of publications. Although existing methods have explored the capabilities of Large Language Models (LLMs) for automated scientific reviewing, their generated contents…

2024

FusedNet: End-to-End Mobile Robot Relocalization in Dynamic Large-Scale Scene

RA-L 2024

To improve robot relocalization accuracy in both static and dynamic environments, we introduce a novel network, FusedNet, which incorporates a cross-attention to fuse global and local image features for end-to-end relocalization. This approach relies solely on a monocular camera sensor that is fixed

Cited by 5SourceScholar
2024

GRID: Scene-Graph-based Instruction-driven Robotic Task Planning

IROS 2024

Recent works have shown that Large Language Models (LLMs) can facilitate the grounding of instructions for robotic task planning. Despite this progress, most existing works have primarily focused on utilizing raw images to aid LLMs in understanding environmental information. However, this approach n

Cited by 45SourcecodeScholar
2024

Mobile Robot Oriented Large-Scale Indoor Dataset for Dynamic Scene Understanding

ICRA 2024poster

Most existing robotic datasets capture static scene data and thus are limited in evaluating robots’ dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University Dynamic) robotic dataset, for training and evaluating their dy…

Cited by 9SourceScholar
2024

ParametricNet++: A 6DoF Pose Estimation Network with Sparse Keypoint Recovery for Parametric Shapes in Stacked Scenarios

IROS 2024poster

Most industrial parts are designed from parametric shapes with the properties of diversity and uncertainty. We propose a 6DoF pose estimation network, ParametricNet++, based on pointwise regression and sparse keypoint recovery, which is extended from ParametricNet to include optimizations of keypoin…

Cited by 1SourceScholar
2024

SD-Net: Symmetric-Aware Keypoint Prediction and Domain Adaptation for 6D Pose Estimation In Bin-picking Scenarios

IROS 2024poster

Despite the success of 6D pose estimation in bin-picking scenarios, existing methods still struggle to produce accurate prediction results for symmetry objects in real-world scenarios. The primary bottlenecks include 1) the ambiguity in keypoints caused by object symmetries; and 2) the domain gap be…

Cited by 6SourcecodeScholar
2023

Domain Adaptation on Point Clouds for 6D Pose Estimation in Bin-Picking Scenarios

IROS 2023poster

Training with simulated data is a common approach in pose estimation research. However, a sim-to-real gap between clean simulated data and noisy real data will seriously weaken the generalization ability of the algorithm, especially for point clouds. To address this problem, this paper proposes a do…

Cited by 4SourceScholar
2023

Reinforcement Learning Based Pushing and Grasping Objects from Ungraspable Poses

ICRA 2023poster

Grasping an object when it is in an ungraspable pose is a challenging task, such as books or other large flat objects placed horizontally on a table. Inspired by human manipulation, we address this problem by pushing the object to the edge of the table and then grasping it from the hanging part. In…

Cited by 18SourceScholar
2022

Class-Aware Contrastive Semi-Supervised Learning

CVPR 2022poster

Pseudo-label-based semi-supervised learning (SSL) has achieved great success on raw data utilization. However, its training procedure suffers from confirmation bias due to the noise contained in self-generated artificial labels. Moreover, the model's judgment becomes noisier in real-world applicatio…

Cited by 139PDFcodeScholar
2021

Efficient SE(3) Reachability Map Generation via Interplanar Integration of Intra-planar Convolutions

ICRA 2021poster

Convolution has been used for fast computation of reachability maps, but it has high computational costs when performing SE(3) convolution operations for general joint arrangements in industrial robots and 3D workspace. Its application is also limited to planar robots, 2D workspace, or robots with s…

Cited by 7SourceScholar
2021

ParametricNet: 6DoF Pose Estimation Network for Parametric Shapes in Stacked Scenarios

ICRA 2021poster

Most industrial parts are parametric and their special properties are not fully explored yet. This paper proposes a new 6DoF pose estimation network for parametric shapes in stacked scenarios (ParametricNet). It treats a parametric shape, instead of a part object, as a category. The keypoints of ind…

Cited by 12SourceScholar
2019

PPR-Net:Point-wise Pose Regression Network for Instance Segmentation and 6D Pose Estimation in Bin-picking Scenarios

IROS 2019poster

Accurate object 6D pose estimation is a core task for robot bin-picking applications, especially when objects are randomly stacked with heavy occlusion. To address this problem, this paper proposes a simple but novel Point-wise Pose Regression Network (PPR-Net). For each point in the point cloud, th…

Cited by 87SourceScholar
2018

SpiderCNN: Deep Learning on Point Sets with Parameterized Convolutional Filters

ECCV 2018poster

Deep neural networks have enjoyed remarkable success for various vision tasks, however it remains challenging to apply CNNs to domains lacking a regular underlying structures such as 3D point clouds. Towards this we propose a novel convolutional architecture, termed SpiderCNN, to efficiently extract…

Cited by 1022SourcePDFScholar