← Search

Pengteng Li

13 accepted papers

2026

Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection

AAAI 2026technical

Source-Free Object Detection (SFOD) aims to adapt a source-pretrained object detector to a target domain without access to source data. However, existing SFOD methods predominantly rely on internal knowledge from the source model, which limits their capacity to generalize across domains and often re

Cited by 0SourcePDFScholar
2026

From Events to Clarity: The Event-Guided Diffusion Framework for Dehazing

CVPR 2026

Clear imaging under hazy conditions is a critical task. Prior-based and neural methods have improved results. However, they operate on RGB frames, which suffer from limited dynamic range. Therefore, dehazing remains ill- posed and can erase structure and illumination details. To address this, we use

Cited by 0SourceScholar
2026

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

ICML 2026poster

We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objects are always visible, leading to brittle and reactive behaviors when targets fall outside the camera’s field of view. S…

Cited by 0SourceScholar
2025

See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model

NeurIPS 2025poster

We introduce See&Trek, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMs) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely…

Cited by 0SourceScholar
2025

Subgraph Invariant Learning Towards Large-Scale Graph Node Classification

AAAI 2025technical

Graph Neural Networks (GNNs) have shown efficacy in graph node classification, but face computational challenges on large-scale graphs. Although existing graph reduction methods address these issues, they still require high computational resources and fail to prioritize robust performance on out-of-…

2024

A Language-Driven Navigation Strategy Integrating Semantic Maps and Large Language Models

IROS 2024poster

Accurate perception of semantic and spatial information is crucial for robots performing language-driven navigation tasks. Existing approaches utilize visual-language models to extract semantic information from the environment and construct maps. However, constrained by the generalization and accura…

Cited by 0SourceScholar
2024

Implicit Grasp Diffusion: Bridging the Gap between Dense Prediction and Sampling-based Grasping

CoRL 2024poster

There are two dominant approaches in modern robot grasp planning: dense prediction and sampling-based methods. Dense prediction calculates viable grasps across the robot’s view but is limited to predicting one grasp per voxel. Sampling-based methods, on the other hand, encode multi-modal grasp distr…

Cited by 2SourceScholar
2024

LLaKey: Follow My Basic Action Instructions to Your Next Key State

IROS 2024poster

In 3D object manipulation, collecting expert data for end-to-end imitation learning becomes a mainstream method. Though successful, previous works neglect the guiding role of language in action execution. These methods lack the understanding of action semantics, in which multiple action sequences ar…

Cited by 0SourceScholar
2024

OTOcc: Optimal Transport for Occupancy Prediction

IJCAI 2024poster

The autonomous driving community is highly interested in 3D occupancy prediction due to its outstanding geometric perception and object recognition capabilities. However, previous methods are limited to existing semantic conversion mechanisms for solving sparse ground truths problem, causing excessi…

2024

Robot Trajectron: Trajectory Prediction-based Shared Control for Robot Manipulation

ICRA 2024poster

We address the problem of (a) predicting the trajectory of an arm reaching motion, based on a few seconds of the motion’s onset, and (b) leveraging this predictor to facilitate shared-control manipulation tasks, by reducing the operator’s cognitive load through assistance in their anticipated direct…

Cited by 24SourcecodeScholar
2023

Bagging R-CNN: Ensemble for Object Detection in Complex Traffic Scenes

ICASSP 2023accepted

Generic object detection methods have achieved preferable results, but it is still challenging to detect objects from complicated traffic scenes like extreme illumination and adverse weather. The existing methods are not robust enough to be extended to new complex traffic scenes. To address this iss…

Cited by 0SourceScholar