← Search

Peizheng Li

7 accepted papers

2026

CareBot-H: Enhancing Patient Transfer with Biomimetic Design and Trajectory Deformation Algorithm

ICRA 2026poster

This paper introduces the CareBot-H Robot, a humanoid nursing robot designed to perform patient transfer tasks in confined environments. The robot is equipped with biomimetic arms that replicate human arm size and function, and distributed tactile sensors that enhance operational safety during physi…

Cited by 0Scholar
2026

RegionCache: Semantic-Aware Region Reuse for Efficient Multi-Turn Image Generation

IJCAI 2026

Real-world image generation generally requires multi-turn editing, where users iteratively refine a small region while the majority of the image remains stable across turns. Despite this strong region-level stability, existing diffusion transformer (DiT)–based editing pipelines recompute the entire

Cited by 0Scholar
2026

Seizure-Semiology-Suite($S^3$): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

ICML 2026spotlight

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporally evolving pathologic motor behaviors such as seizure semiology remains largely untested. To address this gap, we intro…

Cited by 0SourceScholar
2026

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

CVPR 2026

End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-gra

Cited by 0SourcecodeScholar
2025

AGO: Adaptive Grounding for Open World 3D Occupancy Prediction

ICCV 2025poster

Open-world 3D semantic occupancy prediction aims to generate a voxelized 3D representation from sensor inputs while recognizing both known and unknown objects. Transferring open-vocabulary knowledge from vision-language models (VLMs) offers a promising direction but remains challenging. However, met…

2024

SeFlow: A Self-Supervised Scene Flow Method in Autonomous Driving

ECCV 2024poster

"Scene flow estimation predicts the 3D motion at each point in successive LiDAR scans. This detailed, point-level, information can help autonomous vehicles to accurately predict and understand dynamic changes in their surroundings. Current state-of-the-art methods require annotated data to train sce…

2023

PowerBEV: A Powerful Yet Lightweight Framework for Instance Prediction in Bird’s-Eye View

IJCAI 2023poster

Accurately perceiving instances and predicting their future motion are key tasks for autonomous vehicles, enabling them to navigate safely in complex urban traffic. While bird’s-eye view (BEV) representations are commonplace in perception for autonomous driving, their potential in a motion predictio…