← Search

Zhengcai Cao

9 accepted papers

2026

Long-Horizon Planning with Large Language Models for Indoor Assistive Navigation of the Visually Impaired

ICRA 2026poster

For visually impaired individuals, assistive navigation systems play a crucial role in enabling independent mobility. However, long-horizon planning based on natural language (NL) instructions in complex indoor environments remains a significant challenge. Recent studies show the strong potential of…

Cited by 0codeScholar
2025

EDSOD: An Encoder-Decoder, Diffusion-model, and Swin-Transformer-based Small Object Detector

IROS 2025

Small object detection (SOD) given aerial images suffers from an information imbalance across different feature scales. This makes it extremely challenging to perform accurate SOD. Existing methods, e.g., Feature Pyramid Network (FPN)-based algorithms, focus on extracting high-resolution and low-res

Cited by 1SourcecodeScholar
2025

Illumination Adaptation for SAM to Achieve Accurate Segmentation of Images Taken in Low-Light Scenes

ICRA 2025

Achieving accurate segmentation in low-light scenes is challenging due to 1) severe domain shift encountered when models trained on daylight data are applied to such scenes and 2) lack of large-scale fine-grained labels in low-light conditions. A good idea is to use the generalization capabilities o

Cited by 1SourcecodeScholar
2025

Robust and Real-Time Perception and Planning for UGVs in Complex Outdoor Environments

IROS 2025

Large-scale outdoor navigation is essential for unmanned ground vehicles (UGVs), but despite significant advancements, they still face two key challenges in practical applications. The first one is how to ensure safe navigation in environments with dynamic and low-lying obstacles that LiDAR cannot d

Cited by 1SourceScholar
2024

A Context-Enhanced Full-Resolution Floor Plan Segmentation Network for Topological Semantic Mapping

IROS 2024poster

Topological semantic maps provide a practical solution to enhance indoor navigation for the Partially Sighted or Visually Impaired (PSVI). Segmenting indoor floor plans and extracting boundaries are key to constructing these maps. The existing methods exhibit low accuracy in segmentation. To achieve…

Cited by 1SourceScholar
2024

A Lightweight De-confounding Transformer for Image Captioning in Wearable Assistive Navigation Device

IROS 2024poster

Image captioning is a multi-modal task that enables the transformation from scene images to natural language, providing valuable insights for visually impaired individuals to understand their environment. Therefore, its application to wearable navigation devices for visually impaired individuals hol…

Cited by 1SourceScholar
2024

QuerySOD: A Small Object Detection Algorithm Based on Sparse Convolutional Network and Query Mechanism

IROS 2024poster

Although remarkable advances have been achieved in generic object detection, small object detection (SOD) remains challenging owing to small objects’ information loss and noisy representation caused by their non-uniform distribution. Their limited width and height, scale variations, and redundant co…

Cited by 1SourceScholar