← Search

Xiaoguang Zhao

9 accepted papers

2026

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

RSS 2026poster

Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. However, achieving such dexterity in cluttered scenes remains challenging and underexplored, as it requires selectively exploiting contact among multiple interacting objects with inherently co…

Cited by 0SourceScholar
2026

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action (VLA) models benefit from Chain-of-Thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception and control. We propose Latent Reasoning VLA (LaRA-VLA), a unified VLA framework…

Cited by 0SourceScholar
2026

MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning

ICML 2026poster

Multi-objective discrete optimization problems, such as molecular design, pose significant challenges due to their vast and unstructured combinatorial spaces. Traditional evolutionary algorithms often get trapped in local optima, while expert knowledge can provide crucial guidance for accelerating c…

Cited by 0SourceScholar
2026

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

RSS 2026poster

While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and struggle with viewpoint generalization. This limitation largely stems from the reliance on pretrained RGB encoders, which lack explicit geometric cues and prioritize semanti…

Cited by 0SourceScholar
2025

FAWL: Weakly-Supervised Video Corpus Moment Retrieval with Frame-Wise Auxiliary Alignment and Weighted Contrastive Learning

ICASSP 2025accepted

Video Corpus Moment Retrieval (VCMR) is a challenging task that aims to localize query-specified moments from a collection of untrimmed videos. The recent state-of-the-art method, JSG, tries to tackle this task using only video-level annotations in a weakly-supervised setting. However, the late fusi…

Cited by 0SourceScholar
2025

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval

EMNLP 2025

Partially Relevant Video Retrieval (PRVR) is a practical yet challenging task that involves retrieving videos based on queries relevant to only specific segments. While existing works follow the paradigm of developing models to process unimodal features, powerful pretrained vision-language models li

2025

RefCap: Zero-shot Video Corpus Moment Retrieval Based on Refined Dense Video Captioning

ICASSP 2025accepted

Video corpus moment retrieval (VCMR) is a challenging task aimed at localizing specific segments from untrimmed videos within a vast video collection. It has long been addressed using end-to-end supervised or weakly-supervised methods, which often lack explainability and rely on laborious annotation…

Cited by 0SourceScholar
2019

A Novel Development of Robots with Cooperative Strategy for Long-term and Close-proximity Autonomous Transmission-line Inspection

ICRA 2019poster

We develop two cooperative robots for power transmission lines (PTLs) inspection - a light climbing robot (CBR) which can stably move on the overhead ground wire (OGW) for sensor data collection and an unmanned aerial vehicle (UAV) with a grabbing mechanism, which can automatically put the CBR on th…

Cited by 12SourceScholar
2018

A Novel Monocular-Based Navigation Approach for UAV Autonomous Transmission-Line Inspection

IROS 2018poster

This paper proposes a unique and robust UAV autonomous navigation approach along one side of overhead transmission lines for inspection. To this end, we establish a perspective model and develop a novel Pan/Tilt monocular-based navigation scheme. Simultaneously, the following three key issues are ad…

Cited by 53SourceScholar