← Search

Yuhao Zhao

5 accepted papers

2025

A Bio-inspired Spherical Soft Magnetic Millirobot for Gastrointestinal Applications

IROS 2025

Gastroscopy and colonoscopy have become the fundamental tools for gastrointestinal (GI) tract diagnosis and treatment. Conventional tethered devices usually lead to the use of anesthetic agents and patient discomfort. Capsule endoscopy is becoming an ideal alternative, however, the smooth capsule sh

Cited by 0SourceScholar
2025

IoU-Aware Clustering for Anchor Configuration Determination in Efficient Defect Detection

IROS 2025

Deep-learning-based object detection has gained widespread application in surface defect inspection, with anchor-based detectors achieving remarkable success by utilizing dense anchors to align with defects. Determining the optimal anchor configuration, i.e., sizes and aspect ratios of anchor boxes,

Cited by 0SourceScholar
2025

Monocular Depth Estimation and Segmentation for Transparent Object with Iterative Semantic and Geometric Fusion

ICRA 2025

Transparent object perception is indispensable for numerous robotic tasks. However, accurately segmenting and estimating the depth of transparent objects remain challenging due to complex optical properties. Existing methods primarily delve into only one task using extra inputs or specialized sensor

Cited by 8SourcecodeScholar
2024

Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization

ICASSP 2024accepted

Audio-Visual Source Localization (AVSL) is the task of identifying specific sounding objects in the scene given audio cues. In our work, we focus on semi-supervised AVSL with pseudo-labeling. To address the issues with vanilla hard pseudo-labels including bias accumulation, noise sensitivity, and in…

Cited by 0SourceScholar
2023

Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization

NeurIPS 2023poster

Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise…