← Search

Haogang Zhu

9 accepted papers

2026

Great Minds Think Alike: Contextual Tacit Communication for Decentralized LLM-Agent Cooperation

ICML 2026poster

Large language models (LLMs) are increasingly used as planners for cooperative embodied agents, but multi-agent settings amplify inconsistency under partial observability and make explicit communication costly or even unavailable. Many existing approaches rely on online message passing; when communi…

Cited by 0SourceScholar
2026

TANGO: Learning Distribution-wise Foundation Prior Consistency and Instance-wise Style Calibration for Medical Image Generalization

CVPR 2026

Test-time adaptation (TTA) has emerged as a promising solution to address real world domain shifts in medical image segmentation. Current approaches adapt by updating or regularizing a pre-trained source model. However, they face two major issues: (i) the source models on which they rely are prone t

Cited by 0SourceScholar
2026

When Vision-Language Models Meet Fetal Cardiac Ultrasound: Dual-Level Contrastive Learning for Out-of-Distribution Detection

IJCAI 2026

Recent advances in vision-language models (VLMs) have shown remarkable performance in medical image classification tasks. However, applying VLMs to fetal cardiac ultrasound (FCU) remains challenging due to compound distribution shifts, including covariate shifts caused by cross-center heterogeneity

Cited by 0Scholar
2025

Perturbating, Tuning, and Collaborating: Harnessing Vision Foundation Models for Single Domain Generalization on Medical Imaging

AAAI 2025technical

Single Domain Generalization (SDG) is critical in medical imaging applications. Recently, Vision Foundation Models (VFMs) have spearheaded a trend in AI development due to their robust generalizability and versatility. This work aims to fully explore the generalization capabilities of VFMs alongside…

Cited by 0SourcePDFScholar
2025

TinyMIG: Transferring Generalization from Vision Foundation Models to Single-Domain Medical Imaging

ICML 2025poster

Medical imaging faces significant challenges in single-domain generalization (SDG) due to the diversity of imaging devices and the variability among data collection centers. To address these challenges, we propose \textbf{TinyMIG}, a framework designed to transfer generalization capabilities from vi…

Cited by 0SourcePDFScholar
2023

An Open-Source Robotic Chinese Chess Player

IROS 2023poster

Consumer robots can accompany children growing up, improving their abilities while playing and entertaining. This paper presents an open-source, practical, low-cost robotic Chinese chess player. The proposed system includes an elaborate mechanical structure, a simple kinematic solution, a novel robo…

Cited by 1SourcecodeScholar
2022

Deep Tri-Training for Semi-Supervised Image Segmentation

RA-L 2022

Semantic segmentation is of great value to autonomous driving and many robotic applications, while it highly depends on costly and time-consuming pixel-level annotation. To make full use of unlabeled data, this work proposes a deep tri-training framework (dubbed DTT) to utilize labeled along with un

Cited by 11SourceScholar
2022

Dual Regression for Efficient Hand Pose Estimation

ICRA 2022poster

Hand pose estimation constitutes prime attainment for human-machine interaction-based applications. Real-time operation is vital in such tasks. Thus, a reliable estimator should exhibit low computational complexity and high precision at the same time. Previous works have explored the regression tech…

Cited by 9SourceScholar
2021

Real-Time Monocular Human Depth Estimation and Segmentation on Embedded Systems

IROS 2021poster

Estimating a scene’s depth to achieve collision avoidance against moving pedestrians is a crucial and fundamental problem in the robotic field. This paper proposes a novel, low complexity network architecture for fast and accurate human depth estimation and segmentation in indoor environments, aimin…

Cited by 27SourcecodeScholar