← Search

Huan Zhao

36 accepted papers

2026

Anatomical Prior-Driven Framework for Autonomous Robotic Cardiac Ultrasound Standard View Acquisition

ICRA 2026poster

Cardiac ultrasound diagnosis is critical for cardiovascular disease assessment, but acquiring standard views remains highly operator-dependent. Existing medical segmentation models often yield anatomically inconsistent results in images with poor textural differentiation between distinct feature cla…

2026

CoEvoer: Collaborative Evolution Transformer for Upper-Body Expressive Human Pose and Shape Estimation

AAAI 2026technical

Expressive Human Pose and Shape Estimation (EHPS) plays a crucial role in various AR/VR applications and has witnessed significant progress in recent years. However, current state-of-the-art methods still struggle with accurate parameter estimation for facial and hand regions and exhibit limited gen

Cited by 0SourcePDFScholar
2026

Human-Like Robot Action Policy through Game-Based Empathetic Inference for Human-Robot Collaboration

ICRA 2026poster

Harmonious human-robot collaboration requires the robot to behave like a human partner, which raises the critical question of what factors make the robot do so. This paper proposes a series of policies based on empathetic and non-empathetic intent inference, proactive and reactive action planning, a…

Cited by 0Scholar
2026

Learning-Based Adaptive Control for Surgical Robotic Exposure Task on Deformable Tissues

RSS 2026poster

In various surgical procedures, regions of interest (ROIs) such as organs or lesions are often occluded by overlying tissues, requiring surgeons to achieve adequate exposure for precise intervention. However, the irregular geometry, nonlinear biomechanical properties of overlying tissues, and limite…

Cited by 0SourceScholar
2026

Making Visual Dialogue More Engaging: A New Task, Method, and Metric

AAAI 2026technical

Large language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a chal

Cited by 0SourcePDFScholar
2026

MindTracker: Unveiling Implicit Emotions in Long-Horizon Dialogues

IJCAI 2026

Affective computing has achieved notable success in recognizing explicit emotions from short, isolated dialogue segments. However, human emotions are often implicitly expressed, internally regulated, and dynamically evolve over extended interactions. Existing models struggle to disentangle internal

Cited by 0Scholar
2026

PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network

AAAI 2026technical

Existing multimodal representation learning approaches often rely on simple feature concatenation or unified transformations, which fail to effectively disentangle and leverage common and private information across different modalities in a progressive manner. Moreover, they typically lack adaptive

Cited by 0SourcePDFScholar
2026

SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Models

CVPR 2026

Scientific images often require accurate numerical representations and correct object attributes. However, current faithfulness metrics are primarily tailored toward photorealistic, real-life imagery, rendering them ill-suited for scientific image evaluation. To address this gap, we introduce a nove

Cited by 0SourceScholar
2025

Autonomous Bimanual Manipulation of Deformable Objects Using Deep Reinforcement Learning Guided Adaptive Control

ICRA 2025

Deformable object manipulation (DOM) which is a common subtask in various surgical procedures represents an inevitable challenge in robot-assisted surgery (RAS) due to complex nonlinear deformation. This paper proposes a deep reinforcement learning guided adaptive control (RLAC) modelfree framework,

Cited by 1SourceScholar
2025

DSSM: Dual State Space Model For Human Motions Generation

ICASSP 2025accepted

Text-driven human motion generation has attracted considerable critical attention in recent years. The task requires generating movements that are diverse, natural, and comfortable in accordance with the text description. However, while generating the human motion, there is a significant gap in the…

Cited by 0SourceScholar
2025

Dual-View Learning for Conversational Emotion Recognition Through Context and Emotion-Shift Modeling

AAAI 2025technical

Conversational Emotion Recognition (CER) has recently been explored through conversational context modeling to learn the emotion distribution, i.e., the likelihood over emotion categories associated with each utterance. While these methods have shown promising results in emotion classification, they…

Cited by 0SourcePDFScholar
2025

Enhanced Multimodal Emotion Recognition in Conversations via Contextual Filtering and Multi-Frequency Graph Propagation

ICASSP 2025accepted

Multimodal Emotion Recognition in Conversations (ERC) plays a crucial role in understanding human language and behavior in real-world scenarios. However, existing research tends to simply concatenate multimodal representations, failing to capture the complex relationships between modalities. Recent…

Cited by 0SourceScholar
2025

Geometry and Force-Informed Robotic Assembly with Small Relative Initial Deviations for Circular Electrical Connectors

ICRA 2025

Circular electrical connectors (CECs) have a wide range of applications in scenarios that require reliable connections. However, sockets are often located in narrow scenes with random spatial orientations, complex lighting conditions, and obstructions from cables, making it difficult to accurately l

Cited by 0SourceScholar
2025

Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language Models

ACL 2025long

We revisit knowledge-based visual reasoning (KB-VR) in light of modern advances in multimodal large language models (MLLMs), and make the following contributions: (i) We propose Visual Knowledge Card (VKC) – a novel image that incorporates not only internal visual knowledge (e.g., scene-aware inform…

Cited by 0SourcePDFScholar
2025

Large Language Model Should Understand Pinyin for Chinese ASR Error Correction

ICASSP 2025accepted

Large language models (LLMs) can enhance automatic speech recognition (ASR) systems through generative error correction (GEC). In this paper, we propose Pinyin-enhanced GEC (PY-GEC), which leverages Pinyin—the phonetic representation of Mandarin Chinese—as supplementary information to improve Chines…

Cited by 0SourceScholar
2025

Leveraging Surgical Activity Grammar for Primary Intention Prediction in Laparoscopy Procedures

ICRA 2025

Surgical procedures are inherently complex and dynamic, with intricate dependencies and various execution paths. Accurate identification of the intentions behind critical actions, referred to as Primary Intentions (PIs), is crucial to understanding and planning the procedure. This paper presents a n

Cited by 2SourceScholar
2025

Parameter-Efficient Federal-Tuning Enhances Privacy Preserving for Speech Emotion Recognition

ICASSP 2025accepted

The Pre-trained Speech Models (PSMs) generate universal speech representations using self-supervised or weakly-supervised learning from large-scale datasets. It achieves promising performance when fine-tuned for specific tasks such as Speech Emotion Recognition (SER). However, fine-tuning on various…

Cited by 0SourceScholar
2025

Robust Robotic Breast Ultrasound Scanning and Real-Time Lesion Localization

ICRA 2025

The inherent flexibility and real-time deformation of breast tissue pose significant challenges for achieving full coverage and accurate lesion localization in autonomous breast ultrasound scanning. This paper introduces a robust finite state machine-based framework that mimics the decision-making p

Cited by 0SourceScholar
2025

Sensor-Free Strategy for Estimating Guidewire/Catheter Shape and Contact Force in Endovascular Interventions

RA-L 2025

Accurate assessment of guidewire shape and contact forces is critical for autonomous robotic endovascular procedures. However, existing sensor-based approaches often require modifications to standard guidewires or the use of custom-made alternatives, which can hinder integration into conventional su

Cited by 1SourceScholar
2025

XDGesture: An xLSTM-based Diffusion Model for Co-speech Gesture Generation

ICASSP 2025accepted

In multimodal human-computer interaction, generating co-speech gestures is crucial for enhancing interaction naturalness and user experience. However, achieving synchronized and natural gesture sequences remains a significant challenge due to the complexity of modeling temporal dependencies across d…

Cited by 0SourceScholar
2024

Customising General Large Language Models for Specialised Emotion Recognition Tasks

ICASSP 2024accepted

The advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tasks but also in emotion recognition in terms of accuracy, universality, explanation, robustness, few/zero-shot learning,…

Cited by 0SourceScholar
2024

EmoTransKG: An Innovative Emotion Knowledge Graph to Reveal Emotion Transformation

ACL 2024findings

This paper introduces EmoTransKG, an innovative Emotion Knowledge Graph (EKG) that establishes connections and transformations between emotions across diverse open-textual events. Compared to existing EKGs, which primarily focus on linking emotion keywords to related terms or on assigning sentiment…

2024

Esihgnn: Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion Recognition

ICASSP 2024accepted

Conversational Emotion Recognition (CER) aims to predict the emotion expressed by an utterance (referred to as an "event") during a conversation. Existing graph-based methods mainly focus on event interactions to comprehend the conversational context, while overlooking the direct influence of the sp…

Cited by 0SourceScholar
2024

Graph-Based Spatial Reasoning for Tracking Landmarks in Dynamic Laparoscopic Environments

RA-L 2024

Accurate anatomical landmark tracking is crucial yet challenging in laparoscopic surgery due to the changing appearance of landmarks during dynamic tool-anatomy interactions and visual domain shifts between cases. Unlike appearance-based detection methods, this work proposes a novel graph-based appr

Cited by 4SourceScholar
2024

LSTDial: Enhancing Dialogue Generation via Long- and Short-Term Measurement Feedback

NAACL 2024long

Generating high-quality responses is a key challenge for any open domain dialogue systems. However, even though there exist a variety of quality dimensions especially designed for dialogue evaluation (e.g., coherence and diversity scores), current dialogue systems rarely utilize them to guide the re…

2024

Vascular Centerline-Guided Autonomous Navigation Methods for Robot-Lead Endovascular Interventions

ICRA 2024poster

In minimally invasive endovascular interventional surgery, guidewire navigation is an indispensable process. However, even experienced physicians often encounter difficulties in manually manipulating the guidewire for branch selection, while also facing the risk of radiation exposure. In this study,…

Cited by 2SourceScholar
2023

Privacy-Enhanced Federated Learning Against Attribute Inference Attack for Speech Emotion Recognition

ICASSP 2023accepted

Federal learning-based (FL) Speech Emotion Recognition (SER) framework aims to protect data privacy when characterizing emotions. However, previous studies have shown that the framework is vulnerable, because curious servers can indirectly infer user private information. To address this challenge, w…

Cited by 0SourceScholar
2022

Automatic Keyframe Detection for Critical Actions from the Experience of Expert Surgeons

IROS 2022poster

Robot-Assisted Minimally Invasive Surgery (RAMIS), which introduced robot-actuated invasive tools to increase the dexterity and efficiency of traditional MIS, has become popular. Investigations on how to achieve autonomy in RAMIS have drawn vast intention recently, which urges further insights into…

Cited by 6SourceScholar
2022

Autonomous Laparoscope Control for Minimally Invasive Surgery With Intuition and RCM Constraints

RA-L 2022

The automated laparoscope control can ensure the stability of the laparoscopic image and enhance the performance of surgeons during surgical operations. The letter presented a method that autonomously adjusted the laparoscope pose to obtain the optimal field of view (FOV). An optimization problem wa

Cited by 14SourceScholar
2022

Improving Event Representation via Simultaneous Weakly Supervised Contrastive Learning and Clustering

ACL 2022long

Representations of events described in text are important for various tasks. In this work, we present SWCC: a Simultaneous Weakly supervised Contrastive learning and Clustering framework for event representation learning. SWCC learns event representations by making better use of co-occurrence inform…

2022

Mask-then-Fill: A Flexible and Effective Data Augmentation Framework for Event Extraction

EMNLP 2022finding

We present Mask-then-Fill, a flexible and effective data augmentation framework for event extraction. Our approach allows for more flexible manipulation of text and thus can generate more diverse data while keeping the original event structure unchanged as much as possible. Specifically, it first ra…

Cited by 33SourcePDFScholar
2019

Compact Convolutional Recurrent Neural Networks via Binarization for Speech Emotion Recognition

ICASSP 2019accepted

Despite the great advances, most of the recently developed automatic speech recognition systems focus on working in a server-client manner, and thus often require a high computational cost, such as the storage size and memory accesses. This, however, does not satisfy the increasing demand for a succ…

Cited by 0SourceScholar