← Search

Haoyu Zhang

35 accepted papers

2026

A Hybrid Magnetic Actuation System for Hybrid Microrobotic Targeted Delivery

ICRA 2026poster

Magnetic microrobots hold great promise for biomedical applications. However, achieving flexible magnetic field adjustment with a magnetic actuation system (MAS) to actuate diverse microrobots remains a significant challenge. In this work, we propose an Electromagnetic-Permanent Magnet Actuation (EP…

Cited by 0Scholar
2026

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

AAAI 2026technical

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric) vision, overlooking the unique challenges of first-person (egoce

Cited by 0SourcePDFScholar
2026

Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation

AAAI 2026technical

Long-term action anticipation from egocentric video is critical for applications such as human-computer interaction and assistive technologies, where anticipating user intent enables proactive and context-aware AI assistance. However, existing approaches suffer from three key limitations: 1) underut

Cited by 0SourcePDFScholar
2026

Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization

ICLR 2026poster

We introduce Qronos---a new post-training quantization algorithm that not only explicitly corrects errors due to both weight and activation quantization, but also corrects errors accumulated from previously quantized layers. Our iterative algorithm is based on an interpretable and disciplined optimi…

Cited by 0SourcecodeScholar
2026

TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models

CVPR 2026

Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for their safe deployment to prevent the creation of harmful content.This has fostered a dynamic interplay between the development of erasure defenses and the adversarial probes desi

Cited by 0SourcecodeScholar
2026

Task-Aware Retrieval Augmentation for Dynamic Recommendation

AAAI 2026technical

Dynamic recommendation systems aim to provide personalized suggestions by modeling temporal user-item interactions across time-series behavioral data. Recent studies have leveraged pre-trained dynamic graph neural networks (GNNs) to learn user-item representations over temporal snapshot graphs. Howe

Cited by 0SourcePDFScholar
2025

All-in-One: Transferring Vision Foundation Models into Stereo Matching

AAAI 2025technical

As a fundamental vision task, stereo matching has made remarkable progress. While recent iterative optimization-based methods have achieved promising performance, their feature extraction capabilities still have room for improvement. Inspired by the ability of vision foundation models (VFMs) to ext…

Cited by 1SourcePDFScholar
2025

Consistency-aware Self-Training for Iterative-based Stereo Matching

CVPR 2025poster

Iterative-based methods have become mainstream in stereo matching due to their high performance. However, these methods heavily rely on labeled data and face challenges with unlabeled real-world data. To this end, we propose a consistency-aware self-training framework for iterative-based stereo matc…

Cited by 0SourcePDFScholar
2025

Depth-PC: Sim-to-Real Transfer for Zero-Shot Visual Servoing via Cross-Modal Fusion

RA-L 2025

Visual servoing techniques guide robotic motion using visual information to accomplish manipulation tasks, requiring high precision and robustness against noise. Traditional methods often require prior knowledge and are susceptible to external disturbances. Learning-driven alternatives, while promis

Cited by 0SourcecodeScholar
2025

HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Assistant Scenarios

ACL 2025finding

Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In this paper, we introduce HammerBench, a novel benchmark framework for assessing LLMs’ function-calling capabilities in re…

2025

Haptic Feedback Control Strategy for Microswarm Navigation in Flowing Environments

IROS 2025

Swarming microrobots offer great promise for targeted delivery in biofluidic environments. However, current approaches insufficiently utilize the operator’s perceptual awareness and interactive decision-making capabilities. This work proposes a real-time navigation and control strategy with haptic f

Cited by 0SourceScholar
2025

Improving Task-Specific Multimodal Sentiment Analysis with General MLLMs via Prompting

NeurIPS 2025poster

Multimodal Sentiment Analysis (MSA) aims to predict sentiment from diverse data types, such as video, audio, and language. Recent progress in Multimodal Large Language Models (MLLMs) have demonstrated impressive performance across various tasks. However, in MSA, the increase in computational costs d…

Cited by 0SourceScholar
2025

Long-Distance Delivery of Collective Cell Microrobots Driven by Mobile Magnetic Actuation System

IROS 2025

Collective microrobots enable controlled batch delivery, showing promising application in the biomedical field. However, significant challenges remain in achieving long-distance delivery of collective microrobots in dynamic environments. This study proposes a magnetic actuation strategy for deliveri

Cited by 0SourceScholar
2025

Modular Gaussian Splatting: Instance Decomposable Learning and Adaptive Rendering of 3D Scenes via Mixture of Experts

ICASSP 2025accepted

This paper introduces Modular Gaussian Splatting (Modular-GS), a novel method that leverages 3D Gaussian Splatting and Mixture of Experts (MoEs) for decomposing and representing 3D scenes as a combination of Instance Gaussians. Modular-GS achieves scene decomposition by inputting multi-view data and…

Cited by 0SourceScholar
2025

Neural-Lyapunov Fusion: Stable Dynamical System Learning for Robotic Motion Generation

IROS 2025

Point-to-point and periodic motions are ubiquitous in the world of robotics. To master these motions, Autonomous Dynamic System (ADS) based algorithms are fundamental in the domain of Learning from Demonstration (LfD). However, these algorithms face the significant challenge of balancing precision i

Cited by 0SourceScholar
2025

Object-Shot Enhanced Grounding Network for Egocentric Video

CVPR 2025poster

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric videos but often neglect key characteristics of egocentric vid…

2025

Reinforcement Learning-Based Microrobotic Swarm Navigation and Obstacle Avoidance in Partially Observable Environments

IROS 2025

Microrobotic swarms have shown promising features due to their collective and flexible behaviours, while achieving precise swarm control and autonomous navigation in complex environments remains a challenge. Here, we propose a Transformer-based reinforcement learning strategy that integrates Proxima

Cited by 0SourceScholar
2025

STRAP: Spatio-Temporal Pattern Retrieval for Out-of-Distribution Generalization

NeurIPS 2025poster

Spatio-Temporal Graph Neural Networks (STGNNs) have emerged as a powerful tool for modeling dynamic graph-structured data across diverse domains. However, they often fail to generalize in Spatio-Temporal Out-of-Distribution (STOOD) scenarios, where both temporal dynamics and spatial structures evolv…

Cited by 0SourceScholar
2025

Selective Motion Control of Cell Microrobots in Three-Dimensional Space

IROS 2025

Magnetic microrobots are showing great potential in micromanipulation due to the capability of motion control under external fields. However, achieving selective control of magnetic microrobots in three-dimensional (3D) space using global magnetic fields still presents a challenge. In this work, we

Cited by 0SourceScholar
2025

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

NeurIPS 2025spotlight

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial uncertainty and data scarcity, limiting the 3D spatial reasoning capab…

Cited by 0SourceScholar
2025

TC–RAG: Turing–Complete RAG’s Case study on Medical LLM Systems

ACL 2025long

In the pursuit of enhancing domain-specific Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) emerges as a promising solution to mitigate issues such as hallucinations, outdated knowledge, and limited expertise in highly specialized queries. However, existing approaches to RAG fall…

2025

Time Series Supplier Allocation via Deep Black-Litterman Model

AAAI 2025technical

As a typical problem of Spatiotemporal Resource Management, Time Series Supplier Allocation (TSSA) poses a complex NP-hard challenge, aimed at refining future order dispatching strategies to satisfy the trade-off between demands and maximum supply. The Black-Litterman (BL) model, which comes from fi…

2024

Learning a Stable Dynamic System with a Lyapunov Energy Function for Demonstratives Using Neural Networks

ICRA 2024poster

Autonomous Dynamic System (DS)-based algorithms hold a pivotal and foundational role in the field of Learning from Demonstration (LfD). Nevertheless, they confront the formidable challenge of striking a delicate balance between achieving precision in learning and ensuring the overall stability of th…

Cited by 2SourceScholar
2024

MOSE: Monocular Semantic Reconstruction Using NeRF-Lifted Noisy Priors

RA-L 2024

Accurately reconstructing dense and semantically annotated 3D meshes from monocular images remains a challenging task due to the lack of geometry guidance and imperfect view-dependent 2D priors. Though we have witnessed recent advancements in implicit neural scene representations enabling precise 2D

Cited by 0SourceScholar
2024

Multi-Factor Adaptive Vision Selection for Egocentric Video Question Answering

ICML 2024poster

The challenge of interpreting the world from a human perspective in Artificial Intelligence (AI) is particularly evident in egocentric video question answering, which grapples with issues like small object recognition, noise suppression, and spatial-temporal reasoning. To address these challenges, w…

2024

Towards Robust Multimodal Sentiment Analysis with Incomplete Data

NeurIPS 2024poster

The field of Multimodal Sentiment Analysis (MSA) has recently witnessed an emerging direction seeking to tackle the issue of data incompleteness. Recognizing that the language modality typically contains dense sentiment information, we consider it as the dominant modality and present an innovative L…

2023

Dense Depth Completion Based on Multi-Scale Confidence and Self-Attention Mechanism for Intestinal Endoscopy

ICRA 2023poster

Doctors perform limited one-way intestine endoscopy, in which advanced surgical robots with depth sensors, such as stereo and ToF endoscopes, can only provide sparse and incomplete depth information. However, dense, accurate and instant depth estimation during endoscopy is vital for doctors to judge…

Cited by 10SourceScholar
2023

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

EMNLP 2023long main

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (*e.g.,* language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate…

Cited by 0SourcecodeScholar
2023

Learning Robust Point-to-Point Motions Adversarially: A Stochastic Differential Equation Approach

RA-L 2023

This letter proposes a robust stochastic differential equation approach for learning point-to-point motions in an adversarial way. The proposed stochastic dynamical model combines the advantages of the stochastic differential equation and the transformer-like function together to achieve both robust

Cited by 11SourceScholar
2023

LiteG2P: A Fast, Light and High Accuracy Model for Grapheme-to-Phoneme Conversion

ICASSP 2023accepted

As a key component of automated speech recognition (ASR) and the front-end in text-to-speech (TTS), grapheme-to-phoneme (G2P) plays the role of converting letters to their corresponding pronunciations. Existing methods are either slow or poor in performance, and are limited in application scenarios,…

Cited by 0SourceScholar
2022

Divide and Denoise: Learning from Noisy Labels in Fine-Grained Entity Typing with Cluster-Wise Loss Correction

ACL 2022long

Fine-grained Entity Typing (FET) has made great progress based on distant supervision but still suffers from label noise. Existing FET noise learning methods rely on prediction distributions in an instance-independent manner, which causes the problem of confirmation bias. In this work, we propose a…

2022

End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding

ACL 2022long

Natural language spatial video grounding aims to detect the relevant objects in video frames with descriptive sentences as the query. In spite of the great advances, most existing methods rely on dense video frame annotations, which require a tremendous amount of human effort. To achieve effective g…

Cited by 41SourcePDFScholar
2020

Learning with Noise: Improving Distantly-Supervised Fine-grained Entity Typing via Automatic Relabeling

IJCAI 2020poster

Fine-grained entity typing (FET) is a fundamental task for various entity-leveraging applications. Although great success has been made, existing systems still have challenges in handling noisy samples in training data introduced by distant supervision methods. To address these noise, previous studi…

Cited by 0SourcePDFScholar