← Search

Hao Huang

61 accepted papers

2026

A Benchmark for Joint Dialogue Satisfaction, Emotion Recognition, and Emotion State Transition Prediction

ICASSP 2026poster

User satisfaction is closely related to enterprises, as it not only directly reflects users' subjective evaluation of service quality or products, but also affects customer loyalty and long-term business revenue. Monitoring and understanding user emotions during interactions helps predict and improv…

Cited by 0SourcePDFScholar
2026

HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning

ICML 2026poster

Vision–Language–Action (VLA) models have shown strong performance in robotic manipulation, but often struggle in long-horizon or out-of-distribution scenarios due to the lack of explicit mechanisms for multimodal reasoning and anticipating how the world will evolve under action. Recent works introdu…

Cited by 0SourceScholar
2026

IACW: Intent-Aware Controllable Watermarking for Scalable Authorial Intent Attribution

ICML 2026poster

As Large Language Models (LLMs) integrate into writing workflows, precise governance requires distinguishing ''how AI participated'' rather than merely ''whether AI was used.'' Traditional binary detection often misclassifies ``AI-polished'' content as generated, creating fairness risks. We propose …

Cited by 0SourceScholar
2026

Integrating Advantage Actor-Critic in Multi-Robot Collaboration

RA-L 2026

Recent advances in large language models (LLMs) have spurred interest in using these models to coordinate multi-agent robot systems. However, existing approaches often fail to handle dynamic and complex environments effectively. We present A2C-Collab, an <underline xmlns:mml="http://www.w3.org/1998/

Cited by 0SourceScholar
2026

Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding

AAAI 2026technical

Spoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating contex

Cited by 0SourcePDFScholar
2026

LoRAGen: Structure-Aware Weight Space Learning for LoRA Generation

ICLR 2026poster

The widespread adoption of Low-Rank Adaptation (LoRA) for efficient fine-tuning of large language models has created demand for scalable parameter generation methods that can synthesize adaptation weights directly from task descriptions, avoiding costly task-specific training. We present LoRAGen, a…

Cited by 0SourcecodeScholar
2026

MoL: Adaptive Mixture-of-Length Reasoning for Efficient Question Answering with Context

ICLR 2026poster

We present Mixture-of-Length (MoL), an approach for Question Answering (QA) with context that aims to improve the balance between reasoning quality and response efficiency. Our method introduces a principled difficulty assessment based on information-theoretic principles and a dual-objective reward…

Cited by 0SourceScholar
2026

ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation

ICML 2026poster

The prefill stage of long-context Retrieval-Augmented Generation (RAG) is severely bottlenecked by computational overhead. To mitigate this, recent methods assemble pre-calculated KV caches of retrieved RAG documents (by a *user query*) and reprocess selected tokens to recover cross-attention betwee…

Cited by 0SourceScholar
2026

RESA: Bringing Back What Sparse Attention Ignores with Residual Estimation

ICLR 2026poster

Large Language Models (LLM) have gained significant attention. KV cache, stored to avoid quadratic complexity of attention, becomes a bottleneck due to the demands for long-context. Sparse attention (SA) has been proposed to address this by only selecting critical KVs for attention, which ma…

Cited by 0SourceScholar
2026

Vision-Only Gaussian Splatting for Collaborative Semantic Occupancy Prediction

AAAI 2026technical

Collaborative perception enables connected vehicles to share information, overcoming occlusions and extending the limited sensing range inherent in single-agent (non-collaborative) systems. Existing vision-only methods for 3D semantic occupancy prediction commonly rely on dense 3D voxels, which incu

Cited by 0SourcePDFScholar
2025

GADACE: Graph Anomaly Detection Combining Attribute Contrast and Structure Reconstruction

ICASSP 2025accepted

Unsupervised graph anomaly detection aims to identify nodes that deviate from typical behaviors in graphs. Existing approaches can be briefly categorized into two main groups, namely, reconstruction-based approaches that detect anomalies through reconstruction errors, and contrastive learning-based…

Cited by 0SourceScholar
2025

HANet: A Harmonic Attention-Based Network for Singing Melody Extraction from Polyphonic Music

ICASSP 2025accepted

Singing melody extraction from polyphonic music is a complex but important task in music information retrieval. Harmonic relationships have been shown to be crucial in this task, but most existing models based on Convolutional Neural Networks (CNNs) struggle to capture long-range harmonic dependenci…

Cited by 0SourceScholar
2025

INT: Establishing Information Transfer for Multilingual Intent Detection and Slot Filling

ACL 2025finding

Multilingual spoken language understanding (SLU) involves intent detection (ID) and slot filling (SF) across multiple languages. The inherent linguistic diversity presents significant challenges in achieving performance comparable to traditional SLU. Recent studies have attempted to improve multilin…

Cited by 0SourcePDFScholar
2025

Improved Cross-Lingual Speaker Verification Using Speaker Sensitive Feature Guidance and Fine-grained Phonetic Information

ICASSP 2025accepted

Speaker verification performance significantly degrades when there exists a language mismatch between training and evaluation. Domain Adversarial Training (DAT) has shown to be effective in mitigating this gap by incorporating adversarial training with domain information (language id). Inspired by r…

Cited by 0SourceScholar
2025

Multi-Segment Soft Robot Control Via Deep Koopman-Based Model Predictive Control

ICRA 2025

Soft robots, compared to regular rigid robots, as their multiple segments with soft materials bring flexibility and compliance, have the advantages of safe interaction and dexterous operation in the environment. However, due to its characteristics of high dimensional, nonlinearity, time-varying natu

Cited by 0SourcecodeScholar
2025

OpenVIS: Open-vocabulary Video Instance Segmentation

AAAI 2025technical

Open-vocabulary Video Instance Segmentation (OpenVIS) can simultaneously detect, segment, and track arbitrary object categories in a video, without being constrained to categories seen during training. In this work, we propose InstFormer, a carefully designed framework for the OpenVIS task that achi…

2025

Robust and Efficient Text-based Speech Editing using Noise Conditioning and Rectified Flow

ICASSP 2025accepted

Significant advancements have been made in text-based speech editing (TSE) for clear speech, but effectively editing the noise-contaminated speech remains a challenge. Background noise degrades the quality of generated speech, and edited speech that fails to maintain noise context consistency often…

Cited by 0SourceScholar
2025

Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language Models

IROS 2025

Robotic navigation plays a pivotal role in a wide range of real-world applications. While traditional navigation systems focus on efficiency and obstacle avoidance, their inability to model complex human behaviors in shared spaces has underscored the growing need for socially aware navigation. In th

Cited by 6SourcecodeScholar
2025

Utterance as A Bridge: Few-shot Joint Learning of Empathy Detection and Empathy Intent Classification

ICASSP 2025accepted

Empathy detection (ED) and empathy intent classification (EIC) aim to identify the empathy direction expressed in user utterances and the underlying empathy intent behind them. Previous studies show that facilitating information transfer between tasks can enhance model performance. However, the inte…

Cited by 0SourceScholar
2025

Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon Tasks

ICCV 2025poster

Policy learning focuses on devising strategies for agents in embodied artificial intelligence systems to perform optimal actions based on their perceived states. One of the key challenges in policy learning involves handling complex, long-horizon tasks that require managing extensive sequences of ac…

Cited by 0SourcePDFScholar
2024

$\texttt{dattri}$: A Library for Efficient Data Attribution

NeurIPS 2024spotlight

Data attribution methods aim to quantify the influence of individual training samples on the prediction of artificial intelligence (AI) models. As training data plays an increasingly crucial role in the modern development of large-scale AI models, data attribution has found broad applications in imp…

2024

Benchmarking Complex Instruction-Following with Multiple Constraints Composition

NeurIPS 2024poster

Instruction following is one of the fundamental capabilities of large language models (LLMs). As the ability of LLMs is constantly improving, they have been increasingly applied to deal with complex human instructions in real-world scenarios. Therefore, how to evaluate the ability of complex instruc…

2024

ChatMap: A Wearable Platform Based on the Multi-modal Foundation Model to Augment Spatial Cognition for People with Blindness and Low Vision

IROS 2024poster

Spatial cognition refers to the ability to gain knowledge about their surroundings and utilize this information to identify their location, acquire resources, and navigate their way back to familiar places. People with blindness and low vision (pBLV) face significant challenges with spatial cognitio…

Cited by 0SourceScholar
2024

Domain-Slot Aware Contrastive Learning for Improved Dialogue State Tracking

ICASSP 2024accepted

Large-scale pre-trained neural language model has facilitated to achieve the state-of-the-art performance on Dialogue State Tracking (DST) tasks. One of the existing works models the semantic correlation between the dialogue context and (domain, slot) pair encoded by BERT and make the prediction. De…

Cited by 0SourceScholar
2024

Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language Understanding

ICASSP 2024accepted

Multi-intent spoken language understanding consists of two typical subtasks: multi-intent detection and slot filling. Existing approach suffers from two limitations: (1) It fails to explicitly model the information transfer between slots associated within the same intent clause; (2) Using a co-occur…

Cited by 0SourceScholar
2024

Energy Efficient Streaming Time Series Classification with Attentive Power Iteration

AAAI 2024technical

Efficiently processing time series data streams in real-time on resource-constrained devices offers significant advantages in terms of enhanced computational energy efficiency and reduced time-related risks. We introduce an innovative streaming time series classification network that utilizes attent…

2024

Fact-Aware Summarization with Contrastive Learning for Few-Shot Dialogue State Tracking

ICASSP 2024accepted

Dialogue state tracking (DST) is a crucial component of task-oriented dialogue systems, as it aims to accurately track the user’s goals throughout the dialogue history. However, DST models struggle with new domains due to limited annotated data, leading to poor performance. To solve this key challen…

Cited by 0SourceScholar
2024

FairCLIP: Harnessing Fairness in Vision-Language Learning

CVPR 2024poster

Fairness is a critical concern in deep learning especially in healthcare where these models influence diagnoses and treatment decisions. Although fairness has been investigated in the vision-only domain the fairness of medical vision-language (VL) models remains unexplored due to the scarcity of med…

2024

FairDomain: Achieving Fairness in Cross-Domain Medical Image Segmentation and Classification

ECCV 2024poster

"Addressing fairness in artificial intelligence (AI), particularly in medical AI, is crucial for ensuring equitable healthcare outcomes. Recent efforts to enhance fairness have introduced new methodologies and datasets in medical AI. However, the fairness issue under the setting of domain transfer i…

2024

GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance

NeurIPS 2024poster

Zero-Shot Object Goal Navigation (ZS-OGN) enables robots to navigate toward objects of unseen categories without prior training. Traditional approaches often leverage categorical semantic information for navigation guidance, which struggles when only partial objects are observed or detailed and func…

Cited by 4SourcePDFScholar
2024

Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection

ECCV 2024poster

"Video Anomaly Detection (VAD) has been extensively studied under the settings of One-Class Classification (OCC) and Weakly-Supervised learning (WS), which however both require laborious human-annotated normal/abnormal labels. In this paper, we study Unsupervised VAD (UVAD) that does not depend on a…

2024

Introducing Multilingual Phonetic Information to Speaker Embedding for Speaker Verification

ICASSP 2024accepted

Incorporating frame-level phonetic information during the extraction of speaker embeddings has been shown to enhance the performance of speaker verification systems. However, previous studies have primarily relied on phonetic information obtained from pre-trained models of monolingual automatic spee…

Cited by 0SourceScholar
2024

Learning Diffusions under Uncertainty

AAAI 2024technical

To infer a diffusion network based on observations from historical diffusion processes, existing approaches assume that observation data contain exact occurrence time of each node infection, or at least the eventual infection statuses of nodes in each diffusion process. They determine potential infl…

2024

Phase Continuity-Aware Self-Attentive Recurrent Network with Adaptive Feature Selection for Robust VAD

ICASSP 2024accepted

Deep neural network (DNN) applications have significantly progressed in voice activity detection (VAD). Most current DNN-based VAD methods ignore the rich audio information in the phase domain. Therefore, applying this auxiliary information rationally and coping with low signal-to-noise ratio (SNR)…

Cited by 0SourceScholar
2024

QI-IRA: Quantum-Inspired Interactive Ranking Aggregation for Person Re-identification

AAAI 2024technical

Ranking aggregation (RA), the process of aggregating multiple rankings derived from multiple search strategies, has been proved effective in person re-identification (re-ID) because of a single re-ID method can not always achieve consistent superiority for different scenarios. Existing RA research m…

2024

SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual Attention

ICASSP 2024accepted

We propose a deep neural network with spectrogram matching and mutual attention (SMMA-Net) for audio clue-based target speaker extraction (TSE). To effectively use the auxiliary speech, we proposed spectrogram matching (SM) strategy and mutual attention (MA) block. We conducted all experiments on th…

Cited by 0SourceScholar
2023

Efficient Decision-based Black-box Patch Attacks on Video Recognition

ICCV 2023poster

Although Deep Neural Networks (DNNs) have demonstrated excellent performance, they are vulnerable to adversarial patches that introduce perceptible and localized perturbations to the input. Generating adversarial patches on images has received much attention, while adversarial patches on videos have…

Cited by 23PDFScholar
2023

Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech Recognition

ICASSP 2023accepted

Low-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on the hypothesis that similar linguistic units in neighboring languages exhibit c…

Cited by 0SourceScholar
2023

Investigation into Phone-Based Subword Units for Multilingual End-to-End Speech Recognition

ICASSP 2023accepted

Multilingual automatic speech recognition (ASR) models with phones as modeling units have have improved greatly in low-resource and similar-language scenarios, which benefits from shared representation across languages. Meanwhile, subwords have demonstrated their effectiveness for monolingual end-to…

Cited by 0SourceScholar
2023

Mitigating Domain Dependency for Improved Speech Enhancement Via SNR Loss Boosting

ICASSP 2023accepted

Current supervised speech enhancement methods based on deep learning typically utilize amplitude-based loss functions for optimization, such as Mean Absolute Error (MAE) or Mean Square Error (MSE) loss, which measures the difference between the amplitudes of the estimated and clean speech signals. H…

Cited by 0SourceScholar
2023

SRTNET: Time Domain Speech Enhancement via Stochastic Refinement

ICASSP 2023accepted

Diffusion model, as a new generative model which is very popular in image generation and audio synthesis, is rarely used in speech enhancement. In this paper, we use the diffusion model as a module for stochastic refinement. We propose SRTNet, a novel method for speech enhancement via Stochastic Ref…

Cited by 0SourceScholar
2023

Scalable-DSC: A Structural Template Prompt Approach to Scalable Dialogue State Correction

EMNLP 2023long main

Dialogue state error correction has recently been proposed to correct wrong slot values in predicted dialogue states, thereby mitigating the error propagation problem for dialogue state tracking (DST). These approaches, though effective, are heavily intertwined with specific DST models, limiting the…

Cited by 0SourceScholar
2023

Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation

ICASSP 2023accepted

Existing speech separation models based on deep learning typically generalize poorly due to domain mismatch. In this paper, we propose SpeakerAugment (SA), a data augmentation method for generalizable speech separation that aims to increase the diversity of speaker identity in training data, to miti…

Cited by 0SourceScholar
2023

Speech-Text Based Multi-Modal Training with Bidirectional Attention for Improved Speech Recognition

ICASSP 2023accepted

To let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the synchronicity of feature sampling rates between speech and language (aka text data); 2) the homogeneity of the learned re…

Cited by 0SourceScholar
2022

CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating Deepfakes

AAAI 2022technical

Malicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models,…

2022

Correctable-DST: Mitigating Historical Context Mismatch between Training and Inference for Improved Dialogue State Tracking

EMNLP 2022main

Recently proposed dialogue state tracking (DST) approaches predict the dialogue state of a target turn sequentially based on the previous dialogue state. During the training time, the ground-truth previous dialogue state is utilized as the historical context. However, only the previously predicted d…

Cited by 5SourcePDFScholar
2022

Fine-Grained Predicates Learning for Scene Graph Generation

CVPR 2022poster

The performance of current Scene Graph Generation models is severely hampered by some hard-to-distinguish predicates, e.g., "woman-on/standing on/walking on-beach" or "woman-near/looking at/in front of-child". While general SGG models are prone to predict head predicates and existing re-balancing st…

Cited by 63PDFcodeScholar
2022

Minimum Word Error Training For Non-Autoregressive Transformer-Based Code-Switching ASR

ICASSP 2022accepted

Non-autoregressive end-to-end ASR framework might be potentially appropriate for code-switching recognition task thanks to its inherent property that present output token being independent of historical ones. However, it still under-performs the state-of-the-art autoregressive ASR frameworks. In thi…

Cited by 0SourceScholar
2022

Mining Hard Samples Locally And Globally For Improved Speech Separation

ICASSP 2022accepted

Speech separation dataset typically consists of hard and non-hard samples, and the former is minority and latter majority. The data imbalance problem biases the model towards non-hard samples and weakens the generalization capability. Given that the average separation performance is sufficiently goo…

Cited by 0SourceScholar
2022

Understand before Answer: Improve Temporal Reading Comprehension via Precise Question Understanding

NAACL 2022long

This work studies temporal reading comprehension (TRC), which reads a free-text passage and answers temporal ordering questions. Precise question understanding is critical for temporal reading comprehension. For example, the question “What happened before the victory” and “What happened after the vi…

2021

Diffusion Network Inference from Partial Observations

AAAI 2021technical

To infer the structure of a diffusion network from observed diffusion results, existing approaches customarily assume that observed data are complete and contain the final infection status of each node, as well as precise timestamps of node infections. Due to high cost and uncertainties in the monit…

Cited by 6SourcePDFScholar
2021

Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone Classification

ICASSP 2021accepted

We pursue an interpretable pitch tracking model and a jointly trained tone model for Mandarin tone classification. For pitch tracking, present deep learning based pitch model structure seldom considers the Viterbi decoding commonly implemented in prevalent manually designed pitch tracking algorithms…

Cited by 0SourceScholar
2021

Reasoning over Entity-Action-Location Graph for Procedural Text Understanding

ACL 2021long

Procedural text understanding aims at tracking the states (e.g., create, move, destroy) and locations of the entities mentioned in a given paragraph. To effectively track the states and locations, it is essential to capture the rich semantic relations between entities, actions, and locations in the…

2020

RatE: Relation-Adaptive Translating Embedding for Knowledge Graph Completion

COLING 2020main

Many graph embedding approaches have been proposed for knowledge graph completion via link prediction. Among those, translating embedding approaches enjoy the advantages of light-weight structure, high efficiency and great interpretability. Especially when extended to complex vector space, they show…