← Search

Hua Zhang

17 accepted papers

2026

AquaSentinel: Next-Generation AI System Integrating Sensor Networks for Urban Underground Water Pipeline Anomaly Detection via Collaborative MoE-LLM Agent Architecture

AAAI 2026technical

Underground pipeline leaks and infiltrations pose significant threats to water security and environmental safety. Traditional manual inspection methods provide limited coverage and delayed response, often missing critical anomalies. This paper proposes AquaSentinel, a novel physics-informed AI syste

Cited by 0SourcePDFScholar
2026

Deconstructing Positional Information: From Attention Logits to Training Biases

ICLR 2026poster

Positional encodings, a mechanism for incorporating sequential information into the Transformer model, are central to contemporary research on neural architectures. Previous work has largely focused on understanding their function through the principle of distance attenuation, where proximity dictat…

Cited by 0SourceScholar
2026

PhaseWin Search Framework Enable Efficient Object-Level Interpretation

CVPR 2026

Attribution is essential for interpreting object-level foundation models. Recent methods based on submodular subset selection have achieved high faithfulness, but their efficiency limitations hinder practical deployment in real-world scenarios. To address this, we propose PhaseWin, a novel phase-win

Cited by 0SourcecodeScholar
2026

The Emotional Baby Is Truly Deadly: Does Your Multimodal Large Reasoning Model Have Emotional Flattery Towards Humans?

AAAI 2026technical

Multimodal large reasoning models (MLRMs) have advanced visual-textual integration, enabling sophisticated human-AI interaction. While prior work has exposed MLRMs to visual jailbreaks, it remains underexplored how their reasoning capabilities reshape the security landscape under adversarial inputs.

Cited by 0SourcePDFScholar
2026

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

CVPR 2026

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated tokens depend on visual modalities remains poorly understood, limiting interpretability and reliability. In this work, we pre

Cited by 0SourcecodeScholar
2025

Boosting Few-Shot Open-Set Object Detection via Prompt Learning and Robust Decision Boundary

IJCAI 2025

Few-shot Open-set Object Detection (FOOD) poses a challenge in many open-world scenarios. It aims to train an open-set detector to detect known objects while rejecting unknowns with scarce training samples. Existing FOOD methods are subject to limited visual information, and often exhibit an ambiguo

2025

Fair Text-to-Image Diffusion via Fair Mapping

AAAI 2025technical

In this paper, we address the limitations of existing text-to-image diffusion models in generating demographically fair results when given human-related descriptions. These models often struggle to disentangle the target language context from sociocultural biases, resulting in biased image generatio…

Cited by 14SourcePDFScholar
2025

Interpreting Object-level Foundation Models via Visual Precision Search

CVPR 2025highlight

Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. However, interpreting these models' decisions has grown increasingly challenging. Existing interpretable attribution methods…

2024

Less is More: Fewer Interpretable Region via Submodular Subset Selection

ICLR 2024oral

Image attribution algorithms aim to identify important regions that are highly relevant to model decisions. Although existing attribution solutions can effectively assign importance to target elements, they still face the following challenges: 1) existing attribution methods generate inaccurate smal…

2023

Improving Dynamic HDR Imaging with Fusion Transformer

AAAI 2023technical

Reconstructing a High Dynamic Range (HDR) image from several Low Dynamic Range (LDR) images with different exposures is a challenging task, especially in the presence of camera and object motion. Though existing models using convolutional neural networks (CNNs) have made great progress, challenges s…

2021

Progressive Contour Regression for Arbitrary-Shape Scene Text Detection

CVPR 2021poster

State-of-the-art scene text detection methods usually model the text instance with local pixels or components from the bottom-up perspective and, therefore, are sensitive to noises and dependent on the complicated heuristic post-processing especially for arbitrary-shape texts. To relieve these two i…

Cited by 141PDFcodeScholar
2019

Deep CNN for Wideband Mmwave Massive Mimo Channel Estimation Using Frequency Correlation

ICASSP 2019accepted

For millimeter wave (mmWave) systems with large-scale arrays, hybrid processing structure is usually used at both transmitters and receivers to reduce the complexity and cost, which poses a very challenging issue in channel estimation, especially at the low transmit signal-to-noise ratio regime. In…

Cited by 0SourceScholar
2018

Multi-Class Learning: From Theory to Algorithm

NeurIPS 2018poster

In this paper, we study the generalization performance of multi-class classification and obtain a shaper data-dependent generalization error bound with fast convergence rate, substantially improving the state-of-art bounds in the existing data-dependent generalization analysis. The theoretical analy…

Cited by 58SourcePDFScholar
2017

Segment-tree based cost aggregation for stereo matching with enhanced segmentation advantage

ICASSP 2017accepted

Segment-tree (ST) based cost aggregation algorithm for stereo matching successfully integrates the information of segmentation with non-local cost aggregation framework. The tree structure which is generated by the segmentation strategy directly determines the final results for this kind of algorith…

Cited by 0SourceScholar