← Search

Zhu Liu

26 accepted papers

2026

A Hybrid Space Model for Misaligned Multi-modality Image Fusion

AAAI 2026technical

Infrared and visible image fusion aims to integrate complementary information, such as thermal saliency from infrared imagery and fine-grained texture details from visible imagery. However, real-world multi-modal misalignment and geometric deformation often introduce severe artifacts. Most existing

Cited by 1SourcePDFScholar
2026

Diffuse to Detect: Bi-Level Sample Rebalancing with Pseudo-Label Diffusion for Point-Supervised Infrared Small-Target Detection

ICML 2026spotlight

Point supervision has become a scalable solution to address dense annotation for infrared small target detection, but its performance is limited by two coupled bottlenecks: unstable pseudo-label evolution in cluttered, low-contrast infrared imagery and severe sample-distribution imbalance. In this p…

Cited by 0SourceScholar
2026

EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision

AAAI 2026technical

Egocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages episodic appearance- and geometry-aware memory to achieve unified 2D-3D visual query

Cited by 0SourcePDFScholar
2026

HiDRA: Hierarchical Degradation Representation and Adaptation with Generative Priors for Enhancing Infrared Vision

CVPR 2026

Thermal infrared (TIR) imaging enables robust perception in adverse conditions. However, it often suffers from complex degradations (e.g., fixed-pattern noise and low-resolution) due to sensor limitations and environmental dynamics. Existing methods, whether traditional or learning-based, easily fai

Cited by 0SourcecodeScholar
2026

Learning with Semantic Priors: Stabilizing Point-Supervised Infrared Small Target Detection via Hierarchical Knowledge Distillation

IJCAI 2026

Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are expensive. Point supervision with online label evolution reduces annotation cost; however, lightweight CNN detectors often lack sufficient semantics,

Cited by 0Scholar
2026

RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels

AAAI 2026technical

Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their i

Cited by 0SourcePDFScholar
2025

A Top-down Graph-based Tool for Modeling Classical Semantic Maps: A Case Study of Supplementary Adverbs

NAACL 2025long

Semantic map models (SMMs) construct a network-like conceptual space from cross-linguistic instances or forms, based on the connectivity hypothesis. This approach has been widely used to represent similarity and entailment relationships in cross-linguistic concept comparisons. However, most SMMs are…

2025

Beyond Speaker Identity: Text Guided Target Speech Extraction

ICASSP 2025accepted

Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker’s identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided TSE model StyleTSE that uses natural language descriptions of speaking style in…

Cited by 0SourceScholar
2025

Bilevel Optimization for Adversarial Learning Problems: Sharpness, Generation, and Beyond

NeurIPS 2025poster

Adversarial learning is a widely used paradigm in machine learning, often formulated as a min-max optimization problem where the inner maximization imposes adversarial constraints to guide the outer learner toward more robust solutions. This framework underlies methods such as Sharpness-Aware Minimi…

Cited by 0SourceScholar
2025

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation

ACL 2025finding

Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Op…

Cited by 0SourcePDFScholar
2025

DEAL: Data-Efficient Adversarial Learning for High-Quality Infrared Imaging

CVPR 2025poster

Thermal imaging is often compromised by dynamic, complex degradations caused by hardware limitations and unpredictable environmental factors. The scarcity of high-quality infrared data, coupled with the challenges of dynamic, intricate degradations, makes it difficult to recover details using exis…

2025

Detect, Disambiguate, and Translate: On-Demand Visual Reasoning for Multimodal Machine Translation with Large Vision-Language Models

NAACL 2025long

Multimodal machine translation (MMT) aims to leverage additional modalities to assist in language translation. With limited parallel data, current MMT systems rely heavily on monolingual English captioning data. These systems face three key issues: they often overlook that visual signals are unneces…

Cited by 0SourcePDFScholar
2025

Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark

NeurIPS 2025poster

We engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one e…

Cited by 0SourceScholar
2025

GLTW: Joint Improved Graph Transformer and LLM via Three-Word Language for Knowledge Graph Completion

ACL 2025finding

Knowledge Graph Completion (KGC), which aims to infer missing or incomplete facts, is a crucial task for KGs. However, integrating the vital structural information of KGs into Large Language Models (LLMs) and outputting predictions deterministically remains challenging. To address this, we propose a…

Cited by 0SourcePDFScholar
2025

Learning Rich Speech Representations with Acoustic-Semantic Factorization

ICASSP 2025accepted

Self-supervised pretraining has transformed speech representation learning, enabling models to generalize across various downstream tasks. However, empirical studies have highlighted two notable gaps. First, different speech tasks require varying levels of acoustic and semantic information, which ar…

Cited by 0SourceScholar
2025

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing

EMNLP 2025

We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X util

Cited by 0SourcePDFScholar
2024

Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics

ACL 2024findings

Large language models have achieved remarkable success in general language understanding tasks. However, as a family of generative methods with the objective of next token prediction, the semantic evolution with the depth of these models are not fully explored, unlike their predecessors, such as BER…

2024

Local Contrast Prior-Guided Cross Aggregation Model for Effective Infrared Small Target Detection

ICASSP 2024accepted

Infrared small target detection, referring to discovering the precise shapes of dim targets from complex clutter background, has gradually become a hot spot. In recent years, learning-based methods have become the mainstream schemes with high efficiency. However, these methods seldom consider the na…

Cited by 0SourceScholar
2024

Moreau Envelope for Nonconvex Bi-Level Optimization: A Single-Loop and Hessian-Free Solution Strategy

ICML 2024poster

This work focuses on addressing two major challenges in the context of large-scale nonconvex Bi-Level Optimization (BLO) problems, which are increasingly applied in machine learning due to their ability to model nested structures. These challenges involve ensuring computational efficiency and provid…

Cited by 9SourcePDFScholar
2024

Segmentation-Driven Infrared and Visible Image Fusion Via Transformer-Enhanced Architecture Searching

ICASSP 2024accepted

A series of infrared and visible image fusion (IVIF) methods have emerged to improve the performance of segmentation task. However, existing perception-focused IVIF methods take visual effects and semantic information as a unified goal for training, ignoring the task conflicts. Moreover, these metho…

Cited by 0SourceScholar
2024

Where Elegance Meets Precision: Towards a Compact, Automatic, and Flexible Framework for Multi-modality Image Fusion and Applications

IJCAI 2024poster

Multi-modality image fusion aims to integrate images from multiple sensors, producing an image that is visually appealing and offers more comprehensive information than any single one. To ensure high visual quality and facilitate accurate subsequent perception tasks, previous methods have often casc…

2023

Ambiguity Meets Uncertainty: Investigating Uncertainty Estimation for Word Sense Disambiguation

ACL 2023findings

Word sense disambiguation (WSD), which aims to determine an appropriate sense for a target word given its context, is crucial for natural language understanding. Existing supervised methods treat WSD as a classification task and have achieved remarkable performance. However, they ignore uncertainty…

2023

Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and Beyond

IJCAI 2023poster

Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and neglecting others, seldom investigating their underlying co…

2023

Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation

ICCV 2023oral

Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach `Best of Both Worlds'. To overcome this issue, in this paper, we propos…

Cited by 172PDFcodeScholar
2022

ActiveZero: Mixed Domain Learning for Active Stereovision With Zero Annotation

CVPR 2022poster

Traditional depth sensors generate accurate real world depth estimates that surpass even the most advanced learning approaches trained only on simulation domains. Since ground truth depth is readily available in the simulation domain but quite difficult to obtain in the real domain, we propose a met…

Cited by 8PDFcodeScholar