← Search

Jianzong Wang

55 accepted papers

2026

DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement

ICML 2026poster

Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge lies in inductive biases induced by distinct supervision signals: generation branch prefers high-fidelity, fine-grained represen…

Cited by 0SourceScholar
2026

FROM KNOWING TO DOING PRECISELY: A GENERAL SELF-CORRECTION AND TERMINATION FRAMEWORK FOR VLA MODELS

ICASSP 2026poster

While vision-language-action (VLA) models for embodied agents integrate perception, reasoning, and control, they remain constrained by two critical weaknesses: first, during grasping tasks, the action tokens generated by the language model often exhibit subtle spatial deviations from the target obje…

Cited by 0SourcePDFScholar
2026

TRIAGE: HIERARCHICAL VISUAL BUDGETING FOR EFFICIENT VIDEO REASONING IN VISION-LANGUAGE MODELS

ICASSP 2026oral

Vision-Language Models (VLMs) face significant computational challenges in video processing due to massive data redundancy, which creates prohibitively long token sequences. To address this, we introduce Triage, a training-free, plug-and-play framework that reframes video reasoning as a resource all…

Cited by 0SourcePDFScholar
2026

Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc Queries

AAAI 2026technical

Streaming video question answering (Streaming Video QA) poses distinct challenges for multimodal large language models (MLLMs), as video frames arrive sequentially and user queries can be issued at arbitrary timepoints. Existing solutions relying on fixed-size memory or naive compression often suffe

Cited by 0SourcePDFScholar
2025

ACCon: Angle-Compensated Contrastive Regularizer for Deep Regression

AAAI 2025technical

In deep regression, capturing the relationship among continuous labels in feature space is a fundamental challenge that has attracted increasing interest. Addressing this issue can prevent models from converging to suboptimal solutions across various regression tasks, leading to improved performance…

Cited by 0SourcePDFScholar
2025

CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation

ICASSP 2025accepted

Voice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a non-parallel VC scenario, which induces the train-inference…

Cited by 0SourceScholar
2025

EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition

EMNLP 2025

Although large audio-language models (LALMs) have demonstrated remarkable capabilities in audio perception, their performance in affective computing scenarios, particularly in emotion recognition, reasoning, and subtle sentiment differentiation, remains suboptimal. Recent advances in reinforcement l

Cited by 0SourcePDFScholar
2025

Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-Based Planner and Graph-Based Policy

ICRA 2025

Multi-agent systems (MAS) have shown great potential in executing complex tasks, but coordination and safety remain significant challenges. Multi-Agent Reinforcement Learning (MARL) offers a promising framework for agent collaboration, but it faces difficulties in handling complex tasks and designin

Cited by 12SourceScholar
2025

Federated Domain Generalization with Domain-specific Soft Prompts Generation

ICCV 2025poster

Prompt learning has become an efficient paradigm for adapting CLIP to downstream tasks. Compared with traditional fine-tuning, prompt learning optimizes a few parameters yet yields highly competitive results, especially appealing in federated learning for computational efficiency. engendering domain…

Cited by 0SourcePDFScholar
2025

Graph Contrastive Learning with Decoupled Augmentation

ICASSP 2025accepted

Graph contrastive learning based on augmentation strategies has recently demonstrated remarkable performance. Existing methods typically jointly leverage attribute and structural augmentations to generate graph views, learning data invariance information through contrasting sample pairs. However, th…

Cited by 0SourceScholar
2025

Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning

ACL 2025long

Previous continual learning setups for embodied intelligence focused on executing low-level actions based on human commands, neglecting the ability to learn high-level planning and multi-level knowledge. To address these issues, we propose the Hierarchical Embodied Continual Learning Setups (HEC) th…

Cited by 0SourcePDFScholar
2025

Homogeneous Graph Extraction: An Approach to Learning Heterogeneous Graph Embedding

ICASSP 2025accepted

Heterogeneous Graph Neural Networks (HGNNs) aim to embed rich structural and semantic information of heterogeneous graphs into low-dimensional node representations. While HGNNs extend the foundational work of homogeneous Graph Neural Networks, the methodology for effectively transforming heterogeneo…

Cited by 0SourceScholar
2025

MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts

ACL 2025long

One of the primary challenges in optimizing large language models (LLMs) for long-context inference lies in the high memory consumption of the Key-Value (KV) cache. Existing approaches, such as quantization, have demonstrated promising results in reducing memory usage. However, current quantization…

Cited by 0SourcePDFScholar
2025

PointActionCLIP: Preventing Transfer Degradation in Point Cloud Action Recognition with a Triple-Path CLIP

ICASSP 2025accepted

Directly applying CLIP to point cloud action recognition can cause severe accuracy collapse. In this paper, we propose PointActionCLIP, which successfully prevents this transfer degradation with a triplepath CLIP, including the image path, the sequence path, and the label path. Specifically, the ima…

Cited by 0SourceScholar
2025

RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models

ACL 2025finding

Object Navigation (ObjectNav) is a fundamental task in embodied artificial intelligence. Although significant progress has been made in semantic map construction and target direction prediction in current research, redundant exploration and exploration failures remain inevitable. A critical but unde…

Cited by 0SourcePDFScholar
2025

RUNA: Object-Level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations

AAAI 2025technical

Enabling object detectors to recognize out-of-distribution (OOD) objects is vital for building reliable systems. A primary obstacle stems from the fact that models frequently do not receive supervisory signals from unfamiliar data, leading to overly confident predictions regarding OOD objects. Despi…

Cited by 0SourcePDFScholar
2025

VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection

ICASSP 2025accepted

As object detectors are increasingly deployed as black-box cloud services or pre-trained models with restricted access to the original training data, the challenge of zero-shot object-level out-of-distribution (OOD) detection arises. This task becomes crucial in ensuring the reliability of detectors…

Cited by 0SourceScholar
2024

ED-TTS: Multi-Scale Emotion Modeling Using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

ICASSP 2024accepted

Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech synthesis model that leverages Speech Emotion Diarization (…

Cited by 0SourceScholar
2024

EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model

ICASSP 2024accepted

In recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges related to limited generalization, particularly when dealing with…

Cited by 0SourceScholar
2024

From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

NAACL 2024long

In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point. Recognizing this, we introduce a self-guided methodology for LLMs to autonomously discern and select cherry samples from open-source datasets, effectively minimizing manual curat…

2024

IDEAW: Robust Neural Audio Watermarking with Invertible Dual-Embedding

EMNLP 2024main

The audio watermarking technique embeds messages into audio and accurately extracts messages from the watermarked audio. Traditional methods develop algorithms based on expert experience to embed watermarks into the time-domain or transform-domain of signals. With the development of deep neural netw…

2024

INCPrompt: Task-Aware Incremental Prompting for Rehearsal-Free Class-Incremental Learning

ICASSP 2024accepted

This paper introduces INCPrompt, an innovative continual learning solution that effectively addresses catastrophic forgetting. INCPrompt’s key innovation lies in its use of adaptive key-learner and task-aware prompts that capture task-relevant information. This unique combination encapsulates genera…

Cited by 0SourceScholar
2024

Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval

ICASSP 2024accepted

Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio has the potential ability to well represent content. Besides,…

Cited by 0SourceScholar
2024

Leveraging Biases in Large Language Models: "bias-kNN" for Effective Few-Shot Learning

ICASSP 2024accepted

Large Language Models (LLMs) have shown significant promise in various applications, including zero-shot and few-shot learning. However, their performance can be hampered by inherent biases. Instead of traditionally sought methods that aim to minimize or correct these biases, this study introduces a…

Cited by 0SourceScholar
2024

P2DT: Mitigating Forgetting in Task-Incremental Learning with Progressive Prompt Decision Transformer

ICASSP 2024accepted

Catastrophic forgetting poses a substantial challenge for managing intelligent agents controlled by a large model, causing performance degradation when these agents face new tasks. In our work, we propose a novel solution - the Progressive Prompt Decision Transformer (P2DT). This method enhances a t…

Cited by 0SourceScholar
2024

Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning

ACL 2024long

Instruction tuning is critical to improve LLMs but usually suffers from low-quality and redundant data. Data filtering for instruction tuning has proved important in improving both the efficiency and performance of the tuning process. But it also leads to extra cost and computation due to the involv…

2023

Detecting Out-of-Distribution Examples Via Class-Conditional Impressions Reappearing

ICASSP 2023accepted

Out-of-distribution (OOD) detection aims at enhancing standard deep neural networks to distinguish anomalous inputs from original training data. Previous progress has introduced various approaches where the in-distribution training data and even several OOD examples are prerequisites. However, due t…

Cited by 0SourceScholar
2023

Dynamic Alignment Mask CTC: Improved Mask CTC With Aligned Cross Entropy

ICASSP 2023accepted

Because of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autoregressive models. In this work, we present dynamic alignment Mask CTC, introducing two methods: (1) Aligned Cross Entrop…

Cited by 0SourceScholar
2023

Efficient Uncertainty Estimation with Gaussian Process for Reliable Dialog Response Retrieval

ICASSP 2023accepted

Deep neural networks have achieved remarkable performance in retrieval-based dialogue systems, but they are shown to be ill calibrated. Though basic calibration methods like Monte Carlo Dropout and Ensemble can calibrate well, these methods are time-consuming in the training or inference stages. To…

Cited by 0SourceScholar
2023

Feature-Rich Audio Model Inversion for Data-Free Knowledge Distillation Towards General Sound Classification

ICASSP 2023accepted

Data-Free Knowledge Distillation (DFKD) has recently attracted growing attention in the academic community, especially with major breakthroughs in computer vision. Despite promising results, the technique has not been well applied to audio and signal processing. Due to the variable duration of audio…

Cited by 0SourceScholar
2023

FedET: A Communication-Efficient Federated Class-Incremental Learning Framework Based on Enhanced Transformer

IJCAI 2023poster

Federated Learning (FL) has been widely concerned for it enables decentralized learning while ensuring data privacy. However, most existing methods unrealistically assume that the classes encountered by local clients are fixed over time. After learning new classes, this impractical assumption will m…

Cited by 30SourcePDFScholar
2023

GAIA: Delving into Gradient-based Attribution Abnormality for Out-of-distribution Detection

NeurIPS 2023poster

Detecting out-of-distribution (OOD) examples is crucial to guarantee the reliability and safety of deep neural networks in real-world settings. In this paper, we offer an innovative perspective on quantifying the disparities between in-distribution (ID) and OOD data---analyzing the uncertainty that…

2023

Improving EEG-based Emotion Recognition by Fusing Time-Frequency and Spatial Representations

ICASSP 2023accepted

Using deep learning methods to classify EEG signals can accurately identify people’s emotions. However, existing studies have rarely considered the application of the information in another domain’s representations to feature selection in the time-frequency domain. We propose a classification networ…

Cited by 0SourceScholar
2023

Improving Music Genre Classification from multi-modal Properties of Music and Genre Correlations Perspective

ICASSP 2023accepted

Music genre classification has been widely studied in past few years for its various applications in music information retrieval. Previous works tend to perform unsatisfactorily, since those methods only use audio content or jointly use audio content and lyrics content inefficiently. In addition, as…

Cited by 0SourceScholar
2023

Learning Speech Representations with Flexible Hidden Feature Dimensions

ICASSP 2023accepted

Non-parallel many-to-many voice conversion is a kind of style transfer task in speech. Recently, AutoVC has been applied in this field as a popular solution, as it can achieve distribution-matching style transfer by training only the re- construction loss. However, in order to strike a good balance…

Cited by 0SourceScholar
2023

On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval Models

AAAI 2023technical

Deep neural retrieval models have amply demonstrated their power but estimating the reliability of their predictions remains challenging. Most dialog response retrieval models output a single score for a response on how relevant it is to a given question. However, the bad calibration of deep neural…

2023

PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual Adapter

EMNLP 2023long main

The Retrieval Question Answering (ReQA) task employs the retrieval-augmented framework, composed of a retriever and generator. The generators formulate the answer based on the documents retrieved by the retriever. Incorporating Large Language Models (LLMs) as generators is beneficial due to their ad…

Cited by 0SourceScholar
2023

QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis

ICASSP 2023accepted

Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and control intonation to further deliver the speaker’s questioning intention while tran…

Cited by 0SourceScholar
2023

VQ-CL: Learning Disentangled Speech Representations with Contrastive Learning and Vector Quantization

ICASSP 2023accepted

Voice Conversion(VC) refers to converting the voice characteristics of audio to another one as it is said by other people. Recently, more and more studies have focused on disentangle-based VC, which separates the timbre and linguistic content information from an audio signal to effectively achieve V…

Cited by 0SourceScholar
2022

Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive Learning

ICASSP 2022accepted

Voice Conversion(VC) refers to changing the timbre of a speech while retaining the discourse content. Recently, many works have focused on disentangle-based learning techniques to separate the timbre and the linguistic content information from a speech signal. Once successful, voice conversion will…

Cited by 0SourceScholar
2022

DRVC: A Framework of Any-to-Any Voice Conversion with Self-Supervised Learning

ICASSP 2022accepted

Any-to-any voice conversion problem aims to convert voices for source and target speakers, which are out of the training data. Previous works wildly utilize the disentangle-based models. The disentangle-based model assumes the speech consists of content and speaker style information and aims to unta…

Cited by 0SourceScholar
2022

nnSpeech: Speaker-Guided Conditional Variational Autoencoder for Zero-Shot Multi-speaker text-to-speech

ICASSP 2022accepted

Multi-speaker text-to-speech (TTS) using a few adaption data is a challenge in practical applications. To address that, we propose a zero-shot multi-speaker TTS, named nnSpeech, that could synthesis a new speaker voice without fine-tuning and using only one adaption utterance. Compared with using a…

Cited by 0SourceScholar
2022

r-G2P: Evaluating and Enhancing Robustness of Grapheme to Phoneme Conversion by Controlled Noise Introducing and Contextual Information Incorporation

ICASSP 2022accepted

Grapheme-to-phoneme (G2P) conversion is the process of converting the written form of words to their pronunciations. It has an important role for text-to-speech (TTS) synthesis and automatic speech recognition (ASR) systems. In this paper, we aim to evaluate and enhance the robustness of G2P models.…

Cited by 0SourceScholar
2021

Efficient Client Contribution Evaluation for Horizontal Federated Learning

ICASSP 2021accepted

In federated learning (FL), fair and accurate measurement of the contribution of each federated participant is of great significance. The level of contribution not only provides a rational metric for distributing financial benefits among federated participants, but also helps to discover malicious p…

Cited by 0SourceScholar
2021

Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual Interpolation

ICASSP 2021accepted

Knowledge distillation refers to a technique of transferring the knowledge from a large learned model or an ensemble of learned models to a small model. This method relies on access to the original training set, which might not always be available. A possible solution is a data-free adversarial dist…

Cited by 0SourceScholar
2021

Joint Intent Detection and Slot Filling Based on Continual Learning Model

ICASSP 2021accepted

Slot filling and intent detection have become a significant theme in the field of natural language understanding. Even though slot filling is intensively associated with intent detection, the characteristics of the information required for both tasks are different while most of those approaches may…

Cited by 0SourceScholar
2021

LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation

ICASSP 2021accepted

In this paper, we propose a novel conditional convolution network, named location-variable convolution, to model the dependencies of the waveform sequence. Different from the use of unified convolution kernels in WaveNet to capture the dependencies of arbitrary waveform, the location-variable convol…

Cited by 0SourceScholar
2021

Unidirectional Memory-Self-Attention Transducer for Online Speech Recognition

ICASSP 2021accepted

Self-attention models have been successfully applied in end-to-end speech recognition systems, which greatly improve the performance of recognition accuracy. However, such attention-based models cannot be used in online speech recognition, because these models usually have to utilize a whole acousti…

Cited by 0SourceScholar
2020

A Robust Speaker Clustering Method Based on Discrete Tied Variational Autoencoder

ICASSP 2020accepted

Recently, the speaker clustering model based on aggregation hierarchy cluster (AHC) is a common method to solve two main problems: no preset category number clustering and fix category number clustering. In general, model takes features like i-vectors as input of probability and linear discriminant…

Cited by 0SourceScholar
2020

Aligntts: Efficient Feed-Forward Text-to-Speech System Without Explicit Alignment

ICASSP 2020accepted

Targeting at both high efficiency and performance, we propose AlignTTS to predict the mel-spectrum in parallel. AlignTTS is based on a Feed-Forward Transformer which generates mel-spectrum from a sequence of characters, and the duration of each character is determined by a duration predictor. Instea…

Cited by 0SourceScholar
2020

GraphTTS: Graph-to-Sequence Modelling in Neural Text-to-Speech

ICASSP 2020accepted

This paper leverages the graph-to-sequence method in neural text-to-speech (GraphTTS), which maps the graph embedding of the input sequence to spectrograms. The graphical inputs consist of node and edge representations constructed from input texts. The encoding of these graphical inputs incorporates…

Cited by 0SourceScholar