← Search

Wei Zou

23 accepted papers

2026

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

CVPR 2026

Recent advances in text-to-audio (TTA) generation excel at synthesizing short audio clips but struggle with long-form narrative audio, which requires temporal coherence and compositional reasoning. To fill this gap, we propose AudioStory, a unified framework that integrates large language models (LL

Cited by 0SourcecodeScholar
2025

Aligned Better, Listen Better for Audio-Visual Large Language Models

ICLR 2025poster

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric settings. However, existing Video-LLMs and Audio-Visual Larg…

Cited by 2SourcePDFScholar
2025

C3oT: Generating Shorter Chain-of-Thought Without Compromising Effectiveness

AAAI 2025technical

Generating Chain-of-Thought (CoT) before deriving the answer can effectively improve the reasoning capabilities of large language models (LLMs) and significantly improve the accuracy of the generated answer. However, in most cases, the length of the generated CoT is much longer than the desired fina…

Cited by 10SourcePDFScholar
2025

IoU-Aware Clustering for Anchor Configuration Determination in Efficient Defect Detection

IROS 2025

Deep-learning-based object detection has gained widespread application in surface defect inspection, with anchor-based detectors achieving remarkable success by utilizing dense anchors to align with defects. Determining the optimal anchor configuration, i.e., sizes and aspect ratios of anchor boxes,

Cited by 0SourceScholar
2025

Monocular Depth Estimation and Segmentation for Transparent Object with Iterative Semantic and Geometric Fusion

ICRA 2025

Transparent object perception is indispensable for numerous robotic tasks. However, accurately segmenting and estimating the depth of transparent objects remain challenging due to complex optical properties. Existing methods primarily delve into only one task using extra inputs or specialized sensor

Cited by 8SourcecodeScholar
2025

SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models

EMNLP 2025

Large Language Models (LLMs) excel at various natural language processing tasks but remain vulnerable to jailbreaking attacks that induce harmful content generation. In this paper, we reveal a critical safety inconsistency: LLMs can more effectively identify harmful requests as discriminators than d

2025

TRANS-ZERO: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data

ACL 2025finding

The rise of Large Language Models (LLMs) has reshaped machine translation (MT), but multilingual MT still relies heavily on parallel data for supervised fine-tuning (SFT), facing challenges like data scarcity for low-resource languages and catastrophic forgetting. To address these issues, we propose…

2025

Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models

EMNLP 2025

End-to-end Large Speech Language Models (LSLMs) have demonstrated impressive conversational generation abilities, yet consistently fall short of traditional pipeline systems on semantic understanding benchmarks. In this work, we reveal through systematic experimentation that although LSLMs lose some

Cited by 0SourcePDFScholar
2024

Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization

ICASSP 2024accepted

Audio-Visual Source Localization (AVSL) is the task of identifying specific sounding objects in the scene given audio cues. In our work, we focus on semi-supervised AVSL with pseudo-labeling. To address the issues with vanilla hard pseudo-labels including bias accumulation, noise sensitivity, and in…

Cited by 0SourceScholar
2024

CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-Training

CVPR 2024poster

Learning joint and coordinated features across modalities is essential for many audio-visual tasks. Existing pre-training methods primarily focus on global information neglecting fine-grained features and positions leading to suboptimal performance in dense prediction tasks. To address this issue we…

Cited by 5SourcePDFScholar
2024

MAPO: Advancing Multilingual Reasoning through Multilingual-Alignment-as-Preference Optimization

ACL 2024long

Intuitively, reasoning abilities are considered language-agnostic. However, existing LLMs exhibit inconsistent reasoning abilities across different languages, e.g., reasoning in the dominant language like English is superior to other languages due to the imbalance of multilingual training data. To e…

2024

MMCert: Provable Defense against Adversarial Attacks to Multi-modal Models

CVPR 2024poster

Different from a unimodal model whose input is from a single modality the input (called multi-modal input) of a multi-modal model is from multiple modalities such as image 3D points audio text etc. Similar to unimodal models many existing studies show that a multi-modal model is also vulnerable to a…

2023

Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization

NeurIPS 2023poster

Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise…

2023

Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam Search

EMNLP 2023long main

Machine translation (MT) quality estimation (QE) is a crucial task to estimate the quality of MT outputs when reference translations are unavailable. Many studies focus on generating pseudo data using large parallel corpus and achieve remarkable success in the supervised setting. However, pseudo dat…

Cited by 0SourcecodeScholar
2023

Local Interpretation of Transformer Based on Linear Decomposition

ACL 2023long

In recent years, deep neural networks (DNNs) have achieved state-of-the-art performance on a wide range of tasks. However, limitations in interpretability have hindered their applications in the real world. This work proposes to interpret neural networks by linear decomposition and finds that the Re…

Cited by 14SourcePDFScholar
2022

Audio Deepfake Detection System with Neural Stitching for ADD 2022

ICASSP 2022accepted

This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge[1]. The very same system was used for both two rounds of evaluation in Track 3.2 with similar training methodology. The first round of Track 3.2 data is generated from Text-to-Speech…

Cited by 0SourceScholar
2022

Audio-Visual Wake Word Spotting System for MISP Challenge 2021

ICASSP 2022accepted

This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage both audio and video information to improve the environmental robustness of far-field wake word spotting. In the propose…

Cited by 0SourceScholar
2022

Time Domain Adversarial Voice Conversion for ADD 2022

ICASSP 2022accepted

In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) system to convert source speech with arbitrary language content into target speaker’s fake speech. Then the converted spee…

Cited by 0SourceScholar
2021

A Further Study of Unsupervised Pretraining for Transformer Based Speech Recognition

ICASSP 2021accepted

The construction of an effective good speech recognition system typically requires large amounts of transcribed data, which is expensive to collect. To overcome this problem, many unsupervised pretraining methods have been proposed. Among these methods, Masked Predictive Coding achieved significant…

Cited by 0SourceScholar
2021

Didispeech: A Large Scale Mandarin Speech Corpus

ICASSP 2021accepted

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is recorded in quiet environment and is suitable for various speech p…

Cited by 0SourceScholar
2021

KeSpeech: An Open Source Speech Dataset of Mandarin and Its Eight Subdialects

NeurIPS 2021poster

This paper introduces an open source speech dataset, KeSpeech, which involves 1,542 hours of speech signals recorded by 27,237 speakers in 34 cities in China, and the pronunciation includes standard Mandarin and its 8 subdialects. The new dataset possesses several properties. Firstly, the dataset pr…

Cited by 41SourceScholar
2021

Transformer Based Unsupervised Pre-Training for Acoustic Representation Learning

ICASSP 2021accepted

Recently, a variety of acoustic tasks and related applications arised. For many acoustic tasks, the labeled data size may be limited. To handle this problem, we propose an unsupervised pre-training method using Transformer based encoder to learn a general and robust high-level representation for all…

Cited by 0SourceScholar
2018

End-to-End Flow Correlation Tracking With Spatial-Temporal Attention

CVPR 2018poster

Discriminative correlation filters (DCF) with deep convolutional features have achieved favorable performance in recent tracking benchmarks. However, most of existing DCF trackers only consider appearance features of current frame, and hardly benefit from motion and inter-frame information. The lack…