← Search

Weitai Zhang

9 accepted papers

2025

Adversarial Speech-Text Pre-Training for Speech Translation

ICASSP 2025accepted

Large-scale pre-training has been shown to benefit speech translation tasks. However, existing multimodal pre-training efforts rely on parallel corpora for semantic alignment, potentially limiting performance to the scale of available data and causing data imbalance. Hence, we propose an adversarial…

Cited by 0SourceScholar
2025

Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation

ICASSP 2025accepted

End-to-end speech-to-text translation (E2E ST) has increasingly aroused interest and attention recently, attempting to address the problem of data scarcity and modeling burden. Several attempts exploring the combination of Large Speech and Language Models into a unified model to improve E2E ST are c…

Cited by 0SourceScholar
2025

Large Language Models Are Efficient Learners as Zero-Shot Speech Translators

ICASSP 2025accepted

Significant progress has recently been made in combining Speech Foundation Models (SFMs) and Large Language Models (LLMs) into a unified model to tackle Speech-to-Text Translation (ST) tasks. However, fine-tuning LLMs to adapt to specific downstream tasks requires substantial resources, which is oft…

Cited by 0SourceScholar
2025

Leveraging Boolean Directivity Embedding for Binaural Target Speaker Extraction

ICASSP 2025accepted

Direction-based target speaker extraction (TSE) attracts a constant attention due to the convenience of direction acquisition over assistive video or enrollment audio. The direction clue heavily affects the TSE performance, which might be more seriously in the case of binaural setups due to the smal…

Cited by 0SourceScholar
2025

Scalable Data Synthesis through Human-like Cognitive Imitation and Data Recombination

EMNLP 2025

Large language models (LLMs) rely on massive amounts of training data, however, the quantity of empirically observed data is limited. To alleviate this issue, lots of LLMs leverage synthetic data to enhance the quantity of training data. Despite significant advancements in LLMs, the efficiency and s

Cited by 0SourcePDFScholar
2025

Semi-Supervised Multilingual Alignment with Lexical Memory for Massively Parallel Text Mining

ICASSP 2025accepted

Existing state-of-the-art techniques that employ multilingual sentence embeddings for mining parallel texts predominantly rely on extensive supervision, which often results in sub-optimal performance in the absence of large-scale parallel training datasets. In this study, we introduce a novel method…

Cited by 0SourceScholar
2025

Unlocking the Power of Patch: Patch-Based MLP for Long-Term Time Series Forecasting

AAAI 2025technical

Recent studies have attempted to refine the Transformer architecture to demonstrate its effectiveness in Long-Term Time Series Forecasting (LTSF) tasks. Despite surpassing many linear forecasting models with ever-improving performance, we remain skeptical of Transformers as a solution for LTSF. We a…

2024

A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker Extraction

ICASSP 2024accepted

Target speaker extraction (TSE) based on direction of arrival (DOA) has a wide range of applications in e.g., remote conferencing, hearing aids, in-car speech interaction. Due to the inherent phase uncertainty, existing TSE methods usually suffer from speaker confusion within specific frequency band…

Cited by 0SourceScholar
2024

Pre-Trained Acoustic-and-Textual Modeling for End-To-End Speech-To-Text Translation

ICASSP 2024accepted

End-to-end paradigm has aroused more and more interests and attention for improving speech-to-text translation (ST) recently. Existing end-to-end models mainly attributes and attempts to address the problem of modeling burden and data scarcity, while always fail to maintain both cross-modal and cros…

Cited by 0SourceScholar