← Search

Xinyuan Zhou

7 accepted papers

2025

Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation

ICASSP 2025accepted

End-to-end speech-to-text translation (E2E ST) has increasingly aroused interest and attention recently, attempting to address the problem of data scarcity and modeling burden. Several attempts exploring the combination of Large Speech and Language Models into a unified model to improve E2E ST are c…

Cited by 0SourceScholar
2025

SAR Ship Detector Using Cross-stage Feature Fusion and Decoupled Head with Mutual Guidance

ICASSP 2025accepted

Deep learning-based SAR ship detection methods enhance resilience to noise, distortion, and interference in ocean environments, establishing them as the foremost approach for ship detection nowadays. Nonetheless, substantial difficulties persist in separating ships from the complex backgrounds found…

Cited by 0SourceScholar
2025

Scalable Data Synthesis through Human-like Cognitive Imitation and Data Recombination

EMNLP 2025

Large language models (LLMs) rely on massive amounts of training data, however, the quantity of empirically observed data is limited. To alleviate this issue, lots of LLMs leverage synthetic data to enhance the quantity of training data. Despite significant advancements in LLMs, the efficiency and s

Cited by 0SourcePDFScholar
2024

A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker Extraction

ICASSP 2024accepted

Target speaker extraction (TSE) based on direction of arrival (DOA) has a wide range of applications in e.g., remote conferencing, hearing aids, in-car speech interaction. Due to the inherent phase uncertainty, existing TSE methods usually suffer from speaker confusion within specific frequency band…

Cited by 0SourceScholar
2024

Pre-Trained Acoustic-and-Textual Modeling for End-To-End Speech-To-Text Translation

ICASSP 2024accepted

End-to-end paradigm has aroused more and more interests and attention for improving speech-to-text translation (ST) recently. Existing end-to-end models mainly attributes and attempts to address the problem of modeling burden and data scarcity, while always fail to maintain both cross-modal and cros…

Cited by 0SourceScholar
2023

DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation

EMNLP 2023long main

While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial problem. Specifically, due to the low information density of speech data, the transfo…

Cited by 0SourceScholar
2021

Multi-Channel Target Speech Extraction with Channel Decorrelation and Target Speaker Adaptation

ICASSP 2021accepted

The end-to-end approaches for single-channel target speech extraction have attracted widespread attention. However, the studies for end-to-end multi-channel target speech extraction are still relatively limited. In this work, we propose two methods for exploiting the multi-channel spatial informatio…

Cited by 0SourceScholar