← Search

Xiong Wang

8 accepted papers

2026

Dimension-Free Minimax Rates for Learning Pairwise Interactions in Attention-Style Models

ICLR 2026poster

We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a non-linear activation function. We prove that the minimax rate is $M^{-\frac{2\beta}{2\beta+1}}$ with $M$ being the sample size, depending only…

Cited by 0SourceScholar
2025

FedIGL: Federated Invariant Graph Learning for Non-IID Graphs

NeurIPS 2025poster

Federated Graph Learning (FGL) shows superiority in cross-domain graph training while preserving data privacy. Existing approaches usually assume shared generic knowledge (e.g., prototypes, spectral features) via aggregating local structures statistically to alleviate structural heterogeneity. Howev…

Cited by 0SourceScholar
2025

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

ICML 2025poster

The GPT-4o's excellent duplex speech interaction ability has given users an impressive experience. Researchers have recently proposed several multimodal LLMs to achieve user-agent speech-to-speech conversations. In this paper, we propose a novel speech-text multimodal LLM architecture called Freeze-…

Cited by 32SourcePDFScholar
2025

InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training

ACL 2025long

Recent advancements in speech large language models (SpeechLLMs) have attracted considerable attention. Nonetheless, current methods exhibit suboptimal performance in adhering to speech instructions. Notably, the intelligence of models significantly diminishes when processing speech-form input as co…

2025

M-MoE: Mixture of Mixture-of-Expert Model for CTC-based Streaming Multilingual ASR

ICASSP 2025accepted

The Mixture-of-Expert (MoE) structure has been effectively utilized in multilingual ASR tasks. However, the potential of external language information remains underutilized. In this paper, we introduce the Mixture of MoE (M-MoE) structure, featuring multiple language-specific MoEs and a language-unk…

Cited by 0SourceScholar
2025

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

NeurIPS 2025spotlight

Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in bot…

Cited by 0SourcecodeScholar
2019

Adversarial Examples for Improving End-to-end Attention-based Small-footprint Keyword Spotting

ICASSP 2019accepted

In this paper, we explore the use of adversarial examples for improving a neural network based keyword spotting (KWS) system. Specially, in our system, an effective and small-footprint attention-based neural network model is used. Adversarial example is defined as a misclassified example by a model,…

Cited by 0SourceScholar