← Search

Yike Zhang

8 accepted papers

2026

CRAFT: Adapting VLA Models to Contact-Rich Manipulation Via Force-Aware Curriculum Fine-Tuning

ICRA 2026poster

Vision-Language-Action (VLA) models have shown a strong capability in enabling robots to execute general instructions, yet they struggle with contact-rich manipulation tasks, where success requires precise alignment, stable contact maintenance,and effective handling of deformable objects. A fundamen…

2025

FocusLLM: Precise Understanding of Long Context by Dynamic Condensing

ACL 2025long

Empowering LLMs with the ability to precisely understand long contexts is crucial for many downstream applications. However, handling long contexts with conventional transformer architecture requires substantial training and inference resources. Existing context condensing methods cannot accurately…

2025

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

ICML 2025poster

The GPT-4o's excellent duplex speech interaction ability has given users an impressive experience. Researchers have recently proposed several multimodal LLMs to achieve user-agent speech-to-speech conversations. In this paper, we propose a novel speech-text multimodal LLM architecture called Freeze-…

Cited by 32SourcePDFScholar
2025

LeanK: Learnable K Cache Channel Pruning for Efficient Decoding

EMNLP 2025

Large language models (LLMs) enable long-context tasks but face efficiency challenges due to the growing key-value (KV) cache. We propose LeanK, a learning-based method that prunes unimportant key (K) cache channels by leveraging static channel sparsity. LeanK reduces GPU memory and accelerates deco

Cited by 0SourcePDFScholar
2025

M-MoE: Mixture of Mixture-of-Expert Model for CTC-based Streaming Multilingual ASR

ICASSP 2025accepted

The Mixture-of-Expert (MoE) structure has been effectively utilized in multilingual ASR tasks. However, the potential of external language information remains underutilized. In this paper, we introduce the Mixture of MoE (M-MoE) structure, featuring multiple language-specific MoEs and a language-unk…

Cited by 0SourceScholar
2022

Conversational Speech Recognition by Learning Conversation-Level Characteristics

ICASSP 2022accepted

Conversational automatic speech recognition (ASR) is a task to recognize conversational speech including multiple speakers. Unlike sentence-level ASR, conversational ASR can naturally take advantages from specific characteristics of conversation, such as role preference and topical coherence. This p…

Cited by 0SourceScholar
2022

Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models

ICASSP 2022accepted

Recently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2vec2.0 models. Due to the conditional independence assumption, CTC-based models are always weaker than attention-based e…

Cited by 35SourceScholar
2019

Multiple Temporal Scales Based Speaker Embeddings Learning for Text-dependent Speaker Recognition

ICASSP 2019accepted

To extract high speaker-sensitive embeddings from deep neural networks is still a challenge in the field of speaker recognition. This paper proposes a novel network that learns speaker embeddings from multiple temporal scales. This idea comes from the recent biological research that the human audito…

Cited by 0SourceScholar