← Search

Zejun Ma

50 accepted papers

2026

SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?

ICML 2026poster

Code performance optimization is paramount in real-world software engineering and critical for production-level systems. While Large Language Models (LLMs) have demonstrated impressive capabilities in code generation and bug fixing, their proficiency in enhancing code performance at the repository l…

Cited by 0SourceScholar
2026

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

ICLR 2026poster

Large Language Models (LLMs) can enhance their reasoning by interacting with external tools, a paradigm known as Tool-Integrated Reasoning (TIR). However, extending TIR to multi-turn settings using Reinforcement Learning (RL) often exhibits training instability and degraded performance. We attribute…

Cited by 0SourcecodeScholar
2026

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

ICML 2026poster

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhanced streaming audio-visual large language model that processes over 3-hour videos at $1$ FPS and $360$p resolution, outper…

Cited by 0SourceScholar
2025

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

NeurIPS 2025poster

Large Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optimization framework to address this, employing a closed-loop system where LLMs ite…

Cited by 0SourcecodeScholar
2025

Audio-centric Video Understanding Benchmark without Text Shortcut

EMNLP 2025

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly depends on auditory information, as audio offers critical cont

2025

General-Reasoner: Advancing LLM Reasoning Across All Domains

NeurIPS 2025poster

Reinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs). Particularly, the "Zero" reinforcement learning introduced by Deepseek-R1-Zero, enables direct RL training of base LLMs without relying on an intermediate s…

Cited by 0SourceScholar
2025

Improving LLM Video Understanding with 16 Frames Per Second

ICML 2025poster

Human vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) $\leqslant$2, leading to critical visual informat…

2025

LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

ICLR 2025spotlight

Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains less explored. Additionally, prior LMM research separately tac…

Cited by 0SourcePDFScholar
2025

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

CVPR 2025poster

Recent video large language models (Video LLMs) often depend on costly human annotations or proprietary APIs (e.g., GPT-4o) to produce training data, which limits their training at scale. In this paper, we explore large-scale training for Video LLM with cheap automatic speech recognition (ASR) trans…

2025

Robust SuperAlignment: Weak-to-Strong Robustness Generalization for Vision-Language Models

NeurIPS 2025spotlight

Numerous well-established studies have demonstrated the superhuman capabilities of modern Vision-Language Models (VLMs) across a wide range of tasks. However, growing is the doubt about the continuing availability of reliable high-quality labeling (supervision) from human annotators, leading to stag…

Cited by 6SourceScholar
2025

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

NeurIPS 2025poster

Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. However, we observe that most real-world scenarios do not require such an extensive number of visual tokens. While the perf…

Cited by 0SourceScholar
2025

ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear Attention

NeurIPS 2025poster

Linear attention mechanisms deliver significant advantages for Large Language Models (LLMs) by providing linear computational complexity, enabling efficient processing of ultra-long sequences (e.g., 1M context). However, existing Sequence Parallelism (SP) methods, essential for distributing these wo…

Cited by 0SourceScholar
2025

video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

ICML 2025poster

While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in gener…

2024

Connecting Speech Encoder and Large Language Model for ASR

ICASSP 2024accepted

The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative s…

Cited by 0SourceScholar
2024

Extending Large Language Models for Speech and Audio Captioning

ICASSP 2024accepted

Multimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achie…

Cited by 0SourceScholar
2024

Extending Multilingual ASR to New Languages Using Supplementary Encoder and Decoder Components

ICASSP 2024accepted

Extending multilingual automatic speech recognition (mASR) systems to new languages poses challenges, particularly when training data for existing languages is limited or unavailable. To tackle this issue, we suggest utilizing supplementary encoder and decoder components. Specifically, we propose ap…

Cited by 0SourceScholar
2024

Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

ICLR 2024poster

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspe…

2024

PolyVoice: Language Models for Speech to Speech Translation

ICLR 2024poster

With the huge success of GPT models in natural language processing, there is a growing interest in applying language modeling approaches to speech tasks. Currently, the dominant architecture in speech-to-speech translation (S2ST) remains the encoder-decoder paradigm, creating a need to investigate t…

2024

RePOSE: 3D Human Pose Estimation via Spatio-Temporal Depth Relational Consistency

ECCV 2024poster

"We introduce RePOSE, a simple yet effective approach for addressing occlusion challenges in the learning of 3D human pose estimation (HPE) from videos. Conventional approaches typically employ absolute depth signals as supervision, which are adept at discernible keypoints but become less reliable w…

2024

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

ICLR 2024spotlight

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talkin…

2024

SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR

ICASSP 2024accepted

Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among variou…

Cited by 0SourceScholar
2024

SALMONN: Towards Generic Hearing Abilities for Large Language Models

ICLR 2024poster

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a…

2024

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

ICML 2024poster

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video processing, which can understand not only visual frame sequences…

2023

An ASR-Free Fluency Scoring Approach with Self-Supervised Learning

ICASSP 2023accepted

A typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for the subsequent calculation of fluency-related features or directly modeling speech fluency with an end-to-end approach. This paper describes a novel ASR-free ap…

Cited by 0SourceScholar
2023

Bytecover3: Accurate Cover Song Identification On Short Queries

ICASSP 2023accepted

Deep learning based methods have become a paradigm for cover song identification (CSI) in recent years, where the ByteCover systems have achieved state-of-the-art results on all the mainstream datasets of CSI. However, with the burgeon of short videos, many real-world applications require matching s…

Cited by 0SourceScholar
2023

Internal Language Model Estimation Based Adaptive Language Model Fusion for Domain Adaptation

ICASSP 2023accepted

ASR model deployment environment is ever-changing, and the incoming speech can be switched across different domains during a session. This brings a challenge for effective domain adaptation when only target domain text data is available, and our objective is to obtain obviously improved performance…

Cited by 0SourceScholar
2023

Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation Scoring

ICASSP 2023accepted

Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual pronunciation of the target phone as the phone-level pronunciation qu…

Cited by 0SourceScholar
2023

LiteG2P: A Fast, Light and High Accuracy Model for Grapheme-to-Phoneme Conversion

ICASSP 2023accepted

As a key component of automated speech recognition (ASR) and the front-end in text-to-speech (TTS), grapheme-to-phoneme (G2P) plays the role of converting letters to their corresponding pronunciations. Existing methods are either slow or poor in performance, and are limited in application scenarios,…

Cited by 0SourceScholar
2023

Virtual Try-On with Pose-Garment Keypoints Guided Inpainting

ICCV 2023poster

Virtual try-on is an important technology supporting online apparel shopping, which provides consumers with a virtual experience to fit garments without physically wearing them. Recently, the image-based virtual try-on has received growing research attention. However, the synthetic results of existi…

Cited by 32PDFcodeScholar
2022

BiFSMN: Binary Neural Network for Keyword Spotting

IJCAI 2022poster

The deep neural networks, such as the Deep-FSMN, have been widely studied for keyword spotting (KWS) applications. However, computational resources for these networks are significantly constrained since they usually run on-call on edge devices. In this paper, we present BiFSMN, an accurate and extre…

2022

Bytecover2: Towards Dimensionality Reduction of Latent Embedding for Efficient Cover Song Identification

ICASSP 2022accepted

Convolutional neural network (CNN)-based methods have dominated the recent research of cover song identification (CSI). A typical example is the ByteCover system we proposed, which has achieved state-of-the-art results on all the mainstream datasets of CSI. In this paper, we propose an up-graded ver…

Cited by 0SourceScholar
2022

HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection

ICASSP 2022accepted

Audio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted in this field. However, existing audio transformers require large GPU memories and long training time, meanwhile relying…

Cited by 0SourceScholar
2022

Improving Contextual Representation with Gloss Regularized Pre-training

NAACL 2022findings

Though achieving impressive results on many NLP tasks, the BERT-like masked language models (MLM) encounter the discrepancy between pre-training and inference. In light of this gap, we investigate the contextual representation of pre-training and inference from the perspective of word probability di…

2022

Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection

ICASSP 2022accepted

Nowadays, most methods for end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may suffer from the confusion between simil…

Cited by 59SourceScholar
2022

Improving Pseudo-Label Training For End-To-End Speech Recognition Using Gradient Mask

ICASSP 2022accepted

In the recent trend of semi-supervised speech recognition, both self-supervised representation learning and pseudo-labeling have shown promising results. In this paper, we propose a novel approach to combine their ideas for end-to-end speech recognition model. Without any extra loss function, we uti…

Cited by 0SourceScholar
2022

Language Adaptive Cross-Lingual Speech Representation Learning with Sparse Sharing Sub-Networks

ICASSP 2022accepted

Unsupervised cross-lingual speech representation learning (XLSR) has recently shown promising results in speech recognition by leveraging vast amounts of unlabeled data across multiple languages. However, standard XLSR model suffers from language interference problem due to the lack of language spec…

Cited by 0SourceScholar
2022

S3T: Self-Supervised Pre-Training with Swin Transformer For Music Classification

ICASSP 2022accepted

In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S3T introduces a momentum-based paradigm, MoCo, with Swin Transformer as its feat…

Cited by 0SourceScholar
2022

The Volcspeech System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to make the clustering-based speaker diarization system enable to handle overlapped speech. Front-end dereverberation and the direction-of-…

Cited by 0SourceScholar
2022

Towards Using Clothes Style Transfer for Scenario-Aware Person Video Generation

ICASSP 2022accepted

Clothes style transfer for person video generation is a challenging task, due to drastic variations of intra-person appearance and video scenarios. To tackle this problem, most recent AdaIN-based architectures are proposed to extract clothes and scenario features for generation. However, these appro…

Cited by 0SourceScholar
2022

Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled Data

AAAI 2022technical

Deep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a single model to target multiple sources, they have difficulty…

2021

A Chapter-Wise Understanding System for Text-To-Speech in Chinese Novels

ICASSP 2021accepted

In TTS-based audiobook production, multi-role dubbing and emotional expressions can significantly improve the naturalness of audiobooks. However, it requires manual annotation of original novels with explicit speaker and emotion tags in sentence level, which is extremely time-consuming and costly. I…

Cited by 0SourceScholar
2021

An Hrnet-Blstm Model With Two-Stage Training For Singing Melody Extraction

ICASSP 2021accepted

Well-labeled datasets available for melody extraction are scarce, which limits the further advancement of deep learning based methods. To overcome this problem, we propose to use a pitch refinement method to refine the semitone-level pitch sequences decoded from massive melody MIDI files to generate…

Cited by 0SourceScholar
2021

Bytecover: Cover Song Identification Via Multi-Loss Training

ICASSP 2021accepted

We present in this paper ByteCover, which is a new feature learning method for cover song identification (CSI). Byte-Cover is built based on the classical ResNet model, and two major improvements are designed to further enhance the capability of the model for CSI. In the first improvement, we introd…

Cited by 0SourceScholar
2021

Improving RNN Transducer Modeling for Small-Footprint Keyword Spotting

ICASSP 2021accepted

The recurrent neural network transducer (RNN-T) model has been proved effective for keyword spotting (KWS) recently. However, compared with cross-entropy (CE) or connectionist temporal classification (CTC) based models, the additional prediction network in the RNN-T model increases the model size an…

Cited by 0SourceScholar
2021

PPG-Based Singing Voice Conversion with Adversarial Representation Learning

ICASSP 2021accepted

Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion works, we propose a novel model to steadily convert songs while keeping their naturalness and intonation. We build an end-to…

Cited by 0SourceScholar
2021

Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video Streams

ICASSP 2021accepted

Detecting anchor’s voice in live musical streams is an important preprocessing step for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to effectively focus on the target voice in noisy environme…

Cited by 0SourceScholar
2021

Singing Melody Extraction from Polyphonic Music based on Spectral Correlation Modeling

ICASSP 2021accepted

Convolutional neural network (CNN) based methods have achieved state-of-the-art performance for singing melody extraction from polyphonic music. However, most of these methods focus on the learning of local features, while relationships among spectral components locating far apart are often neglecte…

Cited by 0SourceScholar
2020

A Hybrid Text Normalization System Using Multi-Head Self-Attention For Mandarin

ICASSP 2020accepted

In this paper, we propose a hybrid text normalization system using multi-head self-attention. The system combines the advantages of a rule-based model and a neural model for text preprocessing tasks. Previous studies in Mandarin text normalization usually use a set of hand-written rules, which are h…

Cited by 0SourceScholar
2020

A Unified Sequence-to-Sequence Front-End Model for Mandarin Text-to-Speech Synthesis

ICASSP 2020accepted

In Mandarin text-to-speech (TTS) system, the front-end text processing module significantly influences the intelligibility and naturalness of synthesized speech. Building a typical pipeline-based front-end which consists of multiple individual components requires extensive efforts. In this paper, we…

Cited by 0SourceScholar