← Search

Lu Lu

28 accepted papers

2026

DuPO: Enabling Reliable Self-Verification via Dual Preference Optimization

ICLR 2026poster

We present DuPO, a dual learning-based preference optimization framework that generates annotation-free feedback via the generalized duality. DuPO addresses two key limitations: Reinforcement Learning with Verifiable Rewards (RLVR)’s reliance on costly labels and applicability restricted to verifiab…

Cited by 0SourceScholar
2026

EarlyTom: Early Token Compression Completes Fast Video Understanding

CVPR 2026

Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficiency introduced by processing massive amounts of visual tokens. Although recent approaches achieve extremely low token ret

Cited by 0SourceScholar
2026

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

ICLR 2026poster

Speech-to-Speech (S2S) models have shown promising dialogue capabilities, but their ability to handle paralinguistic cues—such as emotion, tone, and speaker attributes—and to respond appropriately in both content and style remains underexplored. Progress is further hindered by the scarcity of high-q…

Cited by 0SourceScholar
2026

Q Cache: Visual Attention Is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model

AAAI 2026technical

Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual tokens engenders a substantial computational load and key-value (KV) cache footprint bottleneck. Existing approaches focus

Cited by 0SourcePDFScholar
2026

RAPID: Reusing Attention Sparsity with Inter-step Adaptation for Efficient Video Diffusion

CVPR 2026

The prohibitive cost of 3D attention hinders high-quality video generation with diffusion models. Existing sparse attention methods either lack content adaptivity (static) or incur excessive overhead from per-step recalculation (dynamic). Our work challenges the necessity of this trade-off, based on

Cited by 0SourceScholar
2025

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

ICASSP 2025accepted

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced audito…

Cited by 0SourceScholar
2025

MSTBI: Head CT Detection and Prognostic Assessment of Traumatic Brain Injury Dataset

ICASSP 2025accepted

A high-quality traumatic brain injury (TBI) dataset is essential to intelligent assisted diagnosis. In this work, we present a new labeled public TBI dataset containing head CT scans and various examination variables. The dataset was collected from 545 patients with moderate to severe TBI and CT ann…

Cited by 0SourceScholar
2025

QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions

ACL 2025long

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing d…

2025

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

NeurIPS 2025poster

In order to enable fluid and natural human-machine speech interaction, existing full-duplex conversational systems often adopt modular architectures with auxiliary components such as voice activity detectors, interrupters, conversation state predictors, or multiple LLMs. These systems, however, suff…

Cited by 0SourcecodeScholar
2025

ST3: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

AAAI 2025technical

Multimodal large language models (MLLMs) enhance their perceptual capabilities by integrating visual and textual information. However, processing the massive number of visual tokens incurs a significant computational cost. Existing analysis of the MLLM attention mechanisms remains shallow, leading t…

Cited by 2SourcePDFScholar
2025

Spy Inside: Scalable Verification of Dependable Transformers for Event Time Series Systems

ICASSP 2025accepted

Event time series appear in many software scenarios and are a necessary data type in data analytics systems. Transformers are the preferred type of sequential neural network for advanced analytics on event time series, particularly due to their significant contributions to the recent surge of large…

Cited by 0SourceScholar
2024

Challenges in Training PINNs: A Loss Landscape Perspective

ICML 2024oral

This paper explores challenges in training Physics-Informed Neural Networks (PINNs), emphasizing the role of the loss landscape in the training process. We examine difficulties in minimizing the PINN loss function, particularly due to ill-conditioning caused by differential operators in the residual…

2024

Connecting Speech Encoder and Large Language Model for ASR

ICASSP 2024accepted

The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative s…

Cited by 0SourceScholar
2024

Extending Large Language Models for Speech and Audio Captioning

ICASSP 2024accepted

Multimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achie…

Cited by 0SourceScholar
2024

Extending Multilingual ASR to New Languages Using Supplementary Encoder and Decoder Components

ICASSP 2024accepted

Extending multilingual automatic speech recognition (mASR) systems to new languages poses challenges, particularly when training data for existing languages is limited or unavailable. To tackle this issue, we suggest utilizing supplementary encoder and decoder components. Specifically, we propose ap…

Cited by 0SourceScholar
2024

PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs

NeurIPS 2024poster

While significant progress has been made on Physics-Informed Neural Networks (PINNs), a comprehensive comparison of these methods across a wide range of Partial Differential Equations (PDEs) is still lacking. This study introduces PINNacle, a benchmarking tool designed to fill this gap. PINNacle pro…

2024

PolyVoice: Language Models for Speech to Speech Translation

ICLR 2024poster

With the huge success of GPT models in natural language processing, there is a growing interest in applying language modeling approaches to speech tasks. Currently, the dominant architecture in speech-to-speech translation (S2ST) remains the encoder-decoder paradigm, creating a need to investigate t…

2024

SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR

ICASSP 2024accepted

Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among variou…

Cited by 0SourceScholar
2024

SALMONN: Towards Generic Hearing Abilities for Large Language Models

ICLR 2024poster

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a…

2024

SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

NeurIPS 2024poster

Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information. This comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction. Chat-Oriented Large Language Models (LLMs), known fo…

2024

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

ICML 2024poster

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video processing, which can understand not only visual frame sequences…

2023

CIF-PT: Bridging Speech and Text Representations for Spoken Language Understanding via Continuous Integrate-and-Fire Pre-Training

ACL 2023findings

Speech or text representation generated by pre-trained models contains modal-specific information that could be combined for benefiting spoken language understanding (SLU) tasks. In this work, we propose a novel pre-training paradigm termed Continuous Integrate-and-Fire Pre-Training (CIF-PT). It rel…

2020

Spatial Attention for Far-Field Speech Recognition with Deep Beamforming Neural Networks

ICASSP 2020accepted

In this paper, we introduce spatial attention for refining the information in multi-direction neural beamformer for far-field automatic speech recognition. Previous approaches of neural beamformers with multiple look directions, such as the factored complex linear projection, have shown promising re…

Cited by 0SourceScholar
2020

Velocity Field based Active-Assistive Control for Upper Limb Rehabilitation Exoskeleton Robot

ICRA 2020poster

There are limitations of conventional active-assistive control for upper limb rehabilitation exoskeleton robot, such as 1). prior time-dependent trajectories are generally required, 2). task-based rehabilitation exercise involving multi-joint motion is hard to implement, and 3). assistive mechanism…

Cited by 8SourceScholar
2018

Robot Composite Learning and the Nunchaku Flipping Challenge

ICRA 2018poster

Advanced motor skills are essential for robots to physically coexist with humans. Much research on robot dynamics and control has achieved success on hyper robot motor capabilities, but mostly through heavily case-specific engineering. Meanwhile, in terms of robot acquiring skills in a ubiquitous ma…

Cited by 5SourceScholar