← Search

Jing LU

33 accepted papers

2026

GENERATING LOCALIZED AUDIBLE ZONES USING A SINGLE-CHANNEL PARAMETRIC LOUDSPEAKER

ICASSP 2026poster

Advanced sound zone control (SZC) techniques typically rely on massive multi-channel loudspeaker arrays to create high-contrast personal sound zones, making single-loudspeaker SZC seem impossible. In this Letter, we challenge this paradigm by introducing the multi-carrier parametric loudspeaker (MCP…

Cited by 0SourcePDFScholar
2026

PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement

AAAI 2026technical

Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches. However, existing generative SE approaches often overlook the risk of hallucination under severe noise, leading to incorrect spoken conten

Cited by 0SourcePDFScholar
2026

Refine3D: Scene-Adaptive Reference Point Refinement for Sparse 3D Object Detection

AAAI 2026technical

Sparse query-based detectors have emerged as the dominant paradigm in camera-only 3D object detection, owing to their exceptional performance and computational efficiency. A central component of these approaches is the use of reference points, which serve as learnable spatial anchors to guide queri

Cited by 0SourcePDFScholar
2026

Rethinking Flow and Diffusion Bridge Models for Speech Enhancement

AAAI 2026technical

Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a fra

Cited by 0SourcePDFScholar
2025

BridgeVoC: Neural Vocoder with Schrödinger Bridge

IJCAI 2025

While previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restora

Cited by 0SourcePDFScholar
2025

DistillW2N: A Lightweight One-Shot Whisper to Normal Voice Conversion Model Using Distillation of Self-Supervised Features

ICASSP 2025accepted

Whisper to Normal voice conversion (W2N) holds great promise for assistive communication and healthcare, making it an exciting area of research and development. Recent advancements in W2N are predominantly driven by self-supervised speech representation learning (SSL) techniques. While effective, SS…

Cited by 0SourceScholar
2025

Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?

NAACL 2025findings

Human understanding of text depends on general semantic concepts of words rather than their superficial forms. To what extent does our human intuition transfer to language models? In this work, we study the degree to which current multilingual language models (mLMs) understand based on subword-level…

Cited by 0SourcePDFScholar
2025

Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution Constraints

AAAI 2025technical

Sentiment analysis is rapidly advancing by utilizing various data modalities (e.g., text, video, and audio). However, most existing techniques only learn the atomic-level features that reflect strong correlations, while ignoring more complex compositions in multimodal data. Moreover, they also negle…

2024

A Light-Weight State Detection Model for Kalman-Filter-Based Acoustic Feedback Cancellation with Rapid Recovery from Abrupt Path Changes

ICASSP 2024accepted

The partitioned block frequency domain Kalman Filter (PBFDKF) has been applied in acoustic feedback cancellation (AFC) due to its fast convergence and low steady-state misalignment. However, in cases where the feedback path experiences abrupt changes, the Kalman filter, once it reaches a steady stat…

Cited by 0SourceScholar
2024

A Lightweight Hybrid Multi-Channel Speech Extraction System with Directional Voice Activity Detection

ICASSP 2024accepted

Although deep learning (DL) based end-to-end models have shown outstanding performance in multi-channel speech extraction, their practical applications on edge devices are restricted due to their high computational complexity. In this paper, we propose a hybrid system that can more effectively integ…

Cited by 0SourceScholar
2024

FDIG: A Fine-Grained Data Integration Approach for Group Recommendation

ICASSP 2024accepted

Effective group recommendation systems play a pivotal role in enriching the information consumption of users from different groups. Existing group recommendation approaches face challenges such as the sparsity of the rating matrix and low specificity between user clusters, leading to cold-start issu…

Cited by 0SourceScholar
2024

GTCRN: A Speech Enhancement Model Requiring Ultralow Computational Resources

ICASSP 2024accepted

While modern deep learning-based models have significantly outperformed traditional methods in the area of speech enhancement, they often necessitate a lot of parameters and extensive computational power, making them impractical to be deployed on edge devices in real-world applications. In this pape…

Cited by 0SourceScholar
2023

A Learnable Spatial Mapping for Decoding the Directional Focus of Auditory Attention Using EEG

ICASSP 2023accepted

One of the important tasks of auditory attention decoding is to identify the attended speaker’s direction from the listener’s EEG signals. Compared to rule-based methods, deep neural networks (DNNs) have recently shown significantly better identification accuracy, especially with short decision wind…

Cited by 0SourceScholar
2023

A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids

ICASSP 2023accepted

This paper summarizes a hybrid multi-channel speech enhancement system for the ICASSP Signal Processing Grand Challenge: Clarity Challenge (Speech Enhancement for Hearing Aids) 2023. The system consists of a rule-based dereverberation module, a multi-channel enhancement module, and a post-processing…

Cited by 0SourceScholar
2023

Few-Shot Class-Incremental Learning via Class-Aware Bilateral Distillation

CVPR 2023poster

Few-Shot Class-Incremental Learning (FSCIL) aims to continually learn novel classes based on only few training samples, which poses a more challenging task than the well-studied Class-Incremental Learning (CIL) due to data scarcity. While knowledge distillation, a prevailing technique in CIL, can al…

2023

HyperMatch: Noise-Tolerant Semi-Supervised Learning via Relaxed Contrastive Constraint

CVPR 2023poster

Recent developments of the application of Contrastive Learning in Semi-Supervised Learning (SSL) have demonstrated significant advancements, as a result of its exceptional ability to learn class-aware cluster representations and the full exploitation of massive unlabeled data. However, mismatched in…

Cited by 10SourcePDFScholar
2023

Learning List-Level Domain-Invariant Representations for Ranking

NeurIPS 2023spotlight

Domain adaptation aims to transfer the knowledge learned on (data-rich) source domains to (low-resource) target domains, and a popular method is invariant representation learning, which matches and aligns the data distributions on the feature space. Although this method is studied extensively and ap…

Cited by 9SourcePDFScholar
2023

Nested Attention Network with Graph Filtering for Visual Question and Answering

ICASSP 2023accepted

Recently, Visual Question Answering(VQA), which is required to generate the answer by understanding both visual and textual content, has attracted considerable research interest. Most existing works extract visual features with the CNN network and learn its feature embedding with an attention mechan…

Cited by 0SourceScholar
2023

Personalized Speech Enhancement Combining Band-Split RNN and Speaker Attentive Module

ICASSP 2023accepted

Target speaker information can be utilized in speech enhancement (SE) models to more effectively extract the desired speech. Previous works introduce the speaker embedding into speech enhancement models by means of concatenation or affine transformation. In this paper, we propose a speaker attentive…

Cited by 0SourceScholar
2023

Promptagator: Few-shot Dense Retrieval From 8 Examples

ICLR 2023poster

Much recent research on information retrieval has focused on how to transfer from one task (typically with abundant supervised data) to various other retrieval tasks where supervision is limited, with the implicit assumption that it is possible to generalize from one task to all the rest. However, t…

Cited by 230SourcePDFScholar
2022

A Priori SNR Estimation for Speech Enhancement Based on PESQ-Induced Reinforcement Learning

ICASSP 2022accepted

Perceptual evaluation of speech quality (PESQ) is widely accepted as an effective objective metric closely related to the speech quality sensed by human listening perception. Due to its evaluation complexity and non-differentiability, PESQ is difficult to include in the cost function for deep learni…

Cited by 0SourceScholar
2022

Distilling Object Detectors with Global Knowledge

ECCV 2022poster

"Knowledge distillation learns a lightweight student model that mimics a cumbersome teacher. Existing methods regard the knowledge as the feature of each instance or their relations, which is the instance-level knowledge only from the teacher model, i.e., the local knowledge. However, the empirical…

2022

ED2LM: Encoder-Decoder to Language Model for Faster Document Re-ranking Inference

ACL 2022findings

State-of-the-art neural models typically encode document-query pairs using cross-attention for re-ranking. To this end, models generally utilize an encoder-only (like BERT) paradigm or an encoder-decoder (like T5) approach. These paradigms, however, are not without flaws, i.e., running the model on…

Cited by 15SourcePDFScholar
2022

Large Dual Encoders Are Generalizable Retrievers

EMNLP 2022main

It has been shown that dual encoders trained on one domain often fail to generalize to other domains for retrieval tasks. One widespread belief is that the bottleneck layer of a dual encoder, where the final score is simply a dot-product between a query vector and a passage vector, is too limited co…

2021

Multi-stage Training with Improved Negative Contrast for Neural Passage Retrieval

EMNLP 2021main

In the context of neural passage retrieval, we study three promising techniques: synthetic data generation, negative sampling, and fusion. We systematically investigate how these techniques contribute to the performance of the retrieval system and how they complement each other. We propose a multi-s…

2020

Active Control of Line Spectral Noise with Simultaneous Secondary Path Modeling Without Auxiliary Noise

ICASSP 2020accepted

Online secondary path modeling is appealing for most active noise control systems due to its benefit of effective tracking of the varying acoustic environment and possible variation of the control sources and sensors. However, the usually utilized additive noise method inevitably leads to the increa…

Cited by 0SourceScholar
2020

Kalman Filtering Attention for User Behavior Modeling in CTR Prediction

NeurIPS 2020spotlight

Click-through rate (CTR) prediction is one of the fundamental tasks for e-commerce search engines. As search becomes more personalized, it is necessary to capture the user interest from rich behavior data. Existing user behavior modeling algorithms develop different attention mechanisms to emphasize…

2019

Sampling Wisely: Deep Image Embedding by Top-K Precision Optimization

ICCV 2019poster

Deep image embedding aims at learning a convolutional neural network (CNN) based mapping function that maps an image to a feature vector. The embedding quality is usually evaluated by the performance in image search tasks. Since very few users bother to open the second page search results, top-k pre…

Cited by 33PDFcodeScholar