← Search

Junjie Li

33 accepted papers

2026

ADDRESSING GRADIENT MISALIGNMENT IN DATA-AUGMENTED TRAINING FOR ROBUST SPEECH DEEPFAKE DETECTION

ICASSP 2026oral

In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during training, the backpropagated gradients from original and augmented inputs may misalign, which can result in conflicting pa…

Cited by 0SourcePDFScholar
2026

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

ICASSP 2026poster

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language model (LLM)-based designs, rely on scaling fixed emotion embeddings or external…

Cited by 0SourcePDFScholar
2026

FluxNet: Learning Capacity-Constrained Local Transport Operators for Conservative and Bounded PDE Surrogates

ICML 2026poster

Autoregressive learning of time-stepping operators offers an effective approach to data-driven PDE simulation on grids. For conservation laws, however, long-horizon rollouts are often destabilized when learned updates violate global conservation and, in many applications, additional state bounds—suc…

Cited by 0SourceScholar
2026

MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR

ICASSP 2026poster

LLM-based ASR overcomes multilingual data scarcity by projecting speech representations into the LLM space to leverage its robust semantic and reasoning capabilities. However, while previous approaches typically enhance performance by scaling data or model parameters, a single projector often strugg…

Cited by 0SourcePDFScholar
2026

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

ICLR 2026poster

Recent advancements in long chain-of-thought (CoT) reasoning, particularly through the Group Relative Policy Optimization algorithm used by DeepSeek-R1, have led to significant interest in the potential of Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs). While…

Cited by 0SourceScholar
2026

Training-free Boosting for Few-shot Segmentation via Generalizing Semantic Mining

AAAI 2026technical

Few-shot Semantic Segmentation (FSS) aims to segment the novel target objects with the guidance of minimal annotated reference examples. The affinity-based method has great advantages in the FSS inference stage for both specialist model and foundation model. However, current affinity calculation me

Cited by 0SourcePDFScholar
2026

UVLM: Benchmarking Video Language Model for Underwater World Understanding

AAAI 2026technical

Recently, video-language models (VidLMs) have gained widespread attention and adoption. However, existing works primarily focus on terrestrial scenarios, overlooking the highly demanding application needs of underwater observation. To overcome this gap, we introduce UVLM, an under water observation

Cited by 0SourcePDFScholar
2025

Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-Based Planner and Graph-Based Policy

ICRA 2025

Multi-agent systems (MAS) have shown great potential in executing complex tasks, but coordination and safety remain significant challenges. Multi-Agent Reinforcement Learning (MARL) offers a promising framework for agent collaboration, but it faces difficulties in handling complex tasks and designin

Cited by 12SourceScholar
2025

Gait-X: Exploring X modality for Generalized Gait Recognition

ICCV 2025poster

Modality exploration has been repeatedly mentioned in gait recognition, evolving from silhouette to parsing, mesh, point clouds, etc. These latest modalities agree that silhouette is less affected by background and clothing noises, but argue it loses too much valuable discriminative information. The…

Cited by 0SourcePDFScholar
2025

MarS: a Financial Market Simulation Engine Powered by Generative Foundation Model

ICLR 2025poster

Generative models aim to simulate realistic effects of various actions across different contexts, from text generation to visual effects. Despite significant efforts to build real-world simulators, the application of generative models to virtual worlds, like financial markets, remains under-explored…

2025

Multi-Level Speaker Representation for Target Speaker Extraction

ICASSP 2025accepted

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In th…

Cited by 0SourceScholar
2025

RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models

ACL 2025finding

Object Navigation (ObjectNav) is a fundamental task in embodied artificial intelligence. Although significant progress has been made in semantic map construction and target direction prediction in current research, redundant exploration and exploration failures remain inevitable. A critical but unde…

Cited by 0SourcePDFScholar
2025

WuKong: Design, Modeling and Control of a Compact Flexible Hybrid Aerial-Aquatic Vehicle

RA-L 2025

The significant differences in the physical properties of air and water pose a substantial challenge for the development of hybrid aerial-aquatic vehicle (HAAV), which leading to increased prototype size, heavier thrusters, and reduced efficiency or under-actuation in one of the mediums. This letter

Cited by 9SourceScholar
2024

A Piecewise-weighted RANSAC Method Utilizing Abandoned Hypothesis Model Information with a New Application on Robot Self-calibration

IROS 2024poster

Industrial robots and collaborative robots are widely employed in industry and are progressively being utilized to assist individuals in their daily routines. To improve their absolute accuracy, self-calibration methods using portable local measurement devices are cost-effective solutions. However,…

Cited by 0SourceScholar
2024

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker Speech

ICASSP 2024accepted

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the interfering speech. However, this scenario only accounts for a small…

Cited by 0SourceScholar
2024

Probabilistic Contrastive Learning for Domain Adaptation

IJCAI 2024poster

Contrastive learning has shown impressive success in enhancing feature discriminability for various visual tasks in a self-supervised manner, but the standard contrastive paradigm (features+l2 normalization) has limited benefits when applied in domain adaptation. We find that this is mainly because…

2024

SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

ICASSP 2024accepted

Zero-shot voice conversion (VC) aims to transfer the source speaker timbre to arbitrary unseen target speaker timbre, while keeping the linguistic content unchanged. Although the voice of generated speech can be controlled by providing the speaker embedding of the target speaker, the speaker similar…

Cited by 0SourceScholar
2024

Semantic-guided Robustness Tuning for Few-Shot Transfer Across Extreme Domain Shift

ECCV 2024poster

"In this work, we focus on the cross-domain few-shot classification (CDFSC), which is mostly challenged by the low-data problem as well as extreme domain shift between base and novel target classes. Current methods always employ a lightweight backbone and continue to use a linear-probe-like traditio…

Cited by 0SourcePDFScholar
2023

Boundary-Enhanced Co-Training for Weakly Supervised Semantic Segmentation

CVPR 2023poster

The existing weakly supervised semantic segmentation (WSSS) methods pay much attention to generating accurate and complete class activation maps (CAMs) as pseudo-labels, while ignoring the importance of training the segmentation networks. In this work, we observe that there is an inconsistency betwe…

2023

Detecting Out-of-Distribution Examples Via Class-Conditional Impressions Reappearing

ICASSP 2023accepted

Out-of-distribution (OOD) detection aims at enhancing standard deep neural networks to distinguish anomalous inputs from original training data. Previous progress has introduced various approaches where the in-distribution training data and even several OOD examples are prerequisites. However, due t…

Cited by 0SourceScholar
2023

Exploit Domain-Robust Optical Flow in Domain Adaptive Video Semantic Segmentation

AAAI 2023technical

Domain adaptive semantic segmentation aims to exploit the pixel-level annotated samples on source domain to assist the segmentation of unlabeled samples on target domain. For such a task, the key is to construct reliable supervision signals on target domain. However, existing methods can only provid…

2023

GAIA: Delving into Gradient-based Attribution Abnormality for Out-of-distribution Detection

NeurIPS 2023poster

Detecting out-of-distribution (OOD) examples is crucial to guarantee the reliability and safety of deep neural networks in real-world settings. In this paper, we offer an innovative perspective on quantifying the disparities between in-distribution (ID) and OOD data---analyzing the uncertainty that…

2023

Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based Approach

ICCV 2023poster

Weakly-supervised temporal action localization aims to localize action instances in videos with only video-level action labels. Existing methods mainly embrace a localization-by-classification pipeline that optimizes the snippet-level prediction with a video classification loss. However, this formul…

Cited by 19PDFcodeScholar
2023

Stream Attention Based U-Net for L3DAS23 Challenge

ICASSP 2023accepted

Machine learning applications of 3D audio are gaining increasing interest in recent years. In this paper, we propose a stream attention based U-Net to remove background noise and reverberation based on ICASSP Signal Processing Grand Challenge 2023: L3DAS23 Challenge<sup xmlns:mml="http://www.w3.org/…

Cited by 0SourceScholar
2023

Towards Effective Instance Discrimination Contrastive Loss for Unsupervised Domain Adaptation

ICCV 2023poster

Domain adaptation (DA) aims to transfer knowledge from a label-rich source domain to a related but label-scarce target domain. Recently, increasing research has focused on exploring data structure of the target domain. In light of the recent success of Instance Discrimination Contrastive (IDCo) loss…

Cited by 16PDFcodeScholar
2022

Generative Cross-Domain Data Augmentation for Aspect and Opinion Co-Extraction

NAACL 2022long

As a fundamental task in opinion mining, aspect and opinion co-extraction aims to identify the aspect terms and opinion terms in reviews. However, due to the lack of fine-grained annotated resources, it is hard to train a robust model for many domains. To alleviate this issue, unsupervised domain ad…

2022

Targeted Multimodal Sentiment Classification based on Coarse-to-Fine Grained Image-Target Matching

IJCAI 2022poster

Targeted Multimodal Sentiment Classification (TMSC) aims to identify the sentiment polarities over each target mentioned in a pair of sentence and image. Existing methods to TMSC failed to explicitly capture both coarse-grained and fine-grained image-target matching, including 1) the relevance betwe…

2021

Improving Neural Text Normalization with Partial Parameter Generator and Pointer-Generator Network

ICASSP 2021accepted

Text Normalization (TN) is an essential part in conversational systems like text-to-speech synthesis (TTS) and automatic speech recognition (ASR). It is a process of transforming non-standard words (NSW) into a representation of how the words are to be spoken. Existing approaches to TN are mainly ru…

Cited by 0SourceScholar
2021

Learning Intact Features by Erasing-Inpainting for Few-shot Classification

AAAI 2021technical

Few-shot classification aims to categorize the samples from unseen classes with only few labeled samples. To address such a challenge, many methods exploit a base set consisting of massive labeled samples to learn an instance embedding function, i.e., image feature extractor, and it is expected to p…

Cited by 68SourcePDFScholar
2021

PHMOSpell: Phonological and Morphological Knowledge Guided Chinese Spelling Check

ACL 2021long

Chinese Spelling Check (CSC) is a challenging task due to the complex characteristics of Chinese characters. Statistics reveal that most Chinese spelling errors belong to phonological or visual errors. However, previous methods rarely utilize phonological and morphological knowledge of Chinese chara…