← Search

Liming Wang

20 accepted papers

2025

Can Diffusion Models Disentangle? A Theoretical Perspective

NeurIPS 2025poster

This paper presents a novel theoretical framework for understanding how diffusion models can learn disentangled representations with commonly used weak supervision such as partial labels and multiple views. Within this framework, we establish identifiability conditions for diffusion models to disent…

Cited by 0SourceScholar
2024

Mutual Information Based Noise Scale Optimization for Gradient Leakage Resistant Federated Learning

ICASSP 2024accepted

Federated learning decentralizes the learning process, yet it does not provide adequate privacy protection. Current countermeasures predominantly rely on Local Differential Privacy (LDP) techniques. While larger noise injection offers stronger privacy, it also leads to a degradation in model perform…

Cited by 0SourceScholar
2024

Rethinking Word-level Adversarial Attack: The Trade-off between Efficiency, Effectiveness, and Imperceptibility

COLING 2024main

Neural language models have demonstrated impressive performance in various tasks but remain vulnerable to word-level adversarial attacks. Word-level adversarial attacks can be formulated as a combinatorial optimization problem, and thus, an attack method can be decomposed into search space and searc…

Cited by 3SourcePDFScholar
2024

Unsupervised Speech Recognition with N-skipgram and Positional Unigram Matching

ICASSP 2024accepted

Training unsupervised speech recognition systems presents challenges due to GAN-associated instability, misalignment between speech and text, and significant memory demands. To tackle these challenges, we introduce a novel ASR system, ESPUM. This system harnesses the power of lower-order N-skipgrams…

Cited by 0SourceScholar
2023

Contrastive Learning with Adversarial Examples for Alleviating Pathology of Language Model

ACL 2023long

Neural language models have achieved superior performance. However, these models also suffer from the pathology of overconfidence in the out-of-distribution examples, potentially making the model difficult to interpret and making the interpretation methods fail to provide faithful attributions. In t…

Cited by 4SourcePDFScholar
2023

Listen, Decipher and Sign: Toward Unsupervised Speech-to-Sign Language Recognition

ACL 2023findings

Existing supervised sign language recognition systems rely on an abundance of well-annotated data. Instead, an unsupervised speech-to-sign language recognition (SSR-U) system learns to translate between spoken and sign languages by observing only non-parallel speech and sign-language corpora. We pro…

2023

Prototypes-oriented Transductive Few-shot Learning with Conditional Transport

ICCV 2023poster

Transductive Few-Shot Learning (TFSL) has recently attracted increasing attention since it typically outperforms its inductive peer by leveraging statistics of query samples.However, previous TFSL methods usually encode uniform prior that all the classes within query samples are equally likely, whic…

Cited by 23PDFcodeScholar
2023

Similarizing the Influence of Words with Contrastive Learning to Defend Word-level Adversarial Text Attack

ACL 2023findings

Neural language models are vulnerable to word-level adversarial text attacks, which generate adversarial examples by directly substituting discrete input words. Previous search methods for word-level attacks assume that the information in the important words is more influential on prediction than un…

Cited by 7SourcePDFScholar
2022

360-Attack: Distortion-Aware Perturbations From Perspective-Views

CVPR 2022poster

The application of deep neural networks (DNNs) on 360-degree images has achieved remarkable progress in the recent years. However, DNNs have been demonstrated to be vulnerable to well-crafted adversarial examples, which may trigger severe safety problems in the real-world applications based on 360-d…

Cited by 6PDFScholar
2022

Mitigating the Inconsistency Between Word Saliency and Model Confidence with Pathological Contrastive Training

ACL 2022findings

Neural networks are widely used in various NLP tasks for their remarkable performance. However, the complexity makes them difficult to interpret, i.e., they are not guaranteed right for the right reason. Besides the complexity, we reveal that the model pathology - the inconsistency between word sali…

Cited by 6SourcePDFScholar
2022

PARSE: An Efficient Search Method for Black-box Adversarial Text Attacks

COLING 2022main

Neural networks are vulnerable to adversarial examples. The adversary can successfully attack a model even without knowing model architecture and parameters, i.e., under a black-box scenario. Previous works on word-level attacks widely use word importance ranking (WIR) methods and complex search met…

Cited by 9SourcePDFScholar
2022

SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional Images

ICASSP 2022accepted

The safety of Deep Neural Networks (DNNs) processing omnidirectional images (ODIs) is an under-researched topic. In this paper, we propose a novel sparse attack, named Single-Perspective (SP) Attack, towards fooling these models by perturbing only one perspective image (PI) rendered from the target…

Cited by 0SourceScholar
2022

Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech Recognition

ACL 2022long

Phonemes are defined by their relationship to words: changing a phoneme changes the word. Learning a phoneme inventory with little supervision has been a longstanding challenge with important applications to under-resourced speech technology. In this paper, we bridge the gap between the linguistic a…

Cited by 4SourcePDFScholar
2021

Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval

ICASSP 2021accepted

Multimodal word discovery (MWD) is often treated as a byproduct of the speech-to-image retrieval problem. However, our theoretical analysis shows that some kind of alignment/attention mechanism is crucial for a MWD system to learn meaningful word-level representation. We verify our theory by conduct…

Cited by 0SourceScholar
2021

Empowering Adaptive Early-Exit Inference with Latency Awareness

AAAI 2021technical

With the capability of trading accuracy for latency on-the-fly, the technique of adaptive early-exit inference has emerged as a promising line of research to accelerate the deep learning inference. However, studies in this line of research commonly use a group of thresholds to control the accuracy-l…

2018

Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop

ICASSP 2018accepted

We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translate…

Cited by 0SourceScholar
2018

Planar Object Tracking in the Wild: A Benchmark

ICRA 2018poster

Planar object tracking is an actively studied problem in vision-based robotic applications. While several benchmarks have been constructed for evaluating state-of-the-art algorithms, there is a lack of video sequences captured in the wild rather than in constrained laboratory environment. In this pa…

Cited by 66SourceScholar
2016

A general framework for reconstruction and classification from compressive measurements with side information

ICASSP 2016accepted

We develop a general framework for compressive linear-projection measurements with side information. Side information is an additional signal correlated with the signal of interest. We investigate the impact of side information on classification and signal recovery from low-dimensional measurements.…

Cited by 0SourceScholar