← Search

Shilin Wang

17 accepted papers

2026

Enhancing the Security of Visual Speaker Authentication Based on Dynamic Lip-Print Analysis

CVPR 2026

In recent years, face-based authentication methods are gradually replacing traditional methods across various applications, offering enhanced security and user convenience. However, these methods are threatened by the continuously evolving DeepFake techniques. In this paper, a novel Visual Speaker A

Cited by 0SourceScholar
2025

Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning

ICASSP 2025accepted

This paper introduces a novel approach to Visual Forced Alignment (VFA), aiming to accurately synchronize utterances with corresponding lip movements, without relying on audio cues. We propose a novel VFA approach that integrates a local context-aware feature extractor and employs multitask learning…

Cited by 0SourceScholar
2025

Personalized Speech Enhancement without User Enrollment for Real-World Audio Replay Scenarios

ICASSP 2025accepted

Many speech enhancement (SE) approaches have been proposed to deal with cocktail party problem. Personalized speech enhancement (PSE) approaches improve SE performance by utilizing user enrollment speech. However, PSE requires users to record additional clean audio for registration, which can be red…

Cited by 0SourceScholar
2025

Towards A Distribution Alignment Framework for Incomplete Data Classification

ICASSP 2025accepted

Missing attribute values frequently affect data classification, reducing accuracy as most models rely on complete datasets. Imputing missing values is typically used to restore data completeness, which is essential for building models. The effectiveness of imputation significantly impacts the classi…

Cited by 0SourceScholar
2024

Revisiting the Information Capacity of Neural Network Watermarks: Upper Bound Estimation and Beyond

AAAI 2024technical

To trace the copyright of deep neural networks, an owner can embed its identity information into its model as a watermark. The capacity of the watermark quantify the maximal volume of information that can be verified from the watermarked model. Current studies on capacity focus on the ownership veri…

Cited by 4SourcePDFScholar
2024

Speaker-Adaptive Lipreading Via Spatio-Temporal Information Learning

ICASSP 2024accepted

Lipreading has been rapidly developed recently with the help of large-scale datasets and large models. Despite the significant progress made, the performance of lipreading models still falls short when dealing with unseen speakers. Therefore, it is necessary to utilize the speaker’s videos for fine-…

Cited by 0SourceScholar
2023

Content-Insensitive Dynamic Lip Feature Extraction for Visual Speaker Authentication Against Deepfake Attacks

ICASSP 2023accepted

Recent research has shown that lip-based speaker authentication system can achieve good authentication performance. However, with emerging deepfake technology, attackers can make high fidelity talking videos of a user, thus posing a great threat to these systems. Confronted with this threat, we prop…

Cited by 0SourceScholar
2022

Few-shot Table-to-text Generation with Prefix-Controlled Generator

COLING 2022main

Neural table-to-text generation approaches are data-hungry, limiting their adaption for low-resource real-world applications. Previous works mostly resort to Pre-trained Language Models (PLMs) to generate fluent summaries of a table. However, they often contain hallucinated contents due to the uncon…

Cited by 12SourcePDFScholar
2022

PPT: Backdoor Attacks on Pre-trained Models via Poisoned Prompt Tuning

IJCAI 2022poster

Recently, prompt tuning has shown remarkable performance as a new learning paradigm, which freezes pre-trained language models (PLMs) and only tunes some soft prompts. A fixed PLM only needs to be loaded with different prompts to adapt different downstream tasks. However, the prompts associated with…

Cited by 55SourcePDFScholar
2022

Stgat-Mad : Spatial-Temporal Graph Attention Network For Multivariate Time Series Anomaly Detection

ICASSP 2022accepted

Anomaly detection in multivariate time series data is challenging due to complex temporal and feature correlations. This paper proposes a novel unsupervised multi-scale stacked spatial-temporal graph attention network for multivariate time series anomaly detection (STGAT-MAD). The core of our framew…

Cited by 0SourceScholar
2016

Detecting double MPEG compression with the same quantiser scale based on MBM feature

ICASSP 2016accepted

Detecting double MPEG compression is of prime significance in video forensics. However, existing methods are effective only when the primary compression and the secondary compression have different quantiser scales (QS). There is a lack of effective methods dealing with double MPEG compression with…

Cited by 0SourceScholar