← Search

Shengeng Tang

17 accepted papers

2026

Accelerating Controllable Generation via Hybrid-grained Cache

AAAI 2026technical

Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid

Cited by 0SourcePDFScholar
2026

LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition

AAAI 2026technical

Visual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguit

Cited by 0SourcePDFScholar
2026

OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL

ICML 2026poster

Existing forgery detection methods are often limited to uni-modal or bi-modal settings, failing to handle the interleaved text, images, and videos prevalent in real-world misinformation. To bridge this gap, we propose **OmniVL-Guard**, a unified framework for omni vision-language forgery detection a…

Cited by 0SourceScholar
2026

Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning

AAAI 2026technical

Open-world 3D scene understanding is fundamentally challenging for vision and robotics, due to the constraints of closed-vocabulary supervision and static annotations. To address this, we propose a unified framework for Open-World 3D Scene Graph Generation with Retrieval-Augmented Reasoning, which e

Cited by 0SourcePDFScholar
2026

Wi-CBR: Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior Recognition

AAAI 2026technical

The challenge in WiFi-based cross-domain Behavior Recognition lies in the significant interference of domain-specific signals on gesture variation. However, previous methods alleviate this interference by mapping the phase from multiple domains into a common feature space. If the Doppler Frequency S

Cited by 0SourcePDFScholar
2025

Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

AAAI 2025technical

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for longer, untrimmed videos. This task seeks to identify and temporal…

2025

Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observations

CVPR 2025poster

Generating continuous sign language videos from discrete segments is challenging due to the need for smooth transitions that preserve natural flow and meaning. Traditional approaches that simply concatenate isolated signs often result in abrupt transitions, disrupting video coherence. To address thi…

Cited by 0SourcePDFScholar
2025

EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with Events

CVPR 2025highlight

Continuous space-time video super-resolution (C-STVSR) endeavors to upscale videos simultaneously at arbitrary spatial and temporal scales, which has recently garnered increasing interest. However, prevailing methods struggle to yield satisfactory videos at out-of-distribution spatial and temporal s…

2025

Knowledge Swapping via Learning and Unlearning

ICML 2025poster

We introduce Knowledge Swapping, a novel task designed to selectively regulate knowledge of a pretrained model by enabling the forgetting of user-specified information, retaining essential knowledge, and acquiring new knowledge simultaneously. By delving into the analysis of knock-on feature hierar…

2025

Linguistics-Vision Monotonic Consistent Network for Sign Language Production

ICASSP 2025accepted

Sign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to the cross-modal semantic gap and the lack of word-action correspondence labels for strong supervision alignment, the SLP…

Cited by 0SourceScholar
2025

Mixture of Multimodal Adapters for Sentiment Analysis

NAACL 2025long

Pre-trained language model (PLM) have achieved great success in text sentiment analysis. However, in practical applications, sentiment is not only conveyed through language but also hidden in other modalities. Therefore, multimodal sentiment analysis (MSA) has attracted increasing research interest.…

2025

Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning

ICML 2025oral

Class-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability remains a significant challenge, particularly when the task ID is unknown. To address this, our study reveals that the…

2025

Patch-level Sounding Object Tracking for Audio-Visual Question Answering

AAAI 2025technical

Answering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding objects related to the question along the timeline. In this paper, we present a new Patch-level Sounding Object Tracking (PS…

Cited by 6SourcePDFScholar
2025

PhysDiff: Physiology-based Dynamicity Disentangled Diffusion Model for Remote Physiological Measurement

AAAI 2025technical

Recent works on remote PhotoPlethysmoGraphy (rPPG) estimation typically use techniques like CNNs and Transformers to encode implicit features from facial videos for prediction. These methods learn to directly map facial videos to the static values of rPPG signals, overlooking the inherent dynamic ch…

2025

Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing

IJCAI 2025

Text-driven video editing powered by generative diffusion models holds significant promise for applications spanning film production, advertising, and beyond. However, the limited expressiveness of pre-trained word embeddings often restricts nuanced edits, especially when targeting novel concepts wi

2025

Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production

AAAI 2025technical

Sign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit the…

2025

Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition

ICASSP 2025accepted

Speech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain analysis, overlooking the valuable envelope structures of the frequency domain that are equally important for robust emo…

Cited by 0SourceScholar