← Search

Tianyi Wang

21 accepted papers

2026

Amplifying Discrepancies: Exploiting Macro and Micro Inconsistencies for Image Manipulation Localization

AAAI 2026technical

The rapid development of image manipulation technologies poses significant challenges to multimedia forensics, especially in accurate localization of manipulated regions. Existing methods often fail to fully explore the intrinsic discrepancies between manipulated and authentic regions, resulting in

Cited by 0SourcePDFScholar
2026

Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) is increasingly viewed as a tree pruning mechanism. However, we identify a systemic pathology termed Recursive Space Contraction (RSC), an irreversible collapse driven by the combined dynamics of positive sharpening and negative squeezing, where …

Cited by 0SourceScholar
2026

RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction

ICML 2026poster

Histopathology whole-slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular architecture and functional programs defining pathological states, whereas RNA sequencing (RNA-seq) provides genome-wide transcriptional profiles at subst…

Cited by 0SourceScholar
2026

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

ICLR 2026oral

While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underexplored. We introduce WAVE (\textbf{u}nified \& \textbf{v}ersatile \textbf{a}udio-\textbf{v}isual \textbf{e}mbeddings), t…

Cited by 0SourcecodeScholar
2025

A Progressive Local Variance-guided Strategy for Improving Data Augmentation Reliability

ICASSP 2025accepted

Recently, CutMix-based augmentation has emerged as a promising strategy for providing regularization to deep neural networks. However, the randomness in cropping may result in uninformative or non-representative regions being selected, resulting in a synthesized image without the desired features. T…

Cited by 0SourceScholar
2025

Can't Slow Me Down: Learning Robust and Hardware-Adaptive Object Detectors against Latency Attacks for Edge Devices

CVPR 2025poster

Object detection is a fundamental enabler for many real-time downstream applications such as autonomous driving, augmented reality and supply chain management. However, the algorithmic backbone of neural networks is brittle to imperceptible perturbations in the system inputs, which were generally kn…

2025

Diffusion Augmentation Sub-center Modeling for Unsupervised Anomalous Sound Detection with Partially Attribute-Unavailable Conditions

ICASSP 2025accepted

Current state-of-the-art unsupervised anomalous sound detection (ASD) methods typically rely on manually annotated attribute information as labels, employing auxiliary classification tasks to learn an embedding space for normal sounds, which helps detect anomalies deviating from this space. However,…

Cited by 0SourceScholar
2025

Fair Deepfake Detectors Can Generalize

NeurIPS 2025poster

Deepfake detection models face two critical challenges: generalization to unseen manipulations and demographic fairness among population groups. However, existing approaches often demonstrate that these two objectives are inherently conflicting, revealing a trade-off between them. In this paper, we,…

Cited by 0SourceScholar
2025

Fed-DFA: Federated Distillation for Heterogeneous Model Fusion Through the Adversarial Lens

AAAI 2025technical

Most of the federated learning techniques are limited to homogeneous model fusion. With the rapid growth of smart applications on resource-constrained edge devices, it becomes a barrier to accommodate their heterogeneous computing power and memory in the real world. Federated Distillation is a promi…

Cited by 1SourcePDFScholar
2025

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

NeurIPS 2025poster

Long-form video understanding poses a significant challenge for video large language models (VideoLLMs) due to prohibitively high computational and memory demands. In this paper, We propose $\textbf{FlexSelect}$, a flexible and efficient token selection strategy for processing long videos. FlexSele…

Cited by 0SourcecodeScholar
2025

From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment

ICCV 2025poster

Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference, potentially omitting crucial visual information. To address the challe…

Cited by 0SourcePDFScholar
2025

HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization

CVPR 2025poster

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional alignment, thematic coherence, and cultural relevance. We propo…

2025

LLM Sensitivity Evaluation Framework for Clinical Diagnosis

COLING 2025main

Large language models (LLMs) have demonstrated impressive performance across various domains. However, for clinical diagnosis, higher expectations are required for LLM’s reliability and sensitivity: thinking like physicians and remaining sensitive to key medical information that affects diagnostic r…

2025

Learning Real Facial Concepts for Independent Deepfake Detection

IJCAI 2025

Deepfake detection models often struggle with generalization to unseen datasets, manifesting as misclassifying real instances as fake in target domains. This is primarily due to an overreliance on forgery artifacts and a limited understanding of real faces. To address this challenge, we propose a no

Cited by 0SourcePDFScholar
2025

NullSwap: Proactive Identity Cloaking Against Deepfake Face Swapping

ICCV 2025poster

Suffering from performance bottlenecks in passively detecting high-quality Deepfake images due to the advancement of generative models, proactive perturbations offer a promising approach to disabling Deepfake manipulations by inserting signals into benign images. However, existing proactive perturba…

2025

Pinco: Position-induced Consistent Adapter for Diffusion Transformer in Foreground-conditioned Inpainting

ICCV 2025poster

Foreground-conditioned inpainting aims to seamlessly fill the background region of an image by utilizing the provided foreground subject and a text description. While existing T2I-based image inpainting methods can be applied to this task, they suffer from issues of subject shape expansion, distorti…

Cited by 0SourcePDFScholar
2025

Self-Bootstrapping for Versatile Test-Time Adaptation

ICML 2025poster

In this paper, we seek to develop a versatile test-time adaptation (TTA) objective for a variety of tasks — classification and regression across image-, object-, and pixel-level predictions. We achieve this through a self-bootstrapping scheme that optimizes prediction consistency between the test im…

Cited by 0SourcePDFScholar
2023

Contrastive Pre-training with Adversarial Perturbations for Check-In Sequence Representation Learning

AAAI 2023technical

A core step of mining human mobility data is to learn accurate representations for user-generated check-in sequences. The learned representations should be able to fully describe the spatial-temporal mobility patterns of users and the high-level semantics of traveling. However, existing check-in seq…

2022

Which side are you on? Insider-Outsider classification in conspiracy-theoretic social media

ACL 2022long

Social media is a breeding ground for threat narratives and related conspiracy theories. In these, an outside group threatens the integrity of an inside group, leading to the emergence of sharply defined group identities: Insiders – agents with whom the authors identify and Outsiders – agents who th…

2021

RepSum: Unsupervised Dialogue Summarization based on Replacement Strategy

ACL 2021long

In the field of dialogue summarization, due to the lack of training data, it is often difficult for supervised summary generation methods to learn vital information from dialogue context with limited data. Several attempts on unsupervised summarization for text by leveraging semantic information sol…

Cited by 15SourcePDFScholar