← Search

Siwei Lyu

53 accepted papers

2026

AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models

ICLR 2026poster

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in complex problem-solving through Chain-of-Thought (CoT) reasoning. However, the multi-step nature of CoT introduces new safety challenges that extend beyond conventional language model alignment. We identify a failure mode in…

Cited by 0SourceScholar
2026

DICE: Distilling Classifier-Free Guidance into Text Embeddings

AAAI 2026technical

Text-to-image diffusion models are capable of generating high-quality images, but suboptimal pre-trained text representations often result in these images failing to align closely with the given text prompts. Classifier-free guidance (CFG) is a popular and effective technique for improving text-imag

Cited by 0SourcePDFScholar
2025

$\mathcal{X}^2$-DFD: A framework for e$\mathcal{X}$plainable and e$\mathcal{X}$tendable Deepfake Detection

NeurIPS 2025poster

This paper proposes **$\mathcal{X}^2$-DFD**, an **e$\mathcal{X}$plainable** and **e$\mathcal{X}$tendable** framework based on multimodal large-language models (MLLMs) for deepfake detection, consisting of three key stages. The first stage, *Model Feature Assessment*, systematically evaluates the de…

Cited by 0SourcecodeScholar
2025

HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models

CVPR 2025poster

We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (e.g. text, objects, partia…

Cited by 1SourcePDFScholar
2025

Knowledge Distillation with Refined Logits

ICCV 2025poster

Recent research on knowledge distillation has increasingly focused on logit distillation because of its simplicity, effectiveness, and versatility in model compression. In this paper, we introduce Refined Logit Distillation (RLD) to address the limitations of current logit distillation methods. Our…

2025

Your Text Encoder Can Be An Object-Level Watermarking Controller

ICCV 2025poster

Invisible watermarking of AI-generated images can help with copyright protection, enabling detection and identification of AI-generated media. In this work, we present a novel approach to watermark images of T2I Latent Diffusion Models (LDMs). By only fine-tuning text token embeddings \mathcal W _*,…

2024

Enhancing Adversarial Robustness of DNNS Via Weight Decorrelation in Training

ICASSP 2024accepted

Deep Neural Networks (DNNs) are vulnerable to adversarial perturbations, raising significant concerns about their security. Numerous methods have been proposed to enhance DNN robustness. However, many methods, including adversarial training and noise injection, improve robustness by incorporating ex…

Cited by 0SourceScholar
2024

Exposing Text-Image Inconsistency Using Diffusion Models

ICLR 2024poster

In the battle against widespread online misinformation, a growing problem is text-image inconsistency, where images are misleadingly paired with texts with different intent or meaning. Existing classification-based methods for text-image inconsistency can identify contextual inconsistencies but fail…

2024

On the Trajectory Regularity of ODE-based Diffusion Sampling

ICML 2024poster

Diffusion-based generative models use stochastic differential equations (SDEs) and their equivalent ordinary differential equations (ODEs) to establish a smooth connection between a complex data distribution and a tractable prior distribution. In this paper, we identify several intriguing trajectory…

2024

ParallelEdits: Efficient Multi-Aspect Text-Driven Image Editing with Attention Grouping

NeurIPS 2024poster

Text-driven image synthesis has made significant advancements with the development of diffusion models, transforming how visual content is generated from text prompts. Despite these advances, text-driven image editing, a key area in computer graphics, faces unique challenges. A major challenge is ma…

Cited by 2SourcePDFScholar
2024

Simple and Fast Distillation of Diffusion Models

NeurIPS 2024poster

Diffusion-based generative models have demonstrated their powerful performance across various tasks, but this comes at a cost of the slow sampling speed. To achieve both efficient and high-quality synthesis, various distillation-based accelerated sampling methods have been developed recently. Howeve…

2024

Transcending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake Detection

CVPR 2024poster

Deepfake detection faces a critical generalization hurdle with performance deteriorating when there is a mismatch between the distributions of training and testing data. A broadly received explanation is the tendency of these detectors to be overfitted to forgery-specific artifacts rather than learn…

Cited by 67SourcePDFScholar
2023

Controlling Neural Style Transfer with Deep Reinforcement Learning

IJCAI 2023poster

Controlling the degree of stylization in the Neural Style Transfer (NST) is a little tricky since it usually needs hand-engineering on hyper-parameters. In this paper, we propose the first deep Reinforcement Learning (RL) based architecture that splits one-step style transfer into a step-wise proces…

Cited by 1SourcePDFScholar
2023

DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection

NeurIPS 2023poster

A critical yet frequently overlooked challenge in the field of deepfake detection is the lack of a standardized, unified, comprehensive benchmark. This issue leads to unfair performance comparisons and potentially misleading results. Specifically, there is a lack of uniformity in data processing pip…

2023

Detection of Real-Time Deepfakes in Video Conferencing with Active Probing and Corneal Reflection

ICASSP 2023accepted

The COVID pandemic has led to the wide adoption of online video calls in recent years. However, the increasing reliance on video calls provides opportunities for new impersonation attacks by fraudsters using the advanced real-time DeepFakes. Real-time DeepFakes pose new challenges to detection metho…

Cited by 0SourceScholar
2023

RMBench: Benchmarking Deep Reinforcement Learning for Robotic Manipulator Control

IROS 2023poster

Reinforcement learning is used to tackle complex tasks with high-dimensional sensory inputs. Over the past decade, a wide range of reinforcement learning algorithms have been developed, with recent progress benefiting from deep learning for raw sensory signal representation. This raises a natural qu…

Cited by 4SourcecodeScholar
2023

Tracking Multiple Deformable Objects in Egocentric Videos

CVPR 2023poster

Most existing multiple object tracking (MOT) methods that solely rely on appearance features struggle in tracking highly deformable objects. Other MOT methods that use motion clues to associate identities across frames have difficulty handling egocentric videos effectively or efficiently. In this wo…

Cited by 14SourcePDFScholar
2022

Adaptive Face Forgery Detection in Cross Domain

ECCV 2022poster

"It is necessary to develop effective face forgery detection methods with constantly evolving technologies in synthesizing realistic faces which raises serious risks on malicious face tampering. A large and growing body of literature has investigated deep learning-based approaches, especially those…

2022

Differentially private SGDA for minimax problems

UAI 2022poster

Stochastic gradient descent ascent (SGDA) and its variants have been the workhorse for solving minimax problems. However, in contrast to the well-studied stochastic gradient descent (SGD) with differential privacy (DP) constraints, there is little work on understanding the generalization (utility…

2022

Eyes Tell All: Irregular Pupil Shapes Reveal GAN-Generated Faces

ICASSP 2022accepted

Generative adversarial network (GAN) generated high-realistic human faces are visually challenging to discern from real ones. They have been used as profile images for fake social media accounts, which leads to high negative social impacts. In this work, we show that GAN-generated faces can be expos…

Cited by 0SourceScholar
2022

Stochastic Planner-Actor-Critic for Unsupervised Deformable Image Registration

AAAI 2022technical

Large deformations of organs, caused by diverse shapes and nonlinear shape changes, pose a significant challenge for medical image registration. Traditional registration methods need to iteratively optimize an objective function via a specific deformation model along with meticulous parameter tuning…

2022

Text-Image De-Contextualization Detection Using Vision-Language Models

ICASSP 2022accepted

Text-image de-contextualization, which uses inconsistent image-text pairs, is an emerging form of misinformation and drawing increasing attention due to the great threat to information authenticity. With real content but semantic mismatch in multiple modalities, the detection of de-contextualization…

Cited by 0SourceScholar
2022

Towards To-a-T Spatio-Temporal Focus for Skeleton-Based Action Recognition

AAAI 2022technical

Graph Convolutional Networks (GCNs) have been widely used to model the high-order dynamic dependencies for skeleton-based action recognition. Most existing approaches do not explicitly embed the high-order spatio-temporal importance to joints’ spatial connection topology and intensity, and they do n…

Cited by 63SourcePDFScholar
2022

Vocbench: A Neural Vocoder Benchmark for Speech Synthesis

ICASSP 2022accepted

Neural vocoders, used for converting the spectral representations of an audio signal to the waveforms, are a commonly used component in speech synthesis pipelines. It focuses on synthesizing waveforms from low-dimensional representation, such as Mel-Spectrograms. In recent years, different approache…

Cited by 0SourceScholar
2021

Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark

CVPR 2021poster

To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33,600 HD frames in various scenarios. Notably, we annotate 20,800 p…

Cited by 133PDFcodeScholar
2021

Invisible Backdoor Attack With Sample-Specific Triggers

ICCV 2021poster

Recently, backdoor attacks pose a new security threat to the training process of deep neural networks (DNNs). Attackers intend to inject hidden backdoors into DNNs, such that the attacked model performs well on benign samples, whereas its prediction will be maliciously changed if hidden backdoors ar…

Cited by 603PDFcodeScholar
2021

Stability and Differential Privacy of Stochastic Gradient Descent for Pairwise Learning with Non-Smooth Loss

AISTATS 2021poster

Pairwise learning has recently received increasing attention since it subsumes many important machine learning tasks (e.g. AUC maximization and metric learning) into a unifying framework. In this paper, we give the first-ever-known stability and generalization analysis of stochastic gradient descent…

Cited by 23SourcePDFScholar
2021

Stochastic Actor-Executor-Critic for Image-to-Image Translation

IJCAI 2021poster

Training a model-free deep reinforcement learning model to solve image-to-image translation is difficult since it involves high-dimensional continuous state and action spaces. In this paper, we draw inspiration from the recent success of the maximum entropy reinforcement learning framework designed…

2020

Cascade Graph Neural Networks for RGB-D Salient Object Detection

ECCV 2020poster

In this paper, we study the problem of salient object detection for RGB-D images by using both color and depth information. A major technical challenge for detecting salient objects in RGB-D images is to fully leverage the two complementary data sources. The existing works either simply distill prio…

2020

Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics

CVPR 2020poster

AI-synthesized face-swapping videos, commonly known as DeepFakes, is an emerging problem threatening the trustworthiness of online information. The need to develop and evaluate DeepFake detection algorithms calls for datasets of DeepFake videos. However, current DeepFake datasets suffer from low vis…

Cited by 1723PDFcodeScholar
2020

Explainable and Efficient Sequential Correlation Network for 3D Single Person Concurrent Activity Detection

IROS 2020poster

We present the sequential correlation network (SCN) to improve concurrent activity detection. SCN combines a recurrent neural network and a correlation model hierarchically to model the complex correlations and temporal dynamics of concurrent activities. SCN has several advantages that enable effect…

Cited by 2SourceScholar
2020

Learning Semantic Neural Tree for Human Parsing

ECCV 2020poster

In this paper, we design a novel semantic neural tree for human parsing, which uses a tree architecture to encode physiological structure of human body, and design a coarse to fine process in a cascade manner to generate accurate results. Specifically, the semantic neural tree is designed to segment…

Cited by 71SourcePDFScholar
2019

Object-Driven Text-To-Image Synthesis via Adversarial Training

CVPR 2019poster

In this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow attention-driven, multi-stage refinement for synthesizing complex images from text descriptions. With a novel object-driven attentive generative network, the Obj-GAN can synthesize salient objects…

Cited by 386PDFScholar
2018

Multi-Scale Structure-Aware Network for Human Pose Estimation

ECCV 2018poster

We develop a robust multi-scale structure-aware neural network for human pose estimation. This method improves the recent deep conv-deconv hourglass models with four key improvements: (1) multi-scale supervision to strengthen contextual feature learning in matching body keypoints by combining featur…

Cited by 382SourcePDFScholar
2018

Tagging Like Humans: Diverse and Distinct Image Annotation

CVPR 2018poster

In this work we propose a new automatic image annotation model, dubbed diverse and distinct image annotation (D2IA). The generative model D2IA is inspired by the ensemble of human annotations, which create semantically relevant, yet distinct and diverse tags. In D2IA, we generate a relevant and dist…

2017

Adaptive RNN Tree for Large-Scale Human Action Recognition

ICCV 2017poster

In this work, we present the RNN Tree (RNN-T), an adaptive learning framework for skeleton based human action recognition. Our method categorizes action classes and uses multiple Recurrent Neural Networks (RNNs) in a tree-like hierarchy. The RNNs in RNN-T are co-trained with the action category hier…

Cited by 140PDFScholar
2016

Fast Convergence of Online Pairwise Learning Algorithms

AISTATS 2016poster

Pairwise learning usually refers to a learning task which involves a loss function depending on pairs of examples, among which most notable ones are bipartite ranking, metric learning and AUC maximization. In this paper, we focus on online learning algorithms for pairwise learning problems without…

Cited by 20SourcePDFScholar
2015

A comparative study of contact models for contact-aware state estimation

IROS 2015poster

We study the contact-aware state estimation (CASE) problem, i.e., the problem of estimating the state of an object while it is being actively manipulated by a robot. Several researchers have developed particle filters for this problem. They estimate the state (pose and velocity) of manipulated objec…

Cited by 11SourceScholar
2015

UniHIST: A Unified Framework for Image Restoration With Marginal Histogram Constraints

CVPR 2015poster

Marginal histograms provide valuable information for various computer vision problems. However, current image restoration methods do not fully exploit the potential of marginal histograms, in particular, their role as ensemble constraints on the marginal statistics of the restored image. In this pap…

Cited by 12SourcePDFScholar