← Search

Ran Yi

52 accepted papers

2026

Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detection

CVPR 2026

We investigate state-of-the-art deepfake detectors that leverage ViT-based vision foundation models and discover that the [CLS] token suffers from the Pre-trained Information Bias (PIB), i.e., it tends to mainly focus on global semantics due to the knowledge dominated by pre-trained model parameters

Cited by 0SourceScholar
2026

Breaking Manifold Continuity: Vector Quantized Modeling for Real-Centric Deepfake Detection

ICML 2026poster

The increasingly realistic and diverse generative data has led some deepfake detection methods to shift towards learning robust real content, \textit{e.g.}, via reconstruction-based tasks. However, most existing approaches rely primarily on prevalent continuous modeling (\textit{e.g.}, GMMs, VAEs, D…

Cited by 0SourceScholar
2026

Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy

CVPR 2026

The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental challenges of the joint diffusion process: (1) Correspondence Drift,

Cited by 0SourcecodeScholar
2026

Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion Generation

CVPR 2026

Text-to-motion generation is a fundamental task in computer vision, aiming to synthesize 3D human motion sequences from natural language descriptions. However, due to the limited scale and diversity of existing datasets, models trained to directly map raw text to motion often struggle to generalize

Cited by 0SourceScholar
2026

PoseAnything: General Pose-guided Video Generation with Part-aware Temporal Coherence

CVPR 2026

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current pose-guided video generation methods are limited to accepting only hu

Cited by 0SourceScholar
2026

Uncovering and Mitigating Transient Blindness in Multimodal Model Editing

AAAI 2026technical

Multimodal Model Editing (MMED) aims to correct erroneous knowledge in multimodal models. Existing evaluation methods, adapted from textual model editing, overstate success by relying on low-similarity or random inputs, obscure overfitting. We propose a comprehensive locality evaluation framework,

Cited by 0SourcePDFScholar
2025

3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial Representations

CVPR 2025poster

Recent studies have combined 3D Gaussian and 3D Morphable Models (3DMM) to construct high-quality 3D head avatars. In this line of research, existing methods either fail to capture the dynamic textures or incur significant overhead in terms of runtime speed or storage space. To this end, we propose…

Cited by 0SourcePDFScholar
2025

ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting

CVPR 2025poster

Image inpainting aims to fill the missing region of an image.Recently, there has been a surge of interest in foreground-conditioned background inpainting, a sub-task that fills the background of an image while the foreground subject and associated text prompt are provided.Existing background inpaint…

Cited by 0SourcePDFScholar
2025

DiffuseFIST: A Fast Image-guided Style Transfer Method for Adapting Large-scale Diffusion Models

ICASSP 2025accepted

Pre-trained text-to-image (T2I) synthesis diffusion models (DM) have shown remarkable capabilities in generating diverse images. However, they struggle to satisfy the user’s requirements due to (i) text’s inherent imprecision in expressing specific styles and (ii) generation is time-consuming due to…

Cited by 0SourceScholar
2025

ID-Sculpt: ID-aware 3D Head Generation from Single In-the-wild Portrait Image

AAAI 2025technical

While recent works have achieved great success on one-shot 3D common object generation, high quality and fidelity 3D head generation from a single image remains a great challenge. Previous text-based methods for generating 3D heads were limited by text descriptions and image-based methods struggled…

Cited by 0SourcePDFScholar
2025

Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction

CVPR 2025poster

Employing LLMs for visual generation has recently become a research focus. However, the existing methods primarily transfer the LLM architecture to visual generation but rarely investigate the fundamental differences between language and vision. This oversight may lead to suboptimal utilization of v…

2025

MV-Adapter: Multi-View Consistent Image Generation Made Easy

ICCV 2025poster

Existing multi-view image generation methods often make invasive modifications to pre-trained text-to-image (T2I) models and require full fine-tuning, leading to high computational costs and degradation in image quality due to scarce high-quality 3D data. This paper introduces MV-Adapter, an efficie…

Cited by 0SourcePDFScholar
2025

MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning

NeurIPS 2025spotlight

The ability of robots to interpret human instructions and execute manipulation tasks necessitates the availability of task-relevant tabletop scenes for training. However, traditional methods for creating these scenes rely on time-consuming manual layout design or purely randomized layouts, which are…

Cited by 0SourceScholar
2025

Pinco: Position-induced Consistent Adapter for Diffusion Transformer in Foreground-conditioned Inpainting

ICCV 2025poster

Foreground-conditioned inpainting aims to seamlessly fill the background region of an image by utilizing the provided foreground subject and a text description. While existing T2I-based image inpainting methods can be applied to this task, they suffer from issues of subject shape expansion, distorti…

Cited by 0SourcePDFScholar
2025

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

NeurIPS 2025poster

Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyVivid, a multi-subject video customization framework that enables flexible and iden…

Cited by 0SourceScholar
2025

Rectified Diffusion Guidance for Conditional Generation

CVPR 2025poster

Classifier-Free Guidance (CFG), which combines the conditional and unconditional score functions with two coefficients summing to one, serves as a practical technique for diffusion model sampling. Theoretically, however, denoising with CFG cannot be expressed as a reciprocal diffusion process, which…

2025

SIGMAN: Scaling 3D Human Gaussian Generation with Millions of Assets

ICCV 2025poster

3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall in…

Cited by 0SourcePDFScholar
2025

SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation

ICLR 2025poster

The development of diffusion models has led to significant progress in image and video generation tasks, with pre-trained models like the Stable Diffusion series playing a crucial role. However, a key challenge remains in downstream task applications: how to effectively and efficiently adapt pre-tra…

2025

SuperMat: Physically Consistent PBR Material Estimation at Interactive Rates

ICCV 2025poster

Decomposing physically-based materials from images into their constituent properties remains challenging, particularly when maintaining both computational efficiency and physical consistency. While recent diffusion-based approaches have shown promise, they face substantial computational overhead due…

2025

WaveAR: Wavelet-Aware Continuous Autoregressive Diffusion for Accurate Human Motion Prediction

NeurIPS 2025poster

This work tackles a challenging problem: stochastic human motion prediction (SHMP), which aims to forecast diverse and physically plausible future pose sequences based on a short history of observed motion. While autoregressive sequence models have excelled in related generation tasks, their relianc…

Cited by 0SourceScholar
2025

Weighted Poisson-disk Resampling on Large-Scale Point Clouds

AAAI 2025technical

For large-scale point cloud processing, resampling takes the important role of controlling the point number and density while keeping the geometric consistency. However, current methods cannot balance such different requirements. Particularly with large-scale point clouds, classical methods often st…

2024

AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model

AAAI 2024technical

Anomaly inspection plays an important role in industrial manufacture. Existing anomaly inspection methods are limited in their performance due to insufficient anomaly data. Although anomaly generation methods have been proposed to augment the anomaly data, they either suffer from poor generation aut…

2024

Continuous Piecewise-Affine Based Motion Model for Image Animation

AAAI 2024technical

Image animation aims to bring static images to life according to driving videos and create engaging visual content that can be used for various purposes such as animation, entertainment, and education. Recent unsupervised methods utilize affine and thin-plate spline transformations based on keypoint…

2024

MMPI: a Flexible Radiance Field Representation by Multiple Multi-plane Images Blending

ICRA 2024poster

This paper presents a flexible representation of neural radiance fields based on multi-plane images (MPI), for high-quality view synthesis of complex scenes. MPI with Normalized Device Coordinate (NDC) parameterization is widely used in NeRF learning for its simple definition, easy calculation, and…

Cited by 4SourceScholar
2024

Re-thinking Data Availability Attacks Against Deep Neural Networks

CVPR 2024poster

The unauthorized use of personal data for commercial purposes and the covert acquisition of private data for training machine learning models continue to raise concerns. To address these issues researchers have proposed availability attacks that aim to render data unexploitable. However many availab…

Cited by 5SourcePDFScholar
2024

SAMVG: A Multi-Stage Image Vectorization Model with the Segment-Anything Model

ICASSP 2024accepted

Vector graphics are widely used in graphical designs and have received more and more attention. However, unlike raster images which can be easily obtained, acquiring high-quality vector graphics, typically through automatically converting from raster images, remains a significant challenge, especial…

Cited by 0SourceScholar
2024

SDPose: Tokenized Pose Estimation via Circulation-Guide Self-Distillation

CVPR 2024poster

Recently transformer-based methods have achieved state-of-the-art prediction quality on human pose estimation(HPE). Nonetheless most of these top-performing transformer-based models are too computation-consuming and storage-demanding to deploy on edge computing platforms. Those transformer-based mod…

2024

SMaRt: Improving GANs with Score Matching Regularity

ICML 2024poster

Generative adversarial networks (GANs) usually struggle in learning from highly diverse data, whose underlying manifold is complex. In this work, we revisit the mathematical foundations of GANs, and theoretically reveal that the native adversarial loss for GAN training is insufficient to fix the pro…

2024

Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner

CVPR 2024poster

A diffusion model which is formulated to produce an image using thousands of denoising steps usually suffers from a slow inference speed. Existing acceleration algorithms simplify the sampling by skipping most steps yet exhibit considerable performance degradation. By viewing the generation of diffu…

2023

Contrastive Pseudo Learning for Open-World DeepFake Attribution

ICCV 2023poster

The challenge in sourcing attribution for forgery faces has gained widespread attention due to the rapid development of generative techniques. While many recent works have taken essential steps on GAN-generated faces, more threatening attacks related to identity swapping or expression transferring a…

Cited by 23PDFcodeScholar
2023

Instance-Aware Domain Generalization for Face Anti-Spoofing

CVPR 2023poster

Face anti-spoofing (FAS) based on domain generalization (DG) has been recently studied to improve the generalization on unseen scenarios. Previous methods typically rely on domain labels to align the distribution of each domain for learning domain-invariant representations. However, artificial domai…

2023

LiDAR-Camera Panoptic Segmentation via Geometry-Consistent and Semantic-Aware Alignment

ICCV 2023poster

3D panoptic segmentation is a challenging perception task that requires both semantic segmentation and instance segmentation. In this task, we notice that images could provide rich texture, color, and discriminative information, which can complement LiDAR data for evident performance improvement, bu…

Cited by 21PDFcodeScholar
2023

Make-It-3D: High-fidelity 3D Creation from A Single Image with Diffusion Prior

ICCV 2023poster

In this work, we investigate the problem of creating high-fidelity 3D content from only a single image. This is inherently challenging: it essentially involves estimating the underlying 3D geometry while hallucinating unseen textures. To address this challenge, we leverage prior knowledge in a well-…

Cited by 301PDFcodeScholar
2023

Multimodal Industrial Anomaly Detection via Hybrid Fusion

CVPR 2023poster

2D-based Industrial Anomaly Detection has been widely discussed, however, multimodal industrial anomaly detection based on 3D point clouds and RGB images still has many untouched fields. Existing multimodal industrial anomaly detection methods directly concatenate the multimodal features, which lead…

2023

Phasic Content Fusing Diffusion Model with Directional Distribution Consistency for Few-Shot Model Adaption

ICCV 2023poster

Training a generative model with limited number of samples is a challenging task. Current methods primarily rely on few-shot model adaption to train the network. However, in scenarios where data is extremely limited (less than 10), the generative network tends to overfit and suffers from content deg…

Cited by 14PDFcodeScholar
2023

RFENet: Towards Reciprocal Feature Evolution for Glass Segmentation

IJCAI 2023poster

Glass-like objects are widespread in daily life but remain intractable to be segmented for most existing methods. The transparent property makes it difficult to be distinguished from background, while the tiny separation boundary further impedes the acquisition of their exact contour. In this paper,…

2023

Remembering Normality: Memory-guided Knowledge Distillation for Unsupervised Anomaly Detection

ICCV 2023poster

Knowledge distillation (KD) has been widely explored in unsupervised anomaly detection (AD). The student is assumed to constantly produce representations of typical patterns within trained data, named "normality", and the representation discrepancy between the teacher and student model is identified…

Cited by 48PDFScholar
2023

Towards Artistic Image Aesthetics Assessment: A Large-Scale Dataset and a New Method

CVPR 2023poster

Image aesthetics assessment (IAA) is a challenging task due to its highly subjective nature. Most of the current studies rely on large-scale datasets (e.g., AVA and AADB) to learn a general model for all kinds of photography images. However, little light has been shed on measuring the aesthetic qual…

2022

Exploiting Fine-Grained Face Forgery Clues via Progressive Enhancement Learning

AAAI 2022technical

With the rapid development of facial forgery techniques, forgery detection has attracted more and more attention due to security concerns. Existing approaches attempt to use frequency information to mine subtle artifacts under high-quality forged faces. However, the exploitation of frequency informa…

Cited by 155SourcePDFScholar
2022

Generative Domain Adaptation for Face Anti-Spoofing

ECCV 2022poster

"Face anti-spoofing (FAS) approaches based on unsupervised domain adaption (UDA) have drawn growing attention due to promising performances for target scenarios. Most existing UDA FAS methods typically fit the trained models to the target domain via aligning the distribution of semantic high-level f…

Cited by 78SourcePDFScholar
2022

ISDNet: Integrating Shallow and Deep Networks for Efficient Ultra-High Resolution Segmentation

CVPR 2022poster

The huge burden of computation and memory are two obstacles in ultra-high resolution image segmentation. To tackle these issues, most of the previous works follow the global-local refinement pipeline, which pays more attention to the memory consumption but neglects the inference speed. In comparison…

Cited by 60PDFcodeScholar
2022

LAKe-Net: Topology-Aware Point Cloud Completion by Localizing Aligned Keypoints

CVPR 2022poster

Point cloud completion aims at completing geometric and topological shapes from a partial observation. However, some topology of the original shape is missing, existing methods directly predict the location of complete points, without predicting structured and topological information of the complete…

Cited by 84PDFScholar
2022

Optimization over Disentangled Encoding: Unsupervised Cross-Domain Point Cloud Completion via Occlusion Factor Manipulation

ECCV 2022poster

"Recently, studies considering domain gaps in shape completion attracted more attention, due to the undesirable performance of supervised methods on real scans. They only noticed the gap in input scans, but ignored the gap in output prediction, which is specific for completion. In this paper, we dis…

2022

Region-Aware Temporal Inconsistency Learning for DeepFake Video Detection

IJCAI 2022poster

The rapid development of face forgery techniques has drawn growing attention due to security concerns. Existing deepfake video detection methods always attempt to capture the discriminative features by directly exploiting static temporal convolution to mine temporal inconsistency, without explicit…

Cited by 24SourcePDFScholar
2019

APDrawingGAN: Generating Artistic Portrait Drawings From Face Photos With Hierarchical GANs

CVPR 2019oral

Significant progress has been made with image stylization using deep learning, especially with generative adversarial networks (GANs). However, existing methods fail to produce high quality artistic portrait drawings. Such drawings have a highly abstract style, containing a sparse set of continuous…

Cited by 200PDFScholar
2019

Fast Computation of Content-Sensitive Superpixels and Supervoxels Using Q-Distances

ICCV 2019poster

State-of-the-art researches model the data of images and videos as low-dimensional manifolds and generate superpixels/supervoxels in a content-sensitive way, which is achieved by computing geodesic centroidal Voronoi tessellation (GCVT) on manifolds. However, computing exact GCVTs is slow due to com…

Cited by 21PDFScholar