← Search

Jun-Cheng Chen

23 accepted papers

2026

GOT-Edit: Geometry-Aware Generic Object Tracking via Online Model Editing

ICLR 2026poster

Human perception for effective object tracking in a 2D video stream arises from the implicit use of prior 3D knowledge combined with semantic reasoning. In contrast, most generic object tracking (GOT) methods primarily rely on 2D features of the target and its surroundings while neglecting 3D geomet…

Cited by 0SourcecodeScholar
2026

Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective

ICML 2026poster

Recent studies have shown that semantic watermarks, which embed information into the initial noise of latent diffusion models (LDMs), are vulnerable to black-box forgery attacks. However, existing methods primarily rely on empirical evidence and lack a rigorous theoretical understanding of the condi…

Cited by 0SourceScholar
2025

FIPER: Factorized Features for Robust Image Super-Resolution and Compression

NeurIPS 2025poster

In this work, we propose using a unified representation, termed **Factorized Features**, for low-level vision tasks, where we test on **Single Image Super-Resolution (SISR)** and **Image Compression**. Motivated by the shared principles between these tasks, they require recovering and preserving fin…

Cited by 0SourceScholar
2025

Generation and Comprehension Hand-in-Hand: Vision-guided Expression Diffusion for Boosting Referring Expression Generation and Comprehension

ICLR 2025poster

Referring expression generation (REG) and comprehension (REC) are vital and complementary in joint visual and textual reasoning. Existing REC datasets typically contain insufficient image-expression pairs for training, hindering the generalization of REC models to unseen referring expressions. More…

Cited by 0SourcePDFScholar
2025

Pixel Is Not a Barrier: An Effective Evasion Attack for Pixel-Domain Diffusion Models

AAAI 2025technical

Diffusion Models have emerged as powerful generative models for high-quality image synthesis, with many subsequent image editing techniques based on them. However, the ease of text-based image editing introduces significant risks, such as malicious editing for scams or intellectual property infringe…

2025

Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation Model

CVPR 2025poster

The current deep generative models have enabled the creation of synthetic facial images with remarkable photorealism, raising significant societal concerns over their potential misuse. Despite rapid advancements in the field of deepfake detection, developing an efficient and effective approach for t…

2024

ACCEPT: Adaptive Codebook for Composite and Efficient Prompt Tuning

EMNLP 2024finding

Prompt Tuning has been a popular Parameter-Efficient Fine-Tuning method attributed to its remarkable performance with few updated parameters on various large-scale pretrained Language Models (PLMs). Traditionally, each prompt has been considered indivisible and updated independently, leading the par…

2024

Blenda: Domain Adaptive Object Detection Through Diffusion-Based Blending

ICASSP 2024accepted

Unsupervised domain adaptation (UDA) aims to transfer a model learned using labeled data from the source domain to unlabeled data in the target domain. To address the large domain gap issue between the source and target domains, we propose a novel regularization method for domain adaptive object det…

Cited by 0SourceScholar
2024

MeDM: Mediating Image Diffusion Models for Video-to-Video Translation with Temporal Correspondence Guidance

AAAI 2024technical

This study introduces an efficient and effective method, MeDM, that utilizes pre-trained image Diffusion Models for video-to-video translation with consistent temporal flow. The proposed framework can render videos from scene position information, such as a normal G-buffer, or perform text-guided ed…

2023

Hearing and Seeing Abnormality: Self-Supervised Audio-Visual Mutual Learning for Deepfake Detection

ICASSP 2023accepted

The recent development of deepfakes has resulted in serious threats to society, such as spreading misinformation, defamation, etc. Although recent deepfake detection methods are capable of achieving satisfactory results for seen forgeries, the performance drops significantly for unseen ones. With pr…

Cited by 0SourceScholar
2022

CLIPCAM: A Simple Baseline For Zero-Shot Text-Guided Object And Action Localization

ICASSP 2022accepted

The key for the contemporary deep learning-based object and action localization algorithms to work is the large-scale annotated data. However, in real-world scenarios, since there are infinite amounts of unlabeled data beyond the categories of publicly available datasets, it is not only time- and ma…

Cited by 0SourceScholar
2022

Continual Learning for Visual Search With Backward Consistent Feature Embedding

CVPR 2022poster

In visual search, the gallery set could be incrementally growing and added to the database in practice. However, existing methods rely on the model trained on the entire dataset, ignoring the continual updating of the model. Besides, as the model updates, the new model must re-extract features for t…

Cited by 28PDFcodeScholar
2021

Naturalistic Physical Adversarial Patch for Object Detectors

ICCV 2021poster

Most prior works on physical adversarial attacks mainly focus on the attack performance but seldom enforce any restrictions over the appearance of the generated adversarial patches. This leads to conspicuous and attention-grabbing patterns for the generated patches which can be easily identified by…

Cited by 192PDFcodeScholar
2020

Face Feature Recovery via Temporal Fusion for Person Search

ICASSP 2020accepted

Searching actors from videos by a single portrait image is a challenging task, due to large variations of video scenes and intra-person appearance. To tackle this problem, most recent works apply deep neural networks for detecting and extracting robust facial features for matching. However, when the…

Cited by 0SourceScholar
2020

The Devil is in the Details: Self-Supervised Attention for Vehicle Re-Identification

ECCV 2020poster

In recent years, the research community has approached the problem of vehicle re-identification (re-id) with attention-based models, specifically focusing on regions of a vehicle containing discriminative information. These re-id methods rely on expensive key-point labels, part annotations, and addi…

2019

A Dual-Path Model With Adaptive Attention for Vehicle Re-Identification

ICCV 2019oral

In recent years, attention models have been extensively used for person and vehicle re-identification. Most re-identification methods are designed to focus attention on key-point locations. However, depending on the orientation, the contribution of each key-point varies. In this paper, we present a…

Cited by 291PDFcodeScholar
2019

Uncertainty Modeling of Contextual-Connections Between Tracklets for Unconstrained Video-Based Face Recognition

ICCV 2019poster

Unconstrained video-based face recognition is a challenging problem due to significant within-video variations caused by pose, occlusion and blur. To tackle this problem, an effective idea is to propagate the identity from high-quality faces to low-quality ones through contextual connections, which…

Cited by 15PDFScholar