← Search

Zexi Jia

13 accepted papers

2026

Evaluating Generative Models via One-Dimensional Code Distributions

CVPR 2026

Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the spac

Cited by 0SourcecodeScholar
2026

Exploring Specular Reflection Inconsistency for Generalizable Face Forgery Detection

ICLR 2026poster

Detecting deepfakes has become increasingly challenging as forgery faces synthesized by AI-generated methods, particularly diffusion models, achieve unprecedented quality and resolution. Existing forgery detection approaches relying on spatial and frequency features demonstrate limited efficacy agai…

Cited by 0SourceScholar
2026

F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model

AAAI 2026technical

Traditional dialogue retrieval aims to select the most appropriate utterance or image from recent dialogue history. However, they often fail to meet users’ actual needs for revisiting semantically coherent content scattered across long-form conversations. To fill this gap, we define the Fine-grained

Cited by 0SourcePDFScholar
2026

GROW: Watermark Generation with Progressive Guidance for Diffusion Models

CVPR 2026

Digital watermarking is a cornerstone for copyright protection. With the rapid advancement of generative models like diffusion models, in-generation and training-free watermarking techniques have garnered more attention for their endogeneity and convenience. These methods typically embed a watermark

Cited by 0SourceScholar
2026

Manifold-Optimal Guidance: A Unified Riemannian Control View of Diffusion Guidance

ICML 2026spotlight

Classifier-Free Guidance (CFG) serves as the de facto control mechanism for conditional diffusion, yet high guidance scales notoriously induce oversaturation, texture artifacts, and structural collapse. We attribute this failure to a geometric mismatch: standard CFG performs Euclidean extrapolation …

Cited by 0SourceScholar
2026

Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity

CVPR 2026

Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authentic to real-world photography remains challenging. This is partly due to biases in existing evaluation paradigms: human ratings and preference-trained metrics often

Cited by 0SourcecodeScholar
2025

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

ICCV 2025poster

Vision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method utilizes…

Cited by 0SourcePDFScholar
2025

From Imitation to Innovation: The Emergence of AI's Unique Artistic Styles and the Challenge of Copyright Protection

ICCV 2025poster

Current legal frameworks consider AI-generated works eligible for copyright protection when they meet originality requirements and involve substantial human intellectual input. However, systematic legal standards and reliable evaluation methods for AI art copyrights are lacking. Through comprehensiv…

Cited by 0SourcePDFScholar
2025

MCID: Multi-aspect Copyright Infringement Detection for Generated Images

ICCV 2025poster

With the rapid advancement of generative models, we can now create highly realistic images. This represents a significant technical breakthrough but also introduces new challenges for copyright protection. Previous methods for detecting copyright infringement in AI-generated images mainly depend on…

Cited by 0SourcePDFScholar
2025

Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution Analysis

CVPR 2025poster

The advancement of Generative Adversarial Networks (GANs) and diffusion models significantly enhances the realism of synthetic images, driving progress in image processing and creative design. However, this progress also necessitates the development of effective detection methods, as synthetic image…

Cited by 0SourcePDFScholar
2025

Semantic to Structure: Learning Structural Representations for Infringement Detection

ICASSP 2025accepted

Structural information in images is crucial for aesthetic assessment, and it is widely recognized in the artistic field that imitating the structure of other works significantly infringes on creators’ rights. The advancement of diffusion models has led to AI-generated content imitating artists’ stru…

Cited by 0SourceScholar
2025

WalkVLM: Aid Visually Impaired People Walking by Vision Language Model

ICCV 2025poster

Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people.With the recent progress of vision-language models (VLMs), applying VLMs to offer walking guidance has bec…

Cited by 0SourcePDFScholar