← Search

Jinchao Zhang

35 accepted papers

2026

Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection

CVPR 2026

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to classify and localize action segments in untrimmed videos for unseen categories. Previous methods rely solely on global alignment between label-level semantics and visual features, which is insufficient to transfer temporal consistent visual

Cited by 0SourceScholar
2026

Evaluating Generative Models via One-Dimensional Code Distributions

CVPR 2026

Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the spac

Cited by 0SourcecodeScholar
2026

Exploring Specular Reflection Inconsistency for Generalizable Face Forgery Detection

ICLR 2026poster

Detecting deepfakes has become increasingly challenging as forgery faces synthesized by AI-generated methods, particularly diffusion models, achieve unprecedented quality and resolution. Existing forgery detection approaches relying on spatial and frequency features demonstrate limited efficacy agai…

Cited by 0SourceScholar
2026

F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model

AAAI 2026technical

Traditional dialogue retrieval aims to select the most appropriate utterance or image from recent dialogue history. However, they often fail to meet users’ actual needs for revisiting semantically coherent content scattered across long-form conversations. To fill this gap, we define the Fine-grained

Cited by 0SourcePDFScholar
2026

GROW: Watermark Generation with Progressive Guidance for Diffusion Models

CVPR 2026

Digital watermarking is a cornerstone for copyright protection. With the rapid advancement of generative models like diffusion models, in-generation and training-free watermarking techniques have garnered more attention for their endogeneity and convenience. These methods typically embed a watermark

Cited by 0SourceScholar
2026

Manifold-Optimal Guidance: A Unified Riemannian Control View of Diffusion Guidance

ICML 2026spotlight

Classifier-Free Guidance (CFG) serves as the de facto control mechanism for conditional diffusion, yet high guidance scales notoriously induce oversaturation, texture artifacts, and structural collapse. We attribute this failure to a geometric mismatch: standard CFG performs Euclidean extrapolation …

Cited by 0SourceScholar
2026

Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

CVPR 2026

Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-training, existing OVOD methods achieve strong detection performance on general scenarios (e.g., OV-COCO) but suffer severe performance drops when transferred

Cited by 0SourceScholar
2026

Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity

CVPR 2026

Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authentic to real-world photography remains challenging. This is partly due to biases in existing evaluation paradigms: human ratings and preference-trained metrics often

Cited by 0SourcecodeScholar
2025

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

ICCV 2025poster

Vision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method utilizes…

Cited by 0SourcePDFScholar
2025

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

NeurIPS 2025poster

Multimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utterances. Despite its significance, little research has investigated the capability of multimodal large language models (…

Cited by 0SourcecodeScholar
2025

DisComp: A Two-Stage Prompt Optimization Framework Combining Task-Agnostic and Task-Aware Compression

NAACL 2025findings

Large language models (LLMs) exhibit exceptional performance across a wide range of natural language processing tasks, often relying on lengthy prompts to harness their full capabilities. However, extended prompts can lead to substantial computational overhead and increased hardware demands, limitin…

Cited by 0SourcePDFScholar
2025

Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment

ICCV 2025poster

While fine-tuning diffusion models with reinforcement learning (RL) has demonstrated effectiveness in directly optimizing downstream objectives, existing RL frameworks are prone to overfitting the rewards, leading to outputs that deviate from the true data distribution and exhibit reduced diversity.…

Cited by 0SourcePDFScholar
2025

From Imitation to Innovation: The Emergence of AI's Unique Artistic Styles and the Challenge of Copyright Protection

ICCV 2025poster

Current legal frameworks consider AI-generated works eligible for copyright protection when they meet originality requirements and involve substantial human intellectual input. However, systematic legal standards and reliable evaluation methods for AI art copyrights are lacking. Through comprehensiv…

Cited by 0SourcePDFScholar
2025

ILDiff: Generate Transparent Animated Stickers by Implicit Layout Distillation

ICASSP 2025accepted

High-quality animated stickers usually contain transparent channels, which are often ignored by current video generation models. To generate fine-grained animated transparency channels, existing methods can be roughly divided into video matting algorithms and diffusion-based algorithms. The methods…

Cited by 0SourceScholar
2025

MCID: Multi-aspect Copyright Infringement Detection for Generated Images

ICCV 2025poster

With the rapid advancement of generative models, we can now create highly realistic images. This represents a significant technical breakthrough but also introduces new challenges for copyright protection. Previous methods for detecting copyright infringement in AI-generated images mainly depend on…

Cited by 0SourcePDFScholar
2025

MedDiT: A Knowledge-Controlled Diffusion Transformer Framework for Dynamic Medical Image Generation in Virtual Simulated Patient

IJCAI 2025

Medical education relies heavily on Simulated Patients (SPs) to provide a safe environment for students to practice clinical skills, including medical image analysis. However, the high cost of recruiting qualified SPs and the lack of diverse medical imaging datasets have presented significant challe

Cited by 0SourcePDFScholar
2025

Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution Analysis

CVPR 2025poster

The advancement of Generative Adversarial Networks (GANs) and diffusion models significantly enhances the realism of synthetic images, driving progress in image processing and creative design. However, this progress also necessitates the development of effective detection methods, as synthetic image…

Cited by 0SourcePDFScholar
2025

Semantic to Structure: Learning Structural Representations for Infringement Detection

ICASSP 2025accepted

Structural information in images is crucial for aesthetic assessment, and it is widely recognized in the artistic field that imitating the structure of other works significantly infringes on creators’ rights. The advancement of diffusion models has led to AI-generated content imitating artists’ stru…

Cited by 0SourceScholar
2025

WalkVLM: Aid Visually Impaired People Walking by Vision Language Model

ICCV 2025poster

Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people.With the recent progress of vision-language models (VLMs), applying VLMs to offer walking guidance has bec…

Cited by 0SourcePDFScholar
2024

Enhancing Note-Level Singing Transcription Model with Unlabeled and Weakly Labeled Data

ICASSP 2024accepted

Note-level automatic singing transcription, involving the extraction of onset, offset, and pitch information from a singing voice, is a crucial process in the field of Music Information Retrieval (MIR), The recent advancements in deep learning models have led to significant progress in this field. H…

Cited by 0SourceScholar
2023

Humming2Music: Being A Composer As Long As You Can Humming

IJCAI 2023poster

Creating a piece of music is difficult for people who have never been trained to compose. We present an automatic music generation system to lower the threshold of creating music. The system takes the user's humming as input and creates full music based on the humming melody. The system consists of…

2023

LingGe: An Automatic Ancient Chinese Poem-to-Song Generation System

IJCAI 2023poster

This paper presents a novel system, named LingGe ("伶歌" in Chinese), to generate songs for ancient Chinese poems automatically. LingGe takes the poem as the lyric, composes music conditioned on the lyric, and finally outputs a full song including the singing and the accompaniment. It consists of four…

Cited by 1SourcePDFScholar
2023

Rephrasing the Reference for Non-autoregressive Machine Translation

AAAI 2023technical

Non-autoregressive neural machine translation (NAT) models suffer from the multi-modality problem that there may exist multiple possible translations of a source sentence, so the reference sentence may be inappropriate for the training when the NAT output is closer to other translations. In response…

2022

AutoCAD: Automatically Generate Counterfactuals for Mitigating Shortcut Learning

EMNLP 2022finding

Recent studies have shown the impressive efficacy of counterfactually augmented data (CAD) for reducing NLU models’ reliance on spurious features and improving their generalizability. However, current methods still heavily rely on human efforts or task-specific designs to generate counterfactuals, t…

2022

Counterfactual Data Augmentation via Perspective Transition for Open-Domain Dialogues

EMNLP 2022main

The construction of open-domain dialogue systems requires high-quality dialogue datasets. The dialogue data admits a wide variety of responses for a given dialogue history, especially responses with different semantics. However, collecting high-quality such a dataset in most scenarios is labor-inten…

2022

High-Resolution and Arbitrary-Sized Chinese Landscape Painting Creation Based on Generative Adversarial Networks

IJCAI 2022poster

This paper outlines an automated creation system for Chinese landscape paintings based on generative adversarial networks. The system consists of three cascaded modules: generation, resizing, and super-resolution. The generation module first generates a square-shaped painting, then the resizing modu…

Cited by 4SourcePDFScholar
2022

Selecting Stickers in Open-Domain Dialogue through Multitask Learning

ACL 2022findings

With the increasing popularity of online chatting, stickers are becoming important in our online communication. Selecting appropriate stickers in open-domain dialogue requires a comprehensive understanding of both dialogues and stickers, as well as the relationship between the two types of modalitie…

2021

An Iterative Multi-Knowledge Transfer Network for Aspect-Based Sentiment Analysis

EMNLP 2021finding

Aspect-based sentiment analysis (ABSA) mainly involves three subtasks: aspect term extraction, opinion term extraction, and aspect-level sentiment classification, which are typically handled in a separate or joint manner. However, previous approaches do not well exploit the interactive relations amo…

2021

Constructing Emotional Consensus and Utilizing Unpaired Data for Empathetic Dialogue Generation

EMNLP 2021finding

Researches on dialogue empathy aim to endow an agent with the capacity of accurate understanding and proper responding for emotions. Existing models for empathetic dialogue generation focus on the emotion flow in one direction, that is, from the context to response. We argue that conducting an empat…

2021

Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue Utterances

ACL 2021long

Nowadays, open-domain dialogue models can generate acceptable responses according to the historical context based on the large-scale pre-trained language models. However, they generally concatenate the dialogue history directly as the model input to predict the response, which we named as the flat p…

2021

Different Strokes for Different Folks: Investigating Appropriate Further Pre-training Approaches for Diverse Dialogue Tasks

EMNLP 2021main

Loading models pre-trained on the large-scale corpus in the general domain and fine-tuning them on specific downstream tasks is gradually becoming a paradigm in Natural Language Processing. Previous investigations prove that introducing a further pre-training phase between pre-training and fine-tuni…

2021

GTM: A Generative Triple-wise Model for Conversational Question Generation

ACL 2021long

Generating some appealing questions in open-domain conversations is an effective way to improve human-machine interactions and lead the topic to a broader or deeper direction. To avoid dull or deviated questions, some researchers tried to utilize answer, the “future” information, to guide question g…

Cited by 16SourcePDFScholar
2021

Text AutoAugment: Learning Compositional Augmentation Policy for Text Classification

EMNLP 2021main

Data augmentation aims to enrich training samples for alleviating the overfitting issue in low-resource or class-imbalanced situations. Traditional methods first devise task-specific operations such as Synonym Substitute, then preset the corresponding parameters such as the substitution rate artific…

2020

Contrastive Zero-Shot Learning for Cross-Domain Slot Filling with Adversarial Attack

COLING 2020main

Zero-shot slot filling has widely arisen to cope with data scarcity in target domains. However, previous approaches often ignore constraints between slot value representation and related slot description representation in the latent space and lack enough model robustness. In this paper, we propose a…

Cited by 44SourcePDFScholar
2020

One Comment from One Perspective: An Effective Strategy for Enhancing Automatic Music Comment

COLING 2020main

The automatic generation of music comments is of great significance for increasing the popularity of music and the music platform’s activity. In human music comments, there exists high distinction and diverse perspectives for the same song. In other words, for a song, different comments stem from di…