← Search

Cheng Yan

17 accepted papers

2026

Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer

CVPR 2026

Motion transfer has emerged as a promising direction for controllable video generation, yet existing methods largely focus on single-object scenarios and struggle when multiple objects require distinct motion patterns. In this work, we present FlexiMMT, the first implicit image-to-video (I2V) motion

Cited by 0SourcecodeScholar
2025

Commonsense Subgraph for Inductive Relation Reasoning with Meta-learning

COLING 2025main

In knowledge graphs (KGs), predicting missing relations is a critical reasoning task. Recent subgraph-based models have delved into inductive settings, which aim to predict relations between newly added entities. While these models have demonstrated the ability for inductive reasoning, they only con…

Cited by 0SourcePDFScholar
2025

Inductive Reasoning on Few-Shot Knowledge Graphs with Task-Aware Language Models

EMNLP 2025

Knowledge graphs are dynamic structures that continuously evolve as new entities emerge, often accompanied by only a handful of associated triples. Current knowledge graph reasoning methods struggle in these few-shot scenarios due to their reliance on extensive structural information.To address this

Cited by 0SourcePDFScholar
2025

Information Bottleneck-guided MLPs for Robust Spatial-temporal Forecasting

ICML 2025poster

Spatial-temporal forecasting (STF) plays a pivotal role in urban planning and computing. Spatial-Temporal Graph Neural Networks (STGNNs) excel at modeling spatial-temporal dynamics, thus being robust against noise perturbations. However, they often suffer from relatively poor computational efficienc…

2025

LCDB 1.1: A Database Illustrating Learning Curves Are More Ill-Behaved Than Previously Thought

NeurIPS 2025poster

Sample-wise learning curves plot performance versus training set size. They are useful for studying scaling laws and speeding up hyperparameter tuning and model selection. Learning curves are often assumed to be well-behaved: monotone (i.e. improving with more data) and convex. By constructing the L…

Cited by 3SourcecodeScholar
2025

Learning Visual Proxy for Compositional Zero-Shot Learning

ICCV 2025poster

Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions by leveraging knowledge from seen compositions. Existing methods typically align textual prototypes with visual features using Vision-Language Models (VLMs), but they face two key limitations: (1) modality…

2025

Priority on High-Quality: Selecting Instruction Data via Consistency Verification of Noise Injection

EMNLP 2025

Large Language Models (LLMs) have demonstrated a remarkable understanding of language nuances through instruction tuning, enabling them to effectively tackle various natural language processing tasks. Recent research has focused on the quality of instruction data rather than the quantity of instruct

2025

ROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving Object

CVPR 2025poster

The availability of large-scale remote sensing video data underscores the importance of high-quality interactive segmentation. However, challenges such as small object sizes, ambiguous features, and limited generalization make it difficult for current methods to achieve this goal. In this work, we p…

2024

Correcting Language Model Bias for Text Classification in True Zero-Shot Learning

COLING 2024main

Combining pre-trained language models (PLMs) and manual templates is a common practice for text classification in zero-shot scenarios. However, the effect of this approach is highly volatile, ranging from random guesses to near state-of-the-art results, depending on the quality of the manual templat…

Cited by 1SourcePDFScholar
2023

Feature Prediction Diffusion Model for Video Anomaly Detection

ICCV 2023poster

Anomaly detection in the video is an important research area and a challenging task in real applications. Due to the unavailability of large-scale annotated anomaly events, most existing video anomaly detection (VAD) methods focus on learning the distribution of normal samples to detect the substant…

Cited by 58PDFScholar
2023

Structure-aware Knowledge Graph-to-text Generation with Planning Selection and Similarity Distinction

EMNLP 2023long main

The knowledge graph-to-text (KG-to-text) generation task aims to synthesize coherent and engaging sentences that accurately convey the complex information derived from an input knowledge graph. One of the primary challenges in this task is bridging the gap between the diverse structures of the KG an…

Cited by 0SourceScholar
2021

BV-Person: A Large-Scale Dataset for Bird-View Person Re-Identification

ICCV 2021poster

Person Re-IDentification (ReID) aims at re-identifying persons from non-overlapping cameras. Existing person ReID studies focus on horizontal-view ReID tasks, in which the person images are captured by the cameras from a (nearly) horizontal view. In this work we introduce a new ReID task, bird-view…

Cited by 24PDFScholar
2021

Biomedical Concept Normalization by Leveraging Hypernyms

EMNLP 2021main

Biomedical Concept Normalization (BCN) is widely used in biomedical text processing as a fundamental module. Owing to numerous surface variants of biomedical concepts, BCN still remains challenging and unsolved. In this paper, we exploit biomedical concept hypernyms to facilitate BCN. We propose Bio…

2021

Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection

NeurIPS 2021poster

Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when on…

2021

Occluded Person Re-Identification With Single-Scale Global Representations

ICCV 2021poster

Occluded person re-identification (ReID) aims at re-identifying occluded pedestrians from occluded or holistic images taken across multiple cameras. Current state-of-the-art (SOTA) occluded ReID models rely on some auxiliary modules, including pose estimation, feature pyramid and graph matching modu…

Cited by 64PDFScholar
2020

Self-Trained Deep Ordinal Regression for End-to-End Video Anomaly Detection

CVPR 2020poster

Video anomaly detection is of critical practical importance to a variety of real applications because it allows human attention to be focused on events that are likely to be of interest, in spite of an otherwise overwhelming volume of video. We show that applying self-trained deep ordinal regression…

Cited by 318PDFScholar