← Search

Yiding Yang

17 accepted papers

2026

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

ICLR 2026poster

We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference s…

Cited by 0SourcecodeScholar
2026

TGT: Text-Grounded Trajectories for Locally Controlled Video Generation

CVPR 2026

Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However,

Cited by 0SourceScholar
2026

VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization

CVPR 2026

Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired data of simple editing operations, which fundamentally limits t

Cited by 0SourcecodeScholar
2022

Learning Graph Neural Networks for Image Style Transfer

ECCV 2022poster

"State-of-the-art parametric and non-parametric style transfer approaches are prone to either distorted local style patterns due to global statistics alignment, or unpleasing artifacts resulting from patch mismatching. In this paper, we study a novel semi-parametric neural style transfer framework t…

Cited by 75SourcePDFScholar
2021

Amalgamating Knowledge From Heterogeneous Graph Neural Networks

CVPR 2021poster

In this paper, we study a novel knowledge transfer task in the domain of graph neural networks (GNNs). We strive to train a multi-talented student GNN, without accessing human annotations, that "amalgamates" knowledge from a couple of teacher GNNs with heterogeneous architectures and handling distin…

Cited by 120PDFcodeScholar
2021

Learning Dynamics via Graph Neural Networks for Human Pose Estimation and Tracking

CVPR 2021poster

Multi-person pose estimation and tracking serve as crucial steps for video understanding. Most state-of-the-art approaches rely on first estimating poses in each frame and only then implementing data association and refinement. Despite the promising results achieved, such a strategy is inevitably pr…

Cited by 98PDFScholar
2021

Meta-Aggregator: Learning To Aggregate for 1-Bit Graph Neural Networks

ICCV 2021poster

In this paper, we study a novel meta aggregation scheme towards binarizing graph neural networks (GNNs). We begin by developing a vanilla 1-bit GNN framework that binarizes both the GNN parameters and the graph features. Despite the lightweight architecture, we observed that this vanilla framework s…

Cited by 52PDFScholar
2021

Scene Essence

CVPR 2021poster

What scene elements, if any, are indispensable for recognizing a scene? We strive to answer this question through the lens of an end-to-end learning scheme. Our goal is to identify a collection of such pivotal elements, which we term as Scene Essence, to be those that would alter scene recognition i…

Cited by 20PDFScholar
2021

Turning Frequency to Resolution: Video Super-Resolution via Event Cameras

CVPR 2021poster

State-of-the-art video super-resolution (VSR) methods focus on exploiting inter- and intra-frame correlations to estimate high-resolution (HR) video frames from low-resolution (LR) ones. In this paper, we study VSR from an exotic perspective, by explicitly looking into the role of temporal frequency…

Cited by 51PDFScholar
2020

Distilling Knowledge From Graph Convolutional Networks

CVPR 2020poster

Existing knowledge distillation methods focus on convolutional neural networks (CNNs), where the input samples like images lie in a grid domain, and have largely overlooked graph convolutional networks (GCN) that handle non-grid data. In this paper, we propose to our best knowledge the first dedicat…

Cited by 314PDFcodeScholar
2020

Learning Propagation Rules for Attribution Map Generation

ECCV 2020poster

Existing gradient-based attribution-map methods rely on hand-crafted propagation rules for the non-linear/activation layers during the backward pass, so as to produce gradients of the input and then the attribution map. Despite the promising results achieved, such methods are sensitive to the non-in…

Cited by 18SourcePDFScholar