← Search

Yutong Feng

20 accepted papers

2026

ACCFormer: Predicting Analog Circuit Performance Metrics via Topology-Aware Transformers

IJCAI 2026

Reusing and migrating analog circuit intellectual property (IP) across process nodes poses a significant challenge in modern chip design. Efficient and generalizable circuit performance prediction methods for analog circuits are crucial to achieving this goal. Current data-driven approaches typicall

Cited by 0Scholar
2025

Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis

NeurIPS 2025poster

Video generative models can be regarded as world simulators due to their ability to capture dynamic, continuous changes inherent in real-world environments. These models integrate high-dimensional information across visual, temporal, spatial, and causal dimensions, enabling predictions of subjects i…

Cited by 0SourcecodeScholar
2025

Mimir: Improving Video Diffusion Models for Precise Text Understanding

CVPR 2025poster

Text serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features from text encoders yet struggle with limited text comprehension. The recent success of large language models (LLMs) show…

Cited by 4SourcePDFScholar
2025

OmniTry: Virtual Try-On Anything without Masks

NeurIPS 2025poster

Virtual Try-ON (VTON) is a practical and widely-applied task, for which most of existing works focus on clothes. This paper presents OmniTry, a unified framework that extends VTON beyond garment to encompass any wearable objects, e.g., jewelries and accessories, with mask-free setting for more pract…

Cited by 0SourcecodeScholar
2025

ROSE: Remove Objects with Side Effects in Videos

NeurIPS 2025poster

Video object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, \textit{e.g.,} their shadows and reflections, existing works struggle to eliminate these effects for the scarcity of paired video data as…

Cited by 0SourceScholar
2024

Check Locate Rectify: A Training-Free Layout Calibration System for Text-to-Image Generation

CVPR 2024poster

Diffusion models have recently achieved remarkable progress in generating realistic images. However challenges remain in accurately understanding and synthesizing the layout requirements in the textual prompts. To align the generated image with layout instructions we present a training-free layout c…

2024

Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation

CVPR 2024poster

This study focuses on a novel task in text-to-image (T2I) generation namely action customization. The objective of this task is to learn the co-existing action from limited data and generalize it to unseen humans or even animals. Experimental results show that existing subject-driven customization m…

2024

LivePhoto: Real Image Animation with Text-guided Motion Control

ECCV 2024poster

"Despite the recent progress in text-to-video generation, existing studies usually overlook the issue that only spatial contents but not temporal motions in synthesized videos are under the control of text. Towards such a challenge, this work presents a practical system, named , which allows users t…

2024

Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following

CVPR 2024poster

Existing text-to-image (T2I) diffusion models usually struggle in interpreting complex prompts especially those with quantity object-attribute binding and multi-subject descriptions. In this work we introduce a semantic panel as the middleware in decoding texts to images supporting the generator to…

Cited by 48SourcePDFScholar
2024

Spatio-Temporal Field Neural Networks for Air Quality Inference

IJCAI 2024poster

The air quality inference problem aims to utilize historical data from a limited number of observation sites to infer the air quality index at an unknown location. Considering the sparsity of data due to the high maintenance cost of the stations, good inference algorithms can effectively save the co…

Cited by 2SourcePDFScholar
2024

Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning

CVPR 2024poster

Recent compositional zero-shot learning (CZSL) methods adapt pre-trained vision-language models (VLMs) by constructing trainable prompts only for composed state-object pairs. Relying on learning the joint representation of seen compositions these methods ignore the explicit modeling of the state and…

2024

UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training

NeurIPS 2024poster

This work presents a unified knowledge protocol, called UKnow, which facilitates knowledge-based studies from the perspective of data. Particularly focusing on visual and linguistic modalities, we categorize data knowledge into five unit types, namely, in-image, in-text, cross-image, cross-text, and…

Cited by 0SourcePDFScholar
2024

Zero-shot Image Editing with Reference Imitation

NeurIPS 2024poster

Image editing serves as a practical yet challenging task considering the diverse demands from users, where one of the hardest parts is to precisely describe how the edited image should look like. In this work, we present a new form of editing, termed imitative editing, to help users exercise their c…

Cited by 24SourcePDFScholar
2023

ViM: Vision Middleware for Unified Downstream Transferring

ICCV 2023poster

Foundation models are pre-trained on massive data and transferred to downstream tasks via fine-tuning. This work presents Vision Middleware (ViM), a new learning paradigm that targets unified transferring from a single foundation model to a variety of downstream tasks. ViM consists of a zoo of light…

Cited by 1PDFScholar
2022

Grow and Merge: A Unified Framework for Continuous Categories Discovery

NeurIPS 2022accept

Although a number of studies are devoted to novel category discovery, most of them assume a static setting where both labeled and unlabeled data are given at once for finding new categories. In this work, we focus on the application scenarios where unlabeled data are continuously fed into the catego…

Cited by 32SourcePDFScholar
2022

Rethinking Supervised Pre-Training for Better Downstream Transferring

ICLR 2022poster

The pretrain-finetune paradigm has shown outstanding performance on many applications of deep learning, where a model is pre-trained on an upstream large dataset (e.g. ImageNet), and is then fine-tuned to different downstream tasks. Though for most cases, the pre-training stage is conducted based on…

Cited by 52SourcePDFScholar
2021

Event Stream Super-Resolution via Spatiotemporal Constraint Learning

ICCV 2021poster

Event cameras are bio-inspired sensors that respond to brightness changes asynchronously and output in the form of event streams instead of frame-based images. They own outstanding advantages compared with traditional cameras: higher temporal resolution, higher dynamic range, and lower power consump…

Cited by 21PDFScholar
2020

Anomaly Detection with Training Data in Hyperspectral Imagery

ICASSP 2020accepted

In this paper, we investigate the anomaly detection problem for multi-pixel targets in hyperspectral imagery when training data are available. We derive the generalized likelihood ratio test and obtain its analytical expressions of the probability of false alarm and probability of detection. The per…

Cited by 0SourceScholar