← Search

Zongxin Yang

42 accepted papers

2026

Beyond Independent Genes: Learning Module-Inductive Representations for Gene Perturbation Prediction

ICML 2026poster

Predicting transcriptional responses to genetic perturbations is a central problem in functional genomics. In practice, perturbation responses are rarely gene-independent but instead manifest as coordinated, program-level transcriptional changes among functionally related genes. However, most existi…

Cited by 0SourceScholar
2026

BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment

ICLR 2026poster

Conditional image generation augments text-to-image synthesis with structural, spatial, or stylistic priors and is used in many domains. However, current methods struggle to harmonize guidance from both sources when conflicts arise: 1) input-level conflict, where the semantics of the conditioning im…

Cited by 0SourcecodeScholar
2026

Insert Anything: Image Insertion via In-Context Editing in DiT

AAAI 2026technical

This work presents Insert Anything, a unified framework for reference-based image insertion that seamlessly integrates objects from reference images into target scenes under flexible, user-specified control guidance. Instead of training separate models for individual tasks, our approach is trained o

Cited by 0SourcePDFScholar
2026

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

ICLR 2026poster

Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challen…

Cited by 0SourceScholar
2026

ReflFlow: Learning Geometry-Guided Ray Tracing for Dynamic Specular Reconstruction

ICML 2026poster

We present ReflFlow, a novel framework for high-fidelity rendering of dynamic specular scenes by addressing two key challenges: precise reflection direction estimation and physically accurate modeling. To achieve this, we propose a Residual Material-Augmented 2D Gaussian Splatting representation tha…

Cited by 0SourceScholar
2026

Spiked-CFR: Causal Representation Learning from LLMs via Wasserstein Projection Pursuit

ICML 2026poster

Estimating treatment effects from observational text is increasingly practical with Large Language Models (LLMs). However, applying causal representation learning directly to high-dimensional LLM embeddings faces a fundamental barrier: empirical Wasserstein matching suffers from the curse of dimensi…

Cited by 0SourceScholar
2026

Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models

ICLR 2026poster

Rigged 3D assets are fundamental to 3D deformation and animation. However, existing 3D generation methods face challenges in generating animatable geometry, while rigging techniques lack fine-grained structural control over skeleton creation. To address these limitations, we introduce Stroke3D, a no…

Cited by 0SourceScholar
2025

3DIS: Depth-Driven Decoupled Image Synthesis for Universal Multi-Instance Generation

ICLR 2025spotlight

The increasing demand for controllable outputs in text-to-image generation has spurred advancements in multi-instance generation (MIG), allowing users to define both instance layouts and attributes. However, unlike image-conditional generation methods such as ControlNet, MIG techniques have not been…

2025

DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models

ICCV 2025poster

Image-conditioned generation methods, such as depth- and canny-conditioned approaches, have demonstrated remarkable abilities for precise image synthesis. However, existing models still struggle to accurately control the content of multiple instances (or regions). Even state-of-the-art models like F…

2025

Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

NeurIPS 2025poster

Instruction-based image editing enables precise modifications via natural language prompts, but existing methods face a precision-efficiency tradeoff: fine-tuning demands massive datasets (>10M) and computational resources, while training-free approaches suffer from weak instruction comprehension.…

Cited by 0SourceScholar
2025

Few-Shot Incremental Learning via Foreground Aggregation and Knowledge Transfer for Audio-Visual Semantic Segmentation

AAAI 2025technical

Audio-Visual Semantic Segmentation (AVSS) has gained significant attention in the multi-modal domain, aiming to segment video objects that produce specific sounds in the corresponding audio. Despite notable progress, existing methods still struggle to handle new classes not included in the original…

Cited by 0SourcePDFScholar
2025

Origin Identification for Text-Guided Image-to-Image Diffusion Models

ICML 2025poster

Text-guided image-to-image diffusion models excel in translating images based on textual prompts, allowing for precise and creative visual modifications. However, such a powerful technique can be misused for *spreading misinformation*, *infringing on copyrights*, and *evading content tracing*. This…

2025

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

ICLR 2025poster

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapti…

2025

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

CVPR 2025poster

Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segme…

2025

X-Field: A Physically Informed Representation for 3D X-ray Reconstruction

NeurIPS 2025spotlight

X-ray imaging is indispensable in medical diagnostics, yet its use is tightly regulated due to radiation exposure. Recent research borrows representations from the 3D reconstruction area to complete two tasks with reduced radiation dose: X-ray Novel View Synthesis (NVS) and Computed Tomography (CT)…

Cited by 0SourceScholar
2024

Controllable 3D Face Generation with Conditional Style Code Diffusion

AAAI 2024technical

Generating photorealistic 3D faces from given conditions is a challenging task. Existing methods often rely on time-consuming one-by-one optimization approaches, which are not efficient for modeling the same distribution content, e.g., faces. Additionally, an ideal controllable 3D face generation mo…

2024

DRIP: Unleashing Diffusion Priors for Joint Foreground and Alpha Prediction in Image Matting

NeurIPS 2024poster

Recovering the foreground color and opacity/alpha matte from a single image (i.e., image matting) is a challenging and ill-posed problem where data priors play a critical role in achieving precise results. Traditional methods generally predict the alpha matte and then extract the foreground through…

Cited by 2SourcePDFScholar
2024

DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)

ICML 2024poster

Recent LLM-driven visual agents mainly focus on solving image-based tasks, which limits their ability to understand dynamic scenes, making it far from real-life applications like guiding students in laboratory experiments and identifying their mistakes. Hence, this paper explores DoraemonGPT, a comp…

2024

HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting

ECCV 2024poster

"Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods struggle to create high-quality and consistent animated avatars efficiently. Previous animatable head models like FLAME have…

2024

SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human Reconstruction

CVPR 2024highlight

Creating high-quality 3D models of clothed humans from single images for real-world applications is crucial. Despite recent advancements accurately reconstructing humans in complex poses or with loose clothing from in-the-wild images along with predicting textures for unseen areas remains a signific…

2023

Efficient Emotional Adaptation for Audio-Driven Talking-Head Generation

ICCV 2023poster

Audio-driven talking-head synthesis is a popular research topic for virtual human-related applications. However, the inflexibility and inefficiency of existing methods, which necessitate expensive end-to-end training to transfer emotions from guidance videos to talking-head predictions, are signific…

Cited by 122PDFcodeScholar
2023

FedSeg: Class-Heterogeneous Federated Learning for Semantic Segmentation

CVPR 2023poster

Federated Learning (FL) is a distributed learning paradigm that collaboratively learns a global model across multiple clients with data privacy-preserving. Although many FL algorithms have been proposed for classification tasks, few works focus on more challenging semantic seg-mentation tasks, espec…

Cited by 53SourcePDFScholar
2023

Global-correlated 3D-decoupling Transformer for Clothed Avatar Reconstruction

NeurIPS 2023poster

Reconstructing 3D clothed human avatars from single images is a challenging task, especially when encountering complex poses and loose clothing. Current methods exhibit limitations in performance, largely attributable to their dependence on insufficient 2D image features and inconsistent query metho…

2023

Global-to-Local Modeling for Video-Based 3D Human Pose and Shape Estimation

CVPR 2023poster

Video-based 3D human pose and shape estimations are evaluated by intra-frame accuracy and inter-frame smoothness. Although these two metrics are responsible for different ranges of temporal consistency, existing state-of-the-art methods treat them as a unified problem and use monotonous modeling str…

2023

Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and Segmentation

ICCV 2023poster

Tracking any given object(s) spatially and temporally is a common purpose in Visual Object Tracking (VOT) and Video Object Segmentation (VOS). Joint tracking and segmentation have been attempted in some studies but they often lack full compatibility of both box and mask in initialization and predict…

Cited by 17PDFcodeScholar
2023

JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh Recovery

ICCV 2023poster

In this study, we focus on the problem of 3D human mesh recovery from a single image under obscured conditions. Most state-of-the-art methods aim to improve 2D alignment technologies, such as spatial averaging and 2D joint sampling. However, they tend to neglect the crucial aspect of 3D alignment by…

Cited by 19PDFcodeScholar
2023

ProD: Prompting-To-Disentangle Domain Knowledge for Cross-Domain Few-Shot Image Classification

CVPR 2023poster

This paper considers few-shot image classification under the cross-domain scenario, where the train-to-test domain gap compromises classification accuracy. To mitigate the domain gap, we propose a prompting-to-disentangle (ProD) method through a novel exploration with the prompting mechanism. ProD a…

Cited by 27SourcePDFScholar
2023

TransHuman: A Transformer-based Human Representation for Generalizable Neural Human Rendering

ICCV 2023poster

In this paper, we focus on the task of generalizable neural human rendering which trains conditional Neural Radiance Fields (NeRF) from multi-view videos of different characters. To handle the dynamic human motion, previous methods have primarily used a SparseConvNet (SPC)-based human representation…

Cited by 25PDFcodeScholar
2022

H2FA R-CNN: Holistic and Hierarchical Feature Alignment for Cross-Domain Weakly Supervised Object Detection

CVPR 2022poster

Cross-domain weakly supervised object detection (CDWSOD) aims to adapt the detection model to a novel target domain with easily acquired image-level annotations. How to align the source and target domains is critical to the CDWSOD accuracy. Existing methods usually focus on partial detection compone…

Cited by 51PDFcodeScholar
2022

Instance As Identity: A Generic Online Paradigm for Video Instance Segmentation

ECCV 2022poster

"Modeling temporal information for both detection and tracking in a unified framework has been proved a promising solution to video instance segmentation (VIS). However, how to effectively incorporate the temporal information into an online model remains an open problem. In this work, we propose a n…

2021

Associating Objects with Transformers for Video Object Segmentation

NeurIPS 2021poster

This paper investigates how to realize better and more efficient embedding learning to tackle the semi-supervised video object segmentation under challenging multi-object scenarios. The state-of-the-art methods learn to decode features with a single positive object and thus have to match and segment…

2020

Collaborative Video Object Segmentation by Foreground-Background Integration

ECCV 2020poster

This paper investigates the principles of embedding learning to tackle the challenging semi-supervised video object segmentation. Different from previous practices that only explore the embedding learning using pixels from foreground object (s), we consider background should be equally treated and t…

2019

Very Long Natural Scenery Image Prediction by Outpainting

ICCV 2019poster

Comparing to image inpainting, image outpainting receives less attention due to two challenges in it. The first challenge is how to keep the spatial and content consistency between generated images and original input. The second challenge is how to maintain high quality in generated results, especia…

Cited by 114PDFcodeScholar