← Search

Ziyan Wu

40 accepted papers

2026

Consistent Instance Field for Dynamic Scene Understanding

CVPR 2026

We introduce Consistent Instance Field, a continuous and probabilistic spatio-temporal representation for dynamic scene understanding.Unlike prior methods that rely on discrete tracking or view-dependent features, our approach disentangles visibility from persistent object identity by modeling each

Cited by 0SourceScholar
2026

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding

CVPR 2026

Large vision-language models struggle with medical video understanding, where spatial precision, temporal reasoning, and clinical semantics are critical. To address this, we first introduce MedVidBench, a large-scale benchmark of 531,850 video-instruction pairs across 8 medical sources spanning vide

Cited by 0SourceScholar
2025

3D Vision-Language Gaussian Splatting

ICLR 2025poster

Recent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches h…

Cited by 16SourcePDFScholar
2025

6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric Rendering

ICLR 2025poster

Novel view synthesis has advanced significantly with the development of neural radiance fields (NeRF) and 3D Gaussian splatting (3DGS). However, achieving high quality without compromising real-time rendering remains challenging, particularly for physically-based rendering using ray/path tracing wit…

2025

7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting

ICCV 2025poster

Real-time rendering of dynamic scenes with view-dependent effects remains a fundamental challenge in computer graphics. While recent advances in Gaussian Splatting have shown promising results separately handling dynamic scenes (4DGS) and view-dependent effects (6DGS), no existing method unifies the…

Cited by 0SourcePDFScholar
2025

CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single Image

ICCV 2025poster

Reconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on the assumption that the human subject is in an occlusion-free…

Cited by 0SourcePDFScholar
2025

Order-aware Interactive Segmentation

ICLR 2025poster

Interactive segmentation aims to accurately segment target objects with minimal user interactions. However, current methods often fail to accurately separate target objects from the background, due to a limited understanding of order, the relative depth between objects in a scene. To address this is…

Cited by 0SourcePDFScholar
2025

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding

CVPR 2025poster

Temporal awareness is essential for video large language models (LLMs) to understand and reason about events within long videos, enabling applications like dense video captioning and temporal video grounding in a unified system. However, the scarcity of long videos with detailed captions and precise…

Cited by 1SourcePDFScholar
2024

DDGS-CT: Direction-Disentangled Gaussian Splatting for Realistic Volume Rendering

NeurIPS 2024poster

Digitally reconstructed radiographs (DRRs) are simulated 2D X-ray images generated from 3D CT volumes, widely used in preoperative settings but limited in intraoperative applications due to computational bottlenecks. Physics-based Monte Carlo simulations provide accurate representations but are extr…

Cited by 6SourcePDFScholar
2024

DaReNeRF: Direction-aware Representation for Dynamic Scenes

CVPR 2024poster

Addressing the intricate challenge of modeling and re-rendering dynamic scenes most recent approaches have sought to simplify these complexities using plane-based explicit representations overcoming the slow training time issues associated with methods like Neural Radiance Fields (NeRF) and implicit…

Cited by 9SourcePDFScholar
2024

Disguise without Disruption: Utility-Preserving Face De-identification

AAAI 2024technical

With the rise of cameras and smart sensors, humanity generates an exponential amount of data. This valuable information, including underrepresented cases like AI in medical settings, can fuel new deep-learning tools. However, data scientists must prioritize ensuring privacy for individuals in these…

Cited by 15SourcePDFScholar
2024

Federated Learning via Input-Output Collaborative Distillation

AAAI 2024technical

Federated learning (FL) is a machine learning paradigm in which distributed local nodes collaboratively train a central model without sharing individually held private data. Existing FL methods either iteratively share local model parameters or deploy co-distillation. However, the former is highly s…

2024

Implicit Modeling of Non-rigid Objects with Cross-Category Signals

AAAI 2024technical

Deep implicit functions (DIFs) have emerged as a potent and articulate means of representing 3D shapes. However, methods modeling object categories or non-rigid entities have mainly focused on single-object scenarios. In this work, we propose MODIF, a multi-object deep implicit function that jointly…

Cited by 1SourcePDFScholar
2024

PBADet: A One-Stage Anchor-Free Approach for Part-Body Association

ICLR 2024poster

The detection of human parts (e.g., hands, face) and their correct association with individuals is an essential task, e.g., for ubiquitous human-machine interfaces and action recognition. Traditional methods often employ multi-stage processes, rely on cumbersome anchor-based systems, or do not scale…

Cited by 1SourcePDFScholar
2023

CMDA: Cross-Modality Domain Adaptation for Nighttime Semantic Segmentation

ICCV 2023poster

Most nighttime semantic segmentation studies are based on domain adaptation approaches and image input. However, limited by the low dynamic range of conventional cameras, images fail to capture structural details and boundary information in low-light conditions. Event cameras, as a new form of visio…

Cited by 38PDFcodeScholar
2023

Progressive Multi-View Human Mesh Recovery with Self-Supervision

AAAI 2023technical

To date, little attention has been given to multi-view 3D human mesh estimation, despite real-life applicability (e.g., motion capture, sport analysis) and robustness to single-view ambiguities. Existing solutions typically suffer from poor generalization performance to new settings, largely due to…

Cited by 16SourcePDFScholar
2022

Forecasting Human Trajectory from Scene History

NeurIPS 2022accept

Predicting the future trajectory of a person remains a challenging problem, due to randomness and subjectivity. However, the moving patterns of human in constrained scenario typically conform to a limited number of regularities to a certain extent, because of the scenario restrictions (\eg, floor pl…

2022

PREF: Predictability Regularized Neural Motion Fields

ECCV 2022poster

"Knowing the 3D motions in a dynamic scene is essential to many vision applications. Recent progress is mainly focused on estimating the activity of some specific elements like humans. In this paper, we leverage a neural motion field for estimating the motion of all points in a multiview setting. Mo…

Cited by 39SourcePDFScholar
2022

PseudoClick: Interactive Image Segmentation with Click Imitation

ECCV 2022poster

"The goal of click-based interactive image segmentation is to obtain precise object segmentation masks with limited user interaction, i.e., by a minimal number of user clicks. Existing methods require users to provide all the clicks: by first inspecting the segmentation mask and then providing point…

Cited by 69SourcePDFScholar
2022

SMPL-A: Modeling Person-Specific Deformable Anatomy

CVPR 2022poster

A variety of diagnostic and therapeutic protocols rely on locating in vivo target anatomical structures, which can be obtained from medical scans. However, organs move and deform as the patient changes his/her pose. In order to obtain accurate target location information, clinicians have to either c…

Cited by 12PDFScholar
2022

Self-supervised Human Mesh Recovery with Cross-Representation Alignment

ECCV 2022poster

"Fully supervised human mesh recovery methods are data-hungry and have poor generalizability due to the limited availability and diversity of 3D-annotated benchmark datasets. Recent progress in self-supervised human mesh recovery has been made using synthetic-data-driven training paradigms where the…

Cited by 15SourcePDFScholar
2021

A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual Concepts

CVPR 2021poster

Despite substantial progress in applying neural networks (NN) to a wide variety of areas, they still largely suffer from a lack of transparency and interpretability. While recent developments in explainable artificial intelligence attempt to bridge this gap (e.g., by visualizing the correlation betw…

Cited by 60PDFScholar
2021

Ensemble Attention Distillation for Privacy-Preserving Federated Learning

ICCV 2021poster

We consider the problem of Federated Learning (FL) where numerous decentralized computational nodes collaborate with each other to train a centralized machine learning model without explicitly sharing their local data samples. Such decentralized training naturally leads to issues of imbalanced or di…

Cited by 148PDFScholar
2021

Spatio-Temporal Representation Factorization for Video-Based Person Re-Identification

ICCV 2021poster

Despite much recent progress in video-based person re-identification (re-ID), the current state-of-the-art still suffers from common real-world challenges such as appearance similarity among various people, occlusions, and frame misalignment. To alleviate these problems, we propose Spatio-Temporal R…

Cited by 93PDFScholar
2020

Hierarchical Kinematic Human Mesh Recovery

ECCV 2020poster

We consider the problem of estimating a parametric model of 3D human mesh from a single image. While there has been substantial recent progress in this area with direct regression of model parameters, these methods only implicitly exploit the human body kinematic structure, leading to sub-optimal us…

Cited by 128SourcePDFScholar
2020

Towards Visually Explaining Variational Autoencoders

CVPR 2020oral

Recent advances in Convolutional Neural Network (CNN) model interpretability have led to impressive progress in visualizing and understanding model predictions. In particular, gradient-based visual attention methods have driven much recent effort in using visual attention maps as a means for visual…

Cited by 301PDFcodeScholar
2019

Incremental Scene Synthesis

NeurIPS 2019poster

We present a method to incrementally generate complete 2D or 3D scenes with the following properties: (a) it is globally consistent at each step according to a learned scene prior, (b) real observations of a scene can be incorporated while observing global consistency, (c) unobserved regions can be…

Cited by 9SourcePDFScholar
2019

Learning Local RGB-to-CAD Correspondences for Object Pose Estimation

ICCV 2019poster

We consider the problem of 3D object pose estimation. While much recent work has focused on the RGB domain, the reliance on accurately annotated images limits generalizability and scalability. On the other hand, the easily available object CAD models are rich sources of data, providing a large numbe…

Cited by 30PDFScholar
2019

Seeing Beyond Appearance - Mapping Real Images into Geometrical Domains for Unsupervised CAD-based Recognition

IROS 2019poster

While convolutional neural networks are dominating the field of computer vision, one usually does not have access to the large amount of domain-relevant data needed for their training. Therefore, it has become common practice to use available synthetic samples along domain adaptation schemes to prep…

Cited by 14SourceScholar
2019

Sharpen Focus: Learning With Attention Separability and Consistency

ICCV 2019poster

Recent developments in gradient-based attention modeling have seen attention maps emerge as a powerful tool for interpreting convolutional neural networks. Despite good localization for an individual class of interest, these techniques produce attention maps with substantially overlapping responses…

Cited by 41PDFScholar
2018

End-to-End Learning of Keypoint Detector and Descriptor for Pose Invariant 3D Matching

CVPR 2018poster

Finding correspondences between images or 3D scans is at the heart of many computer vision and image retrieval applications and is often enabled by matching local keypoint descriptors. Various learning approaches have been applied in the past to different stages of the matching pipeline, considering…

Cited by 71SourcePDFScholar
2018

Learning Compositional Visual Concepts With Mutual Consistency

CVPR 2018poster

Compositionality of semantic concepts in image synthesis and analysis is appealing as it can help in decomposing known and generatively recomposing unknown data. For instance, we may learn concepts of changing illumination, geometry or albedo of a scene, and try to recombine them to generate physica…

Cited by 17SourcePDFScholar
2018

Tell Me Where to Look: Guided Attention Inference Network

CVPR 2018poster

Weakly supervised learning with only coarse labels can obtain visual explanations of deep neural network such as attention maps by back-propagating gradients. These attention maps are then available as priors for tasks such as object localization and semantic segmentation. In one common framework we…

Cited by 719SourcePDFScholar