← Search

Zhongpai Gao

20 accepted papers

2026

Consistent Instance Field for Dynamic Scene Understanding

CVPR 2026

We introduce Consistent Instance Field, a continuous and probabilistic spatio-temporal representation for dynamic scene understanding.Unlike prior methods that rely on discrete tracking or view-dependent features, our approach disentangles visibility from persistent object identity by modeling each

Cited by 0SourceScholar
2026

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding

CVPR 2026

Large vision-language models struggle with medical video understanding, where spatial precision, temporal reasoning, and clinical semantics are critical. To address this, we first introduce MedVidBench, a large-scale benchmark of 531,850 video-instruction pairs across 8 medical sources spanning vide

Cited by 0SourceScholar
2025

3D Vision-Language Gaussian Splatting

ICLR 2025poster

Recent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches h…

Cited by 16SourcePDFScholar
2025

6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric Rendering

ICLR 2025poster

Novel view synthesis has advanced significantly with the development of neural radiance fields (NeRF) and 3D Gaussian splatting (3DGS). However, achieving high quality without compromising real-time rendering remains challenging, particularly for physically-based rendering using ray/path tracing wit…

2025

7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting

ICCV 2025poster

Real-time rendering of dynamic scenes with view-dependent effects remains a fundamental challenge in computer graphics. While recent advances in Gaussian Splatting have shown promising results separately handling dynamic scenes (4DGS) and view-dependent effects (6DGS), no existing method unifies the…

Cited by 0SourcePDFScholar
2025

CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single Image

ICCV 2025poster

Reconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on the assumption that the human subject is in an occlusion-free…

Cited by 0SourcePDFScholar
2025

Order-aware Interactive Segmentation

ICLR 2025poster

Interactive segmentation aims to accurately segment target objects with minimal user interactions. However, current methods often fail to accurately separate target objects from the background, due to a limited understanding of order, the relative depth between objects in a scene. To address this is…

Cited by 0SourcePDFScholar
2025

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding

CVPR 2025poster

Temporal awareness is essential for video large language models (LLMs) to understand and reason about events within long videos, enabling applications like dense video captioning and temporal video grounding in a unified system. However, the scarcity of long videos with detailed captions and precise…

Cited by 1SourcePDFScholar
2024

DDGS-CT: Direction-Disentangled Gaussian Splatting for Realistic Volume Rendering

NeurIPS 2024poster

Digitally reconstructed radiographs (DRRs) are simulated 2D X-ray images generated from 3D CT volumes, widely used in preoperative settings but limited in intraoperative applications due to computational bottlenecks. Physics-based Monte Carlo simulations provide accurate representations but are extr…

Cited by 6SourcePDFScholar
2024

DaReNeRF: Direction-aware Representation for Dynamic Scenes

CVPR 2024poster

Addressing the intricate challenge of modeling and re-rendering dynamic scenes most recent approaches have sought to simplify these complexities using plane-based explicit representations overcoming the slow training time issues associated with methods like Neural Radiance Fields (NeRF) and implicit…

Cited by 9SourcePDFScholar
2024

DiffStega: Towards Universal Training-Free Coverless Image Steganography with Diffusion Models

IJCAI 2024poster

Traditional image steganography focuses on concealing one image within another, aiming to avoid steganalysis by unauthorized entities. Coverless image steganography (CIS) enhances imperceptibility by not using any cover image. Recent works have utilized text prompts as keys in CIS through diffusion…

2024

Disguise without Disruption: Utility-Preserving Face De-identification

AAAI 2024technical

With the rise of cameras and smart sensors, humanity generates an exponential amount of data. This valuable information, including underrepresented cases like AI in medical settings, can fuel new deep-learning tools. However, data scientists must prioritize ensuring privacy for individuals in these…

Cited by 15SourcePDFScholar
2024

Divide and Fuse: Body Part Mesh Recovery from Partially Visible Human Images

ECCV 2024poster

"We introduce a novel bottom-up approach for human body mesh reconstruction, specifically designed to address the challenges posed by partial visibility and occlusion in input images. Traditional top-down methods, relying on whole-body parametric models like SMPL, falter when only a small part of th…

Cited by 2SourcePDFScholar
2024

Implicit Modeling of Non-rigid Objects with Cross-Category Signals

AAAI 2024technical

Deep implicit functions (DIFs) have emerged as a potent and articulate means of representing 3D shapes. However, methods modeling object categories or non-rigid entities have mainly focused on single-object scenarios. In this work, we propose MODIF, a multi-object deep implicit function that jointly…

Cited by 1SourcePDFScholar
2024

PBADet: A One-Stage Anchor-Free Approach for Part-Body Association

ICLR 2024poster

The detection of human parts (e.g., hands, face) and their correct association with individuals is an essential task, e.g., for ubiquitous human-machine interfaces and action recognition. Traditional methods often employ multi-stage processes, rely on cumbersome anchor-based systems, or do not scale…

Cited by 1SourcePDFScholar
2022

Learning Invisible Markers for Hidden Codes in Offline-to-Online Photography

CVPR 2022poster

QR (quick response) codes are widely used as an offline-to-online channel to convey information (e.g., links) from publicity materials (e.g., display and print) to mobile devices. However, QR Codes are not favorable for taking up valuable space of publicity materials. Recent works propose invisible…

Cited by 37PDFScholar
2021

Learning Local Neighboring Structure for Robust 3D Shape Representation

AAAI 2021technical

Mesh is a powerful data structure for 3D shapes. Representation learning for 3D meshes is important in many computer vision and graphics applications. The recent success of convolutional neural networks (CNNs) for structured data (e.g., images) suggests the value of adapting insight from CNN for 3D…

2021

Learning Spectral Dictionary for Local Representation of Mesh

IJCAI 2021poster

For meshes, sharing the topology of a template is a common and practical setting in face-, hand-, and body-related applications. Meshes are irregular since each vertex's neighbors are unordered and their orientations are inconsistent with other vertices. Previous methods use isotropic filters or pre…

2021

Looking Here or There? Gaze Following in 360-Degree Images

ICCV 2021poster

Gaze following, i.e., detecting the gaze target of a human subject, in 2D images has become an active topic in computer vision. However, it usually suffers from the out of frame issue due to the limited field-of-view (FoV) of 2D images. In this paper, we introduce a novel task, gaze following in 360…

Cited by 23PDFScholar