← Search

Shengjun Zhang

8 accepted papers

2026

OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text

ICLR 2026poster

Composed video retrieval presents a complex challenge: retrieving a target video based on a source video and a textual modification instruction. This task demands fine-grained reasoning over multimodal transformations. However, existing benchmarks predominantly focus on vision–text alignment, largel…

Cited by 0SourceScholar
2025

Learning Efficient and Generalizable Human Representation with Human Gaussian Model

ICCV 2025poster

Modeling animatable human avatars from videos is a long-standing and challenging problem. While conventional methods require per-instance optimization, recent feed-forward methods have been proposed to generate 3D Gaussians with a learnable network.However, these methods predict independent Gaussian…

2025

Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model

CVPR 2025poster

In this paper, we propose Scene Splatter, a momentum-based paradigm for video diffusion to generate generic scenes from single image. Existing methods, which employ video generation models to synthesize novel views, suffer from limited video length and scene inconsistency, leading to artifacts and d…

Cited by 1SourcePDFScholar
2025

ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment

ICCV 2025poster

Perpetual 3D scene generation aims to produce long-range and coherent 3D view sequences, which is applicable for long-term video synthesis and 3D scene reconstruction. Existing methods follow a "navigate-and-imagine" fashion and rely on outpainting for successive view expansion. However, the generat…

Cited by 0SourcePDFScholar
2025

SurfelSplat: Learning Efficient and Generalizable Gaussian Surfel Representations for Sparse-View Surface Reconstruction

NeurIPS 2025poster

3D Gaussian Splatting (3DGS) has demonstrated impressive performance in 3D scene reconstruction. Beyond novel view synthesis, it shows great potential for multi-view surface reconstruction. Existing methods employ optimization-based reconstruction pipelines that achieve precise and complete surface…

Cited by 0SourceScholar
2024

Gaussian Graph Network: Learning Efficient and Generalizable Gaussian Representations from Multi-view Images

NeurIPS 2024poster

3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis performance. While conventional methods require per-scene optimization, more recently several feed-forward methods have been proposed to generate pixel-aligned Gaussian representations with a learnable network, which are g…

Cited by 1SourcePDFScholar
2024

GeoAuxNet: Towards Universal 3D Representation Learning for Multi-sensor Point Clouds

CVPR 2024poster

Point clouds captured by different sensors such as RGB-D cameras and LiDAR possess non-negligible domain gaps. Most existing methods design different network architectures and train separately on point clouds from various sensors. Typically point-based methods achieve outstanding performances on eve…