← Search

Zijie Wu

11 accepted papers

2026

Learning Hierarchical Hyperbolic Mixture Model for Part-aware 3D Generation

CVPR 2026

3D shape generation has become increasingly important for graphics and vision applications. Current part-aware 3D generation usually overlooks hierarchical part relations or inefficiently encodes multi-level semantics in Euclidean space. Thus we propose a novel framework for hierarchical and efficie

Cited by 0SourceScholar
2025

AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation

ICCV 2025poster

Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spatio-temporal distributions and the scarcity of 4D training data. In this paper, we present AnimateAnyMesh, the first fee…

Cited by 0SourcePDFScholar
2025

Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generation

CVPR 2025poster

3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex preprocessing or low controllability. In this paper, we introduce a novel framework des…

Cited by 0SourcePDFScholar
2025

Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud Localization

ICCV 2025poster

Text to point cloud cross-modal localization is a crucial vision-language task for future human-robot collaboration. Existing coarse-to-fine frameworks assume that each query text precisely corresponds to the center area of a submap, limiting their applicability in real-world scenarios. This work re…

2025

Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-Localization

AAAI 2025technical

Visual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing…

2024

External Knowledge Enhanced 3D Scene Generation from Sketch

ECCV 2024poster

"Generating realistic 3D scenes is challenging due to the complexity of room layouts and object geometries. We propose a sketch based knowledge enhanced diffusion architecture (SEK) for generating customized, diverse, and plausible 3D scenes. SEK conditions the denoising process with a hand-drawn sk…

Cited by 6SourcePDFScholar
2024

Referring Human Pose and Mask Estimation In the Wild

NeurIPS 2024poster

We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to pr…

2023

3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point Cloud

CVPR 2023poster

In-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide…

2023

Sketch and Text Guided Diffusion Model for Colored Point Cloud Generation

ICCV 2023poster

Diffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently amb…

Cited by 32PDFScholar
2022

CCPL: Contrastive Coherence Preserving Loss for Versatile Style Transfer

ECCV 2022poster

"In this paper, we aim to devise a universally versatile style transfer method capable of performing artistic, photo-realistic, and video style transfer jointly, without seeing videos during training. Previous single-frame methods assume a strong constraint on the whole image to maintain temporal co…