← Search

Zixin Zhu

6 accepted papers

2026

Learning 3D Shape Fidelity Metric from Real-world Distortions

CVPR 2026

3D generation and reconstruction have become essential in many computer vision applications, where the reconstructed or generated 3D shapes need to appear realistic to human perception. However, traditional metrics like Chamfer Distance to compare two 3D shapes focus primarily on matching accuracy o

Cited by 0SourceScholar
2026

Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network

AAAI 2026technical

Textured high-fidelity 3D models are crucial for games, AR/VR, and film, but human-aligned evaluation methods still fall behind despite recent advances in 3D reconstruction and generation. Existing metrics, such as Chamfer Distance, often fail to align with how humans evaluate the fidelity of 3D sha

Cited by 0SourcePDFScholar
2025

CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation

ICCV 2025poster

In text-to-image (T2I) generation, achieving fine-grained control over attributes - such as age or smile - remains challenging, even with detailed text prompts. Slider-based methods offer a solution for precise control of image attributes.Existing approaches typically train individual adapter for ea…

Cited by 0SourcePDFScholar
2025

GeoRemover: Removing Objects and Their Causal Visual Artifacts

NeurIPS 2025spotlight

Towards intelligent image editing, object removal should eliminate both the target object and its causal visual artifacts, such as shadows and reflections. However, existing image appearance-based methods either follow strictly mask-aligned training and fail to remove these casual effects which are…

Cited by 0SourcecodeScholar
2022

Learning Disentangled Classification and Localization Representations for Temporal Action Localization

AAAI 2022technical

A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that thi…

Cited by 20SourcePDFScholar
2021

Enriching Local and Global Contexts for Temporal Action Localization

ICCV 2021poster

Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by…

Cited by 148PDFcodeScholar