← Search

Jiahao Chang

12 accepted papers

2026

EI-Part:Explode for Completion and Implode for Refinement

CVPR 2026

Part-level 3D generation is crucial for various downstream applications, including gaming, film production, and industrial design. However, decomposing a 3D shape into geometrically plausible and meaningful components remains a significant challenge. Previous part-based generation methods often stru

Cited by 0SourceScholar
2026

ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos

CVPR 2026

The ubiquity of monocular videos capturing daily hand-object interactions presents a valuable resource for embodied intelligence. While 3D hand reconstruction from in-the-wild videos has seen significant progress, reconstructing the involved objects remains challenging due to severe occlusions and t

Cited by 0SourcecodeScholar
2026

GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors

ICML 2026poster

Reconstructing 3D scenes using 3D Gaussian Splatting (3DGS) from sparse views is an ill-posed problem due to insufficient information, often resulting in noticeable artifacts. While recent approaches have sought to leverage generative priors to complete information for under-constrained regions, the…

Cited by 0SourceScholar
2026

MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence

ICML 2026poster

Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Despite its importance, this capability remains a significant bottleneck for current multimodal large language models (MLLM…

Cited by 0SourceScholar
2026

ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation

ICLR 2026poster

Existing multi-view 3D object reconstruction methods heavily rely on sufficient overlap between input views, where occlusions and sparse coverage in practice frequently yield severe reconstruction incompleteness. Recent advancements in diffusion-based 3D generative techniques offer the potential to…

Cited by 0SourcecodeScholar
2026

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models

CVPR 2026

Large vision-language models (VLMs) show strong multimodal understanding but still struggle with 3D spatial reasoning, such as distance estimation, size comparison, and cross-view consistency. Existing 3D-aware methods either depend on auxiliary 3D information or enhance RGB-only VLMs with geometry

Cited by 0SourceScholar
2025

Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging

ICCV 2025poster

With the growing demand for high-fidelity 3D models from 2D images, existing methods still face significant challenges in accurately reproducing fine-grained geometric details due to limitations in domain gaps and inherent ambiguities in RGB images. To address these issues, we propose Hi3DGen, a nov…

Cited by 0SourcePDFScholar
2025

Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion

ICCV 2025poster

3D data simulation aims to bridge the gap between simulated and real-captured 3D data, which is a fundamental problem for real-world 3D visual tasks. Most 3D data simulation methods inject predefined physical priors but struggle to capture the full complexity of real data. An optimal approach involv…

Cited by 0SourcePDFScholar
2024

TopoMLP: A Simple yet Strong Pipeline for Driving Topology Reasoning

ICLR 2024poster

Topology reasoning aims to comprehensively understand road scenes and present drivable routes in autonomous driving. It requires detecting road centerlines (lane) and traffic elements, further reasoning their topology relationship, \textit{i.e.}, lane-lane topology, and lane-traffic topology. In thi…

2023

DETRDistill: A Universal Knowledge Distillation Framework for DETR-families

ICCV 2023poster

Transformer-based detectors (DETRs) are becoming popular for their simple framework, but the large model size and heavy time consumption hinder their deployment in the real world. While knowledge distillation (KD) can be an appealing technique to compress giant detectors into small ones for comparab…

Cited by 39PDFScholar
2023

Towards Domain Generalization for Multi-View 3D Object Detection in Bird-Eye-View

CVPR 2023poster

Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D object detection have been continuously proposed, most of them may risk drastic performance degradation when the domain o…

Cited by 25SourcePDFScholar
2023

Visual Recognition-Driven Image Restoration for Multiple Degradation With Intrinsic Semantics Recovery

CVPR 2023poster

Deep image recognition models suffer a significant performance drop when applied to low-quality images since they are trained on high-quality images. Although many studies have investigated to solve the issue through image restoration or domain adaptation, the former focuses on visual quality rather…

Cited by 23SourcePDFScholar