← Search

Panwang Pan

20 accepted papers

2026

Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene Generation

CVPR 2026

We introduce Diff4Splat, a feed-forward framework for dynamic scene generation from a single image. Our method synergizes the powerful generative priors of video diffusion models with geometric and motion constraints learned from a large-scale 4D dataset. Given a single image, a camera trajectory, a

Cited by 0SourcecodeScholar
2026

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

CVPR 2026

Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark

Cited by 0SourceScholar
2026

ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation

CVPR 2026

Significant progress has been achieved in high-fidelity video synthesis, yet current paradigms often fall short in effectively integrating identity information from multiple subjects. This leads to semantic conflicts and suboptimal performance in preserving identities and interactions, limiting cont

Cited by 0SourceScholar
2026

MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second

CVPR 2026

We present MoVieS, a Motion-aware View Synthesis model that reconstructs 4D dynamic scenes from monocular videos in one second. It represents dynamic 3D scenes with pixel-aligned Gaussian primitives and explicitly supervises their time-varying motions. This allows, for the first time, the unified mo

Cited by 0SourcecodeScholar
2026

SAM-Veteran: An MLLM-Based Human-like SAM Agent for Reasoning Segmentation

ICLR 2026poster

Significant progress has been made in reasoning segmentation by combining multi-modal large language models (MLLMs) with the Segment Anything Model (SAM): the former excel in reasoning and vision–language alignment, while the latter offers powerful pixel-level understanding. However, current paradig…

Cited by 0SourceScholar
2026

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

CVPR 2026

Humans inhabit a physical 4D world, where spatial geometry and semantic content evolve over time, forming a dynamic reality. While current Multimodal Large Language Models (MLLMs) demonstrate strong capabilities in understanding static visual inputs, it remains unclear whether they can effectively "

Cited by 0SourcecodeScholar
2025

4K4DGen: Panoramic 4D Generation at 4K Resolution

ICLR 2025spotlight

The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques either focus solely on dynamic objects or perform outpainting from a single perspe…

2025

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation

ICLR 2025poster

Recent advancements in 3D content generation from text or a single image struggle with limited high-quality 3D datasets and inconsistency from 2D multi-view generation. We introduce DiffSplat, a novel 3D generative framework that natively generates 3D Gaussian splats by taming large-scale text-to-im…

Cited by 5SourcePDFScholar
2025

DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling

NeurIPS 2025poster

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human‑like capabilities. Howev…

Cited by 0SourceScholar
2025

HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation

NeurIPS 2025poster

Recent advances in generative models have achieved high-fidelity in 3D human reconstruction, yet their utility for specific tasks (e.g., human 3D segmentation) remains constrained. We propose HumanCrafter, a unified framework that enables the joint modeling of appearance and human-part semantics fro…

Cited by 0SourceScholar
2025

InfoBridge: Balanced Multimodal Integration through Conditional Dependency Modeling

ICCV 2025poster

Developing systems that interpret diverse real-world signals remains a fundamental challenge in multimodal learning. Current approaches face significant obstacles from inherent modal heterogeneity. While existing methods attempt to enhance fusion through cross-modal alignment or interaction mechanis…

2025

InstantSplamp: Fast and Generalizable Stenography Framework for Generative Gaussian Splatting

ICLR 2025poster

With the rapid development of large generative models for 3D, especially the evolution from NeRF representations to more efficient Gaussian Splatting, the synthesis of 3D assets has become increasingly fast and efficient, enabling the large-scale publication and sharing of generated 3D objects. Howe…

2025

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

NeurIPS 2025poster

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based sol…

Cited by 0SourceScholar
2025

JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration

CVPR 2025poster

Vision-centric perception systems often struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To enable robust and autonomous operation in real-world…

Cited by 2SourcePDFScholar
2025

Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions

NeurIPS 2025poster

The synthesis of realistic Martian landscape videos, essential for mission rehearsal and robotic simulation, presents unique challenges. These primarily stem from the scarcity of high-quality Martian data and the significant domain gap relative to terrestrial imagery. To address these challenges, we…

Cited by 0SourceScholar
2025

Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables

NeurIPS 2025oral

Recently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicabil…

Cited by 0SourceScholar
2025

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers

NeurIPS 2025poster

We introduce PartCrafter, the first structured 3D generative model that jointly synthesizes multiple semantically meaningful and geometrically distinct 3D meshes from a single RGB image. Unlike existing methods that either produce monolithic 3D shapes or follow two-stage pipelines, i.e. first segmen…

Cited by 0SourceScholar
2024

HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors

NeurIPS 2024poster

Despite recent advancements in high-fidelity human reconstruction techniques, the requirements for densely captured images or time-consuming per-instance optimization significantly hinder their applications in broader scenarios. To tackle these issues, we present **HumanSplat**, which predicts the 3…

2024

Multi-modal Relation Distillation for Unified 3D Representation Learning

ECCV 2024poster

"Recent advancements in multi-modal pre-training for 3D point clouds have demonstrated promising results by aligning heterogeneous features across 3D shapes and their corresponding 2D images and language descriptions. However, current straightforward solutions often overlook intricate structural rel…

Cited by 0SourcePDFScholar
2023

StegaNeRF: Embedding Invisible Information within Neural Radiance Fields

ICCV 2023poster

Recent advancements in neural rendering have paved the way for a future marked by the widespread distribution of visual data through the sharing of Neural Radiance Field (NeRF) model weights. However, while established techniques exist for embedding ownership or copyright information within conventi…

Cited by 66PDFcodeScholar