← Search

Songyou Peng

30 accepted papers

2026

Do 3D Large Language Models Really Understand 3D Spatial Relationships?

ICLR 2026poster

Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on text-only question-answer pairs can perform comparably or even surpass these methods on the SQA3D benchmark without using…

Cited by 0SourceScholar
2026

Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature Alignment

CVPR 2026

Novel View Synthesis (NVS) has traditionally relied on models with explicit 3D inductive biases combined with known camera parameters from Structure-from-Motion (SfM) beforehand. Recent vision foundation models like VGGT take an orthogonal approach -- 3D knowledge is gained implicitly through traini

Cited by 0SourcecodeScholar
2026

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving

CVPR 2026

Robust training and validation of Autonomous Driving Systems (ADS) require massive, diverse datasets. Proprietary data collected by Autonomous Vehicle (AV) fleets, while high-fidelity, are limited in scale, diversity of sensor configurations, as well as geographic and long-tail-behavioral coverage.

Cited by 0SourceScholar
2026

UFO-4D: Unposed Feedforward 4D reconstruction from Two Images

ICLR 2026poster

Dense 4D reconstruction from unposed images remains a critical challenge, with current methods relying on slow test-time optimization or fragmented, task-specific feedforward models. We introduce UFO-4D, a unified feedforward framework to reconstruct a dense, explicit 4D representation from just a p…

Cited by 0SourceScholar
2025

CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization

ICCV 2025poster

In dynamic 3D environments, accurately updating scene representations over time is crucial for applications in robotics, mixed reality, and embodied AI. As scenes evolve, efficient methods to incorporate changes are needed to maintain up-to-date, high-quality reconstructions without the computationa…

Cited by 0SourcePDFScholar
2025

DepthSplat: Connecting Gaussian Splatting and Depth

CVPR 2025poster

Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present DepthSplat to connect Gaussian splatting and depth estimation and study their interactions. More specifically, we first contribute a robust multi-view depth model by leveraging pre-train…

2025

Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views

CVPR 2025poster

Neural rendering has demonstrated remarkable success in high-quality 3D neural reconstruction and novel view synthesis with dense input views and accurate poses. However, applying it to sparse, unposed views in unbounded 360* scenes remains a challenging problem. In this paper, we propose a novel ne…

2025

LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering

NeurIPS 2025spotlight

In this work, we present a novel level-of-detail (LOD) method for 3D Gaussian Splatting that enables real-time rendering of large-scale scenes on memory-constrained devices. Our approach introduces a hierarchical LOD representation that iteratively selects optimal subsets of Gaussians based on camer…

Cited by 0SourceScholar
2025

No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images

ICLR 2025oral

We introduce NoPoSplat, a feed-forward model capable of reconstructing 3D scenes parameterized by 3D Gaussians from unposed sparse multi-view images. Our model, trained exclusively with photometric loss, achieves real-time 3D Gaussian reconstruction during inference. To eliminate the need for accura…

2025

Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation

CVPR 2025poster

Prompts play a critical role in unleashing the power of language and vision foundation models for specific tasks. For the first time, we introduce prompting into depth foundation models, creating a new paradigm for metric depth estimation termed Prompt Depth Anything. Specifically, we use a low-cost…

2025

SplatTalk: 3D VQA with Gaussian Splatting

ICCV 2025poster

Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D vision-language models (VLMs) have achieved remarkable success in 2D V…

Cited by 0SourcePDFScholar
2025

Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images

ICCV 2025poster

We present a system using Multimodal LLMs (MLLMs) to analyze a large database with tens of millions of images captured at different times, with the aim of discovering patterns in temporal changes. Specifically, we aim to capture frequent co-occurring changes ("trends") across a city over a certain p…

Cited by 0SourcePDFScholar
2025

WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments

CVPR 2025poster

We present WildGS-SLAM, a robust and efficient monocular RGB SLAM system designed to handle dynamic environments by leveraging uncertainty-aware geometric mapping. Unlike traditional SLAM systems, which assume static scenes, our approach integrates depth and uncertainty information to enhance tracki…

2024

NeRF On-the-go: Exploiting Uncertainty for Distractor-free NeRFs in the Wild

CVPR 2024poster

Neural Radiance Fields (NeRFs) have shown remarkable success in synthesizing photorealistic views from multi-view images of static scenes but face challenges in dynamic real-world environments with distractors like moving objects shadows and lighting changes. Existing methods manage controlled envir…

2024

Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization

ICLR 2024poster

Large foundation models are becoming ubiquitous, but training them from scratch is prohibitively expensive. Thus, efficiently adapting these powerful models to downstream tasks is increasingly important. In this paper, we study a principled finetuning paradigm -- Orthogonal Finetuning (OFT) -- for d…

Cited by 57SourcePDFScholar
2024

Renovating Names in Open-Vocabulary Segmentation Benchmarks

NeurIPS 2024poster

Names are essential to both human cognition and vision-language models. Open-vocabulary models utilize class names as text prompts to generalize to categories unseen during training. However, the precision of these names is often overlooked in existing datasets. In this paper, we address this undere…

2024

Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels

ECCV 2024poster

"Current 3D scene segmentation methods are heavily dependent on manually annotated 3D training datasets. Such manual annotations are labor-intensive, and often lack fine-grained details. Furthermore, models trained on this data typically struggle to recognize object classes beyond the annotated trai…

Cited by 34SourcePDFScholar
2024

Ternary-Type Opacity and Hybrid Odometry for RGB NeRF-SLAM

IROS 2024poster

In this work, we address the challenge of deploying Neural Radiance Field (NeRFs) in Simultaneous Localization and Mapping (SLAM) under the condition of lacking depth information, relying solely on RGB inputs. The key to unlocking the full potential of NeRF in such a challenging context lies in the…

Cited by 0SourceScholar
2024

WildGaussians: 3D Gaussian Splatting In the Wild

NeurIPS 2024poster

While the field of 3D scene reconstruction is dominated by NeRFs due to their photorealistic quality, 3D Gaussian Splatting (3DGS) has recently emerged, offering similar quality with real-time rendering speeds. However, both methods primarily excel with well-controlled 3D scenes, while in-the-wild d…

2023

DiffDreamer: Towards Consistent Unsupervised Single-view Scene Extrapolation with Conditional Diffusion Models

ICCV 2023poster

Scene extrapolation---the idea of generating novel views by flying into a given image---is a promising, yet challenging task. For each predicted frame, a joint inpainting and 3D refinement problem has to be solved, which is ill posed and includes a high level of ambiguity. Moreover, training data fo…

Cited by 38PDFcodeScholar
2023

OpenScene: 3D Scene Understanding With Open Vocabularies

CVPR 2023poster

Traditional 3D scene understanding approaches rely on labeled 3D datasets to train a model for a single task with supervision. We propose OpenScene, an alternative approach where a model predicts dense features for 3D scene points that are co-embedded with text and image pixels in CLIP feature space…

Cited by 341SourcePDFScholar
2022

MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction

NeurIPS 2022accept

In recent years, neural implicit surface reconstruction methods have become popular for multi-view 3D reconstruction. In contrast to traditional multi-view stereo methods, these approaches tend to produce smoother and more complete reconstructions due to the inductive smoothness bias of neural netwo…

Cited by 505SourcePDFScholar
2022

NICE-SLAM: Neural Implicit Scalable Encoding for SLAM

CVPR 2022poster

Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over-smoothed scene reconstructions and have difficulty scaling up to large scenes. These l…

Cited by 766PDFcodeScholar
2021

KiloNeRF: Speeding Up Neural Radiance Fields With Thousands of Tiny MLPs

ICCV 2021poster

NeRF synthesizes novel views of a scene with unprecedented quality by fitting a neural radiance field to RGB images. However, NeRF requires querying a deep Multi-Layer Perceptron (MLP) millions of times, leading to slow rendering times, even on modern GPUs. In this paper, we demonstrate that real-ti…

Cited by 855PDFcodeScholar
2021

Shape As Points: A Differentiable Poisson Solver

NeurIPS 2021oral

In recent years, neural implicit representations gained popularity in 3D reconstruction due to their expressiveness and flexibility. However, the implicit nature of neural implicit representations results in slow inference times and requires careful initialization. In this paper, we revisit the clas…

2021

UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction

ICCV 2021poster

Neural implicit 3D representations have emerged as a powerful paradigm for reconstructing surfaces from multi-view images and synthesizing novel views. Unfortunately, existing methods such as DVR or IDR require accurate per-pixel object masks as supervision. At the same time, neural radiance fields…

Cited by 862PDFcodeScholar
2020

Convolutional Occupancy Networks

ECCV 2020poster

Recently, implicit neural representations have gained popularity for learning-based 3D reconstruction. While demonstrating promising results, most implicit approaches are limited to comparably simple geometry of single objects and do not scale to more complicated or large-scale scenes. The key limit…

2020

DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere Tracing

CVPR 2020poster

We propose a differentiable sphere tracing algorithm to bridge the gap between inverse graphics methods and the recently proposed deep learning based implicit signed distance function. Due to the nature of the implicit function, the rendering process requires tremendous function queries, which is pa…

Cited by 350PDFcodeScholar
2019

Calibration Wizard: A Guidance System for Camera Calibration Based on Modelling Geometric and Corner Uncertainty

ICCV 2019oral

It is well known that the accuracy of a calibration depends strongly on the choice of camera poses from which images of a calibration object are acquired. We present a system -- Calibration Wizard -- that interactively guides a user towards taking optimal calibration images. For each new image to be…

Cited by 37PDFcodeScholar