← Search

Ruilong Li

15 accepted papers

2026

3DGS$^2$-TR: A Scalable Second-Order Trust-Region Method for 3D Gaussian Splatting

ICML 2026poster

We propose 3DGS$^2$-TR, a second-order optimizer for accelerating the scene training problem in 3D Gaussian Splatting (3DGS). Unlike existing second-order approaches that rely on explicit or dense curvature representations, such as 3DGS-LM (Höllein et al., 2025) or 3DGS2 (Lan et al., 2025), our meth…

Cited by 0SourceScholar
2026

SimULi: Real-Time LiDAR and Camera Simulation with Unscented Transforms

ICLR 2026poster

Rigorous testing of autonomous robots, such as self-driving vehicles, is essential to ensure their safety in real-world deployments. This requires building high-fidelity simulators to test scenarios beyond those that can be safely or exhaustively collected in the real-world. Existing neural renderin…

Cited by 0SourceScholar
2026

VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale

CVPR 2026

We present a scalable 3D reconstruction model that addresses a critical limitation in offline feed-forward methods: their computational and memory requirements grow quadratically w.r.t. the number of input images. Our approach is built on the key insight that this bottleneck stems from the varying-l

Cited by 0SourcecodeScholar
2023

NeRF-Det: Learning Geometry-Aware Volumetric Representation for Multi-View 3D Object Detection

ICCV 2023poster

We present NeRF-Det, a novel method for indoor 3D detection with posed RGB images as input. Unlike existing indoor 3D detection methods that struggle to model scene geometry, our method makes novel use of NeRF in an end-to-end manner to explicitly estimate 3D geometry, thereby improving 3D detection…

Cited by 51PDFcodeScholar
2022

Monocular Dynamic View Synthesis: A Reality Check

NeurIPS 2022accept

We study the recent progress on dynamic view synthesis (DVS) from monocular video. Though existing approaches have demonstrated impressive results, we show a discrepancy between the practical capture process and the existing experimental protocols, which effectively leaks in multi-view signals durin…

2022

TAVA: Template-Free Animatable Volumetric Actors

ECCV 2022poster

"Coordinate-based volumetric representations have the potential to generate photo-realistic virtual avatars from images. However, virtual avatars need to be controllable and be rendered in novel poses that may not have been observed. Traditional techniques, such as LBS, provide such a controlling fu…

2021

AI Choreographer: Music Conditioned 3D Dance Generation With AIST++

ICCV 2021poster

We present AIST++, a new multi-modal dataset of 3D dance motion and music, along with FACT, a Full-Attention Cross-modal Transformer network for generating 3D dance motion conditioned on music. The proposed AIST++ dataset contains 1.1M frames of 3D dance motion in 1408 sequences, covering 10 dance g…

Cited by 576PDFcodeScholar
2021

PlenOctrees for Real-Time Rendering of Neural Radiance Fields

ICCV 2021poster

We introduce a method to render Neural Radiance Fields (NeRFs) in real time using PlenOctrees, an octree-based 3D representation which supports view-dependent effects. Our method can render 800x800 images at more than 150 FPS, which is over 3000 times faster than conventional NeRFs. We do so without…

Cited by 1173PDFcodeScholar
2020

Learning Formation of Physically-Based Face Attributes

CVPR 2020poster

Based on a combined data set of 4000 high resolution facial scans, we introduce a non-linear morphable face model, capable of producing multifarious face geometry of pore-level resolution, coupled with material attributes for use in physically-based rendering. We aim to maximize the variety of the p…

Cited by 122PDFcodeScholar
2020

Monocular Real-Time Volumetric Performance Capture

ECCV 2020poster

We present the first approach to volumetric performance capture and novel-view rendering at real-time speed from monocular video, eliminating the need for expensive multi-view systems or cumbersome pre-acquisition of a personalized template model. Our system reconstructs a fully textured 3D human fr…

2019

Example-Guided Style-Consistent Image Synthesis From Semantic Labeling

CVPR 2019poster

Example-guided image synthesis aims to synthesize an image from a semantic label map and an exemplary image indicating style. We use the term "style" in this problem to refer to implicit characteristics of images, for example: in portraits "style" includes gender, racial identity, age, hairstyle;…

Cited by 102PDFcodeScholar
2019

Pose2Seg: Detection Free Human Instance Segmentation

CVPR 2019poster

The standard approach to image instance segmentation is to perform the object detection first, and then segment the object from the detection bounding-box. More recently, deep learning methods like Mask R-CNN perform them jointly. However, little research takes into account the uniqueness of the "hu…

Cited by 276PDFcodeScholar