← Search

Linning Xu

21 accepted papers

2026

ARTDECO: Toward High-Fidelity On-the-Fly Reconstruction with Hierarchical Gaussian Structure and Feed-Forward Guidance

ICLR 2026poster

On-the-fly 3D reconstruction from monocular image sequences is a long-standing challenge in computer vision, critical for applications such as real-to-sim, AR/VR, and robotics. Existing methods face a major tradeoff: per-scene optimization yields high fidelity but is computationally expensive, where…

Cited by 0SourcecodeScholar
2026

DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos

RSS 2026poster

Data scarcity fundamentally limits the generalization of bimanual dexterous manipulation, as real-world data collection for dexterous hands is expensive and labor-intensive. Human manipulation videos, as a direct carrier of manipulation knowledge, offer significant potential for scaling up robot lea…

Cited by 0SourceScholar
2026

Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction

RSS 2026poster

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models lack the precision and metric consistency required for phys…

Cited by 0SourceScholar
2026

SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-Body Manipulation

ICML 2026poster

Simulating deformable objects under rich interactions remains a fundamental challenge for real-to-sim robot manipulation, with dynamics jointly driven by environmental effects and robot actions. Existing simulators rely on predefined physics or data-driven dynamics without robot-conditioned control,…

Cited by 0SourceScholar
2025

Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning

NeurIPS 2025poster

Realistic 3D indoor scene synthesis is vital for embodied AI and digital content creation. It can be naturally divided into two subtasks: object generation and layout generation. While recent generative models have significantly advanced object-level quality and controllability, layout generation re…

Cited by 0SourceScholar
2025

FlashGS: Efficient 3D Gaussian Splatting for Large-scale and High-resolution Rendering

CVPR 2025poster

Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated significant potential over traditional rendering techniques, attracting widespread attention from both industry and academia. However, real-time rendering with 3DGS remains a challenging problem, particularly in large-scale, high-reso…

2025

Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes

CVPR 2025poster

Seamless integration of both aerial and street view images remains a significant challenge in neural scene reconstruction and rendering. Existing methods predominantly focus on single domain, limiting their applications in immersive environments, which demand extensive free view exploration with lar…

Cited by 1SourcePDFScholar
2025

MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow Generation

NeurIPS 2025poster

Object compositing offers significant promise for augmented reality (AR) and embodied intelligence applications. Existing approaches predominantly focus on single-image scenarios or intrinsic decomposition techniques, facing challenges with multi-view consistency, complex scenes, and diverse lightin…

Cited by 0SourceScholar
2025

ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting

ICCV 2025poster

3D Gaussian Splatting is renowned for its high-fidelity reconstructions and real-time novel view synthesis, yet its lack of semantic understanding limits object-level perception. In this work, we propose ObjectGS, an object-aware framework that unifies 3D scene reconstruction with semantic understan…

2024

Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text

NeurIPS 2024poster

Recent advancements in 3D generation have leveraged synthetic datasets with ground truth 3D assets and predefined camera trajectories. However, the potential of adopting real-world datasets, which can produce significantly more realistic 3D scenes, remains largely unexplored. In this work, we delve…

2024

GSDF: 3DGS Meets SDF for Improved Neural Rendering and Reconstruction

NeurIPS 2024poster

Representing 3D scenes from multiview images remains a core challenge in computer vision and graphics, requiring both reliable rendering and reconstruction, which often conflicts due to the mismatched prioritization of image quality over precise underlying scene geometry. Although both neural implic…

Cited by 3SourcePDFScholar
2023

AssetField: Assets Mining and Reconfiguration in Ground Feature Plane Representation

ICCV 2023poster

Both indoor and outdoor environments are inherently structured and repetitive. Traditional modeling pipelines keep an asset library storing unique object templates, which is both versatile and memory efficient in practice. Inspired by this observation, we propose AssetField, a novel neural scene rep…

Cited by 12PDFScholar
2023

Grid-Guided Neural Radiance Fields for Large Urban Scenes

CVPR 2023poster

Purely MLP-based neural radiance fields (NeRF-based methods) often suffer from underfitting with blurred renderings on large-scale scenes due to limited model capacity. Recent approaches propose to geographically divide the scene and adopt multiple sub-NeRFs to model each region individually, leadin…

Cited by 94SourcePDFScholar
2023

MatrixCity: A Large-scale City Dataset for City-scale Neural Rendering and Beyond

ICCV 2023poster

Neural radiance fields (NeRF) and its subsequent variants have led to remarkable progress in neural rendering. While most of recent neural rendering works focus on objects and small-scale scenes, developing neural rendering methods for city-scale scenes is of great potential in many real-world appli…

Cited by 78PDFcodeScholar
2023

OmniCity: Omnipotent City Understanding With Multi-Level and Multi-View Images

CVPR 2023poster

This paper presents OmniCity, a new dataset for omnipotent city understanding from multi-level and multi-view images. More precisely, OmniCity contains multi-view satellite images as well as street-level panorama and mono-view images, constituting over 100K pixel-wise annotated images that are well-…

2022

BungeeNeRF: Progressive Neural Radiance Field for Extreme Multi-Scale Scene Rendering

ECCV 2022poster

"Neural Radiance Field (NeRF) has achieved outstanding performance in modeling 3D objects and controlled scenes, usually under a single scale. In this work, we focus on multi-scale cases where large changes in imagery are observed at drastically different scales. This scenario vastly exists in the r…

Cited by 267SourcePDFScholar
2021

BlockPlanner: City Block Generation With Vectorized Graph Representation

ICCV 2021poster

City modeling is the foundation for computational urban planning, navigation, and entertainment. In this work, we present the first generative model of city blocks named BlockPlanner, and showcase its ability to synthesize valid city blocks with varying land lots configurations. We propose a novel v…

Cited by 22PDFScholar
2020

A Local-to-Global Approach to Multi-Modal Movie Scene Segmentation

CVPR 2020poster

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of movies. This is very challenging - compared to the videos st…

Cited by 155PDFcodeScholar
2020

A Unified Framework for Shot Type Classification Based on Subject Centric Lens

ECCV 2020poster

In film making, shot has a profound influence on how the story is delivered and how the audiences are echoed. As different scale and movement types of shots can express different emotions and contents, recognizing shots and their attributes is important to the understanding of movies as well as gene…

Cited by 85SourcePDFScholar
2020

Learn to Propagate Reliably on Noisy Affinity Graphs

ECCV 2020poster

Recent works have shown that exploiting unlabeled data through label propagation can substantially reduce the labeling cost, which has been a critical issue in developing visual recognition models. Yet, how to propagate labels reliably, especially on a dataset with unknown outliers, remains an open…

Cited by 16SourcePDFScholar
2020

Online Multi-modal Person Search in Videos

ECCV 2020poster

The task of searching certain people in videos has seen increasing potential in real-world applications, such as video organization and editing. Most existing approaches are devised to work in an offline manner, where identifies can only be inferred after an entire video is examined. This working ma…

Cited by 35SourcePDFScholar