← Search

Yang-Tian Sun

12 accepted papers

2026

Dynamic Important Example Mining for Reinforcement Finetuning

CVPR 2026

Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and used. Most data-centric RFT methods rely on static or heuristic sample selection, implicitly assuming a sample's value is f

Cited by 0SourcecodeScholar
2026

Repurposing 3D Generative Model for Autoregressive Layout Generation

CVPR 2026

We introduce LaviGen, a framework that repurposes 3D generative models for 3D layout generation. Unlike previous methods that infer object layouts from textual descriptions, LaviGen operates directly in the native 3D space, formulating layout generation as an autoregressive process that explicitly m

Cited by 0SourcecodeScholar
2026

Stabilizing Streaming Video Geometry via Dynamic Feature Normalization

CVPR 2026

Consistent 3D geometry estimation from streaming RGB input is crucial for real-world applications such as autonomous driving, embodied AI, and large-scale reconstruction. While modern monocular geometry foundation models achieve strong single-image accuracy, they exhibit severe temporal inconsistenc

Cited by 0SourcecodeScholar
2026

Stereo World Model: Camera-Guided Stereo Video Generation

CVPR 2026

We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry dire

Cited by 0SourcecodeScholar
2025

Deformable Radial Kernel Splatting

CVPR 2025poster

Recently, Gaussian splatting has emerged as a robust technique for representing 3D scenes, enabling real-time rasterization and high-fidelity rendering. However, Gaussians' inherent radial symmetry and smoothness constraints limit their ability to represent complex shapes, often requiring thousands…

Cited by 1SourcePDFScholar
2024

Let the Avatar Talk using Texts without Paired Training Data

ECCV 2024poster

"This paper introduces text-driven talking avatar generation, a task that uses text to instruct both the generation and animation of an avatar. One significant obstacle in this task is the absence of paired text and talking avatar data for model training, limiting data-driven methodologies. To this…

Cited by 0SourcePDFScholar
2024

SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes

CVPR 2024poster

Novel view synthesis for dynamic scenes is still a challenging problem in computer vision and graphics. Recently Gaussian splatting has emerged as a robust technique to represent static scenes and enable high-quality and real-time novel view synthesis. Building upon this technique we propose a new r…

2024

Spec-Gaussian: Anisotropic View-Dependent Appearance for 3D Gaussian Splatting

NeurIPS 2024poster

The recent advancements in 3D Gaussian splatting (3D-GS) have not only facilitated real-time rendering through modern GPU rasterization pipelines but have also attained state-of-the-art rendering quality. Nevertheless, despite its exceptional rendering quality and performance on standard datasets, 3…

Cited by 45SourcePDFScholar
2024

Splatter a Video: Video Gaussian Representation for Versatile Processing

NeurIPS 2024poster

Video representation is a long-standing problem that is crucial for various downstream tasks, such as tracking, depth prediction, segmentation, view synthesis, and editing. However, current methods either struggle to model complex motions due to the absence of 3D structure or rely on implicit 3D rep…

Cited by 7SourcePDFScholar
2024

Total-Decom: Decomposed 3D Scene Reconstruction with Minimal Interaction

CVPR 2024highlight

Scene reconstruction from multi-view images is a fundamental problem in computer vision and graphics. Recent neural implicit surface reconstruction methods have achieved high-quality results; however editing and manipulating the 3D geometry of reconstructed scenes remains challenging due to the abse…

2022

NeRF-Editing: Geometry Editing of Neural Radiance Fields

CVPR 2022poster

Implicit neural rendering, especially Neural Radiance Field (NeRF), has shown great potential in novel view synthesis of a scene. However, current NeRF-based methods cannot enable users to perform user-controlled shape deformation in the scene. While existing works have proposed some approaches to m…

Cited by 285PDFScholar