← Search

Yutao Tang

5 accepted papers

2026

Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding

CVPR 2026

Recent advances in 3D vision-language models (VLMs) highlight a strong potential for 3D scene understanding and reasoning.However, effectively tokenizing 3D scenes into holistic scene tokens, and leveraging these tokens across diverse 3D understanding tasks, remain highly challenging. We present NDT

Cited by 0SourcecodeScholar
2025

MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild

NeurIPS 2025poster

In-the-wild photo collections often contain limited volumes of imagery and exhibit multiple appearances, e.g., taken at different times of day or seasons, posing significant challenges to scene reconstruction and novel view synthesis. Although recent adaptations of Neural Radiance Field (NeRF) and 3…

Cited by 0SourceScholar
2025

SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction

CVPR 2025poster

Recent efforts in Gaussian-Splat-based Novel View Synthesis can achieve photorealistic rendering; however, such capability is limited in sparse-view scenarios due to sparse initialization and over-fitting floaters. Recent progress in depth estimation and alignment can provide dense point cloud using…

2024

BAGS: Blur Agnostic Gaussian Splatting through Multi-Scale Kernel Modeling

ECCV 2024poster

"Recent efforts in using 3D Gaussians for scene reconstruction and novel view synthesis can achieve impressive results on curated benchmarks; however, images captured in real life are often blurry. In this work, we analyze the robustness of Gaussian-Splatting-based methods against various image blur…

2024

Sod-Uav: Small Object Detection For Unmanned Aerial Vehicle Images Via Improved Yolov7

ICASSP 2024accepted

Detecting small objects in Unmanned Aerial Vehicle (UAV) images is pivotal for a multitude of applications. Given the high-altitude perspective of UAVs, the images they capture often feature intricate backgrounds, pronounced object heterogeneity, and a plethora of sparsely situated small targets. Th…

Cited by 0SourceScholar