← Search

Tianyu Huang

22 accepted papers

2026

AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing

ICML 2026spotlight

Recent studies have extended diffusion-based instruction-driven 2D image editing pipelines to 3D Gaussian Splatting (3DGS), enabling faithful manipulation of 3DGS assets and greatly advancing 3DGS content creation. However, it also exposes these assets to serious risks of unauthorized editing and ma…

Cited by 0SourceScholar
2026

DispViT: Direct Stereo Disparity Regression with a Single-Stream Vision Transformer

ICLR 2026poster

Deep stereo disparity estimation has long been dominated by a \textbf{matching-centric paradigm}, built on constructing cost volumes and iteratively refining local correspondences. Despite its success, this paradigm exhibits an intrinsic vulnerability: visual ambiguities from occlusion or non-Lamber…

Cited by 0SourcecodeScholar
2026

LLM-Guided Semantic Stereo Adaptive Visual Servoing for Precise Peg-In-Hole

ICRA 2026poster

Precision assembly tasks like peg-in-hole remain challenging for robotic manipulation. While visual servoing offers a robust framework, it depends heavily on accurate calibration and manual feature engineering. Learning-based methods, including vision-language models (VLMs), provide strong semantic …

Cited by 0Scholar
2026

ST-TPP: Learning Semi-Transductive Temporal Point Processes with Gromov-Wasserstein Barycentric Regularization

AAAI 2026technical

The generative mechanisms behind real-world event sequences are often heterogeneous, leading to data that possesses inherent clustering structures. However, most existing temporal point processes (TPPs) treat different event sequences independently, without leveraging the clustering structures when

Cited by 0SourcePDFScholar
2026

Towards Solving the Gilbert-Pollak Conjecture via Large Language Models

ICML 2026poster

The Gilbert-Pollak Conjecture, also known as the Steiner Ratio Conjecture, states that for any finite point set in the Euclidean plane, the Steiner minimum tree has length at least $\sqrt{3}/2 \approx 0.866$ times that of the Euclidean minimum spanning tree (the Steiner ratio). A sequence of improve…

Cited by 0SourceScholar
2026

VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the Wild

AAAI 2026technical

Reconstructing topologically consistent facial geometry is crucial for the digital avatar creation pipelines. Existing methods either require tedious manual efforts, lack generalization to in-the-wild data, or are constrained by the limited expressiveness of 3D Morphable Models. To address these lim

Cited by 0SourcePDFScholar
2025

6-DoF Shape Servoing of Deformable Objects in Co-Rotated Space of Modal Graph

ICRA 2025

Shape control of deformable objects under both rotational and translational deformations is important for versatile robotic applications. However, deformation control with full 6-degree-of-freedom (DoF) manipulation is an open problem, since modeling and describing rotational deformations lead to si

Cited by 0SourceScholar
2025

Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

CVPR 2025highlight

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video diffusion model to generate video depth conditioned on the i…

Cited by 14SourcePDFScholar
2025

DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors

AAAI 2025technical

Dynamic 3D interaction has been attracting a lot of attention recently. However, creating such 4D content remains challenging. One solution is to animate 3D scenes with physics-based simulation, which requires manually assigning precise physical properties to the object or the simulated results woul…

2025

TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels

NeurIPS 2025poster

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall short in separating the camera motion from foreground dynamic…

Cited by 0SourceScholar
2024

DreamControl: Control-Based Text-to-3D Generation with 3D Self-Prior

CVPR 2024poster

3D generation has raised great attention in recent years. With the success of text-to-image diffusion models the 2D-lifting technique becomes a promising route to controllable 3D generation. However these methods tend to present inconsistent geometry which is also known as the Janus problem. We obse…

2024

Efficient and Globally Optimal Camera Orientation Estimation With Line Correspondences

RA-L 2024

Given a set of outlier-contaminated 2D–3D line correspondences between the scene and a captured image, we aim to recover the absolute camera pose. This is a fundamental problem in computer vision and robotics, for which many methods have been developed and shown impressive performance, but they fail

Cited by 6SourceScholar
2024

Physically-Based Photometric Bundle Adjustment in Non-Lambertian Environments

IROS 2024poster

Photometric bundle adjustment (PBA) is widely used in estimating the camera pose and 3D geometry by assuming a Lambertian world. However, the assumption of photometric consistency is often violated since the non-diffuse reflection is common in real-world environments. The photometric inconsistency s…

Cited by 0SourceScholar
2024

Scalable 3D Registration via Truncated Entry-wise Absolute Residuals

CVPR 2024poster

Given an input set of 3D point pairs the goal of outlier-robust 3D registration is to compute some rotation and translation that align as many point pairs as possible. This is an important problem in computer vision for which many highly accurate approaches have been recently proposed. Despite their…

2024

TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields

ICLR 2024poster

Recent works learn 3D representation explicitly under text-3D guidance. However, limited text-3D data restricts the vocabulary scale and text control of generations. Generators may easily fall into a stereotype concept for certain text prompts, thus losing open-vocabulary generation ability. To tack…

Cited by 12SourcePDFScholar
2024

UniM2AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving

ECCV 2024poster

"Masked Autoencoders (MAE) play a pivotal role in learning potent representations, delivering outstanding results across various 3D perception tasks essential for autonomous driving. In real-world driving scenarios, it’s commonplace to deploy multiple sensors for comprehensive environment perception…

2023

CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-Training

ICCV 2023poster

Pre-training across 3D vision and language remains under development because of limited training data. Recent works attempt to transfer vision-language (V-L) pre-training methods to 3D vision. However, the domain gap between 3D and images is unsolved, so that V-L pre-trained models are restricted in…

Cited by 167PDFcodeScholar
2023

DDIT: Semantic Scene Completion via Deformable Deep Implicit Templates

ICCV 2023poster

Scene reconstructions are often incomplete due to occlusions and limited viewpoints. There have been efforts to use semantic information for scene completion. However, the completed shapes may be rough and imprecise since respective methods rely on 3D convolution and/or lack effective shape constrai…

Cited by 10PDFScholar
2023

Learning Accurate 3D Shape Based on Stereo Polarimetric Imaging

CVPR 2023poster

Shape from Polarization (SfP) aims to recover surface normal using the polarization cues of light. The accuracy of existing SfP methods is affected by two main problems. First, the ambiguity of polarization cues partially results in false normal estimation. Second, the widely-used assumption about o…

Cited by 12SourcePDFScholar
2023

Symmetry-Aware Transformer-Based Mirror Detection

AAAI 2023technical

Mirror detection aims to identify the mirror regions in the given input image. Existing works mainly focus on integrating the semantic features and structural features to mine specific relations between mirror and non-mirror regions, or introducing mirror properties like depth or chirality to help a…

2022

Learning Monocular Mesh Recovery of Multiple Body Parts Via Synthesis

ICASSP 2022accepted

In this paper, we focus on simultaneously recovering the 3D mesh of multiple body parts from a single RGB image. One of the main challenges is that available datasets with full-body 3D annotations are very limited. This results in poor generalization ability of existing learning-based methods. Exist…

Cited by 0SourceScholar