← Search

Ying He

56 accepted papers

2026

AquaSplatting: A Hybrid 3D Representation for Robust Underwater Scene Reconstruction via Dual-Branch Rendering

AAAI 2026technical

While 3D Gaussian Splatting (3DGS) excels at real-time rendering of standard scenes, it struggles to reconstruct underwater environments due to severe challenges such as light scattering, color attenuation, and sparse coverage of Gaussian kernels in far-field aqueous regions. To address this, we int

Cited by 0SourcePDFScholar
2026

Focusing: View-Consistent Sparse Voxels for Efficient 3D VAE

ICML 2026poster

High-fidelity 3D generation remains difficult. Although some methods have proposed converting raw meshes to SDFs, it remains a lossy process. TripoSF presented a VAE training paradigm based on a rendering loss to circumvent this lossy SDF conversion, achieving high-precision surface reconstruction. …

Cited by 0SourceScholar
2026

ForeDiffusion: Foresight-Conditioned Diffusion Policy via Future View Construction for Robot Manipulation

AAAI 2026technical

Diffusion strategies have advanced visual motor control by progressively denoising high-dimensional action sequences, providing a promising method for robot manipulation. However, as task complexity increases, the success rate of existing baseline models decreases considerably. Analysis indicates th

Cited by 0SourcePDFScholar
2026

From Extrinsic to Intrinsic: Geodesic-Guided Representation Learning for 3D Geometric Data

ICML 2026poster

Geometric analysis fundamentally distinguishes between extrinsic and intrinsic perspectives. The dominant paradigm in current 3D representation learning relies on either extrinsic spatial structures or high-level semantics, struggling to capture the essence of shape identity and underlying manifold …

Cited by 0SourceScholar
2026

GAUSSIAN2SCENE: 3D SCENE REPRESENTATION LEARNING VIA SELF-SUPERVISED LEARNING WITH 3D GAUSSIAN SPLATTING

ICASSP 2026poster

Self-supervised learning (SSL) for point cloud pre-training has become a cornerstone for many 3D vision tasks, enabling effective learning from large-scale unannotated data. At the scene level, existing SSL methods often incorporate volume rendering into the pre-training framework, using RGB-D image…

Cited by 0SourcePDFScholar
2026

LiteGE: Lightweight Geodesic Embedding for Efficient Geodesics Computation and Non-Isometric Shape Correspondence

AAAI 2026technical

Computing geodesic distances on 3D surfaces is fundamental to many tasks in 3D vision and geometry processing, with deep connections to tasks such as shape correspondence. Recent learning-based methods achieve strong performance but rely on large 3D backbones, leading to high memory usage and latenc

Cited by 0SourcePDFScholar
2026

Metric–-Phase Fields: Decoupling Distance and Sign for Thin-Structure Reconstruction from Unoriented Point Clouds

ICML 2026poster

Neural Signed Distance Functions (SDFs) excel at reconstructing watertight manifolds but fail on thin structures and open boundaries due to strict inside-outside constraints. Conversely, Unsigned Distance Fields (UDFs) accommodate general geometries but suffer from gradient singularities at the zero…

Cited by 0SourceScholar
2026

MonoCloth: Reconstruction and Animation of Cloth-Decoupled Human Avatars from Monocular Videos

AAAI 2026technical

Reconstructing realistic 3D human avatars from monocular videos is a challenging task due to the limited geometric information and complex non-rigid motion involved. We present MonoCloth, a new method for reconstructing and animating clothed human avatars from monocular videos. To overcome the limit

Cited by 0SourcePDFScholar
2026

PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization

AAAI 2026technical

Reinforcement Fine-tuning (RFT) methods such as Group Relative Policy Optimization (GRPO) have demonstrated strong capabilities in aligning Large Language Models with human preferences. However, these approaches often suffer from limited data efficiency, necessitating extensive on-policy rollouts to

Cited by 0SourcePDFScholar
2026

SCENERAG: SCENE-LEVEL RETRIEVAL-AUGMENTED GENERATION FOR VIDEO UNDERSTANDING

ICASSP 2026poster

Despite recent advances in retrieval-augmented generation (RAG) for video understanding, effectively understanding long-form video content remains underexplored due to the vast scale and high complexity of video data. Current RAG approaches typically segment videos into fixed-length chunks, which of…

Cited by 0SourcePDFScholar
2026

Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation

ICML 2026poster

Test-time policy adaptation for multi-turn interactions (T$^2$PAM) is essential for aligning Large Language Models (LLMs) with dynamic user needs during inference time. However, existing paradigms commonly treat test-time adaptation as a single-axis problem, either purely refining instructions (Prom…

Cited by 0SourceScholar
2025

3DMambaIPF: A State Space Model for Iterative Point Cloud Filtering via Differentiable Rendering

AAAI 2025technical

Noise is an inevitable aspect of point cloud acquisition, necessitating filtering as a fundamental task within the realm of 3D vision. Existing learning-based filtering methods have shown promising capabilities on commonly used datasets. Nonetheless, the effectiveness of these methods is constrained…

2025

A Lightweight UDF Learning Framework for 3D Reconstruction Based on Local Shape Functions

CVPR 2025poster

Unsigned distance fields (UDFs) provide a versatile framework for representing a diverse array of 3D shapes, encompassing both watertight and non-watertight geometries. Traditional UDF learning methods typically require extensive training on large 3D shape datasets, which is costly and necessitates…

2025

DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language Models

ICASSP 2025accepted

Inderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing,…

Cited by 0SourceScholar
2025

DMF-Net: Image-Guided Point Cloud Completion with Dual-Channel Modality Fusion and Shape-Aware Upsampling Transformer

AAAI 2025technical

In this paper we study the task of a single-view image-guided point cloud completion. Existing methods have got promising results by fusing the information of image into point cloud explicitly or implicitly. However, given that the image has global shape information and the partial point cloud has r…

Cited by 2SourcePDFScholar
2025

Details Enhancement in Unsigned Distance Field Learning for High-fidelity 3D Surface Reconstruction

AAAI 2025technical

While Signed Distance Fields (SDF) are well-established for modeling watertight surfaces, Unsigned Distance Fields (UDF) broaden the scope to include open surfaces and models with complex inner structures. Despite their flexibility, UDFs encounter significant challenges in high-fidelity 3D reconstru…

Cited by 0SourcePDFScholar
2025

Do Not DeepFake Me: Privacy-Preserving Neural 3D Head Reconstruction Without Sensitive Images

AAAI 2025technical

While 3D head reconstruction is widely used for modeling, existing neural reconstruction approaches rely on high-resolution multi-view images, posing notable privacy issues. Individuals are particularly sensitive to facial features, and facial image leakage can lead to many malicious activities, suc…

Cited by 0SourcePDFScholar
2025

Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction

IJCAI 2025

Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive obje

2025

Inverse Rendering using Multi-Bounce Path Tracing and Reservoir Sampling

ICLR 2025poster

We introduce MIRReS, a novel two-stage inverse rendering framework that jointly reconstructs and optimizes explicit geometry, materials, and lighting from multi-view images. Unlike previous methods that rely on implicit irradiance fields or oversimplified ray tracing, our method begins with an initi…

Cited by 0SourcePDFScholar
2025

JAM: Keypoint-Guided Joint Prediction after Classification-Aware Marginal Proposal for Multi-Agent Interaction

IROS 2025

Predicting the future motion of road participants is a critical task in autonomous driving. In this work, we address the challenge of low-quality generation of low-probability modes in multi-agent joint prediction. To tackle this issue, we propose a two-stage multi-agent interactive prediction frame

Cited by 0SourcecodeScholar
2025

MGSR: 2D/3D Mutual-boosted Gaussian Splatting for High-fidelity Surface Reconstruction under Various Light Conditions

ICCV 2025poster

Novel view synthesis (NVS) and surface reconstruction (SR) are essential tasks in 3D Gaussian Splatting (3DGS). Despite recent progress, these tasks are often addressed independently, with GS-based rendering methods struggling under diverse light conditions and failing to produce accurate surfaces,…

2025

MIND: Material Interface Generation from UDFs for Non-Manifold Surface Reconstruction

NeurIPS 2025poster

Unsigned distance fields (UDFs) are widely used in 3D deep learning due to their ability to represent shapes with arbitrary topology. While prior work has largely focused on learning UDFs from point clouds or multi-view images, extracting meshes from UDFs remains challenging, as the learned fields r…

Cited by 0SourcecodeScholar
2025

Resource Allocation for Semantic Segmentation Tasks in Autonomous Driving: A Likelihood Active Inference Approach

ICASSP 2025accepted

The latest Segment Anything Model enables realtime scene annotation and understanding for autonomous driving systems, enhancing driving safety. However, effectively allocating resources for real-time performance and accuracy remains challenging in edge-cloud architectures. Traditional reinforcement…

Cited by 0SourceScholar
2025

SFDM: Robust Decomposition of Geometry and Reflectance for Realistic Face Rendering from Sparse-view Images

CVPR 2025poster

In this study, we introduce a novel two-stage technique for decomposing and reconstructing facial features from sparse-view images, a task made challenging by the unique geometry and complex skin reflectance of each individual. To synthesize 3D facial models more realistically, we endeavor to decoup…

Cited by 0SourcePDFScholar
2025

UNIS: A Unified Framework for Achieving Unbiased Neural Implicit Surfaces in Volume Rendering

ICCV 2025poster

Reconstruction from multi-view images is a fundamental challenge in computer vision that has been extensively studied over the past decades. Recently, neural radiance fields have driven significant advancements, especially through methods using implicit functions and volume rendering, achieving high…

Cited by 0SourcePDFScholar
2025

You Should Learn to Stop Denoising on Point Clouds in Advance

AAAI 2025technical

Point clouds have become the preferred data format for a variety of tasks in 3D vision and graphics. However, raw point clouds often contain significant noise. This paper introduces the Adaptive Stop Denoising Network (ASDN), a novel approach aimed at restoring high-quality point clouds from noisy d…

2024

2S-UDF: A Novel Two-stage UDF Learning Method for Robust Non-watertight Model Reconstruction from Multi-view Images

CVPR 2024poster

Recently building on the foundation of neural radiance field various techniques have emerged to learn unsigned distance fields (UDF) to reconstruct 3D non-watertight models from multi-view images. Yet a central challenge in UDF-based volume rendering is formulating a proper way to convert unsigned d…

2024

A Language-Driven Navigation Strategy Integrating Semantic Maps and Large Language Models

IROS 2024poster

Accurate perception of semantic and spatial information is crucial for robots performing language-driven navigation tasks. Existing approaches utilize visual-language models to extract semantic information from the environment and construct maps. However, constrained by the generalization and accura…

Cited by 0SourceScholar
2024

Flatten Anything: Unsupervised Neural Surface Parameterization

NeurIPS 2024poster

Surface parameterization plays an essential role in numerous computer graphics and geometry processing applications. Traditional parameterization approaches are designed for high-quality meshes laboriously created by specialized 3D modelers, thus unable to meet the processing demand for the current…

2024

From Transparent to Opaque: Rethinking Neural Implicit Surfaces with $\alpha$-NeuS

NeurIPS 2024poster

Traditional 3D shape reconstruction techniques from multi-view images, such as structure from motion and multi-view stereo, face challenges in reconstructing transparent objects. Recent advances in neural radiance fields and its variants primarily address opaque or transparent objects, encountering…

2024

LLaKey: Follow My Basic Action Instructions to Your Next Key State

IROS 2024poster

In 3D object manipulation, collecting expert data for end-to-end imitation learning becomes a mainstream method. Though successful, previous works neglect the guiding role of language in action execution. These methods lack the understanding of action semantics, in which multiple action sequences ar…

Cited by 0SourceScholar
2024

Large Motion Model for Unified Multi-Modal Motion Generation

ECCV 2024poster

"Human motion generation, a cornerstone technique in animation and video production, has widespread applications in various tasks like text-to-motion and music-to-dance. Previous works focus on developing specialist models tailored for each task without scalability. In this work, we present Large Mo…

Cited by 27SourcePDFScholar
2024

OTOcc: Optimal Transport for Occupancy Prediction

IJCAI 2024poster

The autonomous driving community is highly interested in 3D occupancy prediction due to its outstanding geometric perception and object recognition capabilities. However, previous methods are limited to existing semantic conversion mechanisms for solving sparse ground truths problem, causing excessi…

2024

O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion Model

AAAI 2024technical

Occlusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based in-painting model, to reconstruct complete surfaces for the hidd…

2024

PP-TIL: Personalized Planning for Autonomous Driving with Instance-based Transfer Imitation Learning

IROS 2024poster

Personalized motion planning holds significant importance within urban automated driving, catering to the unique requirements of individual users. Nevertheless, prior endeavors have frequently encountered difficulties in simultaneously addressing two crucial aspects: personalized planning within int…

Cited by 0SourcecodeScholar
2023

Bagging R-CNN: Ensemble for Object Detection in Complex Traffic Scenes

ICASSP 2023accepted

Generic object detection methods have achieved preferable results, but it is still challenging to detect objects from complicated traffic scenes like extreme illumination and adverse weather. The existing methods are not robust enough to be extended to new complex traffic scenes. To address this iss…

Cited by 0SourceScholar
2023

Deformable Model-Driven Neural Rendering for High-Fidelity 3D Reconstruction of Human Heads Under Low-View Settings

ICCV 2023poster

Reconstructing 3D human heads in low-view settings presents technical challenges, mainly due to the pronounced risk of overfitting with limited views and high-frequency signals. To address this, we propose geometry decomposition and adopt a two-stage, coarse-to-fine training strategy, allowing for p…

Cited by 9PDFcodeScholar
2023

GeoUDF: Surface Reconstruction from 3D Point Clouds via Geometry-guided Distance Representation

ICCV 2023poster

We present a learning-based method, namely GeoUDF, to tackle the long-standing and challenging problem of reconstructing a discrete surface from a sparse point cloud. To be specific, we propose a geometry-guided learning method for UDF and its gradient estimation that explicitly formulates the unsig…

Cited by 27PDFcodeScholar
2023

IterativePFN: True Iterative Point Cloud Filtering

CVPR 2023poster

The quality of point clouds is often limited by noise introduced during their capture process. Consequently, a fundamental 3D vision task is the removal of noise, known as point cloud filtering or denoising. State-of-the-art learning based methods focus on training neural networks to infer filtered…

2023

NeuroGF: A Neural Representation for Fast Geodesic Distance and Path Queries

NeurIPS 2023poster

Geodesics play a critical role in many geometry processing applications. Traditional algorithms for computing geodesics on 3D mesh models are often inefficient and slow, which make them impractical for scenarios requiring extensive querying of arbitrary point-to-point geodesics. Recently, deep impli…

2023

OPE-SR: Orthogonal Position Encoding for Designing a Parameter-Free Upsampling Module in Arbitrary-Scale Image Super-Resolution

CVPR 2023poster

Arbitrary-scale image super-resolution (SR) is often tackled using the implicit neural representation (INR) approach, which relies on a position encoding scheme to improve its representation ability. In this paper, we introduce orthogonal position encoding (OPE), an extension of position encoding, a…

2023

RePaint-NeRF: NeRF Editting via Semantic Masks and Diffusion Models

IJCAI 2023poster

The emergence of Neural Radiance Fields (NeRF) has promoted the development of synthesized high-fidelity views of the intricate real world. However, it is still a very demanding task to repaint the content in NeRF. In this paper, we propose a novel framework that can take RGB images as input and alt…

2022

Audio-Driven Stylized Gesture Generation with Flow-Based Model

ECCV 2022poster

"Generating stylized audio-driven gestures for robots and virtual avatars has attracted increasing considerations recently. Existing methods require style labels (e.g. speaker identities), or complex preprocessing of the data to obtain style control parameters. In this paper, we propose a new end-to…

Cited by 28SourcePDFScholar
2022

IDEA-Net: Dynamic 3D Point Cloud Interpolation via Deep Embedding Alignment

CVPR 2022poster

This paper investigates the problem of temporally interpolating dynamic 3D point clouds with large non-rigid deformation. We formulate the problem as estimation of point-wise trajectories (i.e., smooth curves) and further reason that temporal irregularity and under-sampling are two major challenges.…

Cited by 23PDFcodeScholar
2022

Multi-Constraint Deep Reinforcement Learning for Smooth Action Control

IJCAI 2022poster

Deep reinforcement learning (DRL) has been studied in a variety of challenging decision-making tasks, e.g., autonomous driving. \textcolor{black}{However, DRL typically suffers from the action shaking problem, which means that agents can select actions with big difference even though states only sli…

2021

CorrNet3D: Unsupervised End-to-End Learning of Dense Correspondence for 3D Point Clouds

CVPR 2021poster

Motivated by the intuition that one can transform two aligned point clouds to each other more easily and meaningfully than a misaligned pair, we propose CorrNet3D -the first unsupervised and end-to-end deep learning-based framework - to drive the learning of dense correspondence between 3D shapes by…

Cited by 97PDFcodeScholar
2020

PUGeo-Net: A Geometry-centric Network for 3D Point Cloud Upsampling

ECCV 2020poster

In this paper, we propose a novel deep neural network based method, called PUGeo-Net, for upsampling 3D point clouds. PUGeo-Net incorporates discrete differential geometry into deep learning elegantly by learning the first and second fundamental forms that are able to fully represent the local geome…

2019

Fast Computation of Content-Sensitive Superpixels and Supervoxels Using Q-Distances

ICCV 2019poster

State-of-the-art researches model the data of images and videos as low-dimensional manifolds and generate superpixels/supervoxels in a content-sensitive way, which is achieved by computing geodesic centroidal Voronoi tessellation (GCVT) on manifolds. However, computing exact GCVTs is slow due to com…

Cited by 21PDFScholar
2017

Sparse representation for colors of 3D point cloud via virtual adaptive sampling

ICASSP 2017accepted

Sparse signal representation has proven to be an extremely powerful tool in a wide range of engineering applications. However, most of the existing techniques are designed for regular data (such as audio signals and images/videos) that uniformly lies in regular Euclidian spaces. This paper aims at e…

Cited by 0SourceScholar