← Search

Jiahe Li

22 accepted papers

2026

Assembling the Mind's Mosaic: Towards EEG Semantic Intent Decoding

ICLR 2026poster

Enabling natural communication through brain–computer interfaces (BCIs) remains one of the most profound challenges in neuroscience and neurotechnology. While existing frameworks offer partial solutions, they are constrained by oversimplified semantic representations and a lack of interpretability.…

Cited by 0SourceScholar
2026

CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual Localization

CVPR 2026

Scene Coordinate Regression (SCR) has emerged as a memory-efficient paradigm for visual localization. While SCR has demonstrated performance comparable to classic feature matching based approaches in small-scale scenes, it has consistently underperformed in large-scale environments. Large-scale loca

Cited by 0SourceScholar
2026

FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM

AAAI 2026technical

We present FoundationSLAM, a learning-based monocular dense SLAM system that addresses the absence of geometric consistency in previous flow-based approaches for accurate and robust tracking and mapping. Our core idea is to bridge flow estimation with geometric reasoning by leveraging the guidance f

Cited by 0SourcePDFScholar
2026

Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization

CVPR 2026

Recent advances in image-to-video (I2V) generation have achieved remarkable progress in synthesizing high-quality, temporally coherent videos from static images. Among all the applications of I2V, human-centric video generation includes a large portion. However, existing I2V models encounter difficu

Cited by 0SourcecodeScholar
2026

Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction

ICML 2026poster

Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have strictly bottlenecked existing approaches. This paper presents AmbiSuR, a framework that explores an intrinsic solution upon Gaussian Splatting for…

Cited by 0SourceScholar
2026

SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddings

CVPR 2026

Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual divergence of estimated scale over long sequences. Existing frame-to-frame methods achieve real-time performance through lo

Cited by 0SourceScholar
2026

SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction

AAAI 2026technical

Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this

Cited by 0SourcePDFScholar
2025

EVPGS: Enhanced View Prior Guidance for Splatting-based Extrapolated View Synthesis

CVPR 2025poster

Gaussian Splatting (GS)-based methods rely on sufficient training view coverage and perform synthesis on interpolated views. In this work, we tackle the more challenging and underexplored Extrapolated View Synthesis (EVS) task. Here we enable GS-based models trained with limited view coverage to gen…

Cited by 0SourcePDFScholar
2025

Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian Splatting

NeurIPS 2025poster

We present Eve3D, a novel framework for dense 3D reconstruction based on 3D Gaussian Splatting (3DGS). While most existing methods rely on imperfect priors derived from pre-trained vision models, Eve3D fully leverages these priors by jointly optimizing both them and the 3DGS backbone. This joint opt…

Cited by 0SourceScholar
2025

GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction

NeurIPS 2025spotlight

Reconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are increasingly constrained by representational bottlenecks. In this paper, we introduce GeoSVR, an explicit voxel-based fra…

Cited by 0SourcecodeScholar
2025

HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data Augmentation

AAAI 2025technical

Understanding of bimanual hand-object interaction plays an important role in robotics and virtual reality. However, due to significant occlusions between hands and object as well as the high degree-of-freedom motions, it is challenging to collect and annotate a high-quality, large-scale dataset, whi…

Cited by 0SourcePDFScholar
2025

InsTaG: Learning Personalized 3D Talking Head from Few-Second Video

CVPR 2025poster

Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces InsTaG, a 3D talking head synthesis framework that allows a f…

2025

Towards Feature-Consistent Parameter Collaboration for Personalized Federated Learning

ICASSP 2025accepted

Personalized federated learning (PFL) aims to improve the performance of the local model on each client with the non-IID data among different clients. This paper introduces FedFPC, a PFL method that allows effective and robust parameter-wise collaboration to achieve outperforming performance. Stem f…

Cited by 0SourceScholar
2024

CoR-GS: Sparse-View 3D Gaussian Splatting via Co-Regularization

ECCV 2024poster

"3D Gaussian Splatting (3DGS) creates a radiance field consisting of 3D Gaussians to represent a scene. With sparse training views, 3DGS easily suffers from overfitting, negatively impacting rendering. This paper introduces a new co-regularization perspective for improving sparse-view 3DGS. When tra…

2024

Con4m: Context-aware Consistency Learning Framework for Segmented Time Series Classification

NeurIPS 2024poster

Time Series Classification (TSC) encompasses two settings: classifying entire sequences or classifying segmented subsequences. The raw time series for segmented TSC usually contain Multiple classes with Varying Duration of each class (MVD). Therefore, the characteristics of MVD pose unique challenge…

2024

DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization

CVPR 2024poster

Radiance fields have demonstrated impressive performance in synthesizing novel views from sparse input views yet prevailing methods suffer from high training costs and slow inference speed. This paper introduces DNGaussian a depth-regularized framework based on 3D Gaussian radiance fields offering r…

2024

Fast Updating Truncated SVD for Representation Learning with Sparse Matrices

ICLR 2024poster

Updating truncated Singular Value Decomposition (SVD) has extensive applications in representation learning. The continuous evolution of massive-scaled data matrices in practical scenarios highlights the importance of aligning SVD-based models with fast-paced updates. Recent methods for updating tru…

Cited by 2SourcePDFScholar
2024

Robust Synthetic-to-Real Transfer for Stereo Matching

CVPR 2024poster

With advancements in domain generalized stereo matching networks models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However few studies have investigated the robustness after fine-tuning them in real-world scenarios during which the domain generalization ability ca…

2024

TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting

ECCV 2024poster

"Radiance fields have demonstrated impressive performance in synthesizing lifelike 3D talking heads. However, due to the difficulty in fitting steep appearance changes, the prevailing paradigm that presents facial motions by directly modifying point appearance may lead to distortions in dynamic regi…

2023

Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis

ICCV 2023poster

This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with small model size. Our idea is to explicitly exploit the unequal…

Cited by 84PDFcodeScholar
2023

STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition

ICCV 2023poster

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Seco…

Cited by 24PDFScholar
2023

Self-supervised Learning of Implicit Shape Representation with Dense Correspondence for Deformable Objects

ICCV 2023poster

Learning 3D shape representation with dense correspondence for deformable objects is a fundamental problem in computer vision. Existing approaches often need additional annotations of specific semantic domain, e.g., skeleton pose for human body or animals, which require extra annotation effort and s…

Cited by 9PDFScholar