← Search

Erwei Yin

18 accepted papers

2026

MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs

CVPR 2026

3D hand pose estimation (HPE) from sparse inertial measurement units (IMUs) has shown great potential in human-computer interaction. However, due to the significant semantic gap between sparse local motion information and structured global pose information, estimating hand poses from sparse IMU sign

Cited by 0SourceScholar
2026

MHED-SLAM: Multi-Scale Hybrid Encoding-Based Decoupled SLAM

AAAI 2026technical

Neural Radiance Fields (NeRF)-based Visual Simultaneous Localization and Mapping (SLAM) achieve superior scene geometric modeling and robust camera tracking by leveraging neural representations. Existing methods typically relied on multi-resolution hash encoding with truncated signed distance field

Cited by 0SourcePDFScholar
2026

OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition

CVPR 2026

Online micro gesture recognition from hand skeletons is critical for VR/AR interaction but faces challenges due to limited public datasets and task-specific algorithms. Micro gestures involve subtle motion patterns, which make constructing datasets with precise skeletons and frame-level annotations

Cited by 0SourceScholar
2026

PURIFICATION BEFORE FUSION: TOWARD MASK-FREE SPEECH ENHANCEMENT FOR ROBUST AUDIO-VISUAL SPEECH RECOGNITION

ICASSP 2026poster

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse interference into the feature fusion process. To mitigate this, rece…

Cited by 0SourcePDFScholar
2026

Seeing the Unseen: Physics-as-Representation for Generalizable Gaze Perception

ICML 2026poster

We introduce physics-as-representation, a learning paradigm that encodes physical structure and geometric laws into visual representations, enabling models to see the unseen—the underlying 3D geometry and motion dynamics not apparent in raw pixels. We instantiate this paradigm in gaze perception by …

Cited by 0SourceScholar
2026

UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

CVPR 2026

Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between input images and rendered outputs, or multi-view consistency constraints, as the driving force to optimize networks and p

Cited by 0SourceScholar
2025

A Multi-Prior Fusion Network for Video-based Micro-Expression Recognition

ICASSP 2025accepted

The analysis of facial micro-expressions (MEs) has emerged as a significant application and topic within the field of image and video processing. However, challenges persist due to the brief duration and subtle intensity of these spontaneous expressions. This paper presents a novel multi-prior fusio…

Cited by 0SourceScholar
2025

De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimation

CVPR 2025poster

3D Gaze estimation is a challenging task due to two main issues. First, existing methods focus on analyzing dense features (e.g., large pixel regions), which are sensitive to local noise (e.g., light spots, blurs) and result in increased computational complexity. Second, an eyeball model can corresp…

Cited by 0SourcePDFScholar
2025

Hierarchical-aware Orthogonal Disentanglement Framework for Fine-grained Skeleton-based Action Recognition

ICCV 2025poster

In recent years, skeleton-based action recognition has gained significant attention due to its robustness in varying environmental conditions. However, most existing methods struggle to distinguish fine-grained actions due to subtle motion features, minimal inter-class variation, and they often fail…

Cited by 0SourcePDFScholar
2025

Learning Neural Vocoder from Range-Null Space Decomposition

IJCAI 2025

Despite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F) domain-based neural vocoder to resolve the above-mentioned chal

2025

LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition

ICASSP 2025accepted

Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have significantly enhanced the performance of lip reading models. Despi…

Cited by 0SourceScholar
2025

Lunar Twins: We Choose to Go to the Moon with Large Language Models

ACL 2025finding

In recent years, the rapid advancement of large language models (LLMs) has significantly reshaped the landscape of scientific research. While LLMs have achieved notable success across various domains, their application in specialized fields such as lunar exploration remains underdeveloped, and their…

2025

M2EIT: Multi-Domain Mixture of Experts for Robust Neural Inertial Tracking

ICCV 2025poster

Inertial tracking (IT), independent of the environment and external infrastructure, has long been the ideal solution for providing location services to humans. Despite significant strides in inertial tracking empowered by deep learning, prevailing neural inertial tracking predominantly utilizes conv…

Cited by 0SourcePDFScholar
2025

Self-supervised Contrastive Pre-training for Dry Electrode EEG Emotion Recognition via Cross Device Representation Consistency

ICASSP 2025accepted

The use of dry electrode electroencephalography (EEG) systems holds significant importance in advancing the everyday application of emotion recognition. However, adapting it to real-world applications faces unique challenges due to low signal-to-noise ratios and unreliable emotion labels. To address…

Cited by 0SourceScholar
2024

Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization

COLING 2024main

Lip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip reading in cross-speaker scenarios where the speaker identity…

Cited by 1SourcePDFScholar
2024

Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments

RA-L 2024

The task of vision-and-language navigation in continuous environments (VLN-CE) aims at training an autonomous agent to perform low-level actions to navigate through 3D continuous surroundings using visual observations and language instructions. The significant potential of VLN-CE for mobile robots h

Cited by 26SourceScholar
2023

Grounded Entity-Landmark Adaptive Pre-Training for Vision-and-Language Navigation

ICCV 2023oral

Cross-modal alignment is one key challenge for Vision-and-Language Navigation (VLN). Most existing studies concentrate on mapping the global instruction or single sub-instruction to the corresponding trajectory. However, another critical problem of achieving fine-grained alignment at the entity leve…

Cited by 21PDFcodeScholar