← Search

Yutong Chen

8 accepted papers

2025

EgoM2P: Egocentric Multimodal Multitask Pretraining

ICCV 2025accepted

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the camera wearer's actions, intentions, and surrounding environ…

Cited by 0SourcePDFScholar
2025

SplatFormer: Point Transformer for Robust 3D Gaussian Splatting

ICLR 2025spotlight

3D Gaussian Splatting (3DGS) has recently transformed photorealistic reconstruction, achieving high visual fidelity and real-time performance. However, rendering quality significantly deteriorates when test views deviate from the camera angles used during training, posing a major challenge for appli…

2024

Within the Dynamic Context: Inertia-aware 3D Human Modeling with Pose Sequence

ECCV 2024poster

"Neural rendering techniques have significantly advanced 3D human body modeling. However, previous approaches overlook dynamics induced by factors such as motion inertia, leading to challenges in scenarios where the pose remains static while the appearance changes, such as abrupt stops after spinnin…

Cited by 7SourcePDFScholar
2022

A Simple Multi-Modality Transfer Learning Baseline for Sign Language Translation

CVPR 2022poster

This paper proposes a simple transfer learning baseline for sign language translation. Existing sign language datasets (e.g. PHOENIX-2014T, CSL-Daily) contain only about 10K-20K pairs of sign videos, gloss annotations and texts, which are an order of magnitude smaller than typical parallel data for…

Cited by 182PDFcodeScholar
2022

Two-Stream Network for Sign Language Recognition and Translation

NeurIPS 2022accept

Sign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existing approaches directly encode RGB videos into hidden representations. RGB videos, however, are raw signals with substanti…