← Search

Yusuke Sugano

7 accepted papers

2026

Robust Long-Term Test-Time Adaptation for 3D Human Pose Estimation Through Motion Discretization

AAAI 2026technical

Online test-time adaptation addresses the train-test domain gap by adapting the model on unlabeled streaming test inputs before making the final prediction. However, online adaptation for 3D human pose estimation suffers from error accumulation when relying on self-supervision with imperfect predict

Cited by 0SourcePDFScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Interact Before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action Recognition

CVPR 2022poster

Unsupervised domain adaptive video action recognition aims to recognize actions of a target domain using a model trained with only out-of-domain (source) annotations. The inherent complexity of videos makes this task challenging but also provides ground for leveraging multi-modal inputs (e.g., RGB,…

Cited by 49PDFScholar
2018

Light Structure from Pin Motion: Simple and Accurate Point Light Calibration for Physics-based Modeling

ECCV 2018poster

We present a practical method for geometric point light source calibration. Unlike in prior works that use Lambertian spheres, mirror spheres, or mirror planes, our calibration target consists of a Lambertian plane and small shadow casters at unknown positions above the plane. Due to their small siz…

Cited by 8SourcePDFScholar
2015

Rendering of Eyes for Eye-Shape Registration and Gaze Estimation

ICCV 2015poster

Images of the eye are key in several computer vision problems, such as shape registration and gaze estimation. Recent large-scale supervised methods for these problems require time-consuming data collection and manual annotation, which can be unreliable. We propose synthesizing perfectly labelled ph…

Cited by 422PDFScholar