← Search

Yoichi Sato

41 accepted papers

2026

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

CVPR 2026

Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing benchmarks rarely provide the fine-grained, grounded evidence needed to rigorously

Cited by 0SourceScholar
2026

EgoBrain: Synergizing Minds and Eyes For Human Action Understanding

ICLR 2026poster

The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human cognition and behavior from neural signals. In particular, the rise of multimodal AI models have brought new possibilities…

Cited by 0SourcecodeScholar
2026

HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

CVPR 2026

Hand-object interaction (HOI) involves dynamics where human manipulations produce spatio-temporal effects on objects. However, existing semantic HOI benchmarks focus on either manipulation or effects at a coarse level, lacking fine-grained spatio-temporal reasoning to capture HOI dynamics. We introd

Cited by 0SourceScholar
2026

Multi-speaker Attention Alignment for Multimodal Social Interaction

CVPR 2026

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures.While Multimodal Large Language Models (MLLMs) are natural candidates, simply adding visual inputs yields surprisingly inconsi

Cited by 0SourcecodeScholar
2025

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance

ICCV 2025poster

This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individual within a 3D point cloud. Human inertial localization is challenging due to IM…

Cited by 0SourcePDFScholar
2025

Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities

CVPR 2025poster

We address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze signals. While traditional methods rely on manually labeled mistakes, our approach does not require mistake annotations, hence overcoming the need of domain-specifi…

Cited by 3SourcePDFScholar
2025

Generative Modeling of Shape-Dependent Self-Contact Human Poses

ICCV 2025poster

One can hardly model self-contact of human poses without considering underlying body shapes. For example, the pose of rubbing a belly for a person with a low BMI leads to penetration of the hand into the belly for a person with a high BMI. Despite its relevance, existing self-contact datasets lack t…

Cited by 0SourcePDFScholar
2025

SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training

ICLR 2025poster

We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SiMHand. Pre-training with large-scale images achieves promising results in various tasks, but prior methods for 3D hand pose pre-training have not fully…

2024

Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects

ECCV 2024poster

"We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic understanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion generation. Accurately reconstructing such interactions in is c…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition

ECCV 2024poster

"Compared with visual signals, Inertial Measurement Units (IMUs) placed on human limbs can capture accurate motion signals while being robust to lighting variation and occlusion. While these characteristics are intuitively valuable to help egocentric action recognition, the potential of IMUs remains…

Cited by 9SourcePDFScholar
2024

Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation

CVPR 2024poster

The pursuit of accurate 3D hand pose estimation stands as a keystone for understanding human activity in the realm of egocentric vision. The majority of existing estimation methods still rely on single-view images as input leading to potential limitations e.g. limited field-of-view and ambiguity in…

2024

WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding

ECCV 2024poster

"In this paper, we address the challenge of fine-grained video event understanding in traffic scenarios, vital for autonomous driving and safety. Traditional datasets focus on driver or vehicle behavior, often neglecting pedestrian perspectives. To fill this gap, we introduce the WTS dataset, highli…

2023

DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-To-Fine Contrastive Ranking

CVPR 2023poster

Understanding dense action in videos is a fundamental challenge towards the generalization of vision models. Several works show that compositionality is key to achieving generalization by combining known primitive elements, especially for handling novel composited structures. Compositional temporal…

Cited by 15SourcePDFScholar
2023

Structural Multiplane Image: Bridging Neural View Synthesis and 3D Reconstruction

CVPR 2023poster

The Multiplane Image (MPI), containing a set of fronto-parallel RGBA layers, is an effective and efficient representation for view synthesis from sparse inputs. Yet, its fixed structure limits the performance, especially for surfaces imaged at oblique angles. We introduce the Structural MPI (S-MPI),…

Cited by 10SourcePDFScholar
2023

Weakly Supervised Temporal Sentence Grounding With Uncertainty-Guided Self-Training

CVPR 2023poster

The task of weakly supervised temporal sentence grounding aims at finding the corresponding temporal moments of a language description in the video, given video-language correspondence only at video-level. Most existing works select mismatched video-language pairs as negative samples and train the m…

Cited by 31SourcePDFScholar
2022

CompNVS: Novel View Synthesis with Scene Completion

ECCV 2022poster

"We introduce a scalable framework for novel view synthesis from RGB-D images with largely incomplete scene coverage. While generative neural approaches have demonstrated spectacular results on 2D images, they have not yet achieved similar photorealistic results in combination with scene completion…

Cited by 8SourcePDFScholar
2022

Domain Adaptive Hand Keypoint and Pixel Localization in the Wild

ECCV 2022poster

"We aim to improve the performance of regressing hand keypoints and segmenting pixel-level hand masks under new imaging conditions (e.g., outdoors) when we only have labeled images taken under very different conditions (e.g., indoors). In the real world, it is important that the model trained for bo…

Cited by 23SourcePDFScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Interact Before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action Recognition

CVPR 2022poster

Unsupervised domain adaptive video action recognition aims to recognize actions of a target domain using a model trained with only out-of-domain (source) annotations. The inherent complexity of videos makes this task challenging but also provides ground for leveraging multi-modal inputs (e.g., RGB,…

Cited by 49PDFScholar
2021

Unsupervised Common Particular Object Discovery and Localization by Analyzing a Match Graph

ICASSP 2021accepted

Although the unsupervised discovery and localization of common objects from within a set of images has received considerable attention, the difficulty of this task means that current methods are not sufficiently accurate. This paper describes an unsupervised method that more accurately discovers and…

Cited by 0SourceScholar
2020

Generalizing Hand Segmentation in Egocentric Videos With Uncertainty-Guided Model Adaptation

CVPR 2020poster

Although the performance of hand segmentation in egocentric videos has been significantly improved by using CNNs, it still remains a challenging issue to generalize the trained models to new domains, e.g., unseen environments. In this work, we solve the hand segmentation generalization problem witho…

Cited by 63PDFcodeScholar
2018

Future Person Localization in First-Person Videos

CVPR 2018poster

We present a new task that predicts future locations of people observed in first-person videos. Consider a first-person video stream continuously recorded by a wearable camera. Given a short clip of a person that is extracted from the complete stream, we aim to predict that person's location in futu…

2018

Predicting Gaze in Egocentric Video by Learning Task-dependent Attention Transition

ECCV 2018poster

We present a new computational model for gaze prediction in egocentric videos by exploring patterns in temporal shift of gaze fixations (attention transition) that are dependent on egocentric manipulation tasks. Our assumption is that the high-level context of how a task is completed in a certain wa…

2017

From RGB to Spectrum for Natural Scenes via Manifold-Based Mapping

ICCV 2017poster

Spectral analysis of natural scenes can provide much more detailed information about the scene than an ordinary RGB camera. The richer information provided by hyperspectral images has been beneficial to numerous applications, such as understanding natural environmental changes and classifying plants…

Cited by 125PDFScholar
2017

Privacy-Preserving Visual Learning Using Doubly Permuted Homomorphic Encryption

ICCV 2017poster

We propose a privacy-preserving framework for learning visual classifiers by leveraging distributed private image data. This framework is designed to aggregate multiple classifiers updated locally using private data and to ensure that no private information about the data is exposed during and after…

Cited by 70PDFScholar
2016

Exploiting Spectral-Spatial Correlation for Coded Hyperspectral Image Restoration

CVPR 2016poster

Conventional scanning and multiplexing techniques for hyperspectral imaging suffer from limited temporal and/or spatial resolution. To resolve this issue, coding techniques are becoming increasingly popular in developing snapshot systems for high-resolution hyperspectral imaging. For such systems,…

Cited by 121PDFScholar
2016

Hierarchical Gaussian Descriptor for Person Re-Identification

CVPR 2016poster

Describing the color and textural information of a person image is one of the most crucial aspects of person re-identification. In this paper, we present a novel descriptor based on a hierarchical distribution of pixel features. A hierarchical covariance descriptor has been successfully applied for…

Cited by 744PDFScholar
2016

Joint Recovery of Dense Correspondence and Cosegmentation in Two Images

CVPR 2016poster

We propose a new technique to jointly recover cosegmentation and dense per-pixel correspondence in two images. Our method parameterizes the correspondence field using piecewise similarity transformations and recovers a mapping between the estimated common "foreground" regions in the two images allow…

Cited by 126PDFcodeScholar
2016

Understanding Hand-Object Manipulation with Grasp Types and Object Attributes

RSS 2016poster

Our goal is to automate the understanding of natural hand-object manipulation by developing computer vision- based techniques. Our hypothesis is that it is necessary to model the grasp types of hands and the attributes of manipulated objects in order to accurately recognize manipulation actions. Spe…

Cited by 124SourcePDFScholar
2015

Adaptive Spatial-Spectral Dictionary Learning for Hyperspectral Image Denoising

ICCV 2015poster

Hyperspectral imaging is beneficial in a diverse range of applications from diagnostic medicine, to agriculture, to surveillance to name a few. However, hyperspectral images often times suffer from degradation due to the limited light, which introduces noise into the imaging process. In this paper,…

Cited by 49PDFScholar
2015

Illumination and Reflectance Spectra Separation of a Hyperspectral Image Meets Low-Rank Matrix Factorization

CVPR 2015poster

This paper addresses the illumination and reflectance spectra separation (IRSS) problem of a hyperspectral image captured under general spectral illumination. The huge amount of pixels in a hypersepctral image poses tremendous challenges on computational efficiency, yet in turn offers greater color…

Cited by 46SourcePDFScholar
2015

Separating Fluorescent and Reflective Components by Using a Single Hyperspectral Image

ICCV 2015poster

This paper introduces a novel method to separate fluorescent and reflective components in the spectral domain. In contrast to existing methods, which require to capture two or more images under varying illuminations, we aim to achieve this separation task by using a single hyperspectral image. After…

Cited by 12PDFScholar
2015

Uncalibrated Photometric Stereo Based on Elevation Angle Recovery From BRDF Symmetry of Isotropic Materials

CVPR 2015poster

This paper addresses the problem of uncalibrated photometric stereo with isotropic reflectances. Existing methods face difficulty in solving for the elevation angles of surface normals when the light sources only cover the visible hemisphere. Here, we introduce the notion of "constrained half-vector…

Cited by 36SourcePDFScholar