← Search

Ryosuke Furuta

9 accepted papers

2026

EgoBrain: Synergizing Minds and Eyes For Human Action Understanding

ICLR 2026poster

The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human cognition and behavior from neural signals. In particular, the rise of multimodal AI models have brought new possibilities…

Cited by 0SourcecodeScholar
2026

Multi-speaker Attention Alignment for Multimodal Social Interaction

CVPR 2026

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures.While Multimodal Large Language Models (MLLMs) are natural candidates, simply adding visual inputs yields surprisingly inconsi

Cited by 0SourcecodeScholar
2025

Generative Modeling of Shape-Dependent Self-Contact Human Poses

ICCV 2025poster

One can hardly model self-contact of human poses without considering underlying body shapes. For example, the pose of rubbing a belly for a person with a low BMI leads to penetration of the hand into the belly for a person with a high BMI. Despite its relevance, existing self-contact datasets lack t…

Cited by 0SourcePDFScholar
2025

SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training

ICLR 2025poster

We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SiMHand. Pre-training with large-scale images achieves promising results in various tasks, but prior methods for 3D hand pose pre-training have not fully…

2024

ActionVOS: Actions as Prompts for Video Object Segmentation

ECCV 2024oral

"Delving into the realm of egocentric vision, the advancement of referring video object segmentation (RVOS) stands as pivotal in understanding human activities. However, existing RVOS task primarily relies on static attributes such as object names to segment target objects, posing challenges in dist…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2022

Domain Adaptive Hand Keypoint and Pixel Localization in the Wild

ECCV 2022poster

"We aim to improve the performance of regressing hand keypoints and segmenting pixel-level hand masks under new imaging conditions (e.g., outdoors) when we only have labeled images taken under very different conditions (e.g., indoors). In the real world, it is important that the model trained for bo…

Cited by 23SourcePDFScholar
2018

Cross-Domain Weakly-Supervised Object Detection Through Progressive Domain Adaptation

CVPR 2018poster

Can we detect common objects in a variety of image domains without instance-level annotations? In this paper, we present a framework for a novel task, cross-domain weakly supervised object detection, which addresses this question. For this paper, we have access to images with instance-level annotati…

2017

Object detection refinement using Markov random field based pruning and learning based rescoring

ICASSP 2017accepted

Contextual information such as the co-occurrence of objects and the location of objects has played an important role in object detection. We present candidate pruning and object rescoring methods that leverage contextual information and that can improve the state-of-the-art CNN-based object detectio…

Cited by 0SourceScholar