← Search

Shugao Ma

17 accepted papers

2026

OSMO: Open-vocabulary Self-eMOtion Tracking

CVPR 2026

We introduce the novel task of egocentric self-emotion tracking, which aims to infer an individual's evolving emotions from egocentric multimodal streams such as voice, visual surroundings, semantic subtext, and eye-tracking signals. To establish this research direction, we present: (1) OSMO dataset

Cited by 0SourcecodeScholar
2025

HuMoCon: Concept Discovery for Human Motion Understanding

CVPR 2025poster

We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract semantically meaningful and generalizable features. HuMoCon add…

Cited by 0SourcePDFScholar
2025

LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand Diffusion.

CVPR 2025poster

Current research on generating 3D hand-object interaction motion primarily focuses on in-domain objects. Generalization to unseen objects is essential for practical applications, yet it remains both challenging and largely unexplored.In this paper, we propose LatentHOI, a novel approach designed to…

Cited by 0SourcePDFScholar
2025

Streaming VideoLLMs for Real-Time Procedural Video Understanding

ICCV 2025poster

We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens -- verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded…

Cited by 9SourcePDFScholar
2024

On the Utility of 3D Hand Poses for Action Recognition

ECCV 2024poster

"3D hand pose is an underexplored modality for action recognition. Poses are compact yet informative and can greatly benefit applications with limited compute budgets. However, poses alone offer an incomplete understanding of actions, as they cannot fully capture objects and environments with which…

2024

POET: Prompt Offset Tuning for Continual Human Action Adaptation

ECCV 2024oral

"As extended reality (XR) is redefining how users interact with computing devices, research in human action recognition is gaining prominence. Typically, models deployed on immersive computing devices are static and limited to their default set of classes. The goal of our research is to provide user…

2024

X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization

CVPR 2024poster

Lately there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However the adaptation of these models to egocentric videos has been largely unexplored. To address this gap we propose a sim…

2023

Data-Free Class-Incremental Hand Gesture Recognition

ICCV 2023poster

This paper investigates data-free class-incremental learning (DFCIL) for hand gesture recognition from 3D skeleton sequences. In this class-incremental learning (CIL) setting, while incrementally registering the new classes, we do not have access to the training samples (i.e. data-free) of t…

Cited by 9PDFcodeScholar
2023

Opening the Vocabulary of Egocentric Actions

NeurIPS 2023poster

Human actions in egocentric videos often feature hand-object interactions composed of a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations — sparsity of action compositions and a closed set of interacting objects. Thi…

2022

LiP-Flow: Learning Inference-Time Priors for Codec Avatars via Normalizing Flows in Latent Space

ECCV 2022poster

"Neural face avatars that are trained from multi-view data captured in camera domes can produce photo-realistic 3D reconstructions. However, at inference time, they must be driven by limited inputs such as partial views recorded by headset-mounted cameras or a front-facing camera, and sparse facial…

Cited by 1SourcePDFScholar
2020

Expressive Telepresence via Modular Codec Avatars

ECCV 2020poster

VR telepresence consists of interacting with another human in a virtual space represented by an avatar. Today most avatars are cartoon-like, but soon the technology will allow video-realistic ones. This paper aims in this direction and presents Modular Codec Avatars (MCA), a method to generate hyper…

Cited by 40SourcePDFScholar
2019

Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and Synthesis

ICCV 2019poster

We present a 16.2-million frame (50-hour) multimodal dataset of two-person face-to-face spontaneous conversations. Our dataset features synchronized body and finger motion as well as audio data. To the best of our knowledge, it represents the largest motion capture and audio dataset of natural conve…

Cited by 118PDFScholar
2015

Salient Object Subitizing

CVPR 2015poster

People can immediately and precisely identify 1, 2, 3 or 4 items by a simple glance. The phenomenon, known as Subitizing, inspires us to pursue the task of Salient Object Subitizing (SOS), i.e. predicting the existence and the number of salient objects in a scene using holistic cues. To study this p…

Cited by 138SourcePDFScholar