← Search

Mingfang Zhang

12 accepted papers

2026

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

CVPR 2026

Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing benchmarks rarely provide the fine-grained, grounded evidence needed to rigorously

Cited by 0SourceScholar
2026

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding

CVPR 2026

Current vision-language pre-training (VLP) paradigms excel at global scene understanding but struggle with instance-level reasoning due to global-only supervision. We introduce InstAP, an Instance-Aware Pre-training framework that jointly optimizes global vision-text alignment and fine-grained, inst

Cited by 0SourceScholar
2026

Multi-speaker Attention Alignment for Multimodal Social Interaction

CVPR 2026

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures.While Multimodal Large Language Models (MLLMs) are natural candidates, simply adding visual inputs yields surprisingly inconsi

Cited by 0SourcecodeScholar
2026

Parameter-Efficient Adaptation for MLLMs via Implicit Modality Decomposition

CVPR 2026

Parameter-efficient fine-tuning (PEFT) has become a compelling approach for adapting large language models (LLMs) into multimodal large language models (MLLMs), enabling them to handle diverse modalities with substantially lower memory and computational costs. However, most existing PEFT methods neg

Cited by 0SourcecodeScholar
2025

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance

ICCV 2025poster

This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individual within a 3D point cloud. Human inertial localization is challenging due to IM…

Cited by 0SourcePDFScholar
2025

SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training

ICLR 2025poster

We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SiMHand. Pre-training with large-scale images achieves promising results in various tasks, but prior methods for 3D hand pose pre-training have not fully…

2024

EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

CVPR 2024poster

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability we introduce EgoExoLearn a large-scale dataset that emulates the human demonstration following process in which ind…

2024

Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition

ECCV 2024poster

"Compared with visual signals, Inertial Measurement Units (IMUs) placed on human limbs can capture accurate motion signals while being robust to lighting variation and occlusion. While these characteristics are intuitively valuable to help egocentric action recognition, the potential of IMUs remains…

Cited by 9SourcePDFScholar
2024

Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation

CVPR 2024poster

The pursuit of accurate 3D hand pose estimation stands as a keystone for understanding human activity in the realm of egocentric vision. The majority of existing estimation methods still rely on single-view images as input leading to potential limitations e.g. limited field-of-view and ambiguity in…

2023

Structural Multiplane Image: Bridging Neural View Synthesis and 3D Reconstruction

CVPR 2023poster

The Multiplane Image (MPI), containing a set of fronto-parallel RGBA layers, is an effective and efficient representation for view synthesis from sparse inputs. Yet, its fixed structure limits the performance, especially for surfaces imaged at oblique angles. We introduce the Structural MPI (S-MPI),…

Cited by 10SourcePDFScholar
2020

Optical Flow in the Dark

CVPR 2020poster

Many successful optical flow estimation methods have been proposed, but they become invalid when tested in dark scenes because low-light scenarios are not considered when they are designed and current optical flow benchmark datasets lack low-light samples. Even if we preprocess to enhance the dark i…

Cited by 69PDFScholar