← Search

Lifeng Fan

8 accepted papers

2026

Read the Room: Video Social Reasoning with Mental-Physical Causal Chains

ICLR 2026poster

``Read the room,'' or the ability to infer others' mental states from subtle social cues, is a hallmark of human social intelligence but remains a major challenge for current AI systems. Existing social reasoning datasets are limited in complexity, scale, and coverage of mental states, falling short…

Cited by 0SourcecodeScholar
2024

Learning Concept-Based Causal Transition and Symbolic Reasoning for Visual Planning

IROS 2024poster

Visual planning simulates how humans make decisions to achieve desired goals in the form of searching for visual causal transitions between an initial visual state and a final visual goal state. It has become increasingly important in egocentric vision with its advantages in guiding agents to perfor…

Cited by 1SourcecodeScholar
2022

Emergent Graphical Conventions in a Visual Communication Game

NeurIPS 2022accept

Humans communicate with graphical sketches apart from symbolic languages. Primarily focusing on the latter, recent studies of emergent communication overlook the sketches; they do not account for the evolution process through which symbolic sign systems emerge in the trade-off between iconicity and…

Cited by 19SourcePDFScholar
2021

Learning Triadic Belief Dynamics in Nonverbal Communication From Videos

CVPR 2021poster

Humans possess a unique social cognition capability; nonverbal communication can convey rich social information among agents. In contrast, such crucial social characteristics are mostly missing in the existing scene understanding literature. In this paper, we incorporate different nonverbal communic…

Cited by 27PDFcodeScholar
2020

Joint Inference of States, Robot Knowledge, and Human (False-)Beliefs

ICRA 2020poster

Aiming to understand how human (false-)belief— a core socio-cognitive ability—would affect human interactions with robots, this paper proposes to adopt a graphical model to unify the representation of object states, robot knowledge, and human (false-)beliefs. Specifically, a parse graph (pg) is lear…

Cited by 27SourceScholar
2019

Understanding Human Gaze Communication by Spatio-Temporal Graph Reasoning

ICCV 2019poster

This paper addresses a new problem of understanding human gaze communication in social videos from both atomic-level and event-level, which is significant for studying human social interactions. To tackle this novel and challenging problem, we contribute a large-scale video dataset, VACATION, which…

Cited by 145PDFcodeScholar