← Search

Vuong Le

8 accepted papers

2023

Persistent-Transient Duality: A Multi-Mechanism Approach for Modeling Human-Object Interaction

ICCV 2023poster

Humans are highly adaptable, swiftly switching between different modes to progressively handle different tasks, situations and contexts. In Human-object interaction (HOI) activities, these modes can be attributed to two mechanisms: (1) the large-scale consistent plan for the whole activity and (2) t…

Cited by 2PDFcodeScholar
2022

Video Dialog As Conversation about Objects Living in Space-Time

ECCV 2022poster

"It would be a technological feat to be able to create a system that can hold a meaningful conversation with humans about what they watch. A setup toward that goal is presented as a video dialog task, where the system is asked to generate natural utterances in response to a question in an ongoing di…

2021

Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering

IJCAI 2021poster

Video Question Answering (Video QA) is a powerful testbed to develop new AI capabilities. This task necessitates learning to reason about objects, relations, and events across visual and linguistic domains in space-time. High-level reasoning demands lifting from associative visual pattern recognitio…

2021

Learning Asynchronous and Sparse Human-Object Interaction in Videos

CVPR 2021poster

Human activities can be learned from video. With effective modeling it is possible to discover not only the action labels but also the temporal structure of the activities, such as the progression of the sub-activities. Automatically recognizing such structure from raw video signal is a new capabili…

Cited by 47PDFScholar
2020

Hierarchical Conditional Relation Networks for Video Question Answering

CVPR 2020oral

Video question answering (VideoQA) is challenging as it requires modeling capacity to distill dynamic visual artifacts and distant relations and to associate them with linguistic concepts. We introduce a general-purpose reusable neural unit called Conditional Relation Network (CRN) that serves as a…

Cited by 334PDFcodeScholar
2019

Learning Regularity in Skeleton Trajectories for Anomaly Detection in Videos

CVPR 2019poster

Appearance features have been widely used in video anomaly detection even though they contain complex entangled factors. We propose a new method to model the normal patterns of human movements in surveillance video for anomaly detection using dynamic skeleton features. We decompose the skeletal move…

Cited by 379PDFcodeScholar
2019

Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection

ICCV 2019poster

Deep autoencoder has been extensively used for anomaly detection. Training on the normal data, the autoencoder is expected to produce higher reconstruction error for the abnormal inputs than the normal ones, which is adopted as a criterion for identifying anomalies. However, this assumption does not…

Cited by 1738PDFScholar