← Search

Zelun Luo

8 accepted papers

2025

Perception Tokens Enhance Visual Reasoning in Multimodal Language Models

CVPR 2025poster

Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit from depth estimation, and reasoning about 2D object instances benefits from object detection. Yet, MLMs can not produc…

2022

MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing

NeurIPS 2022accept

Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabilities. While complex human activities are often hierarchical and compositional,…

Cited by 23SourcePDFScholar
2021

MOMA: Multi-Object Multi-Actor Activity Parsing

NeurIPS 2021poster

Complex activities often involve multiple humans utilizing different objects to complete actions (e.g., in healthcare settings, physicians, nurses, and patients interact with each other and various medical devices). Recognizing activities poses a challenge that requires a detailed understanding of a…

Cited by 32SourcePDFScholar
2018

DF-Net: Unsupervised Joint Learning of Depth and Flow using Cross-Task Consistency

ECCV 2018poster

We present an unsupervised learning framework for simultaneously training single-view depth prediction and optical flow estimation models using unlabeled video sequences. Existing unsupervised methods often exploit brightness constancy and spatial smoothness priors to train depth or flow models. In…

2018

Graph Distillation for Action Detection with Privileged Modalities

ECCV 2018poster

We propose a technique that tackles action detection in multimodal videos under a realistic and challenging condition in which only limited training data and partially observed modalities are available. Common methods in transfer learning do not take advantage of the extra modalities potentially ava…

2017

Label Efficient Learning of Transferable Representations acrosss Domains and Tasks

NeurIPS 2017poster

We propose a framework that learns a representation transferable across different domains and tasks in a data efficient manner. Our approach battles domain shift with a domain adversarial loss, and generalizes the embedding to novel task using a metric learning-based approach. Our model is simultane…

Cited by 360SourcePDFScholar
2017

Unsupervised Learning of Long-Term Motion Dynamics for Videos

CVPR 2017poster

We present an unsupervised representation learning approach that compactly encodes the motion dependencies in videos. Given a pair of images from a video clip, our framework learns to predict the long-term 3D motions. To reduce the complexity of the learning framework, we propose to describe the mot…

Cited by 254PDFScholar