← Search

Honglu Zhou

14 accepted papers

2025

Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D

EMNLP 2025

Real-world decision-making often begins with identifying which modality contains the most relevant information for a given query. While recent multimodal models have made impressive progress in processing diverse inputs, it remains unclear whether they can reason contrastively across multiple modali

Cited by 0SourcePDFScholar
2025

Unifying Specialized Visual Encoders for Video Language Models

ICML 2025poster

Recent advances in vision backbones have yielded powerful and diverse visual and video encoders. Yet, current Video Large Language Models encode visual inputs using an encoder from a single backbone family, limiting the amount and type of visual information they can process. We propose MERV, a Multi…

2025

ViUniT: Visual Unit Tests for More Robust Visual Programming

CVPR 2025poster

Programming based approaches to reasoning tasks have substantially expanded the types of questions models can answer about visual scenes.Yet on benchmark visual reasoning data, when answering correctly, such models produce incorrect programs 33% of the time. These models are often right for the wron…

2024

Learning from Synthetic Human Group Activities

CVPR 2024poster

The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation we introduce M3…

2024

Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos

CVPR 2024poster

In this paper we explore the capability of an agent to construct a logical sequence of action steps thereby assembling a strategic procedural plan. This plan is crucial for navigating from an initial visual observation to a target visual outcome as depicted in real-life instructional videos. Existin…

2023

Harnessing Neighborhood Modeling and Asymmetry Preservation for Digraph Representation Learning

IJCAI 2023poster

Digraph Representation Learning aims to learn representations for directed homogeneous graphs (digraphs). Prior work is largely constrained or has poor generalizability across tasks. Most Graph Neural Networks exhibit poor performance on digraphs due to the neglect of modeling neighborhoods and pres…

Cited by 0SourcePDFScholar
2023

MSI: Maximize Support-Set Information for Few-Shot Segmentation

ICCV 2023poster

FSS (Few-shot segmentation) aims to segment a target class using a small number of labeled images (support set). To extract the information relevant to target class, a dominant approach in best performing FSS methods removes background features using a support mask. We observe that this feature exci…

Cited by 32PDFcodeScholar
2023

Procedure-Aware Pretraining for Instructional Video Understanding

CVPR 2023poster

Our goal is to learn a video representation that is useful for downstream procedure understanding tasks in instructional videos. Due to the small amount of available annotations, a key challenge in procedure understanding is to be able to extract from unlabeled videos the procedural knowledge such a…

2022

COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality

ECCV 2022poster

"Group Activity Recognition detects the activity collectively performed by a group of actors, which requires compositional reasoning of actors and objects. We approach the task by modeling the video as tokens that represent the multi-scale semantic concepts in the video. We propose COMPOSER, a Multi…

2022

HM: Hybrid Masking for Few-Shot Segmentation

ECCV 2022poster

"We study few-shot semantic segmentation that aims to segment a target object from a query image when provided with a few annotated support images of the target class. Several recent methods resort to a feature masking (FM) technique to discard irrelevant feature activations which eventually facilit…

2022

Harnessing Fourier Isovists and Geodesic Interaction for Long-Term Crowd Flow Prediction

IJCAI 2022poster

With the rise in popularity of short-term Human Trajectory Prediction (HTP), Long-Term Crowd Flow Prediction (LTCFP) has been proposed to forecast crowd movement in large and complex environments. However, the input representations, models, and datasets for LTCFP are currently limited. To this end,…

2021

Hopper: Multi-hop Transformer for Spatiotemporal Reasoning

ICLR 2021poster

This paper considers the problem of spatiotemporal object-centric reasoning in videos. Central to our approach is the notion of object permanence, i.e., the ability to reason about the location of objects as they move through the video while being occluded, contained or carried by other objects. Exi…

2020

HID: Hierarchical Multiscale Representation Learning for Information Diffusion

IJCAI 2020poster

Multiscale modeling has yielded immense success on various machine learning tasks. However, it has not been properly explored for the prominent task of information diffusion, which aims to understand how information propagates along users in online social networks. For a specific user, whether and w…

2020

Laying the Foundations of Deep Long-Term Crowd Flow Prediction

ECCV 2020poster

Predicting the crowd behavior in complex environments is a key requirement for crowd and disaster management, architectural design, and urban planning. Given a crowd's immediate state, current approaches must be successively repeated over multiple time-steps for long-term predictions, leading to com…