← Search

Zi-Yuan Hu

4 accepted papers

2025

Fine-grained Spatiotemporal Grounding on Egocentric Videos

ICCV 2025poster

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite its growing importance in applications such as augmented re…

2024

Beyond Embeddings: The Promise of Visual Table in Visual Reasoning

EMNLP 2024main

Visual representation learning has been a cornerstone in computer vision, involving typical forms such as visual embeddings, structural symbols, and text-based representations. Despite the success of CLIP-type visual embeddings, they often lack access to world knowledge critical for visual reasoning…

2024

Enhancing Temporal Modeling of Video LLMs via Time Gating

EMNLP 2024finding

Video Large Language Models (Video LLMs) have achieved impressive performance on video-and-language tasks, such as video question answering. However, most existing Video LLMs neglect temporal information in video data, leading to struggles with temporal-aware video understanding. To address this gap…

2023

VL-PET: Vision-and-Language Parameter-Efficient Tuning via Granularity Control

ICCV 2023poster

As the model size of pre-trained language models (PLMs) grows rapidly, full fine-tuning becomes prohibitively expensive for model training and storage. In vision-and-language (VL), parameter-efficient tuning (PET) techniques are proposed to integrate modular modifications (e.g., Adapter) into encode…

Cited by 19PDFcodeScholar