← Search

Zhanning Gao

6 accepted papers

2025

Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving

ICCV 2025poster

In light of the dynamic nature of autonomous driving environments and stringent safety requirements, general MLLMs combined with CLIP alone often struggle to accurately represent driving-specific scenarios, particularly in complex interactions and long-tail cases. To address this, we propose the Hin…

Cited by 0SourcePDFScholar
2021

Noise-Resistant Deep Metric Learning With Ranking-Based Instance Selection

CVPR 2021poster

The existence of noisy labels in real-world data negatively impacts the performance of deep learning models. Although much research effort has been devoted to improving robustness to noisy labels in classification tasks, the problem of noisy labels in deep metric learning (DML) remains open. In this…

Cited by 52PDFcodeScholar
2020

An AI-empowered Visual Storyline Generator

IJCAI 2020poster

Video editing is currently a highly skill- and time-intensive process. One of the most important tasks in video editing is to compose the visual storyline. This paper outlines Visual Storyline Generator (VSG), an artificial intelligence (AI)-empowered system that automatically generates visual story…

Cited by 0SourcePDFScholar
2019

Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation Networks

ICCV 2019poster

Weakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video…

Cited by 135PDFcodeScholar
2018

Adding Attentiveness to the Neurons in Recurrent Neural Networks

ECCV 2018poster

Recurrent neural networks (RNNs) are capable of modeling the temporal dynamics of complex sequential information. However, the structures of existing RNN neurons mainly focus on controlling the contributions of current and historical information but do not explore the different importance levels of…

Cited by 105SourcePDFScholar
2017

ER3: A Unified Framework for Event Retrieval, Recognition and Recounting

CVPR 2017poster

We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames an…

Cited by 28PDFScholar