← Search

Shuai Fu

4 accepted papers

2025

Vision-text Enhancement Network For Weakly Supervised Video Anomaly Detection

ICASSP 2025accepted

The recent vision-language pre-training model ImageBind has shown significant success in a wide range of visual tasks, demonstrating excellent ability of joint embedding space across different modalities in visual or textual representation. A worthwhile problem is utilizing such a strong model for w…

Cited by 0SourceScholar
2024

GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning

NeurIPS 2024poster

Large Language Models (LLMs) are increasingly used for various tasks with graph structures. Though LLMs can process graph information in a textual format, they overlook the rich vision modality, which is an intuitive way for humans to comprehend structural information and conduct general graph reaso…

2024

Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models

ICLR 2024spotlight

With the prevalence of large-scale pretrained vision-language models (VLMs), such as CLIP, soft-prompt tuning has become a popular method for adapting these models to various downstream tasks. However, few works delve into the inherent properties of learnable soft-prompt vectors, specifically the im…

2023

Learning Retrieval Augmentation for Personalized Dialogue Generation

EMNLP 2023long main

Personalized dialogue generation, focusing on generating highly tailored responses by leveraging persona profiles and dialogue context, has gained significant attention in conversational AI applications. However, persona profiles, a prevalent setting in current personalized dialogue datasets, typica…

Cited by 0SourcecodeScholar