← Search

Feifei Zhang

8 accepted papers

2026

OAD-Promoter: Enhancing Zero-Shot VQA Using Large Language Models with Object Attribute Description

AAAI 2026technical

Large Language Models (LLMs) have become a crucial tool in Visual Question Answering (VQA) for handling knowledge-intensive questions in few-shot or zero-shot scenarios. However, their reliance on massive training datasets often causes them to inherit language biases during the acquisition of knowle

Cited by 0SourcePDFScholar
2025

Overcoming Dual Drift for Continual Long-Tailed Visual Question Answering

ICCV 2025poster

Visual Question Answering (VQA) is a widely explored multimodal task aimed at answering questions based on images. Recently, a few studies have started to investigate continual learning in VQA to cope with evolving multimodal data streams. However, these studies fall short of tackling another critic…

Cited by 0SourcePDFScholar
2025

When Open-Vocabulary Visual Question Answering Meets Causal Adapter: Benchmark and Approach

AAAI 2025technical

Visual Question Answering (VQA) is a multifaceted task that integrates computer vision and natural language processing to produce textual answers from images and questions. Existing VQA benchmarks predominantly adhere to a closed-set paradigm, limiting their ability to address arbitrary, unseen answ…

Cited by 0SourcePDFScholar
2023

VQACL: A Novel Visual Question Answering Continual Learning Setting

CVPR 2023poster

Research on continual learning has recently led to a variety of work in unimodal community, however little attention has been paid to multimodal tasks like visual question answering (VQA). In this paper, we establish a novel VQA Continual Learning setting named VQACL, which contains two key componen…

2018

Joint Pose and Expression Modeling for Facial Expression Recognition

CVPR 2018poster

Facial expression recognition (FER) is a challenging task due to different expressions under arbitrary poses. Most conventional approaches either perform face frontalization on a non-frontal facial image or learn separate classifiers for each pose. Different from existing methods, in this paper, we…

2016

Domain adaptation for speech emotion recognition by sharing priors between related source and target classes

ICASSP 2016accepted

In speech emotion recognition (SER), speech data is usually captured from different scenarios, which often leads to significant performance degradation due to the inherent mismatch between training and test set. To cope with this problem, we propose a domain adaptation method called Sharing Priors b…

Cited by 0SourceScholar