← Search

Chengyang Fang

3 accepted papers

2024

Prompting Large Language Models with Fine-Grained Visual Relations from Scene Graph for Visual Question Answering

ICASSP 2024accepted

Visual Question Answering (VQA) is a task that requires models to comprehend both questions and images. An increasing number of works are leveraging the strong reasoning capabilities of Large Language Models (LLMs) to address VQA. These methods typically utilize image captions as visual text descrip…

Cited by 0SourceScholar
2024

Segment then Match: Find the Carrier before Reasoning in Scene-Text VQA

ICASSP 2024accepted

Text-based Visual Question Answering (TextVQA) requires models to answer questions about the scene text in images by reasoning the context between the scene text and the question. Previous works demonstrated that clustering the scene text could help the model understand the context between different…

Cited by 0SourceScholar
2023

CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High Quality

ACL 2023long

There are three problems existing in the popular data-to-text datasets. First, the large-scale datasets either contain noise or lack real application scenarios. Second, the datasets close to real applications are relatively small in size. Last, current datasets bias in the English language while lea…

Cited by 2SourcePDFScholar