← Search

Dayong Hu

3 accepted papers

2024

Prompting Large Language Models with Fine-Grained Visual Relations from Scene Graph for Visual Question Answering

ICASSP 2024accepted

Visual Question Answering (VQA) is a task that requires models to comprehend both questions and images. An increasing number of works are leveraging the strong reasoning capabilities of Large Language Models (LLMs) to address VQA. These methods typically utilize image captions as visual text descrip…

Cited by 0SourceScholar
2024

Segment then Match: Find the Carrier before Reasoning in Scene-Text VQA

ICASSP 2024accepted

Text-based Visual Question Answering (TextVQA) requires models to answer questions about the scene text in images by reasoning the context between the scene text and the question. Previous works demonstrated that clustering the scene text could help the model understand the context between different…

Cited by 0SourceScholar
2021

Improving Encoder by Auxiliary Supervision Tasks for Table-to-Text Generation

ACL 2021long

Table-to-text generation aims at automatically generating natural text to help people conveniently obtain salient information in tables. Although neural models for table-to-text have achieved remarkable progress, some problems are still overlooked. Previous methods cannot deduce the factual results…