← Search

Phu-Vinh Nguyen

2 accepted papers

2025

SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization

EMNLP 2025

Visual Language Models have demonstrated remarkable capabilities across various tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness in natural human-machine interactions. Moreover, the quality of language

Cited by 0SourcePDFScholar
2024

ViGLUE: A Vietnamese General Language Understanding Benchmark and Analysis of Vietnamese Language Models

NAACL 2024findings

As the number of language models has increased, various benchmarks have been suggested to assess the proficiency of the models in natural language understanding. However, there is a lack of such a benchmark in Vietnamese due to the difficulty in accessing natural language processing datasets or the…