← Search

Thanh Tam Nguyen

5 accepted papers

2025

Localizing Before Answering: A Benchmark for Grounded Medical Visual Question Answering

IJCAI 2025

Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in

Cited by 0SourcePDFScholar
2025

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

NeurIPS 2025poster

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to…

Cited by 0SourceScholar
2023

Fast Yet Effective Speech Emotion Recognition with Self-Distillation

ICASSP 2023accepted

Speech emotion recognition (SER) is the task of recognising humans’ emotional states from speech. SER is extremely prevalent in helping dialogue systems to truly understand our emotions and become a trustworthy human conversational partner. Due to the lengthy nature of speech, SER also suffers from…

Cited by 0SourceScholar
2023

Knowledge Transfer for on-Device Speech Emotion Recognition With Neural Structured Learning

ICASSP 2023accepted

Speech emotion recognition (SER) has been a popular research topic in human-computer interaction (HCI). As edge devices are rapidly springing up, applying SER to edge devices is promising for a huge number of HCI applications. Although deep learning has been investigated to improve the performance o…

Cited by 0SourceScholar