← Search

Thanh-Dat Truong

13 accepted papers

2026

$\phi$-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models

CVPR 2026

Fairness in Continual Learning for Large Multimodal Models (LMMs) is an emerging yet underexplored challenge, particularly in the presence of imbalanced data distributions that can lead to biased model updates and suboptimal performance across tasks. While recent continual learning studies have made

Cited by 0SourceScholar
2025

Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models

NeurIPS 2025poster

Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some fundamental limitations related to robustness and generalization due to the alignment and correlation between visual and…

Cited by 0SourceScholar
2025

FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene Understanding

CVPR 2025poster

Continual Learning in semantic scene segmentation aims to continually learn new unseen classes in dynamic environments while maintaining previously learned knowledge. Prior studies focused on modeling the catastrophic forgetting and background shift challenges in continual learning. However, fairnes…

Cited by 0SourcePDFScholar
2025

MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning

NeurIPS 2025poster

Multimodal learning has gained much success in recent years. However, current multimodal fusion methods adopt the attention mechanism of Transformers to implicitly learn the underlying correlation of multimodal features. As a result, the multimodal model cannot capture the essential features of each…

Cited by 0SourceScholar
2024

EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view Understanding

NeurIPS 2024poster

Unsupervised Domain Adaptation has been an efficient approach to transferring the semantic segmentation model across data distributions. Meanwhile, the recent Open-vocabulary Semantic Scene understanding based on large-scale vision language models is effective in open-set settings because it can lea…

Cited by 1SourcePDFScholar
2024

Insect-Foundation: A Foundation Model and Large-scale 1M Dataset for Visual Insect Understanding

CVPR 2024highlight

In precision agriculture the detection and recognition of insects play an essential role in the ability of crops to grow healthy and produce a high-quality yield. The current machine vision model requires a large volume of data to achieve high performance. However there are approximately 5.5 million…

Cited by 19SourcePDFScholar
2023

FREDOM: Fairness Domain Adaptation Approach to Semantic Scene Understanding

CVPR 2023poster

Although Domain Adaptation in Semantic Scene Segmentation has shown impressive improvement in recent years, the fairness concerns in the domain adaptation have yet to be well defined and addressed. In addition, fairness is one of the most critical aspects when deploying the segmentation models into…

2023

Fairness Continual Learning Approach to Semantic Scene Understanding in Open-World Environments

NeurIPS 2023poster

Continual semantic segmentation aims to learn new classes while maintaining the information from the previous classes. Although prior studies have shown impressive progress in recent years, the fairness concern in the continual semantic segmentation needs to be better addressed. Meanwhile, fairness…

Cited by 14SourcePDFScholar
2022

DirecFormer: A Directed Attention in Transformer Approach to Robust Action Recognition

CVPR 2022poster

Human action recognition has recently become one ofthe popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in thetask of video action recognition with competitive results.However, these methods…

Cited by 75PDFcodeScholar
2021

BiMaL: Bijective Maximum Likelihood Approach to Domain Adaptation in Semantic Scene Segmentation

ICCV 2021poster

Semantic segmentation aims to predict pixel-level labels. It has become a popular task in various computer vision applications. While fully supervised segmentation methods have achieved high accuracy on large-scale vision datasets, they are unable to generalize on a new test environment or a new dom…

Cited by 43PDFcodeScholar
2021

DyGLIP: A Dynamic Graph Model With Link Prediction for Accurate Multi-Camera Multiple Object Tracking

CVPR 2021poster

Multi-Camera Multiple Object Tracking (MC-MOT) is a significant computer vision problem due to its emerging applicability in several real-world applications. Despite a large number of existing works, solving the data association problem in any MC-MOT pipeline is arguably one of the most challenging…

Cited by 71PDFcodeScholar
2021

The Right To Talk: An Audio-Visual Transformer Approach

ICCV 2021poster

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker's utterances) remains a challenging task. Al…

Cited by 46PDFcodeScholar
2020

Vec2Face: Unveil Human Faces From Their Blackbox Features in Face Recognition

CVPR 2020oral

Unveiling face images of a subject given his/her high-level representations extracted from a blackbox Face Recognition engine is extremely challenging. It is because the limitations of accessible information from that engine including its structure and uninterpretable extracted features. This paper…

Cited by 62PDFScholar