← Search

Khoa Luu

28 accepted papers

2026

$\phi$-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models

CVPR 2026

Fairness in Continual Learning for Large Multimodal Models (LMMs) is an emerging yet underexplored challenge, particularly in the presence of imbalanced data distributions that can lead to biased model updates and suboptimal performance across tasks. While recent continual learning studies have made

Cited by 0SourceScholar
2025

Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models

NeurIPS 2025poster

Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some fundamental limitations related to robustness and generalization due to the alignment and correlation between visual and…

Cited by 0SourceScholar
2025

FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene Understanding

CVPR 2025poster

Continual Learning in semantic scene segmentation aims to continually learn new unseen classes in dynamic environments while maintaining previously learned knowledge. Prior studies focused on modeling the catastrophic forgetting and background shift challenges in continual learning. However, fairnes…

Cited by 0SourcePDFScholar
2025

HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation

CVPR 2025poster

Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames. However, prior methods rely on pairwise connections, limiting their…

Cited by 1SourcePDFScholar
2025

MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning

NeurIPS 2025poster

Multimodal learning has gained much success in recent years. However, current multimodal fusion methods adopt the attention mechanism of Transformers to implicitly learn the underlying correlation of multimodal features. As a result, the multimodal model cannot capture the essential features of each…

Cited by 0SourceScholar
2024

CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial Videos

NeurIPS 2024poster

Video scene graph generation (VidSGG) has emerged as a transformative approach to capturing and interpreting the intricate relationships among objects and their temporal dynamics in video sequences. In this paper, we introduce the new AeroEye dataset that focuses on multi-object relationship modelin…

Cited by 3SourcePDFScholar
2024

DINTR: Tracking via Diffusion-based Interpolation

NeurIPS 2024poster

Object tracking is a fundamental task in computer vision, requiring the localization of objects of interest across video frames. Diffusion models have shown remarkable capabilities in visual generation, making them well-suited for addressing several requirements of the tracking problem. This work pr…

Cited by 0SourcePDFScholar
2024

EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view Understanding

NeurIPS 2024poster

Unsupervised Domain Adaptation has been an efficient approach to transferring the semantic segmentation model across data distributions. Meanwhile, the recent Open-vocabulary Semantic Scene understanding based on large-scale vision language models is effective in open-set settings because it can lea…

Cited by 1SourcePDFScholar
2024

HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding

CVPR 2024poster

Visual interactivity understanding within visual scenes presents a significant challenge in computer vision. Existing methods focus on complex interactivities while leveraging a simple relationship model. These methods however struggle with a diversity of appearance situation position interaction an…

Cited by 14SourcePDFScholar
2024

Insect-Foundation: A Foundation Model and Large-scale 1M Dataset for Visual Insect Understanding

CVPR 2024highlight

In precision agriculture the detection and recognition of insects play an essential role in the ability of crops to grow healthy and produce a high-quality yield. The current machine vision model requires a large volume of data to achieve high performance. However there are approximately 5.5 million…

Cited by 19SourcePDFScholar
2024

Z-GMOT: Zero-shot Generic Multiple Object Tracking

NAACL 2024findings

Despite recent significant progress, Multi-Object Tracking (MOT) faces limitations such as reliance on prior knowledge and predefined categories and struggles with unseen objects. To address these issues, Generic Multiple Object Tracking (GMOT) has emerged as an alternative approach, requiring less…

2023

FREDOM: Fairness Domain Adaptation Approach to Semantic Scene Understanding

CVPR 2023poster

Although Domain Adaptation in Semantic Scene Segmentation has shown impressive improvement in recent years, the fairness concerns in the domain adaptation have yet to be well defined and addressed. In addition, fairness is one of the most critical aspects when deploying the segmentation models into…

2023

Fairness Continual Learning Approach to Semantic Scene Understanding in Open-World Environments

NeurIPS 2023poster

Continual semantic segmentation aims to learn new classes while maintaining the information from the previous classes. Although prior studies have shown impressive progress in recent years, the fairness concern in the continual semantic segmentation needs to be better addressed. Meanwhile, fairness…

Cited by 14SourcePDFScholar
2023

Micron-BERT: BERT-Based Facial Micro-Expression Recognition

CVPR 2023poster

Micro-expression recognition is one of the most challenging topics in affective computing. It aims to recognize tiny facial movements difficult for humans to perceive in a brief period, i.e., 0.25 to 0.5 seconds. Recent advances in pre-training deep Bidirectional Transformers (BERT) have significant…

2023

Type-to-Track: Retrieve Any Object via Prompt-based Tracking

NeurIPS 2023poster

One of the recent trends in vision problems is to use natural language captions to describe the objects of interest. This approach can overcome some limitations of traditional methods that rely on bounding boxes or category annotations. This paper introduces a novel paradigm for Multiple Object Trac…

Cited by 24SourcePDFScholar
2022

DirecFormer: A Directed Attention in Transformer Approach to Robust Action Recognition

CVPR 2022poster

Human action recognition has recently become one ofthe popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in thetask of video action recognition with competitive results.However, these methods…

Cited by 75PDFcodeScholar
2021

BiMaL: Bijective Maximum Likelihood Approach to Domain Adaptation in Semantic Scene Segmentation

ICCV 2021poster

Semantic segmentation aims to predict pixel-level labels. It has become a popular task in various computer vision applications. While fully supervised segmentation methods have achieved high accuracy on large-scale vision datasets, they are unable to generalize on a new test environment or a new dom…

Cited by 43PDFcodeScholar
2021

Clusformer: A Transformer Based Clustering Approach to Unsupervised Large-Scale Face and Visual Landmark Recognition

CVPR 2021poster

The research in automatic unsupervised visual clustering has received considerable attention over the last couple years. It aims at explaining distributions of unlabeled visual images by clustering them via a parameterized model of appearance. Graph Convolutional Neural Networks (GCN) have recently…

Cited by 56PDFcodeScholar
2021

DyGLIP: A Dynamic Graph Model With Link Prediction for Accurate Multi-Camera Multiple Object Tracking

CVPR 2021poster

Multi-Camera Multiple Object Tracking (MC-MOT) is a significant computer vision problem due to its emerging applicability in several real-world applications. Despite a large number of existing works, solving the data association problem in any MC-MOT pipeline is arguably one of the most challenging…

Cited by 71PDFcodeScholar
2021

The Right To Talk: An Audio-Visual Transformer Approach

ICCV 2021poster

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker's utterances) remains a challenging task. Al…

Cited by 46PDFcodeScholar
2020

Vec2Face: Unveil Human Faces From Their Blackbox Features in Face Recognition

CVPR 2020oral

Unveiling face images of a subject given his/her high-level representations extracted from a blackbox Face Recognition engine is extremely challenging. It is because the limitations of accessible information from that engine including its structure and uninterpretable extracted features. This paper…

Cited by 62PDFScholar
2019

Automatic Face Aging in Videos via Deep Reinforcement Learning

CVPR 2019poster

This paper presents a novel approach for synthesizing automatically age-progressed facial images in video sequences using Deep Reinforcement Learning. The proposed method models facial structures and the longitudinal face-aging process of given subjects coherently across video frames. The approach i…

Cited by 46PDFScholar
2017

Faster Than Real-Time Facial Alignment: A 3D Spatial Transformer Network Approach in Unconstrained Poses

ICCV 2017poster

Facial alignment involves finding a set of landmark points on an image with a known semantic meaning. However, this semantic meaning of landmark points is often lost in 2D approaches where landmarks are either moved to visible boundaries or ignored as the pose of the face changes. In order to extrac…

Cited by 149PDFScholar
2017

Temporal Non-Volume Preserving Approach to Facial Age-Progression and Age-Invariant Face Recognition

ICCV 2017oral

Modeling the long-term facial aging process is extremely challenging due to the presence of large and non-linear variations during the face development stages. In order to efficiently address the problem, this work first decomposes the aging process into multiple short-term stages. Then, a novel gen…

Cited by 97PDFScholar
2016

Longitudinal Face Modeling via Temporal Deep Restricted Boltzmann Machines

CVPR 2016poster

Modeling the face aging process is a challenging task due to large and non-linear variations present in different stages of face development. This paper presents a deep model approach for face age progression that can efficiently capture the non-linear aging process and automatically synthesize a se…

Cited by 71PDFScholar
2015

Beyond Principal Components: Deep Boltzmann Machines for Face Modeling

CVPR 2015poster

The "interpretation through synthesis", i.e. Active Appearance Models (AAMs) method, has received considerable attention over the past decades. It aims at "explaining" face images by synthesizing them via a parameterized model of appearance. It is quite challenging due to appearance variations of hu…

Cited by 60SourcePDFScholar