← Search

Van Nguyen

12 accepted papers

2026

CLIP-FMoE: Scalable CLIP via Fused Mixture-of-Experts with Enforced Specialization

ICLR 2026poster

Mixture-of-Experts (MoE) architectures have emerged as a promising approach for scaling deep learning models while maintaining computational efficiency. However, existing MoE adaptations for Contrastive Language-Image Pre-training (CLIP) models suffer from significant computational overhead during s…

Cited by 0SourceScholar
2026

MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources

ICML 2026poster

We present MixtureVitae, an open‑access pretraining corpus built to minimize legal risk while providing strong downstream performance. MixtureVitae follows a permissive‑first, risk‑mitigated sourcing strategy that combines public‑domain and permissively licensed text (e.g., CC‑BY/Apache) with carefu…

Cited by 0SourcecodeScholar
2026

SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images

CVPR 2026

Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface fluid migration. Understanding these phenomena is crucial for assessing hydrocarbon potential and avoiding drilling hazards

Cited by 0SourcecodeScholar
2025

AI2TALE: An Innovative Information Theory-based Approach for Learning to Localize Phishing Attacks

ICLR 2025poster

Phishing attacks remain a significant challenge for detection, explanation, and defense, despite over a decade of research on both technical and non-technical solutions. AI-based phishing detection methods are among the most effective approaches for defeating phishing attacks, providing predictions…

Cited by 0SourcePDFScholar
2025

Effective Context Modeling Framework for Emotion Recognition in Conversations

ICASSP 2025accepted

Emotion Recognition in Conversations (ERC) facilitates a deeper understanding of the emotions conveyed by speakers in each utterance within a conversation. Recently, Graph Neural Networks (GNNs) have demonstrated their strengths in capturing data relationships, particularly in contextual information…

Cited by 0SourceScholar
2025

EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba

ICCV 2025poster

Estimating human dance motion is a challenging task with various industrial applications. Recently, many efforts have focused on predicting human dance motion using either egocentric video or music as input. However, the task of jointly estimating human motion from both egocentric video and music re…

Cited by 0SourcePDFScholar
2025

More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning

ICCV 2025poster

Multi-label learning is a challenging computer vision task that requires assigning multiple categories to each image. However, fully annotating large-scale datasets is often impractical due to high costs and effort, motivating the study of learning from partially annotated data. In the extreme case…

Cited by 0SourcePDFScholar
2025

OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching

ACL 2025long

Text-to-speech (TTS) systems have seen significant advancements in recent years, driven by improvements in deep learning and neural network architectures. Viewing the output speech as a data distribution, previous approaches often employ traditional speech representations, such as waveforms or spect…

2023

An Additive Instance-Wise Approach to Multi-class Model Interpretation

ICLR 2023poster

Interpretable machine learning offers insights into what factors drive a certain prediction of a black-box system. A large number of interpreting methods focus on identifying explanatory input features, which generally fall into two main categories: attribution and selection. A popular attribution-b…

2022

Cycle class consistency with distributional optimal transport and knowledge distillation for unsupervised domain adaptation

UAI 2022poster

Unsupervised domain adaptation (UDA) aims to transfer knowledge from a model trained on a labeled source domain to an unlabeled target domain. To this end, we propose in this paper a novel cycle class-consistent model based on optimal transport (OT) and knowledge distillation. The model consists of…

Cited by 14SourcePDFScholar
2022

Multi-Task Voice Activated Framework Using Self-Supervised Learning

ICASSP 2022accepted

Self-supervised learning methods such as wav2vec 2.0 have shown promising results in learning speech representations from unlabelled and untranscribed speech data that are useful for speech recognition. Since these representations are learned without any task-specific supervision, they can also be u…

Cited by 0SourceScholar