← Search

Hoang Tran Vuong

5 accepted papers

2026

MCW-KD: Multi-Cost Wasserstein Knowledge Distillation for Large Language Models

AAAI 2026technical

Knowledge distillation (KD) is widely recognized as an effective approach for compressing large language models (LLMs). However, standard KD methods often falter when confronted with architectural or tokenization heterogeneity between teacher and student models, which creates a mismatch in their rep

Cited by 0SourcePDFScholar
2025

HiCOT: Improving Neural Topic Models via Optimal Transport and Contrastive Learning

ACL 2025finding

Recent advances in neural topic models (NTMs) have improved topic quality but still face challenges: weak document-topic alignment, high inference costs due to large pretrained language models (PLMs), and limited modeling of hierarchical topic structures. To address these issues, we introduce HiCOT…

2025

Multi-Surrogate-Objective Optimization for Neural Topic Models

EMNLP 2025

Neural topic modeling has substantially improved topic quality and document topic distribution compared to traditional probabilistic methods. These models often incorporate multiple loss functions. However, the disparate magnitudes of these losses can make hyperparameter tuning for these loss functi

2025

Sharpness-Aware Minimization for Topic Models with High-Quality Document Representations

NAACL 2025long

Recent advanced frameworks in topic models have significantly enhanced the performance compared to conventional probabilistic approaches. Such models, mostly constructed from neural network architecture together with other advanced techniques such as contextual embedding, optimal transport distance…

2025

Token-Level Self-Play with Importance-Aware Guidance for Large Language Models

NeurIPS 2025poster

Leveraging the power of Large Language Models (LLMs) through preference optimization is crucial for aligning model outputs with human values. Direct Preference Optimization (DPO) has recently emerged as a simple yet effective method by directly optimizing on preference data without the need for expl…

Cited by 0SourceScholar