← Search

Nguyen Thi Ngoc Diep

7 accepted papers

2026

GloCTM: Cross-Lingual Topic Modeling via a Global Context Space

AAAI 2026technical

Cross-lingual topic modeling seeks to uncover coherent and semantically aligned topics across languages—a task central to multilingual understanding. Yet most existing models learn topics in disjoint, language-specific spaces and rely on alignment mechanisms (e.g., bilingual dictionaries) that often

Cited by 0SourcePDFScholar
2025

EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport Alignments

EMNLP 2025

Knowledge distillation (KD) is crucial for compressing large text embedding models, but faces challenges when teacher and student models use different tokenizers (Cross-Tokenizer KD - CTKD). Vocabulary mismatches impede the transfer of relational knowledge encoded in deep representations, such as hi

Cited by 0SourcePDFScholar
2025

Enhancing Discriminative Representation in Similar Relation Clusters for Few-Shot Continual Relation Extraction

NAACL 2025long

Few-shot Continual Relation Extraction (FCRE) has emerged as a significant challenge in information extraction, necessitating that relation extraction (RE) systems can sequentially identify new relations with limited labeled samples. While existing studies have demonstrated promising results in FCRE…

Cited by 0SourcePDFScholar
2025

MaGiX: A Multi-Granular Adaptive Graph Intelligence Framework for Enhancing Cross-Lingual RAG

EMNLP 2025

Retrieval-Augmented Generation (RAG) enhances large language models by grounding their outputs in external knowledge. Recent advances in Graph-based RAG (GRAG) frameworks, such as GraphRAG, LightRAG, and HippoRAG2, integrate knowledge graphs into the retrieval process to improve multi-hop reasoning

Cited by 0SourcePDFScholar
2025

Mitigating Non-Representative Prototypes and Representation Bias in Few-Shot Continual Relation Extraction

ACL 2025long

To address the phenomenon of similar classes, existing methods in few-shot continual relation extraction (FCRE) face two main challenges: non-representative prototypes and representation bias, especially when the number of available samples is limited. In our work, we propose Minion to address these…

Cited by 0SourcePDFScholar
2025

Mutual-pairing Data Augmentation for Fewshot Continual Relation Extraction

NAACL 2025long

Data scarcity is a major challenge in Few-shot Continual Relation Extraction (FCRE), where models must learn new relations from limited data while retaining past knowledge. Current methods, restricted by minimal data streams, struggle with catastrophic forgetting and overfitting. To overcome this, w…

Cited by 0SourcePDFScholar
2025

ToVo: Toxicity Taxonomy via Voting

NAACL 2025findings

Existing toxic detection models face significant limitations, such as lack of transparency, customization, and reproducibility. These challenges stem from the closed-source nature of their training data and the paucity of explanations for their evaluation mechanism. To address these issues, we propo…

Cited by 0SourcePDFScholar