← Search

Yiyi Chen

7 accepted papers

2026

SPICE: Submodular Penalized Information–Conflict Selection for Efficient Large Language Model Training

ICLR 2026poster

Information-based data selection for instruction tuning is compelling: maximizing the log-determinant of the Fisher information yields a monotone submodular objective, enabling greedy algorithms to achieve a $(1-1/e)$ approximation under a cardinality budget. In practice, however, we identify allevi…

Cited by 0SourceScholar
2025

ALGEN: Few-shot Inversion Attacks on Textual Embeddings via Cross-Model Alignment and Generation

ACL 2025long

With the growing popularity of Large Language Models (LLMs) and vector databases, private textual data is increasingly processed and stored as numerical embeddings. However, recent studies have proven that such embeddings are vulnerable to inversion attacks, where original text is reconstructed to r…

2025

Against All Odds: Overcoming Typology, Script, and Language Confusion in Multilingual Embedding Inversion Attacks

AAAI 2025technical

Large Language Models (LLMs) are susceptible to malicious influence by cyber attackers through intrusions such as adversarial, backdoor, and embedding inversion attacks. In response, the burgeoning field of LLM Security aims to study and defend against such threats. Thus far, the majority of works i…

2025

Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis

NAACL 2025findings

Language Confusion is a phenomenon where Large Language Models (LLMs) generate text that is neither in the desired language, nor in a contextually appropriate language. This phenomenon presents a critical challenge in text generation by LLMs, often appearing as erratic and unpredictable behavior. We…

2025

Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities

EMNLP 2025

We present the first comprehensive study of Memorization in Multilingual Large Language Models (MLLMs), analyzing 95 languages using models across diverse model scales, architectures, and memorization definitions. As MLLMs are increasingly deployed, understanding their memorization behavior has beco

2023

Cross-view Semantic Alignment for Livestreaming Product Recognition

ICCV 2023poster

Live commerce is the act of selling products online through livestreaming. The customer's diverse demands for online products introduces more challenges to Livestreaming Product Recognition. Previous works are either focus on fashion clothing data or subject to single-modal input, thus inconsistent…

Cited by 4PDFcodeScholar