← Search

Zheli Liu

6 accepted papers

2026

DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning

ICLR 2026poster

Real-world large language model deployments (e.g., conversational AI systems, code generation assistants) naturally generate abundant implicit user dissatisfaction (DSAT) signals, as users iterate toward better answers through refinements, corrections, and expressed preferences, while explicit satis…

Cited by 0SourceScholar
2026

TSFAdv: Frequency-Guided Black-Box Adversarial Attacks on Time Series Forecasting

ICML 2026poster

While deep neural network-based long-term time series forecasting (LTSF) has become indispensable for critical infrastructures such as smart grids and IoT platforms, the deployment of these models as black-box APIs introduces severe security vulnerabilities that remain largely underexplored. In this…

Cited by 0SourceScholar
2025

Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models

ICLR 2025poster

Backdoor unalignment attacks against Large Language Models (LLMs) enable the stealthy compromise of safety alignment using a hidden trigger while evading normal safety auditing. These attacks pose significant threats to the applications of LLMs in the real-world Large Language Model as a Service (LL…

2025

Prompt-Guided Internal States for Hallucination Detection of Large Language Models

ACL 2025long

Large Language Models (LLMs) have demonstrated remarkable capabilities across a variety of tasks in different domains. However, they sometimes generate responses that are logically coherent but factually incorrect or misleading, which is known as LLM hallucinations. Data-driven supervised methods tr…

2025

Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark

EMNLP 2025

Embedding-as-a-Service (EaaS) has emerged as a successful business pattern but faces significant challenges related to various forms of copyright infringement, particularly the API misuse and model extraction attacks. Various studies have proposed backdoor-based watermarking schemes to protect the c

2024

BadActs: A Universal Backdoor Defense in the Activation Space

ACL 2024findings

Backdoor attacks pose an increasingly severe security threat to Deep Neural Networks (DNNs) during their development stage. In response, backdoor sample purification has emerged as a promising defense mechanism, aiming to eliminate backdoor triggers while preserving the integrity of the clean conten…