← Search

Lu Lin

20 accepted papers

2026

Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks

ICML 2026poster

Interpretable time series deep learning systems are often assessed by checking temporal consistency on explanations, implicitly treating this as evidence of robustness. We show that this assumption can fail: Predictions and explanations can be adversarially decoupled, enabling targeted misclassifica…

Cited by 0SourceScholar
2025

"Why Is There a Tumor?": Tell Me the Reason, Show Me the Evidence

ICML 2025poster

Medical AI models excel at tumor detection and segmentation. However, their latent representations often lack explicit ties to clinical semantics, producing outputs less trusted in clinical practice. Most of the existing models generate either segmentation masks/labels (localizing where without why)…

Cited by 0SourcePDFScholar
2025

AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion Models

ICML 2025poster

Recent advances in diffusion models have significantly enhanced the quality of image synthesis, yet they have also introduced serious safety concerns, particularly the generation of Not Safe for Work (NSFW) content. Previous research has demonstrated that adversarial prompts can be used to generate…

2025

JoPA: Explaining Large Language Model’s Generation via Joint Prompt Attribution

ACL 2025long

Large Language Models (LLMs) have demonstrated impressive performances in complex text generation tasks. However, the contribution of the input prompt to the generated content still remains obscure to humans, underscoring the necessity of understanding the causality between input and output pairs. E…

2025

Monitoring Decoding: Mitigating Hallucination via Evaluating the Factuality of Partial Response during Generation

ACL 2025finding

While large language models have demonstrated exceptional performance across a wide range of tasks, they remain susceptible to hallucinations – generating plausible yet factually incorrect contents. Existing methods to mitigating such risk often rely on sampling multiple full-length generations, whi…

Cited by 0SourcePDFScholar
2025

Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time

EMNLP 2025

Recently, Multimodal Large Language Models (MLLMs) have gained significant attention across various domains. However, their widespread adoption has also raised serious safety concerns.In this paper, we uncover a new safety risk of MLLMs: the output preference of MLLMs can be arbitrarily manipulated

2025

WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response

NAACL 2025findings

The recent breakthrough in large language models (LLMs) such as ChatGPT has revolutionized every industry at an unprecedented pace. Alongside this progress also comes mounting concerns about LLMs’ susceptibility to jailbreaking attacks, which leads to the generation of harmful or unsafe content. Whi…

Cited by 11SourcePDFScholar
2024

Backdoor Contrastive Learning via Bi-level Trigger Optimization

ICLR 2024poster

Contrastive Learning (CL) has attracted enormous attention due to its remarkable capability in unsupervised representation learning. However, recent works have revealed the vulnerability of CL to backdoor attacks: the feature extractor could be misled to embed backdoored data close to an attack targ…

2024

Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

ACL 2024long

Recently, Large Language Models (LLMs) have made significant advancements and are now widely used across various domains. Unfortunately, there has been a rising concern that LLMs can be misused to generate harmful or malicious content. Though a line of research has focused on aligning LLMs with huma…

2024

Graph Adversarial Diffusion Convolution

ICML 2024poster

This paper introduces a min-max optimization formulation for the Graph Signal Denoising (GSD) problem. In this formulation, we first maximize the second term of GSD by introducing perturbations to the graph structure based on Laplacian distance and then minimize the overall loss of the GSD. By solvi…

2024

Jailbreak Open-Sourced Large Language Models via Enforced Decoding

ACL 2024long

Large Language Models (LLMs) have achieved unprecedented performance in Natural Language Generation (NLG) tasks. However, many existing studies have shown that they could be misused to generate undesired content. In response, before releasing LLMs for public access, model developers usually align th…

Cited by 14SourcePDFScholar
2024

Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization

NeurIPS 2024poster

Researchers have been studying approaches to steer the behavior of Large Language Models (LLMs) and build personalized LLMs tailored for various applications. While fine-tuning seems to be a direct solution, it requires substantial computational resources and may significantly affect the utility of…

2024

SCTrans: Multi-scale scRNA-seq Sub-vector Completion Transformer for Gene-selective Cell Type Annotation

IJCAI 2024poster

Cell type annotation is pivotal to single-cell RNA sequencing data (scRNA-seq)-based biological and medical analysis, e.g., identifying biomarkers, exploring cellular heterogeneity, and understanding disease mechanisms. The previous annotation methods typically learn a nonlinear mapping to infer cel…

Cited by 0SourcePDFScholar
2023

A3FL: Adversarially Adaptive Backdoor Attacks to Federated Learning

NeurIPS 2023poster

Federated Learning (FL) is a distributed machine learning paradigm that allows multiple clients to train a global model collaboratively without sharing their local training data. Due to its distributed nature, many studies have shown that it is vulnerable to backdoor attacks. However, existing studi…

2023

FusionRetro: Molecule Representation Fusion via In-Context Learning for Retrosynthetic Planning

ICML 2023poster

Retrosynthetic planning aims to devise a complete multi-step synthetic route from starting materials to a target molecule. Current strategies use a decoupled approach of single-step retrosynthesis models and search algorithms, taking only the product as the input to predict the reactants for each pl…

2022

Communication-Compressed Adaptive Gradient Method for Distributed Nonconvex Optimization

AISTATS 2022poster

Due to the explosion in the size of the training datasets, distributed learning has received growing interest in recent years. One of the major bottlenecks is the large communication cost between the central server and the local workers. While error feedback compression has been proven to be success…

Cited by 19SourcePDFScholar