← Search

Haodong Zhao

17 accepted papers

2026

GhostEI-Bench: Do Mobile Agent Resilience to Environmental Injection in Dynamic On-Device Environments?

ICLR 2026poster

Vision-Language Models (VLMs) are increasingly deployed as autonomous agents to navigate mobile Graphical User Interfaces (GUIs). However, their operation within dynamic on-device ecosystems, which include notifications, pop-ups, and inter-app interactions, exposes them to a unique and underexplored…

Cited by 0SourceScholar
2026

IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding

CVPR 2026

Recent advances in vision-language models (VLMs) have significantly enhanced the visual grounding task, which involves locating objects in an image based on natural language queries. Despite these advancements, the security of VLM-based grounding systems has not been thoroughly investigated. This pa

Cited by 0SourcecodeScholar
2026

LLM DNA: Tracing Model Evolution via Functional Representations

ICLR 2026oral

The explosive growth of large language models (LLMs) has created a vast but opaque landscape: millions of models exist, yet their evolutionary relationships through fine-tuning, distillation, or adaptation are often undocumented or unclear, complicating LLM management. Existing methods are limited b…

Cited by 0SourcecodeScholar
2025

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models

ACL 2025finding

Existing multi-objective preference alignment methods for large language models (LLMs) face limitations: (1) the inability to effectively balance various preference dimensions, and (2) reliance on auxiliary reward/reference models introduces computational complexity. To address these challenges, we…

2025

Conditional-Balanced Adversarial Delta Tuning for Cross-Domain Implicit Discourse Relation Recognition

ICASSP 2025accepted

Implicit discourse relation recognition (IDRR) is faced with a domain dilemma. Recent studies have achieved breakthroughs in standard datasets, while they are not appropriate in domains with insufficient data, such as bio-medicine. In this paper, we treat this problem as a cross-domain IDRR task, wh…

Cited by 0SourceScholar
2025

Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models

ICML 2025poster

Multi-Modal Large Language Models (MLLMs) have exhibited remarkable performance on various vision-language tasks such as Visual Question Answering (VQA). Despite accumulating evidence of privacy concerns associated with task-relevant content, it remains unclear whether MLLMs inadvertently memorize p…

2025

When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning

EMNLP 2025

Large reasoning models (LRMs) achieve remarkable performance via long reasoning chains, but often incur excessive computational overhead due to redundant reasoning, especially on simple tasks. In this work, we systematically quantify the upper bounds of LRMs under both Long-Thinking and No-Thinking

Cited by 0SourcePDFScholar
2024

Revisiting the Information Capacity of Neural Network Watermarks: Upper Bound Estimation and Beyond

AAAI 2024technical

To trace the copyright of deep neural networks, an owner can embed its identity information into its model as a watermark. The capacity of the watermark quantify the maximal volume of information that can be verified from the watermarked model. Current studies on capacity focus on the ownership veri…

Cited by 4SourcePDFScholar
2024

UOR: Universal Backdoor Attacks on Pre-trained Language Models

ACL 2024findings

Task-agnostic and transferable backdoors implanted in pre-trained language models (PLMs) pose a severe security threat as they can be inherited to any downstream task. However, existing methods rely on manual selection of triggers and backdoor representations, hindering their effectiveness and unive…

Cited by 20SourcePDFScholar
2023

FedPrompt: Communication-Efficient and Privacy-Preserving Prompt Tuning in Federated Learning

ICASSP 2023accepted

Federated learning (FL) has enabled global model training on decentralized data in a privacy-preserving way. However, for tasks that utilize pre-trained language models (PLMs) with massive parameters, there are considerable communication costs. Prompt tuning, which tunes soft prompts without modifyi…

Cited by 0SourceScholar
2023

Infusing Hierarchical Guidance into Prompt Tuning: A Parameter-Efficient Framework for Multi-level Implicit Discourse Relation Recognition

ACL 2023long

Multi-level implicit discourse relation recognition (MIDRR) aims at identifying hierarchical discourse relations among arguments. Previous methods achieve the promotion through fine-tuning PLMs. However, due to the data scarcity and the task gap, the pre-trained feature space cannot be accurately tu…

2023

Is Continuous Prompt a Combination of Discrete Prompts? Towards a Novel View for Interpreting Continuous Prompts

ACL 2023findings

The broad adoption of continuous prompts has brought state-of-the-art results on a diverse array of downstream natural language processing (NLP) tasks. Nonetheless, little attention has been paid to the interpretability and transferability of continuous prompts. Faced with the challenges, we investi…

Cited by 6SourcePDFScholar
2023

PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models

AAAI 2023technical

The huge training overhead, considerable commercial value, and various potential security risks make it urgent to protect the intellectual property (IP) of Deep Neural Networks (DNNs). DNN watermarking has become a plausible method to meet this need. However, most of the existing watermarking scheme…

Cited by 52SourcePDFScholar