← Search

Kaile Wang

9 accepted papers

2026

Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models

ICML 2026poster

As frontier AI systems become increasingly capable, concerns about deceptive behaviors have intensified. Unlike hallucinations, which stem from capability limitations, deception involves strategically misleading responses despite correct internal representations. While prior work has primarily studi…

Cited by 0SourceScholar
2026

FedRD: Reducing Divergences for Generalized Federated Learning via Heterogeneity-aware Parameter Guidance

ICASSP 2026oral

Heterogeneous federated learning (HFL) aims to ensure effective and privacy-preserving collaboration among different entities. As newly joined clients require significant adjustments and additional training to align with the existing system, the problem of generalizing federated learning models to u…

Cited by 0SourcePDFScholar
2025

A Task-Oriented Real-Time and Robust Feature Compression and Selection Method in Collaborative Intelligence System

ICASSP 2025accepted

The emerging autonomous driving has stringent requirements for latency and reliability. In this paper, we propose a task-oriented real-time and robust feature compression and selection method in collaborative intelligence system. Our design, consisting of a three-dimensional channel compression (TDC…

Cited by 0SourceScholar
2025

InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback

NeurIPS 2025spotlight

As multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: \textbf{\textit{What essential capabilities are still missing? }} A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving…

Cited by 0SourceScholar
2025

Language Models Resist Alignment: Evidence From Data Compression

ACL 2025long

Large language models (LLMs) may exhibit unintended or undesirable behaviors. Recent works have concentrated on aligning LLMs to mitigate harmful outputs. Despite these efforts, some anomalies indicate that even a well-conducted alignment process can be easily circumvented, whether intentionally or…

2025

PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

ACL 2025long

In this work, we introduce the PKU-SafeRLHF dataset, designed to promote research on safety alignment in large language models (LLMs). As a sibling project to SafeRLHF and BeaverTails, we separate annotations of helpfulness and harmlessness for question-answering pairs, providing distinct perspectiv…

2025

Reward Generalization in RLHF: A Topological Perspective

ACL 2025finding

Existing alignment methods share a common topology of information flow, where reward information is collected from humans, modeled with preference learning, and used to tune language models. However, this shared topology has not been systematically characterized, nor have its alternatives been thoro…

Cited by 0SourcePDFScholar
2025

Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback

NeurIPS 2025poster

Multimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is critical to explore how to fine-tune MLLMs to preserve capab…

Cited by 0SourceScholar
2025

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction

AAAI 2025technical

The rapid advancement of large language models (LLMs) has led to significant improvements in their capabilities, but also to increased concerns about their alignment with human values and intentions. Current alignment strategies, including adaptive training and inference-time methods, have demonstra…