← Search

Mohan Li

14 accepted papers

2026

AT-Field: Rethinking the Games in Adversarial Training

AAAI 2026technical

Adversarial training is often modeled as a two-player zero-sum game, relying on strong assumptions that limit its practical guidance. In this paper, we instead analyze the interactions between training samples and show that even the fundamental objective—minimizing training loss—may not converge. To

Cited by 0SourcePDFScholar
2026

Federated Learning with Profile Mapping under Distribution Shifts and Drifts

ICLR 2026poster

Federated Learning (FL) enables decentralized model training across clients without sharing raw data, but its performance degrades under real-world data heterogeneity. Existing methods often fail to address distribution shift across clients and distribution drift over time, or they rely on unrealist…

Cited by 0SourcecodeScholar
2026

ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation

ICASSP 2026poster

Large Language Models (LLMs) in multi-turn conversations often suffer from a ``lost-in-conversation'' phenomenon, where they struggle to recover from early incorrect assumptions, particularly when users provide ambiguous initial instructions. We find that standard post-training techniques like Reinf…

Cited by 0SourcePDFScholar
2026

Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks

ICML 2026poster

Triggerable watermarking enables model owners to assert ownership against model extraction attacks. However, most existing approaches require additional training, which limits post-deployment flexibility, and the lack of clear theoretical foundations makes them vulnerable to adaptive attacks. In thi…

Cited by 0SourceScholar
2026

Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation

ICASSP 2026poster

Entropy-based inference methods have gained traction for improving the reliability of Large Language Models (LLMs). However, many existing approaches, such as entropy minimization techniques, suffer from high computational overhead and fail to leverage historical token context effectively. To addres…

Cited by 0SourcePDFScholar
2025

FLUX: Efficient Descriptor-Driven Clustered Federated Learning under Arbitrary Distribution Shifts

NeurIPS 2025poster

Federated Learning (FL) enables collaborative model training across multiple clients while preserving data privacy. Traditional FL methods often use a global model to fit all clients, assuming that clients' data are independent and identically distributed (IID). However, when this assumption does no…

Cited by 0SourceScholar
2025

Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow

AAAI 2025technical

Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e., the generated image descriptions contain objects that do not exist in the image. In this paper, we reveal that object…

2025

PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement

EMNLP 2025

Addressing the intellectual property protection challenges in commercial deployment of large language models (LLMs), existing black-box fingerprinting techniques face dual challenges from incremental fine-tuning erasure and feature-space defense due to their reliance on overfitting high-perplexity t

2024

DiaLoc: An Iterative Approach to Embodied Dialog Localization

CVPR 2024poster

Multimodal learning has advanced the performance for many vision-language tasks. However most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing dialog-based localization approaches assume the availability of entire dialog pr…

Cited by 3SourcePDFScholar
2024

LT-Defense: Searching-free Backdoor Defense via Exploiting the Long-tailed Effect

NeurIPS 2024poster

Language models have shown vulnerability against backdoor attacks, threatening the security of services based on them. To mitigate the threat, existing solutions attempted to search for backdoor triggers, which can be time-consuming when handling a large search space. Looking into the attack process…

Cited by 1SourcePDFScholar
2023

Cumulative Attention Based Streaming Transformer ASR with Internal Language Model Joint Training and Rescoring

ICASSP 2023accepted

This paper presents an approach to improve the performance of streaming Transformer ASR by introducing an internal language model (ILM) as a part of the decoder layers. In the recently pro- posed cumulative attention (CA) based streaming ASR system, only the last or top few decoder layers are equipp…

Cited by 0SourceScholar