← Search

Xiaoming Zhang

19 accepted papers

2026

Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MAS

AAAI 2026technical

Large language model-based multi-agent systems (LLM-MAS) effectively accomplish complex and dynamic tasks through inter-agent communication, but this reliance introduces substantial safety vulnerabilities. Existing attack methods targeting LLM-MAS either compromise agent internals or rely on direct

Cited by 10SourcePDFScholar
2026

LLM-MatLogic: Executable Exchange Contracts for Knowledge-Graph Query Answering with Scoped Negation

ICML 2026poster

LLM-to-KG systems frequently fail on exclusion-rich questions because natural-language negation is both scope-sensitive and evidence-dependent: it may constrain only one subgoal/branch and only certain supporting paths, yet such attachment is rarely explicit in text. We propose the Executable Exchan…

Cited by 0SourceScholar
2026

MergeDNA: Context-Aware Genome Modeling with Dynamic Tokenization Through Token Merging

AAAI 2026technical

Modeling genomic sequences faces two unsolved challenges: the information density varies widely across different regions, while there is no clearly defined minimum vocabulary unit. Relying on either four primitive bases or independently designed DNA tokenizers, existing approaches with naive masked

Cited by 0SourcePDFScholar
2026

TrinityDNA: A Bio-Inspired Foundational Model for Efficient Long-Sequence DNA Modeling

AAAI 2026technical

The modeling of genomic sequences presents unique challenges due to their long length and structural complexity. Traditional sequence models struggle to capture long-range dependencies and biological features inherent in DNA. In this work, we propose TrinityDNA, a novel DNA foundational model design

Cited by 0SourcePDFScholar
2025

Collaborative Evolution: Multi-Round Learning Between Large and Small Language Models for Emergent Fake News Detection

AAAI 2025technical

The proliferation of fake news on social media platforms has exerted a substantial influence on society, leading to discernible impacts and deleterious consequences. Conventional deep learning methodologies employing small language models (SLMs) suffer from the necessity for extensive supervised tra…

Cited by 0SourcePDFScholar
2025

Diff-Shadow: Global-guided Diffusion Model for Shadow Removal

AAAI 2025technical

We propose Diff-Shadow, a global-guided diffusion model for high-quality shadow removal. Previous transformer-based approaches can utilize global information to relate shadow and non-shadow regions but are limited in their synthesis ability and recover images with obvious boundaries. In contrast, di…

2025

LlmFixer: Fix the Helpfulness of Defensive Large Language Models

EMNLP 2025

Defense strategies of large language models besides alignment are introduced to defend against jailbreak attacks, and they have managed to decrease the success rate of jailbreak attacks. However, these defense strategies weakened the helpfulness of large language models. In this work, we propose a u

Cited by 0SourcePDFScholar
2025

M-MoE: Mixture of Mixture-of-Expert Model for CTC-based Streaming Multilingual ASR

ICASSP 2025accepted

The Mixture-of-Expert (MoE) structure has been effectively utilized in multilingual ASR tasks. However, the potential of external language information remains underutilized. In this paper, we introduce the Mixture of MoE (M-MoE) structure, featuring multiple language-specific MoEs and a language-unk…

Cited by 0SourceScholar
2025

TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis

EMNLP 2025

Multimodal Sentiment Analysis (MSA) with missing modalities has attracted increasing attention recently. While current Transformer-based methods leverage dense text information to maintain model robustness, their quadratic complexity hinders efficient long-range modeling and multimodal fusion. To th

2025

What Is a Good Question? Assessing Question Quality via Meta-Fact Checking

AAAI 2025technical

Knowledge-based questions are typically employed to evaluate LLM's knowledge boundaries; meanwhile, numerous studies focus on question generation as a means to enhance the capabilities of both models and individuals. However, there is a lack of in-depth exploration about what constitutes a good ques…

2024

Enhancing Cognitive Diagnosis Using Un-interacted Exercises: A Collaboration-Aware Mixed Sampling Approach

AAAI 2024technical

Cognitive diagnosis is a crucial task in computer-aided education, aimed at evaluating students' proficiency levels across various knowledge concepts through exercises. Current models, however, primarily rely on students' answered exercises, neglecting the complex and rich information contained in u…

2024

Hyperbolic Graph Neural Network for Temporal Knowledge Graph Completion

COLING 2024main

Temporal Knowledge Graphs (TKGs) represent a crucial source of structured temporal information and exhibit significant utility in various real-world applications. However, TKGs are susceptible to incompleteness, necessitating Temporal Knowledge Graph Completion (TKGC) to predict missing facts. Exist…

2024

MDS: A Fine-Grained Dataset for Multi-Modal Dialogue Summarization

COLING 2024main

Due to the explosion of various dialogue scenes, summarizing the dialogue into a short message has drawn much attention recently. In the multi-modal dialogue scene, people tend to use tone and body language to illustrate their intentions. While traditional dialogue summarization has predominantly fo…

2024

Reinforced Adaptive Knowledge Learning for Multimodal Fake News Detection

AAAI 2024technical

Nowadays, detecting multimodal fake news has emerged as a foremost concern since the widespread dissemination of fake news may incur adverse societal impact. Conventional methods generally focus on capturing the linguistic and visual semantics within the multimodal content, which fall short in effe…

Cited by 27SourcePDFScholar
2022

Group-based Interleaved Pipeline Parallelism for Large-scale DNN Training

ICLR 2022poster

The recent trend of using large-scale deep neural networks (DNN) to boost performance has propelled the development of the parallel pipelining technique for efficient DNN training, which has resulted in the development of several prominent pipelines such as GPipe, PipeDream, and PipeDream-2BW. Howev…

2022

Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization

AAAI 2022technical

Multimodal summarization with multimodal output (MSMO) generates a summary with both textual and visual content. Multimodal news report contains heterogeneous contents, which makes MSMO nontrivial. Moreover, it is observed that different modalities of data in the news report correlate hierarchically…

2021

Matching Distributions between Model and Data: Cross-domain Knowledge Distillation for Unsupervised Domain Adaptation

ACL 2021long

Unsupervised Domain Adaptation (UDA) aims to transfer the knowledge of source domain to the unlabeled target domain. Existing methods typically require to learn to adapt the target model by exploiting the source data and sharing the network architecture across domains. However, this pipeline makes t…

Cited by 22SourcePDFScholar