← Search

Xiaodan Zhang

15 accepted papers

2026

Knowledge-Enhanced Image Captioning with Adaptive Graph-based Multimodal Alignment and LLM

AAAI 2026technical

Image captioning is crucial for multimodal understanding, bridging visual content and natural language. Despite recent advancements in Large Multimodal Models (LMMs), when faced with unseen entities or scenes in the open world, even when attempting to leverage learned knowledge, models still struggl

Cited by 0SourcePDFScholar
2026

MARE: Multimodal Analogical Reasoning for Disease Evolution-Aware Radiology Report Generation

AAAI 2026technical

Radiology report generation from longitudinal medical data is critical for assessing disease progression and automating diagnostic workflows. While recent methods incorporate longitudinal information, they primarily rely on multimodal feature fusion, with limited capacity for explicit disease evolut

Cited by 0SourcePDFScholar
2026

MoEA-Net: Modality-Incremental Expert Aggregation Network for Retinal Prognostic Prediction

AAAI 2026technical

Automated analysis of temporal changes in multimodal retinal images is critical for the prognostic assessment of ophthalmic diseases. Unlike traditional single-timepoint diagnosis, tracking longitudinal changes across multiple imaging modalities introduces significant data bias challenges: (1) Imbal

Cited by 0SourcePDFScholar
2026

State Proficiency-Based Adaptive Fine-Tuning for Offline-to-Online Reinforcement Learning

AAAI 2026technical

In offline-to-online (O2O) reinforcement learning, achieving efficient performance improvement while maintaining training stability remains a critical challenge for effective fine-tuning. Existing O2O methods usually focus on the balance between policy improvement and policy constraint during online

Cited by 0SourcePDFScholar
2025

A New Federated Learning Framework Against Gradient Inversion Attacks

AAAI 2025technical

Federated Learning (FL) aims to protect data privacy by enabling clients to collectively train machine learning models without sharing their raw data. However, recent studies demonstrate that information exchanged during FL is subject to Gradient Inversion Attacks (GIA) and, consequently, a variety…

2025

Emotion-aware Structural Enhancement Graph Auto-Encoder for Rumor Detection

ICASSP 2025accepted

Social media is a key channel for information dissemination, making effective rumor detection essential to mitigate misinformation’s societal impact. Although large language models excel in inference and text generation, they struggle with understanding propagation relationships and complex reasonin…

Cited by 0SourceScholar
2025

MEPNet: Medical Entity-Balanced Prompting Network for Brain CT Report Generation

AAAI 2025technical

The automatic generation of brain CT reports has gained widespread attention, given its potential to assist radiologists in diagnosing cranial diseases. However, brain CT scans involve extensive medical entities, such as diverse anatomy regions and lesions, exhibiting highly inconsistent spatial pat…

2025

Resolution Attack: Exploiting Image Compression to Deceive Deep Neural Networks

ICLR 2025poster

Model robustness is essential for ensuring the stability and reliability of machine learning systems. Despite extensive research on various aspects of model robustness, such as adversarial robustness and label noise robustness, the exploration of robustness towards different resolutions, remains les…

2025

Semantic Reshuffling with LLM and Heterogeneous Graph Auto-Encoder for Enhanced Rumor Detection

COLING 2025main

Social media is crucial for information spread, necessitating effective rumor detection to curb misinformation’s societal effects. Current methods struggle against complex propagation influenced by bots, coordinated accounts, and echo chambers, which fragment information and increase risks of misjud…

Cited by 1SourcePDFScholar
2024

General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level

ICASSP 2024accepted

The social biases and unwelcome stereotypes revealed by pretrained language models are becoming obstacles to their application. Compared to numerous debiasing methods targeting word level, there has been relatively less attention on biases present at phrase level, limiting the performance of debiasi…

Cited by 0SourceScholar
2024

See Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning

EMNLP 2024finding

Brain CT report generation is significant to aid physicians in diagnosing cranial diseases.Recent studies concentrate on handling the consistency between visual and textual pathological features to improve the coherence of report.However, there exist some challenges: 1) Redundant visual representing…

2023

Adversarial Text Generation by Search and Learning

EMNLP 2023long findings

Recent research has shown that evaluating the robustness of natural language processing models using textual attack methods is significant. However, most existing text attack methods only use heuristic replacement strategies or language models to generate replacement words at the word level. The bli…

Cited by 0SourceScholar
2023

Granularity Matters: Pathological Graph-driven Cross-modal Alignment for Brain CT Report Generation

EMNLP 2023long main

The automatic Brain CT reports generation can improve the efficiency and accuracy of diagnosing cranial diseases. However, current methods are limited by 1) coarse-grained supervision: the training data in image-text format lacks detailed supervision for recognizing subtle abnormalities, and 2) coup…

Cited by 0SourceScholar
2022

Cross-modal Contrastive Attention Model for Medical Report Generation

COLING 2022main

Medical report automatic generation has gained increasing interest recently as a way to help radiologists write reports more efficiently. However, this image-to-text task is rather challenging due to the typical data biases: 1) Normal physiological structures dominate the images, with only tiny abno…

2017

Delving Into Salient Object Subitizing and Detection

ICCV 2017poster

Subitizing (i.e., instant judgement on the number) and detection of salient objects are human inborn abilities. These two tasks influence each other in the human visual system. In this paper, we delve into the complementarity of these two tasks. We propose a multi-task deep neural network with weigh…

Cited by 68PDFScholar