← Search

Xiaoyan Cai

7 accepted papers

2026

RefleXNet: Targeted Self-Reflection for Accurate Chest X-ray Reporting

AAAI 2026technical

Automated interpretation and reporting of chest X-rays (CXRs) hold significant promise in reducing diagnostic errors and supporting radiologists under heavy clinical workloads. However, existing methods typically rely on global visual features and token-level supervision, limiting their sensitivity

Cited by 0SourcePDFScholar
2025

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning

ACL 2025finding

Direct Preference Optimization (DPO) has gained significant attention for its simplicity and computational efficiency in aligning large language models (LLMs). Recent advancements have extended DPO to multimodal scenarios, achieving strong performance. However, traditional DPO relies on binary prefe…

Cited by 0SourcePDFScholar
2025

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models

ICASSP 2025accepted

The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks. However, current MLLM often struggles to effectively address fine-grained multi-modal challenges. We argue that this lim…

Cited by 0SourceScholar
2025

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

AAAI 2025technical

Existing Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their effectiveness in capturing coarse-grained semantic correlations. However, their capability for fine-grained understanding, which is critical for many…

Cited by 2SourcePDFScholar
2024

MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task Learning

EMNLP 2024main

The growing demand for larger-scale models in the development of Large Language Models (LLMs) poses challenges for efficient training within limited computational resources. Traditional fine-tuning methods often exhibit instability in multi-task learning and rely heavily on extensive training resour…

Cited by 1SourcePDFScholar
2024

Self-Renewal Prompt Optimizing with Implicit Reasoning

EMNLP 2024finding

The effectiveness of Large Language Models (LLMs) relies on their capacity to understand instructions and generate human-like responses. However, aligning LLMs with complex human preferences remains a significant challenge due to the potential misinterpretation of user prompts. Current methods for a…

Cited by 0SourcePDFScholar
2022

An Adaptive Logical Rule Embedding Model for Inductive Reasoning over Temporal Knowledge Graphs

EMNLP 2022main

Temporal knowledge graphs (TKGs) extrapolation reasoning predicts future events based on historical information, which has great research significance and broad application value. Existing methods can be divided into embedding-based methods and logical rule-based methods. Embedding-based methods rel…

Cited by 20SourcePDFScholar