← Search

Kejing Yin

7 accepted papers

2026

Learning Self-Critiquing Mechanisms for Region-Guided Chest X-Ray Report Generation

ICLR 2026poster

Automatic radiology reporting assists radiologists in diagnosing abnormalities in radiology images, where grounding the automatic diagnosis with abnormality locations is important for the report interpretability. However, existing supervised-learning methods could lead to learning the superficial st…

Cited by 0SourceScholar
2026

MASAM: Multimodal Adaptive Sharpness-Aware Minimization for Heterogeneous Data Fusion

ICLR 2026poster

Multimodal learning requires integrating heterogeneous modalities, such as structured records, visual imagery, and temporal signals. It has been revealed that this heterogeneity causes modality encoders to converge at different rates, making the multimodal learning imbalanced. We empirically observe…

Cited by 0SourceScholar
2025

CURV: Coherent Uncertainty-Aware Reasoning in Vision-Language Models for X-Ray Report Generation

NeurIPS 2025poster

Vision-language models have been explored for radiology report generation with promising results. Yet, uncertainty elaborated in findings and the reasoning process for reaching clinical impressions are seldom explicitly modeled, reducing the clinical accuracy and trustworthiness of the generated rep…

Cited by 0SourceScholar
2025

Multimodal Disease Progression Modeling via Spatiotemporal Disentanglement and Multiscale Alignment

NeurIPS 2025spotlight

Longitudinal multimodal data, including electronic health records (EHR) and sequential chest X-rays (CXRs), is critical for modeling disease progression, yet remains underutilized due to two key challenges: (1) redundancy in consecutive CXR sequences, where static anatomical regions dominate over cl…

Cited by 0SourceScholar
2024

Addressing Asynchronicity in Clinical Multimodal Fusion via Individualized Chest X-ray Generation

NeurIPS 2024poster

Integrating multi-modal clinical data, such as electronic health records (EHR) and chest X-ray images (CXR), is particularly beneficial for clinical prediction tasks. However, in a temporal setting, multi-modal data are often inherently asynchronous. EHR can be continuously collected but CXR is gene…

2024

DrFuse: Learning Disentangled Representation for Clinical Multi-Modal Fusion with Missing Modality and Modal Inconsistency

AAAI 2024technical

The combination of electronic health records (EHR) and medical images is crucial for clinicians in making diagnoses and forecasting prognoses. Strategically fusing these two data modalities has great potential to improve the accuracy of machine learning models in clinical prediction tasks. However,…

2021

SWIFT: Scalable Wasserstein Factorization for Sparse Nonnegative Tensors

AAAI 2021technical

Existing tensor factorization methods assume that the input tensor follows some specific distribution (i.e. Poisson, Bernoulli, and Gaussian), and solve the factorization by minimizing some empirical loss functions defined based on the corresponding distribution. However, it suffers from several dra…

Cited by 18SourcePDFScholar